System

The system addresses high costs and risks in conventional marketing by generating personalized content through image, language, and voice generation, enhancing consumer engagement.

JP2026018043APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119104
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional marketing methods face high contract fees, filming costs, and crime risks when hiring influencers, and consumers lack personalized information, making it difficult to attract their interest.

Method used

A system incorporating image, language, and voice generation means, along with data analysis and feedback mechanisms, to create highly personalized marketing content efficiently and risk-free.

Benefits of technology

The system generates personalized marketing content that reduces costs and risks, providing efficient and effective communication with consumers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018043000001_ABST
    Figure 2026018043000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: image generation means; language generation means; audio generation means; input means for accepting user input; and display means for displaying user-generated content.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] With conventional marketing methods, companies often face issues such as high contract fees, filming costs, and crime risks when hiring influencers. Another issue is that consumers are not provided with enough personalized information, making it difficult to attract their interest. The present invention aims to reduce these costs and risks and provide a new means of communication for providing highly personalized information to each individual consumer. [Means for solving the problem]

[0005] The present invention provides a system including an image generation means, a language generation means, a voice generation means, an input means for accepting user input, and a display means for displaying generated content to the user. Furthermore, by providing a data analysis means for analyzing the input user information and converting it into a format suitable for the content to be generated, and a means for accepting feedback, correcting and regenerating the generated content, companies can obtain an efficient and risk-free marketing tool and provide highly personalized information to consumers.

[0006] "Image generation means" refers to a device or program that has the function of generating visual content based on data provided by a user.

[0007] A "language generation means" is a device or program that has the function of analyzing data provided by a user and generating sentences or text based on that data.

[0008] The "voice generation means" is a device or program that has the function of converting the generated text into voice and generating realistic speech.

[0009] An "input means" is a device or program that provides an interface for a user to input information or feedback.

[0010] The "display means" is a device or program that visually presents the generated content to the user.

[0011] The "data analysis means" is a device or program that has the function of analyzing data input by a user and converting it into a format suitable for the content to be generated.

[0012] The "modifying and regenerating means" refers to a device or program that has the function of modifying the generated content based on feedback from users and regenerating it if necessary. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] The present invention is embodied in a system including an image generating means, a language generating means, a voice generating means, an input means, and a display means, which supports corporate marketing and provides personalized information to consumers.

[0035] Server Processing

[0036] The server receives data provided by users and companies and creates content using the following three generation methods.

[0037] 1. Image generation method

[0038] The server generates visual content based on product information and target audience data provided by the user. The image generation means generates high-resolution images, for example, using a generative AI model.

[0039] Example: A server generates promotional images for "sportswear." The data includes the color, design, and background image of the clothing, and the server automatically generates an appropriate image based on this data.

[0040] 2. Language generation means

[0041] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in the language expression requested by the user.

[0042] Example: The server generates text describing the features of "new sportswear," such as "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[0043] 3. Voice Generation Method

[0044] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[0045] Example: Generate an audio file that reads the description of the generated sportswear in a professional voice.

[0046] Terminal handling

[0047] The terminal provides an interface for smooth communication between the user and the server.

[0048] 1. Data entry and customization

[0049] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[0050] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "unisex moisture-wicking, quick-drying sportswear, price, main target demographic").

[0051] 2. Product Labeling

[0052] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[0053] Example: Check the promotional sportswear images, introductory text, and audio files displayed on the device.

[0054] User operations

[0055] The user performs the following operations through the terminal.

[0056] 1. Input and Request Submission

[0057] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[0058] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[0059] 2. Feedback and correction requests

[0060] The user can check the generated content and make corrections or additions as needed.

[0061] Example: A user reviews the generated images and text and submits correction requests such as "make the image background a little brighter" or "make the text tone more casual."

[0062] In this way, the system of the present invention can efficiently generate and provide highly personalized marketing content that meets user requests.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "latest running shoes," along with their features, target age group, and lifestyle.

[0066] Step 2:

[0067] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[0068] Step 3:

[0069] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[0070] Step 4:

[0071] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[0072] Step 5:

[0073] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme "latest running shoes."

[0074] Step 6:

[0075] The server generates text content using language generation tools. It uses a generative AI model to create product descriptions and reviews. For example, it generates a description such as, "These running shoes are lightweight and have excellent cushioning."

[0076] Step 7:

[0077] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[0078] Step 8:

[0079] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[0080] Step 9:

[0081] The device displays the result to the user. The device displays the result received from the server on the screen so that the user can review it. The screen displays an image of the promotional running shoes, introductory text, and an option to play the audio file.

[0082] Step 10:

[0083] Users can review the generated content and provide feedback and correction requests. Users can check the generated content, enter feedback such as "make the background of the image brighter" or "make the tone of the text more casual," and submit it.

[0084] Step 11:

[0085] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the image of the running shoes with a lighter background and adjusting the tone of the text.

[0086] Step 12:

[0087] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[0088] Step 13:

[0089] The device stores the approved content and prepares it for distribution. It stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[0090] Example 1

[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0092] Conventional marketing support systems lack adequate means for efficiently generating highly personalized content that meets user requests. Furthermore, they lack the functionality to incorporate user feedback into the generated content, making it difficult to quickly and accurately revise the content. The present invention aims to solve these problems and more effectively support corporate marketing activities.

[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0094] In this invention, the server includes an image generation means, a language generation means, and a voice generation means. This enables efficient generation of personalized marketing content in response to user requests. The server further includes a means for generating visual content using a generative AI model, a means for generating text using the generative AI model, a means for generating voice based on the generated text, a means for integrating the generated content and sending it to the user, and a means for receiving user feedback on the generated content, correcting it, and regenerating it. This enables content to be quickly and accurately corrected based on user feedback, thereby more effectively supporting corporate marketing activities.

[0095] "Image generation means" is a function for generating visual content (photographs, illustrations, etc.) based on information provided by the user.

[0096] "Language generation means" is a function for generating relevant text based on information and instructions provided by the user.

[0097] The "speech generation means" is a function for generating speech based on the generated text.

[0098] An "input means" is an interface through which a user inputs information into the system.

[0099] The "display means" is an interface for displaying the generated content (images, text, audio) to the user.

[0100] A "generative AI model" is an algorithm or platform that uses machine learning to generate content (images, text, audio, etc.).

[0101] A "prompt" is text containing commands or instructions that are input to a generative AI model to generate specific content.

[0102] The "data analysis means" is a function that analyzes input user information and converts it into a format suitable for the content to be generated.

[0103] The "feedback means" is a function that receives feedback from the user and modifies or regenerates the generated content.

[0104] "Integration means" is a function that combines images, text, and audio into a single piece of content.

[0105] "User feedback" refers to opinions and correction requests from users regarding the generated content.

[0106] The present invention is an advanced content generation system for effectively supporting corporate marketing activities. This system includes image generation means, language generation means, voice generation means, input means, and display means. The following describes in detail the use of specific hardware and software, as well as methods for data processing and data calculation.

[0107] Hardware and Software

[0108] The system is implemented using specific hardware and software as follows:

[0109] Server: A server with high-performance data processing capabilities is used to receive, analyze, generate content, and integrate data.

[0110] Terminal: The user operates the device using a PC, tablet, or other device. This terminal communicates with the server and provides an interface to the user.

[0111] Generative AI models: AI models such as DALL-E and MidJourney are used for image generation, GPT-4 for text generation, and Google Text-to-Speech and Amazon Polly for voice generation.

[0112] Image Generation Means

[0113] The server generates visual content based on product information and target audience data provided by the user. Specifically, it generates images using prompts from a generative AI model (e.g., DALL-E).

[0114] Example: "Generate a scene of a man and woman running in sportswear, moisture-wicking, with a bright background."

[0115] language generation means

[0116] The server has the functionality to generate text about products and services. It generates text using prompt sentences from a generative AI model (e.g., GPT-4).

[0117] Example: "Please explain the features and benefits of your new sportswear in 200 words or less."

[0118] Voice generation means

[0119] The server generates speech based on the generated text, and uses a speech generation AI (e.g., Google Text-to-Speech) to convert the generated text into a voice with a narration that sounds close to a human voice.

[0120] Example: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made with the latest moisture-wicking, quick-drying materials for maximum comfort during exercise.'"

[0121] Data processing and calculation

[0122] The server analyzes the received user data and converts it into a format suitable for the generative AI model. This includes classifying and filtering the data. It also modifies and regenerates the generated content based on feedback provided by the user. To do this, it inputs the prompt sentence into the generative AI model again to generate new content.

[0123] User operations

[0124] The user uses the terminal to input the necessary information through the interface and send a content generation request to the server. The generated content is displayed on the terminal, and the user can check it and send feedback if necessary. The terminal then sends the user's feedback to the server, and the server regenerates the content based on that feedback.

[0125] In this way, the system of the present invention can efficiently generate personalized marketing content in response to user requests and support the marketing activities of companies.

[0126] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0127] Step 1:

[0128] The user uses a terminal to input the necessary data, including product information (e.g., "sportswear"), target audience information (e.g., "runners in their 30s"), and other details required for content creation. The input information is then sent to the server.

[0129] Input: Product information, target audience information

[0130] Output: Data sent to the server

[0131] Step 2:

[0132] The terminal transmits the data entered by the user to the server, which receives the data and prepares it for analysis.

[0133] Input: Data from the user

[0134] Output: Data sent to the server

[0135] Step 3:

[0136] The server analyzes the received data, which may include categorizing and filtering the data, and converts it into a format suitable for the generative AI model.

[0137] Input: Data from the terminal

[0138] Output: Parsed data

[0139] Step 4:

[0140] The server activates the image generation means based on the analyzed data, sending specific prompts to the generative AI model (e.g., DALL-E) to generate visual content.

[0141] Input: Parsed data, prompt statement

[0142] Output: The generated image

[0143] Specifically, the prompt text includes, "Generate a scene of a man and woman running in sportswear, moisture-wicking and quick-drying, with a bright background."

[0144] Step 5:

[0145] The server then activates a language generation mechanism based on the analyzed data, sending specific prompts to a generative AI model (e.g., GPT-4) to generate text.

[0146] Input: Parsed data, prompt statement

[0147] Output: The generated text

[0148] Specific prompts include, "Please explain the features and benefits of your new sportswear in 200 words or less."

[0149] Step 6:

[0150] The server activates a voice generation means based on the generated text, sends the text to a voice generation AI (e.g., Google Text-to-Speech), and generates an audio file.

[0151] Input: Generated text

[0152] Output: Generated audio file

[0153] Specifically, the instructions include: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made from the latest moisture-wicking, quick-drying material to maximize comfort during exercise.'"

[0154] Step 7:

[0155] The server integrates the generated images, text, and audio, and sends the integrated product to the terminal in a format that is easy for the user to view.

[0156] Input: Generated images, text, audio

[0157] Output: The integrated product

[0158] Step 8:

[0159] The server sends the integrated product to the terminal, which receives the product and displays it to the user.

[0160] Input: The integrated product

[0161] Output: The product sent to the terminal

[0162] Step 9:

[0163] The user uses the device to review the generated content and, if desired, enter feedback, which may include requests for image corrections or text adjustments.

[0164] Input: Generated content

[0165] Output: Feedback

[0166] Step 10:

[0167] The terminal sends feedback from the user to the server, which receives this feedback and prepares to modify and regenerate the content.

[0168] Input: User feedback

[0169] Output: Feedback sent to the server

[0170] Step 11:

[0171] The server modifies and regenerates the content based on the feedback, which may involve re-analyzing the data and adjusting the prompts.

[0172] Input: Feedback

[0173] Output: Modified and regenerated content

[0174] Through this process, the system can efficiently generate highly personalized marketing content that meets user requests and supports a company's marketing activities.

[0175] (Application example 1)

[0176] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0177] In today's online shopping environment, consumers cannot actually try on products, making it difficult to experience the feel and comfort of the products. Furthermore, it takes time and effort for consumers to select the right product for themselves. Therefore, there is a need for an easier and more intuitive way for consumers to try on products and understand information about them.

[0178] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0179] In this invention, the server includes an image generation means, a language generation means, a voice generation means, an input means for receiving user input, a display means for displaying the generated content to the user, a means for combining a user image with a product image, and a means for allowing the user to check the generated content, thereby enabling consumers to virtually try on products from the comfort of their own homes and check the product features visually and audibly.

[0180] The "image generation means" is a means for generating visual content based on data and product information provided by the user.

[0181] The "language generation means" is a means for generating text information about products and services.

[0182] The "voice generation means" is a means for generating voice based on the generated text, and creates narration that resembles a human voice.

[0183] The "input means for accepting user input" is an interface for the user to input necessary data.

[0184] The "display means for displaying the generated content to the user" is an interface for displaying the generated content such as images, text, and audio to the user.

[0185] The "means for combining a user image with a product image" refers to a means for combining an acquired user image with a product image, making it appear as if the user is trying on the product.

[0186] The "means for allowing the user to confirm the generated content" refers to a means for presenting the generated content to the user and receiving confirmation of the content and feedback therefrom.

[0187] In order to implement this invention, it is necessary to construct a system that includes an image generation means, a language generation means, a voice generation means, an input means for accepting user input, a display means for displaying the generated content to the user, a means for synthesizing the user image with a product image, and a means for allowing the user to confirm the generated content.

[0188] First, the server receives and analyzes data provided by the user. Specifically, it acquires the user's image data and product information and generates visual content based on this. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate images. This model is capable of generating high-resolution images. For example, a user can take a picture of themselves using a smartphone and send it to the server.

[0189] Next, the server uses language generation tools based on the product information to create a description of the product. A generative AI model (e.g., GPT-3) is used for language generation. For example, when the server generates text to describe the features of "moisture-wicking, quick-drying sportswear," it generates information such as, "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[0190] In addition, the server generates audio based on the generated text. The audio is generated using a speech synthesis AI model (e.g., WaveNet). This model allows for the creation of narration that sounds close to a human voice. Based on the generated description, an audio file introducing the product is generated.

[0191] The user's device displays these artifacts (images, text, and audio). Users can virtually try on items from the comfort of their own home using their smartphone, smart glasses, or head-mounted display. For example, a user can select a particular sportswear item, view its image and description, and listen to an audio description.

[0192] For example:

[0193] 1. The user enters product information for sportswear.

[0194] 2. The server receives image data and product information from the user and creates a synthetic image using a generative AI model.

[0195] 3. Automatically generate product descriptions using a language generation AI model.

[0196] 4. Convert the text generated by the speech synthesis AI model into speech.

[0197] 5. The device displays the generated content to the user, allowing them to virtually try on the clothes at home.

[0198] 6. An example of a prompt sentence is, "Please explain the characteristics of moisture-wicking, quick-drying sportswear."

[0199] Through these steps, consumers can virtually try on products and visually and audibly check the product's features from the comfort of their own home, helping them make product selections more easily and intuitively.

[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0201] Step 1:

[0202] The user takes a picture of themselves using a smartphone or HMD and inputs detailed product information (e.g., the name, features, and size of the moisture-wicking, quick-drying sportswear). This data is sent to the server via the input means. The input data includes the user's image data and product information. The server analyzes this data and converts it into an appropriate format.

[0203] Step 2:

[0204] The server uses an image generation tool to generate visual content based on the image data and product information sent by the user. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate a high-resolution image that synthesizes the user and the product. This image makes it appear as if the user is actually trying on the product. The input is the user image and product information, and the output is the synthesized image.

[0205] Step 3:

[0206] The server uses language generation tools to create product descriptions. It uses a generative AI model (e.g., GPT-3) to generate promotional text. For example, it uses a prompt such as "Please describe the features of moisture-wicking, quick-drying sportswear" to create a detailed product description. The input is product information, and the output is the product description.

[0207] Step 4:

[0208] The server uses a voice generation tool to create a narration based on the generated product description. It uses a speech synthesis AI model (e.g., WaveNet) to generate an audio file that sounds similar to a human voice. This is what the user hears when listening to the product description. The input is the product description, and the output is an audio file.

[0209] Step 5:

[0210] The server sends the generated content (synthetic images, product descriptions, and audio files) to the user's terminal. The terminal displays these contents so that the user can confirm them. Through the display means, the user can have a virtual try-on experience and visually and aurally confirm the generated content. The input is the generated content, and the output is the user's confirmation.

[0211] Step 6:

[0212] The user checks the provided content and inputs feedback as necessary. For example, they send requests for corrections such as "I want the background of the image changed" or "I want the tone of the description to be more casual." The server receives this and corrects and regenerates the content using the regeneration means. The input is the user feedback, and the output is the corrected content.

[0213] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0214] The present invention is embodied in a system including an image generation means, a language generation means, a voice generation means, an input means, a display means, a data analysis means, a feedback response means, and an emotion engine, which recognizes a user's emotions and personalizes content based on the emotions, thereby effectively supporting corporate marketing.

[0215] Server Processing

[0216] The server receives data provided by users and companies and creates content using the following generation methods:

[0217] 1. Image generation method

[0218] The server generates visual content based on user-provided product information and target audience data, for example by using generative AI models to generate high-resolution images.

[0219] Example: A server generates promotional images for the "latest smartphone." The data includes design details and background images, and the appropriate images are automatically generated based on this.

[0220] 2. Language generation means

[0221] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in linguistic expressions that correspond to the user's requests and emotions.

[0222] Example: The server generates text describing the features of a smartphone based on the user's emotions recognized by the emotion engine. The generated content might be something like, "This smartphone has an innovative design and a high-performance camera. It will enrich your daily life."

[0223] 3. Voice Generation Method

[0224] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[0225] Example: Generate an audio file that reads the generated smartphone description in a professional voice.

[0226] Terminal handling

[0227] The terminal provides an interface for smooth communication between the user and the server.

[0228] 1. Data entry and customization

[0229] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[0230] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "latest smartphone, price, target demographic").

[0231] 2. Product Labeling

[0232] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[0233] Example: Check the smartphone promotional images, introductory text, and audio files displayed on the device.

[0234] User operations

[0235] The user performs the following operations through the terminal.

[0236] 1. Input and Request Submission

[0237] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[0238] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[0239] 2. Feedback and correction requests

[0240] The user can check the generated content and make corrections or additions as needed.

[0241] Example: A user reviews the generated images and text and provides feedback such as, "I wish the background of the images was a little brighter" or "I wish the tone of the text was more friendly."

[0242] Use of emotion engine

[0243] The emotion engine has the ability to recognize the user's emotions and reflect them in the generated content.

[0244] 1. Emotion analysis

[0245] The emotion engine analyzes emotions from user input and reactions, such as text input, facial expressions, and tone of voice.

[0246] Example: Determine whether a user has positive or negative emotions based on the comments and feedback they enter into their device.

[0247] 2. Adjust the tone of your content

[0248] The tone and style of the generated content is adjusted based on emotional data analyzed by the emotion engine.

[0249] Example: If the user is feeling positive, the tone of the text will change to be optimistic and uplifting, and if they are feeling negative, the tone will change to be calming and reassuring.

[0250] 3. Real-time emotional response

[0251] The emotion engine analyzes user emotions in real time and instantly modifies and regenerates content accordingly.

[0252] Example: If a user's emotions change (for example, they may initially show interest, but then become disappointed), detect that change and immediately change the way you present content or information.

[0253] In this way, the system of the present invention can provide highly personalized marketing content that takes into account the user's emotions, and effectively support a company's marketing activities.

[0254] The processing flow will be explained below.

[0255] Step 1:

[0256] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "the latest smartphone," along with its features, target age group, and lifestyle.

[0257] Step 2:

[0258] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[0259] Step 3:

[0260] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[0261] Step 4:

[0262] The server uses an emotion engine to analyze the user's emotions. The server analyzes the user's text input, facial expressions, tone of voice, etc. to recognize the emotion. For example, it can detect positive emotions from the user's input.

[0263] Step 5:

[0264] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[0265] Step 6:

[0266] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme of "latest smartphones."

[0267] Step 7:

[0268] The server generates text content using language generation tools. It uses a generative AI model to create product introductions and review text. For example, if the emotion engine detects positive emotions, it generates an introduction such as, "This smartphone features an innovative design and a high-performance camera. It will enrich your daily life."

[0269] Step 8:

[0270] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[0271] Step 9:

[0272] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[0273] Step 10:

[0274] The device displays the product to the user. The device displays the product received from the server on the screen so that the user can review it. The screen displays an image of the promotional smartphone, introductory text, and an option to play the audio file.

[0275] Step 11:

[0276] Users can review the generated content and provide feedback and correction requests. Users can review the generated content and provide feedback such as "I would like the background of the images to be brighter" or "I would like the tone of the text to be more friendly."

[0277] Step 12:

[0278] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the smartphone image with a lighter background and adjusting the tone of the text.

[0279] Step 13:

[0280] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[0281] Step 14:

[0282] The device stores the approved content and prepares it for distribution. It also stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[0283] Step 15:

[0284] The server monitors the user's reaction after distribution and evaluates it in real time using an emotion engine. For example, it analyzes how the user felt about the content and uses the feedback to generate the next content if necessary.

[0285] Example 2

[0286] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0287] Conventional marketing support systems have had difficulty generating highly personalized content that takes user emotions into account, and have been unable to respond to changes in user emotions in real time, resulting in reduced marketing effectiveness.

[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0289] In this invention, the server includes means for receiving data provided by a user, image generation means for generating an image based on the received data, language generation means for generating text based on the received data, voice generation means for generating voice based on the generated text, input means for accepting user input, display means for displaying the generated content to the user, and an emotion engine for analyzing the user's emotions. This makes it possible to analyze the user's emotions in real time and generate and adjust images, text, and voice based on the analyzed emotions.

[0290] "Means for receiving data provided by users" refers to the interface and mechanism for receiving product information and target audience information sent by users and companies.

[0291] "Image generation means" refers to a device or software that creates visual content based on received data, specifically a mechanism that generates high-resolution images using a generative AI model.

[0292] "Language generation means" refers to a device or software that creates text based on received data and has the ability to generate descriptions and comments about products and services using generative AI models.

[0293] "Speech generation means" refers to a device or software that creates an audio file based on the generated text and has a mechanism for generating narration that resembles a human voice.

[0294] "Input means for accepting user input" refers to an interface for a user to input information and a mechanism for receiving data.

[0295] The "display means for displaying generated content to the user" refers to an interface and its display mechanism that allows the user to check generated content such as images, text, and audio.

[0296] An "emotion engine that analyzes user emotions" refers to a device or software that analyzes a user's input and reactions to infer their emotions, and has the ability to adjust the tone and style of the content it generates based on this.

[0297] "Data analysis means" refers to a device or software that analyzes input user information and converts it into a format suitable for the content to be generated.

[0298] "Means for generating images, text, and audio based on prompts using a generative AI model" refers to a generative AI model and its operating mechanism for inputting prompts to generate high-resolution images, appropriate text, natural-sounding audio, etc.

[0299] "Means for analyzing user emotions in real time and modifying and regenerating content in response to emotional changes" refers to a device or software that monitors a user's facial expressions, tone of voice, text input, etc. in real time and instantly adjusts the generated content in response to those emotional changes.

[0300] The present invention relates to a system that analyzes user emotions and generates highly personalized content based on those emotions. This system is composed of a server, a terminal, and a user, and each element has a specific function.

[0301] Server configuration and functions

[0302] The server contains the following main facilities:

[0303] 1. Means of receiving user-provided data

[0304] The server receives product information and target audience information provided by users and businesses through a specific interface, which is implemented using technologies such as HTTP requests and API calls.

[0305] 2. Image generation method

[0306] Based on the received data, the server uses a generative AI model to generate high-resolution images. For example, to generate images of the "latest smartphone" for product promotion, the following prompt sentence is used: "Generate a high-resolution promotional image of the latest smartphone."

[0307] 3. Language generation means

[0308] The server uses a generative AI model to generate text about products and services. It selects appropriate linguistic expressions based on the emotional data analyzed by the emotion engine. For example, when generating promotional text, the prompt is: "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera."

[0309] 4. Voice Generation Method

[0310] Based on the generated text, a voice generation AI is used to generate narration that sounds similar to a human voice, allowing users to create professional audio content. The voice generation prompt is "Please generate a natural voice based on the generated text."

[0311] 5. Emotion Engine

[0312] The emotion engine analyzes user input and responses to infer emotions, and this data is used to adjust the tone and style of the generated content.

[0313] Device configuration and functions

[0314] The terminal provides the following features:

[0315] 1. Data entry and customization

[0316] It provides an interface where users can input necessary information. For example, a company's marketing staff can input information about a new product and set specifications for the content to be generated.

[0317] 2. Product Labeling

[0318] It provides an interface that displays and allows users to check the generated content (images, text, audio) sent from the server. Users can check the generated content and enter feedback.

[0319] User operations

[0320] The user performs the following actions through the terminal:

[0321] 1. Input and Request Submission

[0322] A marketer inputs detailed information about a product or service and requests the server to generate content. For example, the marketer inputs basic product information and target information and clicks the submit button.

[0323] 2. Feedback and correction requests

[0324] Users can review the generated content and submit corrections or additions as needed. For example, they can use the device's feedback interface to submit requests such as "make the background brighter" or "make the text more friendly in tone."

[0325] Specific examples

[0326] For example, when generating promotional content for a new smartphone, a marketer might use prompts like this:

[0327] Image generation prompt: "Generate a high-resolution promotional image for the latest smartphone."

[0328] Text generation prompt: "Based on your emotions, create a text advertising the smartphone's innovative design and high-performance camera."

[0329] Speech generation prompt: "Generate natural-sounding speech based on the generated text."

[0330] By using this system, companies can quickly create highly personalized marketing content that responds to user emotions, thereby maximizing the effectiveness of their marketing activities.

[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0332] Step 1:

[0333] Receiving user data

[0334] The server receives product information and target audience information provided by users and businesses. This data is obtained through HTTP requests and API calls.

[0335] Input: Product information (design details, target demographic, sales price, etc.)

[0336] Output: Received user data

[0337] Specific operation: The server accepts requests to the API endpoint and saves user data in the database.

[0338] Step 2:

[0339] Execution of image generation means

[0340] The server generates images based on the received data, using a generative AI model and prompt text to create high-resolution images.

[0341] Input: User data (product information)

[0342] Output: Generated image data

[0343] Specific operation: The generative AI model is given the prompt "Generate a high-resolution promotional image for the latest smartphone," and the image is generated and saved.

[0344] Step 3:

[0345] Execution of language generation methods

[0346] The server generates text about products and services by inputting prompt sentences into the generative AI model based on the received data and the analysis results of the emotion engine.

[0347] Input: User data (product information), sentiment analysis results

[0348] Output: The generated text

[0349] Specific operation: The generative AI model is given the prompt, "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera," and the generated text is saved in a database.

[0350] Step 4:

[0351] Execution of the speech generation means

[0352] The server generates speech based on the generated text, using a speech generation AI to input prompts and create narration.

[0353] Input: Generated text

[0354] Output: Generated audio file

[0355] Specific operation: Enter a prompt into the voice generation AI saying, "Please generate natural-sounding speech based on the generated text," and save the generated voice file.

[0356] Step 5:

[0357] Data Entry and Customization

[0358] The terminal provides an interface where the user can input the necessary information, such as product details and target audience information.

[0359] Input: Product information, target information

[0360] Output: User data sent

[0361] Specific operation: When the user enters the required information into the terminal interface and clicks the "Submit" button, the data is sent to the server.

[0362] Step 6:

[0363] Display of product

[0364] The terminal displays the products (images, text, and audio) sent from the server to the user for confirmation.

[0365] Input: Generated images, text, and audio files

[0366] Output: The product displayed to the user

[0367] Specific behavior: The device interface is updated to display the generated content to the user.

[0368] Step 7:

[0369] Input and Request Submission

[0370] The user inputs the necessary information through the terminal and sends a request for content generation to the server.

[0371] Input: Product information, target information

[0372] Output: Request to server

[0373] Specific operation: When the user enters information into the terminal and clicks the "Generate" button, the request is sent to the server.

[0374] Step 8:

[0375] Feedback and correction requests

[0376] The user checks the generated content and sends requests for corrections or additions as necessary.

[0377] Input: Feedback

[0378] Output: Submitting a fix request

[0379] Specific operation: The user uses the feedback interface on the device to enter correction requests or addition requests and clicks the submit button.

[0380] Step 9:

[0381] Sentiment analysis and content moderation

[0382] The server uses an emotion engine to analyze emotions based on user input and responses, and adjusts the tone and style of the generated content accordingly.

[0383] Input: User-entered data

[0384] Output: Reconciled content

[0385] Specific operation: The emotion engine analyzes the user's data and, based on the results, adjusts and inputs the prompt sentences into the generative AI model to regenerate the content.

[0386] Step 10:

[0387] Real-time emotional response

[0388] The server analyzes the user's emotions in real time and instantly modifies and regenerates content accordingly.

[0389] Input: Real-time user response data

[0390] Output: Instantly modified content

[0391] Specific operation: Monitors user reactions in real time, and if a change in emotion is detected, inputs new prompt sentences into the generative AI model to modify and regenerate content.

[0392] Through the above steps, highly personalized content based on the user's emotions can be generated quickly.

[0393] (Application example 2)

[0394] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0395] Conventional advertising content is uniform and difficult to fully reflect the diverse emotions and needs of users. As a result, ads that do not elicit a positive response from users are sometimes delivered, preventing the effectiveness of marketing activities from being maximized. Another issue is that modifying and regenerating content to reflect user feedback is time-consuming and laborious.

[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0397] In this invention, the server includes emotion analysis means for analyzing user emotions, image generation means for generating visual advertising content, language generation means for generating text for the advertising content, and voice generation means for generating voice based on the generated text, thereby enabling personalized advertising content to be generated and delivered in real time according to the user's emotions.

[0398] "Emotion analysis means" refers to a device or program that analyzes the user's input information and behavioral data and recognizes the user's emotional state.

[0399] The "image generation means" is a device or program that generates visual advertising content based on specific conditions and input data.

[0400] The "language generation means" is a device or program that automatically generates text for advertising content based on specific conditions and input data.

[0401] The "voice generation means" is a device or program that generates natural voice based on the generated text data.

[0402] "Input means" refers to a device or interface for accepting user information or requests.

[0403] "Display means" refers to a device or interface for visually presenting and displaying the generated advertising content to the user.

[0404] "Data analysis means" refers to a device or program that analyzes input user information and emotion data and converts them into an appropriate format.

[0405] The "feedback receiving means" refers to a device or program for receiving feedback from users and correcting and regenerating the generated advertising content.

[0406] The "advertising content distribution means" refers to a device or program for distributing the generated advertising content to a user's device.

[0407] The system for implementing the present invention operates effectively by the mutual cooperation of the server, terminals, and users.

[0408] Server Processing

[0409] The server automatically generates advertising content using multiple generation means having the following functions:

[0410] Emotion analysis means

[0411] The server receives the user's input information and behavioral data and analyzes the user's emotional state using an emotion analysis tool. For this purpose, it uses a specialized emotion analysis library. For example, it analyzes input text and voice data to determine whether the user is in a positive, negative, excited, or other emotional state.

[0412] Image Generation Means

[0413] The server generates visual ad content based on data from the sentiment analysis method, for example using a generative AI model to generate high-resolution images based on prompts such as:

[0414] Example prompt: 'Create an image of a new smartphone that evokes excitement.'

[0415] language generation means

[0416] The server uses a language generation model to generate text for the advertising content, taking into account the sentiment analysis data and generating the text based on the following prompt:

[0417] Example prompt: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[0418] Voice generation means

[0419] The server generates narration audio based on the generated text, using an AI model specialized in voice generation.

[0420] Terminal handling

[0421] The terminal provides an interface to facilitate interaction between the user and the server.

[0422] Data Entry and Customization

[0423] Users input information about new products and target audiences through their devices, which then sends specific data to the server for generating advertising content.

[0424] Display of product

[0425] The content of the generated products (images, text, audio) sent from the server is displayed on the terminal, allowing the user to check the content.

[0426] User operations

[0427] The user performs the following operations through the terminal.

[0428] Input and Request Submission

[0429] The user enters the necessary information and sends a content generation request to the server.

[0430] Feedback and correction requests

[0431] The user checks the generated content and requests corrections or additions as necessary. The server regenerates the content based on that feedback.

[0432] Hardware and software used

[0433] Sentiment analysis method: EmotionDetector library

[0434] Image generation method: Generative AI model using Tensorflow

[0435] Language generation: transformers (GPT2)

[0436] Sound generation method: Tacotron2

[0437] Example prompt sentence:

[0438] Image Generation: 'Create an image of a new smartphone that evokes excitement.'

[0439] Language generation: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[0440] The above system makes it possible to generate personalized advertising content in real time according to the user's emotions, supporting effective marketing activities that attract the user's attention.

[0441] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0442] Step 1:

[0443] The user inputs information about the new product into the terminal and sends a creation request.

[0444] Input: New product details (e.g. latest smartphone, price, target demographic).

[0445] Output: Generated request data.

[0446] Specific operation: The user uses the input interface of the terminal to input information about the new product, and then clicks the generation request button to send the data to the server.

[0447] Step 2:

[0448] The server analyzes the input information received from the user using a data analysis means and converts it into a required format.

[0449] Input: Generation request data (detailed information about new products, sentiment analysis data).

[0450] Output: The input data required to generate the advertising content.

[0451] Specific operation: The server uses data analysis means to analyze the input data and convert it into a standard format. It also uses emotion analysis means to determine the user's emotional state.

[0452] Step 3:

[0453] The server uses emotion analysis means to recognize the emotion of the user.

[0454] Input: User input information, emotion data.

[0455] Output: The user's emotional state.

[0456] Specific operation: The server analyzes the emotional state based on the received data using an emotion analysis means and recognizes emotions such as positive, negative, and excitement.

[0457] Step 4:

[0458] The server generates visual advertising content using an image generating means.

[0459] Input: Generate request data, user's emotional state.

[0460] Output: The generated ad image.

[0461] What it does: It uses a generative AI model to generate high-resolution advertising images based on a prompt (e.g., "Create an image of a new smartphone that evokes excitement").

[0462] Step 5:

[0463] The server generates text for the advertisement content using a language generation means.

[0464] Input: Generate request data, user's emotional state.

[0465] Output: The generated ad text.

[0466] How it works: Using the GPT2 model, we input an emotion-based prompt (e.g., "Generate an advertisement text for a new smartphone that makes the user feel excited") and generate advertisement text.

[0467] Step 6:

[0468] The server generates a voice based on the text generated by the voice generating means.

[0469] Input: The generated ad text.

[0470] Output: The generated audio file.

[0471] What it does: Uses the Tacotron2 model to convert the generated ad text into natural-sounding speech.

[0472] Step 7:

[0473] The server sends the generated advertising content (images, text, audio) to the terminal.

[0474] Input: Generated ad images, text and audio files.

[0475] Output: The ad content sent.

[0476] Specific operation: The server packages the generated content and sends it to the terminal.

[0477] Step 8:

[0478] The terminal displays the received advertisement content to the user.

[0479] Input: Submitted ad content.

[0480] Output: The ad content displayed to the user.

[0481] Specific operation: The terminal displays and plays the received advertising images, text, and audio to the user.

[0482] Step 9:

[0483] The user submits feedback and the server modifies and regenerates the advertising content based on the feedback.

[0484] Input: User feedback.

[0485] Output: The modified and regenerated ad content.

[0486] How it works: The user sends feedback via their device, and the server then modifies and regenerates the ad content as needed.

[0487] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0488] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0489] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0490] [Second embodiment]

[0491] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0492] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0493] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0494] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0495] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0496] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0497] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0498] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0499] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0500] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0501] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0502] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0503] The present invention is embodied in a system including an image generating means, a language generating means, a voice generating means, an input means, and a display means, which supports corporate marketing and provides personalized information to consumers.

[0504] Server Processing

[0505] The server receives data provided by users and companies and creates content using the following three generation methods.

[0506] 1. Image generation method

[0507] The server generates visual content based on product information and target audience data provided by the user. The image generation means generates high-resolution images, for example, using a generative AI model.

[0508] Example: A server generates promotional images for "sportswear." The data includes the color, design, and background image of the clothing, and the server automatically generates an appropriate image based on this data.

[0509] 2. Language generation means

[0510] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in the language expression requested by the user.

[0511] Example: The server generates text describing the features of "new sportswear," such as "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[0512] 3. Voice Generation Method

[0513] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[0514] Example: Generate an audio file that reads the description of the generated sportswear in a professional voice.

[0515] Terminal handling

[0516] The terminal provides an interface for smooth communication between the user and the server.

[0517] 1. Data entry and customization

[0518] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[0519] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "unisex moisture-wicking, quick-drying sportswear, price, main target demographic").

[0520] 2. Product Labeling

[0521] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[0522] Example: Check the promotional sportswear images, introductory text, and audio files displayed on the device.

[0523] User operations

[0524] The user performs the following operations through the terminal.

[0525] 1. Input and Request Submission

[0526] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[0527] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[0528] 2. Feedback and correction requests

[0529] The user can check the generated content and make corrections or additions as needed.

[0530] Example: A user reviews the generated images and text and submits correction requests such as "make the image background a little brighter" or "make the text tone more casual."

[0531] In this way, the system of the present invention can efficiently generate and provide highly personalized marketing content that meets user requests.

[0532] The processing flow will be explained below.

[0533] Step 1:

[0534] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "latest running shoes," along with their features, target age group, and lifestyle.

[0535] Step 2:

[0536] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[0537] Step 3:

[0538] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[0539] Step 4:

[0540] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[0541] Step 5:

[0542] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme "latest running shoes."

[0543] Step 6:

[0544] The server generates text content using language generation tools. It uses a generative AI model to create product descriptions and reviews. For example, it generates a description such as, "These running shoes are lightweight and have excellent cushioning."

[0545] Step 7:

[0546] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[0547] Step 8:

[0548] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[0549] Step 9:

[0550] The device displays the result to the user. The device displays the result received from the server on the screen so that the user can review it. The screen displays an image of the promotional running shoes, introductory text, and an option to play the audio file.

[0551] Step 10:

[0552] Users can review the generated content and provide feedback and correction requests. Users can check the generated content, enter feedback such as "make the background of the image brighter" or "make the tone of the text more casual," and submit it.

[0553] Step 11:

[0554] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the image of the running shoes with a lighter background and adjusting the tone of the text.

[0555] Step 12:

[0556] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[0557] Step 13:

[0558] The device stores the approved content and prepares it for distribution. It stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[0559] Example 1

[0560] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0561] Conventional marketing support systems lack adequate means for efficiently generating highly personalized content that meets user requests. Furthermore, they lack the functionality to incorporate user feedback into the generated content, making it difficult to quickly and accurately revise the content. The present invention aims to solve these problems and more effectively support corporate marketing activities.

[0562] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0563] In this invention, the server includes an image generation means, a language generation means, and a voice generation means. This enables efficient generation of personalized marketing content in response to user requests. The server further includes a means for generating visual content using a generative AI model, a means for generating text using the generative AI model, a means for generating voice based on the generated text, a means for integrating the generated content and sending it to the user, and a means for receiving user feedback on the generated content, correcting it, and regenerating it. This enables content to be quickly and accurately corrected based on user feedback, thereby more effectively supporting corporate marketing activities.

[0564] "Image generation means" is a function for generating visual content (photographs, illustrations, etc.) based on information provided by the user.

[0565] "Language generation means" is a function for generating relevant text based on information and instructions provided by the user.

[0566] The "speech generation means" is a function for generating speech based on the generated text.

[0567] An "input means" is an interface through which a user inputs information into the system.

[0568] The "display means" is an interface for displaying the generated content (images, text, audio) to the user.

[0569] A "generative AI model" is an algorithm or platform that uses machine learning to generate content (images, text, audio, etc.).

[0570] A "prompt" is text containing commands or instructions that are input to a generative AI model to generate specific content.

[0571] The "data analysis means" is a function that analyzes input user information and converts it into a format suitable for the content to be generated.

[0572] The "feedback means" is a function that receives feedback from the user and modifies or regenerates the generated content.

[0573] "Integration means" is a function that combines images, text, and audio into a single piece of content.

[0574] "User feedback" refers to opinions and correction requests from users regarding the generated content.

[0575] The present invention is an advanced content generation system for effectively supporting corporate marketing activities. This system includes image generation means, language generation means, voice generation means, input means, and display means. The following describes in detail the use of specific hardware and software, as well as methods for data processing and data calculation.

[0576] Hardware and Software

[0577] The system is implemented using specific hardware and software as follows:

[0578] Server: A server with high-performance data processing capabilities is used to receive, analyze, generate content, and integrate data.

[0579] Terminal: The user operates the device using a PC, tablet, or other device. This terminal communicates with the server and provides an interface to the user.

[0580] Generative AI models: AI models such as DALL-E and MidJourney are used for image generation, GPT-4 for text generation, and Google Text-to-Speech and Amazon Polly for voice generation.

[0581] Image Generation Means

[0582] The server generates visual content based on product information and target audience data provided by the user. Specifically, it generates images using prompts from a generative AI model (e.g., DALL-E).

[0583] Example: "Generate a scene of a man and woman running in sportswear, moisture-wicking, with a bright background."

[0584] language generation means

[0585] The server has the functionality to generate text about products and services. It generates text using prompt sentences from a generative AI model (e.g., GPT-4).

[0586] Example: "Please explain the features and benefits of your new sportswear in 200 words or less."

[0587] Voice generation means

[0588] The server generates speech based on the generated text, and uses a speech generation AI (e.g., Google Text-to-Speech) to convert the generated text into a voice with a narration that sounds close to a human voice.

[0589] Example: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made with the latest moisture-wicking, quick-drying materials for maximum comfort during exercise.'"

[0590] Data processing and calculation

[0591] The server analyzes the received user data and converts it into a format suitable for the generative AI model. This includes classifying and filtering the data. It also modifies and regenerates the generated content based on feedback provided by the user. To do this, it inputs the prompt sentence into the generative AI model again to generate new content.

[0592] User operations

[0593] The user uses the terminal to input the necessary information through the interface and send a content generation request to the server. The generated content is displayed on the terminal, and the user can check it and send feedback if necessary. The terminal then sends the user's feedback to the server, and the server regenerates the content based on that feedback.

[0594] In this way, the system of the present invention can efficiently generate personalized marketing content in response to user requests and support the marketing activities of companies.

[0595] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0596] Step 1:

[0597] The user uses a terminal to input the necessary data, including product information (e.g., "sportswear"), target audience information (e.g., "runners in their 30s"), and other details required for content creation. The input information is then sent to the server.

[0598] Input: Product information, target audience information

[0599] Output: Data sent to the server

[0600] Step 2:

[0601] The terminal transmits the data entered by the user to the server, which receives the data and prepares it for analysis.

[0602] Input: Data from the user

[0603] Output: Data sent to the server

[0604] Step 3:

[0605] The server analyzes the received data, which may include categorizing and filtering the data, and converts it into a format suitable for the generative AI model.

[0606] Input: Data from the terminal

[0607] Output: Parsed data

[0608] Step 4:

[0609] The server activates the image generation means based on the analyzed data, sending specific prompts to the generative AI model (e.g., DALL-E) to generate visual content.

[0610] Input: Parsed data, prompt statement

[0611] Output: The generated image

[0612] Specifically, the prompt text includes, "Generate a scene of a man and woman running in sportswear, moisture-wicking and quick-drying, with a bright background."

[0613] Step 5:

[0614] The server then activates a language generation mechanism based on the analyzed data, sending specific prompts to a generative AI model (e.g., GPT-4) to generate text.

[0615] Input: Parsed data, prompt statement

[0616] Output: The generated text

[0617] Specific prompts include, "Please explain the features and benefits of your new sportswear in 200 words or less."

[0618] Step 6:

[0619] The server activates a voice generation means based on the generated text, sends the text to a voice generation AI (e.g., Google Text-to-Speech), and generates an audio file.

[0620] Input: Generated text

[0621] Output: Generated audio file

[0622] Specifically, the instructions include: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made from the latest moisture-wicking, quick-drying material to maximize comfort during exercise.'"

[0623] Step 7:

[0624] The server integrates the generated images, text, and audio, and sends the integrated product to the terminal in a format that is easy for the user to view.

[0625] Input: Generated images, text, audio

[0626] Output: The integrated product

[0627] Step 8:

[0628] The server sends the integrated product to the terminal, which receives the product and displays it to the user.

[0629] Input: The integrated product

[0630] Output: The product sent to the terminal

[0631] Step 9:

[0632] The user uses the device to review the generated content and, if desired, enter feedback, which may include requests for image corrections or text adjustments.

[0633] Input: Generated content

[0634] Output: Feedback

[0635] Step 10:

[0636] The terminal sends feedback from the user to the server, which receives this feedback and prepares to modify and regenerate the content.

[0637] Input: User feedback

[0638] Output: Feedback sent to the server

[0639] Step 11:

[0640] The server modifies and regenerates the content based on the feedback, which may involve re-analyzing the data and adjusting the prompts.

[0641] Input: Feedback

[0642] Output: Modified and regenerated content

[0643] Through this process, the system can efficiently generate highly personalized marketing content that meets user requests and supports a company's marketing activities.

[0644] (Application example 1)

[0645] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0646] In today's online shopping environment, consumers cannot actually try on products, making it difficult to experience the feel and comfort of the products. Furthermore, it takes time and effort for consumers to select the right product for themselves. Therefore, there is a need for an easier and more intuitive way for consumers to try on products and understand information about them.

[0647] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0648] In this invention, the server includes an image generation means, a language generation means, a voice generation means, an input means for receiving user input, a display means for displaying the generated content to the user, a means for combining a user image with a product image, and a means for allowing the user to check the generated content, thereby enabling consumers to virtually try on products from the comfort of their own homes and check the product features visually and audibly.

[0649] The "image generation means" is a means for generating visual content based on data and product information provided by the user.

[0650] The "language generation means" is a means for generating text information about products and services.

[0651] The "voice generation means" is a means for generating voice based on the generated text, and creates narration that resembles a human voice.

[0652] The "input means for accepting user input" is an interface for the user to input necessary data.

[0653] The "display means for displaying the generated content to the user" is an interface for displaying the generated content such as images, text, and audio to the user.

[0654] The "means for combining a user image with a product image" refers to a means for combining an acquired user image with a product image, making it appear as if the user is trying on the product.

[0655] The "means for allowing the user to confirm the generated content" refers to a means for presenting the generated content to the user and receiving confirmation of the content and feedback therefrom.

[0656] In order to implement this invention, it is necessary to construct a system that includes an image generation means, a language generation means, a voice generation means, an input means for accepting user input, a display means for displaying the generated content to the user, a means for synthesizing the user image with a product image, and a means for allowing the user to confirm the generated content.

[0657] First, the server receives and analyzes data provided by the user. Specifically, it acquires the user's image data and product information and generates visual content based on this. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate images. This model is capable of generating high-resolution images. For example, a user can take a picture of themselves using a smartphone and send it to the server.

[0658] Next, the server uses language generation tools based on the product information to create a description of the product. A generative AI model (e.g., GPT-3) is used for language generation. For example, when the server generates text to describe the features of "moisture-wicking, quick-drying sportswear," it generates information such as, "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[0659] In addition, the server generates audio based on the generated text. The audio is generated using a speech synthesis AI model (e.g., WaveNet). This model allows for the creation of narration that sounds close to a human voice. Based on the generated description, an audio file introducing the product is generated.

[0660] The user's device displays these artifacts (images, text, and audio). Users can virtually try on items from the comfort of their own home using their smartphone, smart glasses, or head-mounted display. For example, a user can select a particular sportswear item, view its image and description, and listen to an audio description.

[0661] For example:

[0662] 1. The user enters product information for sportswear.

[0663] 2. The server receives image data and product information from the user and creates a synthetic image using a generative AI model.

[0664] 3. Automatically generate product descriptions using a language generation AI model.

[0665] 4. Convert the text generated by the speech synthesis AI model into speech.

[0666] 5. The device displays the generated content to the user, allowing them to virtually try on the clothes at home.

[0667] 6. An example of a prompt sentence is, "Please explain the characteristics of moisture-wicking, quick-drying sportswear."

[0668] Through these steps, consumers can virtually try on products and visually and audibly check the product's features from the comfort of their own home, helping them make product selections more easily and intuitively.

[0669] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0670] Step 1:

[0671] The user takes a picture of themselves using a smartphone or HMD and inputs detailed product information (e.g., the name, features, and size of the moisture-wicking, quick-drying sportswear). This data is sent to the server via the input means. The input data includes the user's image data and product information. The server analyzes this data and converts it into an appropriate format.

[0672] Step 2:

[0673] The server uses an image generation tool to generate visual content based on the image data and product information sent by the user. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate a high-resolution image that synthesizes the user and the product. This image makes it appear as if the user is actually trying on the product. The input is the user image and product information, and the output is the synthesized image.

[0674] Step 3:

[0675] The server uses language generation tools to create product descriptions. It uses a generative AI model (e.g., GPT-3) to generate promotional text. For example, it uses a prompt such as "Please describe the features of moisture-wicking, quick-drying sportswear" to create a detailed product description. The input is product information, and the output is the product description.

[0676] Step 4:

[0677] The server uses a voice generation tool to create a narration based on the generated product description. It uses a speech synthesis AI model (e.g., WaveNet) to generate an audio file that sounds similar to a human voice. This is what the user hears when listening to the product description. The input is the product description, and the output is an audio file.

[0678] Step 5:

[0679] The server sends the generated content (synthetic images, product descriptions, and audio files) to the user's terminal. The terminal displays these contents so that the user can confirm them. Through the display means, the user can have a virtual try-on experience and visually and aurally confirm the generated content. The input is the generated content, and the output is the user's confirmation.

[0680] Step 6:

[0681] The user checks the provided content and inputs feedback as necessary. For example, they send requests for corrections such as "I want the background of the image changed" or "I want the tone of the description to be more casual." The server receives this and corrects and regenerates the content using the regeneration means. The input is the user feedback, and the output is the corrected content.

[0682] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0683] The present invention is embodied in a system including an image generation means, a language generation means, a voice generation means, an input means, a display means, a data analysis means, a feedback response means, and an emotion engine, which recognizes a user's emotions and personalizes content based on the emotions, thereby effectively supporting corporate marketing.

[0684] Server Processing

[0685] The server receives data provided by users and companies and creates content using the following generation methods:

[0686] 1. Image generation method

[0687] The server generates visual content based on user-provided product information and target audience data, for example by using generative AI models to generate high-resolution images.

[0688] Example: A server generates promotional images for the "latest smartphone." The data includes design details and background images, and the appropriate images are automatically generated based on this.

[0689] 2. Language generation means

[0690] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in linguistic expressions that correspond to the user's requests and emotions.

[0691] Example: The server generates text describing the features of a smartphone based on the user's emotions recognized by the emotion engine. The generated content might be something like, "This smartphone has an innovative design and a high-performance camera. It will enrich your daily life."

[0692] 3. Voice Generation Method

[0693] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[0694] Example: Generate an audio file that reads the generated smartphone description in a professional voice.

[0695] Terminal handling

[0696] The terminal provides an interface for smooth communication between the user and the server.

[0697] 1. Data entry and customization

[0698] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[0699] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "latest smartphone, price, target demographic").

[0700] 2. Product Labeling

[0701] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[0702] Example: Check the smartphone promotional images, introductory text, and audio files displayed on the device.

[0703] User operations

[0704] The user performs the following operations through the terminal.

[0705] 1. Input and Request Submission

[0706] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[0707] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[0708] 2. Feedback and correction requests

[0709] The user can check the generated content and make corrections or additions as needed.

[0710] Example: A user reviews the generated images and text and provides feedback such as, "I wish the background of the images was a little brighter" or "I wish the tone of the text was more friendly."

[0711] Use of emotion engine

[0712] The emotion engine has the ability to recognize the user's emotions and reflect them in the generated content.

[0713] 1. Emotion analysis

[0714] The emotion engine analyzes emotions from user input and reactions, such as text input, facial expressions, and tone of voice.

[0715] Example: Determine whether a user has positive or negative emotions based on the comments and feedback they enter into their device.

[0716] 2. Adjust the tone of your content

[0717] The tone and style of the generated content is adjusted based on emotional data analyzed by the emotion engine.

[0718] Example: If the user is feeling positive, the tone of the text will change to be optimistic and uplifting, and if they are feeling negative, the tone will change to be calming and reassuring.

[0719] 3. Real-time emotional response

[0720] The emotion engine analyzes user emotions in real time and instantly modifies and regenerates content accordingly.

[0721] Example: If a user's emotions change (for example, they may initially show interest, but then become disappointed), detect that change and immediately change the way you present content or information.

[0722] In this way, the system of the present invention can provide highly personalized marketing content that takes into account the user's emotions, and effectively support a company's marketing activities.

[0723] The processing flow will be explained below.

[0724] Step 1:

[0725] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "the latest smartphone," along with its features, target age group, and lifestyle.

[0726] Step 2:

[0727] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[0728] Step 3:

[0729] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[0730] Step 4:

[0731] The server uses an emotion engine to analyze the user's emotions. The server analyzes the user's text input, facial expressions, tone of voice, etc. to recognize the emotion. For example, it can detect positive emotions from the user's input.

[0732] Step 5:

[0733] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[0734] Step 6:

[0735] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme of "latest smartphones."

[0736] Step 7:

[0737] The server generates text content using language generation tools. It uses a generative AI model to create product introductions and review text. For example, if the emotion engine detects positive emotions, it generates an introduction such as, "This smartphone features an innovative design and a high-performance camera. It will enrich your daily life."

[0738] Step 8:

[0739] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[0740] Step 9:

[0741] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[0742] Step 10:

[0743] The device displays the product to the user. The device displays the product received from the server on the screen so that the user can review it. The screen displays an image of the promotional smartphone, introductory text, and an option to play the audio file.

[0744] Step 11:

[0745] Users can review the generated content and provide feedback and correction requests. Users can review the generated content and provide feedback such as "I would like the background of the images to be brighter" or "I would like the tone of the text to be more friendly."

[0746] Step 12:

[0747] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the smartphone image with a lighter background and adjusting the tone of the text.

[0748] Step 13:

[0749] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[0750] Step 14:

[0751] The device stores the approved content and prepares it for distribution. It also stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[0752] Step 15:

[0753] The server monitors the user's reaction after distribution and evaluates it in real time using an emotion engine. For example, it analyzes how the user felt about the content and uses the feedback to generate the next content if necessary.

[0754] Example 2

[0755] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0756] Conventional marketing support systems have had difficulty generating highly personalized content that takes user emotions into account, and have been unable to respond to changes in user emotions in real time, resulting in reduced marketing effectiveness.

[0757] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0758] In this invention, the server includes means for receiving data provided by a user, image generation means for generating an image based on the received data, language generation means for generating text based on the received data, voice generation means for generating voice based on the generated text, input means for accepting user input, display means for displaying the generated content to the user, and an emotion engine for analyzing the user's emotions. This makes it possible to analyze the user's emotions in real time and generate and adjust images, text, and voice based on the analyzed emotions.

[0759] "Means for receiving data provided by users" refers to the interface and mechanism for receiving product information and target audience information sent by users and companies.

[0760] "Image generation means" refers to a device or software that creates visual content based on received data, specifically a mechanism that generates high-resolution images using a generative AI model.

[0761] "Language generation means" refers to a device or software that creates text based on received data and has the ability to generate descriptions and comments about products and services using generative AI models.

[0762] "Speech generation means" refers to a device or software that creates an audio file based on the generated text and has a mechanism for generating narration that resembles a human voice.

[0763] "Input means for accepting user input" refers to an interface for a user to input information and a mechanism for receiving data.

[0764] The "display means for displaying generated content to the user" refers to an interface and its display mechanism that allows the user to check generated content such as images, text, and audio.

[0765] An "emotion engine that analyzes user emotions" refers to a device or software that analyzes a user's input and reactions to infer their emotions, and has the ability to adjust the tone and style of the content it generates based on this.

[0766] "Data analysis means" refers to a device or software that analyzes input user information and converts it into a format suitable for the content to be generated.

[0767] "Means for generating images, text, and audio based on prompts using a generative AI model" refers to a generative AI model and its operating mechanism for inputting prompts to generate high-resolution images, appropriate text, natural-sounding audio, etc.

[0768] "Means for analyzing user emotions in real time and modifying and regenerating content in response to emotional changes" refers to a device or software that monitors a user's facial expressions, tone of voice, text input, etc. in real time and instantly adjusts the generated content in response to those emotional changes.

[0769] The present invention relates to a system that analyzes user emotions and generates highly personalized content based on those emotions. This system is composed of a server, a terminal, and a user, and each element has a specific function.

[0770] Server configuration and functions

[0771] The server contains the following main facilities:

[0772] 1. Means of receiving user-provided data

[0773] The server receives product information and target audience information provided by users and businesses through a specific interface, which is implemented using technologies such as HTTP requests and API calls.

[0774] 2. Image generation method

[0775] Based on the received data, the server uses a generative AI model to generate high-resolution images. For example, to generate images of the "latest smartphone" for product promotion, the following prompt sentence is used: "Generate a high-resolution promotional image of the latest smartphone."

[0776] 3. Language generation means

[0777] The server uses a generative AI model to generate text about products and services. It selects appropriate linguistic expressions based on the emotional data analyzed by the emotion engine. For example, when generating promotional text, the prompt is: "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera."

[0778] 4. Voice Generation Method

[0779] Based on the generated text, a voice generation AI is used to generate narration that sounds similar to a human voice, allowing users to create professional audio content. The voice generation prompt is "Please generate a natural voice based on the generated text."

[0780] 5. Emotion Engine

[0781] The emotion engine analyzes user input and responses to infer emotions, and this data is used to adjust the tone and style of the generated content.

[0782] Device configuration and functions

[0783] The terminal provides the following features:

[0784] 1. Data entry and customization

[0785] It provides an interface where users can input necessary information. For example, a company's marketing staff can input information about a new product and set specifications for the content to be generated.

[0786] 2. Product Labeling

[0787] It provides an interface that displays and allows users to check the generated content (images, text, audio) sent from the server. Users can check the generated content and enter feedback.

[0788] User operations

[0789] The user performs the following actions through the terminal:

[0790] 1. Input and Request Submission

[0791] A marketer inputs detailed information about a product or service and requests the server to generate content. For example, the marketer inputs basic product information and target information and clicks the submit button.

[0792] 2. Feedback and correction requests

[0793] Users can review the generated content and submit corrections or additions as needed. For example, they can use the device's feedback interface to submit requests such as "make the background brighter" or "make the text more friendly in tone."

[0794] Specific examples

[0795] For example, when generating promotional content for a new smartphone, a marketer might use prompts like this:

[0796] Image generation prompt: "Generate a high-resolution promotional image for the latest smartphone."

[0797] Text generation prompt: "Based on your emotions, create a text advertising the smartphone's innovative design and high-performance camera."

[0798] Speech generation prompt: "Generate natural-sounding speech based on the generated text."

[0799] By using this system, companies can quickly create highly personalized marketing content that responds to user emotions, thereby maximizing the effectiveness of their marketing activities.

[0800] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0801] Step 1:

[0802] Receiving user data

[0803] The server receives product information and target audience information provided by users and businesses. This data is obtained through HTTP requests and API calls.

[0804] Input: Product information (design details, target demographic, sales price, etc.)

[0805] Output: Received user data

[0806] Specific operation: The server accepts requests to the API endpoint and saves user data in the database.

[0807] Step 2:

[0808] Execution of image generation means

[0809] The server generates images based on the received data, using a generative AI model and prompt text to create high-resolution images.

[0810] Input: User data (product information)

[0811] Output: Generated image data

[0812] Specific operation: The generative AI model is given the prompt "Generate a high-resolution promotional image for the latest smartphone," and the image is generated and saved.

[0813] Step 3:

[0814] Execution of language generation methods

[0815] The server generates text about products and services by inputting prompt sentences into the generative AI model based on the received data and the analysis results of the emotion engine.

[0816] Input: User data (product information), sentiment analysis results

[0817] Output: The generated text

[0818] Specific operation: The generative AI model is given the prompt, "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera," and the generated text is saved in a database.

[0819] Step 4:

[0820] Execution of the speech generation means

[0821] The server generates speech based on the generated text, using a speech generation AI to input prompts and create narration.

[0822] Input: Generated text

[0823] Output: Generated audio file

[0824] Specific operation: Enter a prompt into the voice generation AI saying, "Please generate natural-sounding speech based on the generated text," and save the generated voice file.

[0825] Step 5:

[0826] Data Entry and Customization

[0827] The terminal provides an interface where the user can input the necessary information, such as product details and target audience information.

[0828] Input: Product information, target information

[0829] Output: User data sent

[0830] Specific operation: When the user enters the required information into the terminal interface and clicks the "Submit" button, the data is sent to the server.

[0831] Step 6:

[0832] Display of product

[0833] The terminal displays the products (images, text, and audio) sent from the server to the user for confirmation.

[0834] Input: Generated images, text, and audio files

[0835] Output: The product displayed to the user

[0836] Specific behavior: The device interface is updated to display the generated content to the user.

[0837] Step 7:

[0838] Input and Request Submission

[0839] The user inputs the necessary information through the terminal and sends a request for content generation to the server.

[0840] Input: Product information, target information

[0841] Output: Request to server

[0842] Specific operation: When the user enters information into the terminal and clicks the "Generate" button, the request is sent to the server.

[0843] Step 8:

[0844] Feedback and correction requests

[0845] The user checks the generated content and sends requests for corrections or additions as necessary.

[0846] Input: Feedback

[0847] Output: Submitting a fix request

[0848] Specific operation: The user uses the feedback interface on the device to enter correction requests or addition requests and clicks the submit button.

[0849] Step 9:

[0850] Sentiment analysis and content moderation

[0851] The server uses an emotion engine to analyze emotions based on user input and responses, and adjusts the tone and style of the generated content accordingly.

[0852] Input: User-entered data

[0853] Output: Reconciled content

[0854] Specific operation: The emotion engine analyzes the user's data and, based on the results, adjusts and inputs the prompt sentences into the generative AI model to regenerate the content.

[0855] Step 10:

[0856] Real-time emotional response

[0857] The server analyzes the user's emotions in real time and instantly modifies and regenerates content accordingly.

[0858] Input: Real-time user response data

[0859] Output: Instantly modified content

[0860] Specific operation: Monitors user reactions in real time, and if a change in emotion is detected, inputs new prompt sentences into the generative AI model to modify and regenerate content.

[0861] Through the above steps, highly personalized content based on the user's emotions can be generated quickly.

[0862] (Application example 2)

[0863] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0864] Conventional advertising content is uniform and difficult to fully reflect the diverse emotions and needs of users. As a result, ads that do not elicit a positive response from users are sometimes delivered, preventing the effectiveness of marketing activities from being maximized. Another issue is that modifying and regenerating content to reflect user feedback is time-consuming and laborious.

[0865] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0866] In this invention, the server includes emotion analysis means for analyzing user emotions, image generation means for generating visual advertising content, language generation means for generating text for the advertising content, and voice generation means for generating voice based on the generated text, thereby enabling personalized advertising content to be generated and delivered in real time according to the user's emotions.

[0867] "Emotion analysis means" refers to a device or program that analyzes the user's input information and behavioral data and recognizes the user's emotional state.

[0868] The "image generation means" is a device or program that generates visual advertising content based on specific conditions and input data.

[0869] The "language generation means" is a device or program that automatically generates text for advertising content based on specific conditions and input data.

[0870] The "voice generation means" is a device or program that generates natural voice based on the generated text data.

[0871] "Input means" refers to a device or interface for accepting user information or requests.

[0872] "Display means" refers to a device or interface for visually presenting and displaying the generated advertising content to the user.

[0873] "Data analysis means" refers to a device or program that analyzes input user information and emotion data and converts them into an appropriate format.

[0874] The "feedback receiving means" refers to a device or program for receiving feedback from users and correcting and regenerating the generated advertising content.

[0875] The "advertising content distribution means" refers to a device or program for distributing the generated advertising content to a user's device.

[0876] The system for implementing the present invention operates effectively by the mutual cooperation of the server, terminals, and users.

[0877] Server Processing

[0878] The server automatically generates advertising content using multiple generation means having the following functions:

[0879] Emotion analysis means

[0880] The server receives the user's input information and behavioral data and analyzes the user's emotional state using an emotion analysis tool. For this purpose, it uses a specialized emotion analysis library. For example, it analyzes input text and voice data to determine whether the user is in a positive, negative, excited, or other emotional state.

[0881] Image Generation Means

[0882] The server generates visual ad content based on data from the sentiment analysis method, for example using a generative AI model to generate high-resolution images based on prompts such as:

[0883] Example prompt: 'Create an image of a new smartphone that evokes excitement.'

[0884] language generation means

[0885] The server uses a language generation model to generate text for the advertising content, taking into account the sentiment analysis data and generating the text based on the following prompt:

[0886] Example prompt: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[0887] Voice generation means

[0888] The server generates narration audio based on the generated text, using an AI model specialized in voice generation.

[0889] Terminal handling

[0890] The terminal provides an interface to facilitate interaction between the user and the server.

[0891] Data Entry and Customization

[0892] Users input information about new products and target audiences through their devices, which then sends specific data to the server for generating advertising content.

[0893] Display of product

[0894] The content of the generated products (images, text, audio) sent from the server is displayed on the terminal, allowing the user to check the content.

[0895] User operations

[0896] The user performs the following operations through the terminal.

[0897] Input and Request Submission

[0898] The user enters the necessary information and sends a content generation request to the server.

[0899] Feedback and correction requests

[0900] The user checks the generated content and requests corrections or additions as necessary. The server regenerates the content based on that feedback.

[0901] Hardware and software used

[0902] Sentiment analysis method: EmotionDetector library

[0903] Image generation method: Generative AI model using Tensorflow

[0904] Language generation: transformers (GPT2)

[0905] Sound generation method: Tacotron2

[0906] Example prompt sentence:

[0907] Image Generation: 'Create an image of a new smartphone that evokes excitement.'

[0908] Language generation: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[0909] The above system makes it possible to generate personalized advertising content in real time according to the user's emotions, supporting effective marketing activities that attract the user's attention.

[0910] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0911] Step 1:

[0912] The user inputs information about the new product into the terminal and sends a creation request.

[0913] Input: New product details (e.g. latest smartphone, price, target demographic).

[0914] Output: Generated request data.

[0915] Specific operation: The user uses the input interface of the terminal to input information about the new product, and then clicks the generation request button to send the data to the server.

[0916] Step 2:

[0917] The server analyzes the input information received from the user using a data analysis means and converts it into a required format.

[0918] Input: Generation request data (detailed information about new products, sentiment analysis data).

[0919] Output: The input data required to generate the advertising content.

[0920] Specific operation: The server uses data analysis means to analyze the input data and convert it into a standard format. It also uses emotion analysis means to determine the user's emotional state.

[0921] Step 3:

[0922] The server uses emotion analysis means to recognize the emotion of the user.

[0923] Input: User input information, emotion data.

[0924] Output: The user's emotional state.

[0925] Specific operation: The server analyzes the emotional state based on the received data using an emotion analysis means and recognizes emotions such as positive, negative, and excitement.

[0926] Step 4:

[0927] The server generates visual advertising content using an image generating means.

[0928] Input: Generate request data, user's emotional state.

[0929] Output: The generated ad image.

[0930] What it does: It uses a generative AI model to generate high-resolution advertising images based on a prompt (e.g., "Create an image of a new smartphone that evokes excitement").

[0931] Step 5:

[0932] The server generates text for the advertisement content using a language generation means.

[0933] Input: Generate request data, user's emotional state.

[0934] Output: The generated ad text.

[0935] How it works: Using the GPT2 model, we input an emotion-based prompt (e.g., "Generate an advertisement text for a new smartphone that makes the user feel excited") and generate advertisement text.

[0936] Step 6:

[0937] The server generates a voice based on the text generated by the voice generating means.

[0938] Input: The generated ad text.

[0939] Output: The generated audio file.

[0940] What it does: Uses the Tacotron2 model to convert the generated ad text into natural-sounding speech.

[0941] Step 7:

[0942] The server sends the generated advertising content (images, text, audio) to the terminal.

[0943] Input: Generated ad images, text and audio files.

[0944] Output: The ad content sent.

[0945] Specific operation: The server packages the generated content and sends it to the terminal.

[0946] Step 8:

[0947] The terminal displays the received advertisement content to the user.

[0948] Input: Submitted ad content.

[0949] Output: The ad content displayed to the user.

[0950] Specific operation: The terminal displays and plays the received advertising images, text, and audio to the user.

[0951] Step 9:

[0952] The user submits feedback and the server modifies and regenerates the advertising content based on the feedback.

[0953] Input: User feedback.

[0954] Output: The modified and regenerated ad content.

[0955] How it works: The user sends feedback via their device, and the server then modifies and regenerates the ad content as needed.

[0956] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0957] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0958] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0959] [Third embodiment]

[0960] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0961] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0962] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0963] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0964] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0965] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0966] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0967] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0968] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0969] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0970] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0971] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0972] The present invention is embodied in a system including an image generating means, a language generating means, a voice generating means, an input means, and a display means, which supports corporate marketing and provides personalized information to consumers.

[0973] Server Processing

[0974] The server receives data provided by users and companies and creates content using the following three generation methods.

[0975] 1. Image generation method

[0976] The server generates visual content based on product information and target audience data provided by the user. The image generation means generates high-resolution images, for example, using a generative AI model.

[0977] Example: A server generates promotional images for "sportswear." The data includes the color, design, and background image of the clothing, and the server automatically generates an appropriate image based on this data.

[0978] 2. Language generation means

[0979] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in the language expression requested by the user.

[0980] Example: The server generates text describing the features of "new sportswear," such as "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[0981] 3. Voice Generation Method

[0982] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[0983] Example: Generate an audio file that reads the description of the generated sportswear in a professional voice.

[0984] Terminal handling

[0985] The terminal provides an interface for smooth communication between the user and the server.

[0986] 1. Data entry and customization

[0987] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[0988] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "unisex moisture-wicking, quick-drying sportswear, price, main target demographic").

[0989] 2. Product Labeling

[0990] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[0991] Example: Check the promotional sportswear images, introductory text, and audio files displayed on the device.

[0992] User operations

[0993] The user performs the following operations through the terminal.

[0994] 1. Input and Request Submission

[0995] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[0996] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[0997] 2. Feedback and correction requests

[0998] The user can check the generated content and make corrections or additions as needed.

[0999] Example: A user reviews the generated images and text and submits correction requests such as "make the image background a little brighter" or "make the text tone more casual."

[1000] In this way, the system of the present invention can efficiently generate and provide highly personalized marketing content that meets user requests.

[1001] The processing flow will be explained below.

[1002] Step 1:

[1003] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "latest running shoes," along with their features, target age group, and lifestyle.

[1004] Step 2:

[1005] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[1006] Step 3:

[1007] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[1008] Step 4:

[1009] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[1010] Step 5:

[1011] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme "latest running shoes."

[1012] Step 6:

[1013] The server generates text content using language generation tools. It uses a generative AI model to create product descriptions and reviews. For example, it generates a description such as, "These running shoes are lightweight and have excellent cushioning."

[1014] Step 7:

[1015] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[1016] Step 8:

[1017] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[1018] Step 9:

[1019] The device displays the result to the user. The device displays the result received from the server on the screen so that the user can review it. The screen displays an image of the promotional running shoes, introductory text, and an option to play the audio file.

[1020] Step 10:

[1021] Users can review the generated content and provide feedback and correction requests. Users can check the generated content, enter feedback such as "make the background of the image brighter" or "make the tone of the text more casual," and submit it.

[1022] Step 11:

[1023] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the image of the running shoes with a lighter background and adjusting the tone of the text.

[1024] Step 12:

[1025] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[1026] Step 13:

[1027] The device stores the approved content and prepares it for distribution. It stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[1028] Example 1

[1029] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1030] Conventional marketing support systems lack adequate means for efficiently generating highly personalized content that meets user requests. Furthermore, they lack the functionality to incorporate user feedback into the generated content, making it difficult to quickly and accurately revise the content. The present invention aims to solve these problems and more effectively support corporate marketing activities.

[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1032] In this invention, the server includes an image generation means, a language generation means, and a voice generation means. This enables efficient generation of personalized marketing content in response to user requests. The server further includes a means for generating visual content using a generative AI model, a means for generating text using the generative AI model, a means for generating voice based on the generated text, a means for integrating the generated content and sending it to the user, and a means for receiving user feedback on the generated content, correcting it, and regenerating it. This enables content to be quickly and accurately corrected based on user feedback, thereby more effectively supporting corporate marketing activities.

[1033] "Image generation means" is a function for generating visual content (photographs, illustrations, etc.) based on information provided by the user.

[1034] "Language generation means" is a function for generating relevant text based on information and instructions provided by the user.

[1035] The "speech generation means" is a function for generating speech based on the generated text.

[1036] An "input means" is an interface through which a user inputs information into the system.

[1037] The "display means" is an interface for displaying the generated content (images, text, audio) to the user.

[1038] A "generative AI model" is an algorithm or platform that uses machine learning to generate content (images, text, audio, etc.).

[1039] A "prompt" is text containing commands or instructions that are input to a generative AI model to generate specific content.

[1040] The "data analysis means" is a function that analyzes input user information and converts it into a format suitable for the content to be generated.

[1041] The "feedback means" is a function that receives feedback from the user and modifies or regenerates the generated content.

[1042] "Integration means" is a function that combines images, text, and audio into a single piece of content.

[1043] "User feedback" refers to opinions and correction requests from users regarding the generated content.

[1044] The present invention is an advanced content generation system for effectively supporting corporate marketing activities. This system includes image generation means, language generation means, voice generation means, input means, and display means. The following describes in detail the use of specific hardware and software, as well as methods for data processing and data calculation.

[1045] Hardware and Software

[1046] The system is implemented using specific hardware and software as follows:

[1047] Server: A server with high-performance data processing capabilities is used to receive, analyze, generate content, and integrate data.

[1048] Terminal: The user operates the device using a PC, tablet, or other device. This terminal communicates with the server and provides an interface to the user.

[1049] Generative AI models: AI models such as DALL-E and MidJourney are used for image generation, GPT-4 for text generation, and Google Text-to-Speech and Amazon Polly for voice generation.

[1050] Image Generation Means

[1051] The server generates visual content based on product information and target audience data provided by the user. Specifically, it generates images using prompts from a generative AI model (e.g., DALL-E).

[1052] Example: "Generate a scene of a man and woman running in sportswear, moisture-wicking, with a bright background."

[1053] language generation means

[1054] The server has the functionality to generate text about products and services. It generates text using prompt sentences from a generative AI model (e.g., GPT-4).

[1055] Example: "Please explain the features and benefits of your new sportswear in 200 words or less."

[1056] Voice generation means

[1057] The server generates speech based on the generated text, and uses a speech generation AI (e.g., Google Text-to-Speech) to convert the generated text into a voice with a narration that sounds close to a human voice.

[1058] Example: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made with the latest moisture-wicking, quick-drying materials for maximum comfort during exercise.'"

[1059] Data processing and calculation

[1060] The server analyzes the received user data and converts it into a format suitable for the generative AI model. This includes classifying and filtering the data. It also modifies and regenerates the generated content based on feedback provided by the user. To do this, it inputs the prompt sentence into the generative AI model again to generate new content.

[1061] User operations

[1062] The user uses the terminal to input the necessary information through the interface and send a content generation request to the server. The generated content is displayed on the terminal, and the user can check it and send feedback if necessary. The terminal then sends the user's feedback to the server, and the server regenerates the content based on that feedback.

[1063] In this way, the system of the present invention can efficiently generate personalized marketing content in response to user requests and support the marketing activities of companies.

[1064] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1065] Step 1:

[1066] The user uses a terminal to input the necessary data, including product information (e.g., "sportswear"), target audience information (e.g., "runners in their 30s"), and other details required for content creation. The input information is then sent to the server.

[1067] Input: Product information, target audience information

[1068] Output: Data sent to the server

[1069] Step 2:

[1070] The terminal transmits the data entered by the user to the server, which receives the data and prepares it for analysis.

[1071] Input: Data from the user

[1072] Output: Data sent to the server

[1073] Step 3:

[1074] The server analyzes the received data, which may include categorizing and filtering the data, and converts it into a format suitable for the generative AI model.

[1075] Input: Data from the terminal

[1076] Output: Parsed data

[1077] Step 4:

[1078] The server activates the image generation means based on the analyzed data, sending specific prompts to the generative AI model (e.g., DALL-E) to generate visual content.

[1079] Input: Parsed data, prompt statement

[1080] Output: The generated image

[1081] Specifically, the prompt text includes, "Generate a scene of a man and woman running in sportswear, moisture-wicking and quick-drying, with a bright background."

[1082] Step 5:

[1083] The server then activates a language generation mechanism based on the analyzed data, sending specific prompts to a generative AI model (e.g., GPT-4) to generate text.

[1084] Input: Parsed data, prompt statement

[1085] Output: The generated text

[1086] Specific prompts include, "Please explain the features and benefits of your new sportswear in 200 words or less."

[1087] Step 6:

[1088] The server activates a voice generation means based on the generated text, sends the text to a voice generation AI (e.g., Google Text-to-Speech), and generates an audio file.

[1089] Input: Generated text

[1090] Output: Generated audio file

[1091] Specifically, the instructions include: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made from the latest moisture-wicking, quick-drying material to maximize comfort during exercise.'"

[1092] Step 7:

[1093] The server integrates the generated images, text, and audio, and sends the integrated product to the terminal in a format that is easy for the user to view.

[1094] Input: Generated images, text, audio

[1095] Output: The integrated product

[1096] Step 8:

[1097] The server sends the integrated product to the terminal, which receives the product and displays it to the user.

[1098] Input: The integrated product

[1099] Output: The product sent to the terminal

[1100] Step 9:

[1101] The user uses the device to review the generated content and, if desired, enter feedback, which may include requests for image corrections or text adjustments.

[1102] Input: Generated content

[1103] Output: Feedback

[1104] Step 10:

[1105] The terminal sends feedback from the user to the server, which receives this feedback and prepares to modify and regenerate the content.

[1106] Input: User feedback

[1107] Output: Feedback sent to the server

[1108] Step 11:

[1109] The server modifies and regenerates the content based on the feedback, which may involve re-analyzing the data and adjusting the prompts.

[1110] Input: Feedback

[1111] Output: Modified and regenerated content

[1112] Through this process, the system can efficiently generate highly personalized marketing content that meets user requests and supports a company's marketing activities.

[1113] (Application example 1)

[1114] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1115] In today's online shopping environment, consumers cannot actually try on products, making it difficult to experience the feel and comfort of the products. Furthermore, it takes time and effort for consumers to select the right product for themselves. Therefore, there is a need for an easier and more intuitive way for consumers to try on products and understand information about them.

[1116] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1117] In this invention, the server includes an image generation means, a language generation means, a voice generation means, an input means for receiving user input, a display means for displaying the generated content to the user, a means for combining a user image with a product image, and a means for allowing the user to check the generated content, thereby enabling consumers to virtually try on products from the comfort of their own homes and check the product features visually and audibly.

[1118] The "image generation means" is a means for generating visual content based on data and product information provided by the user.

[1119] The "language generation means" is a means for generating text information about products and services.

[1120] The "voice generation means" is a means for generating voice based on the generated text, and creates narration that resembles a human voice.

[1121] The "input means for accepting user input" is an interface for the user to input necessary data.

[1122] The "display means for displaying the generated content to the user" is an interface for displaying the generated content such as images, text, and audio to the user.

[1123] The "means for combining a user image with a product image" refers to a means for combining an acquired user image with a product image, making it appear as if the user is trying on the product.

[1124] The "means for allowing the user to confirm the generated content" refers to a means for presenting the generated content to the user and receiving confirmation of the content and feedback therefrom.

[1125] In order to implement this invention, it is necessary to construct a system that includes an image generation means, a language generation means, a voice generation means, an input means for accepting user input, a display means for displaying the generated content to the user, a means for synthesizing the user image with a product image, and a means for allowing the user to confirm the generated content.

[1126] First, the server receives and analyzes data provided by the user. Specifically, it acquires the user's image data and product information and generates visual content based on this. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate images. This model is capable of generating high-resolution images. For example, a user can take a picture of themselves using a smartphone and send it to the server.

[1127] Next, the server uses language generation tools based on the product information to create a description of the product. A generative AI model (e.g., GPT-3) is used for language generation. For example, when the server generates text to describe the features of "moisture-wicking, quick-drying sportswear," it generates information such as, "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[1128] In addition, the server generates audio based on the generated text. The audio is generated using a speech synthesis AI model (e.g., WaveNet). This model allows for the creation of narration that sounds close to a human voice. Based on the generated description, an audio file introducing the product is generated.

[1129] The user's device displays these artifacts (images, text, and audio). Users can virtually try on items from the comfort of their own home using their smartphone, smart glasses, or head-mounted display. For example, a user can select a particular sportswear item, view its image and description, and listen to an audio description.

[1130] For example:

[1131] 1. The user enters product information for sportswear.

[1132] 2. The server receives image data and product information from the user and creates a synthetic image using a generative AI model.

[1133] 3. Automatically generate product descriptions using a language generation AI model.

[1134] 4. Convert the text generated by the speech synthesis AI model into speech.

[1135] 5. The device displays the generated content to the user, allowing them to virtually try on the clothes at home.

[1136] 6. An example of a prompt sentence is, "Please explain the characteristics of moisture-wicking, quick-drying sportswear."

[1137] Through these steps, consumers can virtually try on products and visually and audibly check the product's features from the comfort of their own home, helping them make product selections more easily and intuitively.

[1138] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1139] Step 1:

[1140] The user takes a picture of themselves using a smartphone or HMD and inputs detailed product information (e.g., the name, features, and size of the moisture-wicking, quick-drying sportswear). This data is sent to the server via the input means. The input data includes the user's image data and product information. The server analyzes this data and converts it into an appropriate format.

[1141] Step 2:

[1142] The server uses an image generation tool to generate visual content based on the image data and product information sent by the user. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate a high-resolution image that synthesizes the user and the product. This image makes it appear as if the user is actually trying on the product. The input is the user image and product information, and the output is the synthesized image.

[1143] Step 3:

[1144] The server uses language generation tools to create product descriptions. It uses a generative AI model (e.g., GPT-3) to generate promotional text. For example, it uses a prompt such as "Please describe the features of moisture-wicking, quick-drying sportswear" to create a detailed product description. The input is product information, and the output is the product description.

[1145] Step 4:

[1146] The server uses a voice generation tool to create a narration based on the generated product description. It uses a speech synthesis AI model (e.g., WaveNet) to generate an audio file that sounds similar to a human voice. This is what the user hears when listening to the product description. The input is the product description, and the output is an audio file.

[1147] Step 5:

[1148] The server sends the generated content (synthetic images, product descriptions, and audio files) to the user's terminal. The terminal displays these contents so that the user can confirm them. Through the display means, the user can have a virtual try-on experience and visually and aurally confirm the generated content. The input is the generated content, and the output is the user's confirmation.

[1149] Step 6:

[1150] The user checks the provided content and inputs feedback as necessary. For example, they send requests for corrections such as "I want the background of the image changed" or "I want the tone of the description to be more casual." The server receives this and corrects and regenerates the content using the regeneration means. The input is the user feedback, and the output is the corrected content.

[1151] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1152] The present invention is embodied in a system including an image generation means, a language generation means, a voice generation means, an input means, a display means, a data analysis means, a feedback response means, and an emotion engine, which recognizes a user's emotions and personalizes content based on the emotions, thereby effectively supporting corporate marketing.

[1153] Server Processing

[1154] The server receives data provided by users and companies and creates content using the following generation methods:

[1155] 1. Image generation method

[1156] The server generates visual content based on user-provided product information and target audience data, for example by using generative AI models to generate high-resolution images.

[1157] Example: A server generates promotional images for the "latest smartphone." The data includes design details and background images, and the appropriate images are automatically generated based on this.

[1158] 2. Language generation means

[1159] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in linguistic expressions that correspond to the user's requests and emotions.

[1160] Example: The server generates text describing the features of a smartphone based on the user's emotions recognized by the emotion engine. The generated content might be something like, "This smartphone has an innovative design and a high-performance camera. It will enrich your daily life."

[1161] 3. Voice Generation Method

[1162] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[1163] Example: Generate an audio file that reads the generated smartphone description in a professional voice.

[1164] Terminal handling

[1165] The terminal provides an interface for smooth communication between the user and the server.

[1166] 1. Data entry and customization

[1167] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[1168] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "latest smartphone, price, target demographic").

[1169] 2. Product Labeling

[1170] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[1171] Example: Check the smartphone promotional images, introductory text, and audio files displayed on the device.

[1172] User operations

[1173] The user performs the following operations through the terminal.

[1174] 1. Input and Request Submission

[1175] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[1176] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[1177] 2. Feedback and correction requests

[1178] The user can check the generated content and make corrections or additions as needed.

[1179] Example: A user reviews the generated images and text and provides feedback such as, "I wish the background of the images was a little brighter" or "I wish the tone of the text was more friendly."

[1180] Use of emotion engine

[1181] The emotion engine has the ability to recognize the user's emotions and reflect them in the generated content.

[1182] 1. Emotion analysis

[1183] The emotion engine analyzes emotions from user input and reactions, such as text input, facial expressions, and tone of voice.

[1184] Example: Determine whether a user has positive or negative emotions based on the comments and feedback they enter into their device.

[1185] 2. Adjust the tone of your content

[1186] The tone and style of the generated content is adjusted based on emotional data analyzed by the emotion engine.

[1187] Example: If the user is feeling positive, the tone of the text will change to be optimistic and uplifting, and if they are feeling negative, the tone will change to be calming and reassuring.

[1188] 3. Real-time emotional response

[1189] The emotion engine analyzes user emotions in real time and instantly modifies and regenerates content accordingly.

[1190] Example: If a user's emotions change (for example, they may initially show interest, but then become disappointed), detect that change and immediately change the way you present content or information.

[1191] In this way, the system of the present invention can provide highly personalized marketing content that takes into account the user's emotions, and effectively support a company's marketing activities.

[1192] The processing flow will be explained below.

[1193] Step 1:

[1194] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "the latest smartphone," along with its features, target age group, and lifestyle.

[1195] Step 2:

[1196] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[1197] Step 3:

[1198] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[1199] Step 4:

[1200] The server uses an emotion engine to analyze the user's emotions. The server analyzes the user's text input, facial expressions, tone of voice, etc. to recognize the emotion. For example, it can detect positive emotions from the user's input.

[1201] Step 5:

[1202] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[1203] Step 6:

[1204] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme of "latest smartphones."

[1205] Step 7:

[1206] The server generates text content using language generation tools. It uses a generative AI model to create product introductions and review text. For example, if the emotion engine detects positive emotions, it generates an introduction such as, "This smartphone features an innovative design and a high-performance camera. It will enrich your daily life."

[1207] Step 8:

[1208] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[1209] Step 9:

[1210] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[1211] Step 10:

[1212] The device displays the product to the user. The device displays the product received from the server on the screen so that the user can review it. The screen displays an image of the promotional smartphone, introductory text, and an option to play the audio file.

[1213] Step 11:

[1214] Users can review the generated content and provide feedback and correction requests. Users can review the generated content and provide feedback such as "I would like the background of the images to be brighter" or "I would like the tone of the text to be more friendly."

[1215] Step 12:

[1216] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the smartphone image with a lighter background and adjusting the tone of the text.

[1217] Step 13:

[1218] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[1219] Step 14:

[1220] The device stores the approved content and prepares it for distribution. It also stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[1221] Step 15:

[1222] The server monitors the user's reaction after distribution and evaluates it in real time using an emotion engine. For example, it analyzes how the user felt about the content and uses the feedback to generate the next content if necessary.

[1223] Example 2

[1224] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1225] Conventional marketing support systems have had difficulty generating highly personalized content that takes user emotions into account, and have been unable to respond to changes in user emotions in real time, resulting in reduced marketing effectiveness.

[1226] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1227] In this invention, the server includes means for receiving data provided by a user, image generation means for generating an image based on the received data, language generation means for generating text based on the received data, voice generation means for generating voice based on the generated text, input means for accepting user input, display means for displaying the generated content to the user, and an emotion engine for analyzing the user's emotions. This makes it possible to analyze the user's emotions in real time and generate and adjust images, text, and voice based on the analyzed emotions.

[1228] "Means for receiving data provided by users" refers to the interface and mechanism for receiving product information and target audience information sent by users and companies.

[1229] "Image generation means" refers to a device or software that creates visual content based on received data, specifically a mechanism that generates high-resolution images using a generative AI model.

[1230] "Language generation means" refers to a device or software that creates text based on received data and has the ability to generate descriptions and comments about products and services using generative AI models.

[1231] "Speech generation means" refers to a device or software that creates an audio file based on the generated text and has a mechanism for generating narration that resembles a human voice.

[1232] "Input means for accepting user input" refers to an interface for a user to input information and a mechanism for receiving data.

[1233] The "display means for displaying generated content to the user" refers to an interface and its display mechanism that allows the user to check generated content such as images, text, and audio.

[1234] An "emotion engine that analyzes user emotions" refers to a device or software that analyzes a user's input and reactions to infer their emotions, and has the ability to adjust the tone and style of the content it generates based on this.

[1235] "Data analysis means" refers to a device or software that analyzes input user information and converts it into a format suitable for the content to be generated.

[1236] "Means for generating images, text, and audio based on prompts using a generative AI model" refers to a generative AI model and its operating mechanism for inputting prompts to generate high-resolution images, appropriate text, natural-sounding audio, etc.

[1237] "Means for analyzing user emotions in real time and modifying and regenerating content in response to emotional changes" refers to a device or software that monitors a user's facial expressions, tone of voice, text input, etc. in real time and instantly adjusts the generated content in response to those emotional changes.

[1238] The present invention relates to a system that analyzes user emotions and generates highly personalized content based on those emotions. This system is composed of a server, a terminal, and a user, and each element has a specific function.

[1239] Server configuration and functions

[1240] The server contains the following main facilities:

[1241] 1. Means of receiving user-provided data

[1242] The server receives product information and target audience information provided by users and businesses through a specific interface, which is implemented using technologies such as HTTP requests and API calls.

[1243] 2. Image generation method

[1244] Based on the received data, the server uses a generative AI model to generate high-resolution images. For example, to generate images of the "latest smartphone" for product promotion, the following prompt sentence is used: "Generate a high-resolution promotional image of the latest smartphone."

[1245] 3. Language generation means

[1246] The server uses a generative AI model to generate text about products and services. It selects appropriate linguistic expressions based on the emotional data analyzed by the emotion engine. For example, when generating promotional text, the prompt is: "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera."

[1247] 4. Voice Generation Method

[1248] Based on the generated text, a voice generation AI is used to generate narration that sounds similar to a human voice, allowing users to create professional audio content. The voice generation prompt is "Please generate a natural voice based on the generated text."

[1249] 5. Emotion Engine

[1250] The emotion engine analyzes user input and responses to infer emotions, and this data is used to adjust the tone and style of the generated content.

[1251] Device configuration and functions

[1252] The terminal provides the following features:

[1253] 1. Data entry and customization

[1254] It provides an interface where users can input necessary information. For example, a company's marketing staff can input information about a new product and set specifications for the content to be generated.

[1255] 2. Product Labeling

[1256] It provides an interface that displays and allows users to check the generated content (images, text, audio) sent from the server. Users can check the generated content and enter feedback.

[1257] User operations

[1258] The user performs the following actions through the terminal:

[1259] 1. Input and Request Submission

[1260] A marketer inputs detailed information about a product or service and requests the server to generate content. For example, the marketer inputs basic product information and target information and clicks the submit button.

[1261] 2. Feedback and correction requests

[1262] Users can review the generated content and submit corrections or additions as needed. For example, they can use the device's feedback interface to submit requests such as "make the background brighter" or "make the text more friendly in tone."

[1263] Specific examples

[1264] For example, when generating promotional content for a new smartphone, a marketer might use prompts like this:

[1265] Image generation prompt: "Generate a high-resolution promotional image for the latest smartphone."

[1266] Text generation prompt: "Based on your emotions, create a text advertising the smartphone's innovative design and high-performance camera."

[1267] Speech generation prompt: "Generate natural-sounding speech based on the generated text."

[1268] By using this system, companies can quickly create highly personalized marketing content that responds to user emotions, thereby maximizing the effectiveness of their marketing activities.

[1269] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1270] Step 1:

[1271] Receiving user data

[1272] The server receives product information and target audience information provided by users and businesses. This data is obtained through HTTP requests and API calls.

[1273] Input: Product information (design details, target demographic, sales price, etc.)

[1274] Output: Received user data

[1275] Specific operation: The server accepts requests to the API endpoint and saves user data in the database.

[1276] Step 2:

[1277] Execution of image generation means

[1278] The server generates images based on the received data, using a generative AI model and prompt text to create high-resolution images.

[1279] Input: User data (product information)

[1280] Output: Generated image data

[1281] Specific operation: The generative AI model is given the prompt "Generate a high-resolution promotional image for the latest smartphone," and the image is generated and saved.

[1282] Step 3:

[1283] Execution of language generation methods

[1284] The server generates text about products and services by inputting prompt sentences into the generative AI model based on the received data and the analysis results of the emotion engine.

[1285] Input: User data (product information), sentiment analysis results

[1286] Output: The generated text

[1287] Specific operation: The generative AI model is given the prompt, "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera," and the generated text is saved in a database.

[1288] Step 4:

[1289] Execution of the speech generation means

[1290] The server generates speech based on the generated text, using a speech generation AI to input prompts and create narration.

[1291] Input: Generated text

[1292] Output: Generated audio file

[1293] Specific operation: Enter a prompt into the voice generation AI saying, "Please generate natural-sounding speech based on the generated text," and save the generated voice file.

[1294] Step 5:

[1295] Data Entry and Customization

[1296] The terminal provides an interface where the user can input the necessary information, such as product details and target audience information.

[1297] Input: Product information, target information

[1298] Output: User data sent

[1299] Specific operation: When the user enters the required information into the terminal interface and clicks the "Submit" button, the data is sent to the server.

[1300] Step 6:

[1301] Display of product

[1302] The terminal displays the products (images, text, and audio) sent from the server to the user for confirmation.

[1303] Input: Generated images, text, and audio files

[1304] Output: The product displayed to the user

[1305] Specific behavior: The device interface is updated to display the generated content to the user.

[1306] Step 7:

[1307] Input and Request Submission

[1308] The user inputs the necessary information through the terminal and sends a request for content generation to the server.

[1309] Input: Product information, target information

[1310] Output: Request to server

[1311] Specific operation: When the user enters information into the terminal and clicks the "Generate" button, the request is sent to the server.

[1312] Step 8:

[1313] Feedback and correction requests

[1314] The user checks the generated content and sends requests for corrections or additions as necessary.

[1315] Input: Feedback

[1316] Output: Submitting a fix request

[1317] Specific operation: The user uses the feedback interface on the device to enter correction requests or addition requests and clicks the submit button.

[1318] Step 9:

[1319] Sentiment analysis and content moderation

[1320] The server uses an emotion engine to analyze emotions based on user input and responses, and adjusts the tone and style of the generated content accordingly.

[1321] Input: User-entered data

[1322] Output: Reconciled content

[1323] Specific operation: The emotion engine analyzes the user's data and, based on the results, adjusts and inputs the prompt sentences into the generative AI model to regenerate the content.

[1324] Step 10:

[1325] Real-time emotional response

[1326] The server analyzes the user's emotions in real time and instantly modifies and regenerates content accordingly.

[1327] Input: Real-time user response data

[1328] Output: Instantly modified content

[1329] Specific operation: Monitors user reactions in real time, and if a change in emotion is detected, inputs new prompt sentences into the generative AI model to modify and regenerate content.

[1330] Through the above steps, highly personalized content based on the user's emotions can be generated quickly.

[1331] (Application example 2)

[1332] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1333] Conventional advertising content is uniform and difficult to fully reflect the diverse emotions and needs of users. As a result, ads that do not elicit a positive response from users are sometimes delivered, preventing the effectiveness of marketing activities from being maximized. Another issue is that modifying and regenerating content to reflect user feedback is time-consuming and laborious.

[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1335] In this invention, the server includes emotion analysis means for analyzing user emotions, image generation means for generating visual advertising content, language generation means for generating text for the advertising content, and voice generation means for generating voice based on the generated text, thereby enabling personalized advertising content to be generated and delivered in real time according to the user's emotions.

[1336] "Emotion analysis means" refers to a device or program that analyzes the user's input information and behavioral data and recognizes the user's emotional state.

[1337] The "image generation means" is a device or program that generates visual advertising content based on specific conditions and input data.

[1338] The "language generation means" is a device or program that automatically generates text for advertising content based on specific conditions and input data.

[1339] The "voice generation means" is a device or program that generates natural voice based on the generated text data.

[1340] "Input means" refers to a device or interface for accepting user information or requests.

[1341] "Display means" refers to a device or interface for visually presenting and displaying the generated advertising content to the user.

[1342] "Data analysis means" refers to a device or program that analyzes input user information and emotion data and converts them into an appropriate format.

[1343] The "feedback receiving means" refers to a device or program for receiving feedback from users and correcting and regenerating the generated advertising content.

[1344] The "advertising content distribution means" refers to a device or program for distributing the generated advertising content to a user's device.

[1345] The system for implementing the present invention operates effectively by the mutual cooperation of the server, terminals, and users.

[1346] Server Processing

[1347] The server automatically generates advertising content using multiple generation means having the following functions:

[1348] Emotion analysis means

[1349] The server receives the user's input information and behavioral data and analyzes the user's emotional state using an emotion analysis tool. For this purpose, it uses a specialized emotion analysis library. For example, it analyzes input text and voice data to determine whether the user is in a positive, negative, excited, or other emotional state.

[1350] Image Generation Means

[1351] The server generates visual ad content based on data from the sentiment analysis method, for example using a generative AI model to generate high-resolution images based on prompts such as:

[1352] Example prompt: 'Create an image of a new smartphone that evokes excitement.'

[1353] language generation means

[1354] The server uses a language generation model to generate text for the advertising content, taking into account the sentiment analysis data and generating the text based on the following prompt:

[1355] Example prompt: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[1356] Voice generation means

[1357] The server generates narration audio based on the generated text, using an AI model specialized in voice generation.

[1358] Terminal handling

[1359] The terminal provides an interface to facilitate interaction between the user and the server.

[1360] Data Entry and Customization

[1361] Users input information about new products and target audiences through their devices, which then sends specific data to the server for generating advertising content.

[1362] Display of product

[1363] The content of the generated products (images, text, audio) sent from the server is displayed on the terminal, allowing the user to check the content.

[1364] User operations

[1365] The user performs the following operations through the terminal.

[1366] Input and Request Submission

[1367] The user enters the necessary information and sends a content generation request to the server.

[1368] Feedback and correction requests

[1369] The user checks the generated content and requests corrections or additions as necessary. The server regenerates the content based on that feedback.

[1370] Hardware and software used

[1371] Sentiment analysis method: EmotionDetector library

[1372] Image generation method: Generative AI model using Tensorflow

[1373] Language generation: transformers (GPT2)

[1374] Sound generation method: Tacotron2

[1375] Example prompt sentence:

[1376] Image Generation: 'Create an image of a new smartphone that evokes excitement.'

[1377] Language generation: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[1378] The above system makes it possible to generate personalized advertising content in real time according to the user's emotions, supporting effective marketing activities that attract the user's attention.

[1379] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1380] Step 1:

[1381] The user inputs information about the new product into the terminal and sends a creation request.

[1382] Input: New product details (e.g. latest smartphone, price, target demographic).

[1383] Output: Generated request data.

[1384] Specific operation: The user uses the input interface of the terminal to input information about the new product, and then clicks the generation request button to send the data to the server.

[1385] Step 2:

[1386] The server analyzes the input information received from the user using a data analysis means and converts it into a required format.

[1387] Input: Generation request data (detailed information about new products, sentiment analysis data).

[1388] Output: The input data required to generate the advertising content.

[1389] Specific operation: The server uses data analysis means to analyze the input data and convert it into a standard format. It also uses emotion analysis means to determine the user's emotional state.

[1390] Step 3:

[1391] The server uses emotion analysis means to recognize the emotion of the user.

[1392] Input: User input information, emotion data.

[1393] Output: The user's emotional state.

[1394] Specific operation: The server analyzes the emotional state based on the received data using an emotion analysis means and recognizes emotions such as positive, negative, and excitement.

[1395] Step 4:

[1396] The server generates visual advertising content using an image generating means.

[1397] Input: Generate request data, user's emotional state.

[1398] Output: The generated ad image.

[1399] What it does: It uses a generative AI model to generate high-resolution advertising images based on a prompt (e.g., "Create an image of a new smartphone that evokes excitement").

[1400] Step 5:

[1401] The server generates text for the advertisement content using a language generation means.

[1402] Input: Generate request data, user's emotional state.

[1403] Output: The generated ad text.

[1404] How it works: Using the GPT2 model, we input an emotion-based prompt (e.g., "Generate an advertisement text for a new smartphone that makes the user feel excited") and generate advertisement text.

[1405] Step 6:

[1406] The server generates a voice based on the text generated by the voice generating means.

[1407] Input: The generated ad text.

[1408] Output: The generated audio file.

[1409] What it does: Uses the Tacotron2 model to convert the generated ad text into natural-sounding speech.

[1410] Step 7:

[1411] The server sends the generated advertising content (images, text, audio) to the terminal.

[1412] Input: Generated ad images, text and audio files.

[1413] Output: The ad content sent.

[1414] Specific operation: The server packages the generated content and sends it to the terminal.

[1415] Step 8:

[1416] The terminal displays the received advertisement content to the user.

[1417] Input: Submitted ad content.

[1418] Output: The ad content displayed to the user.

[1419] Specific operation: The terminal displays and plays the received advertising images, text, and audio to the user.

[1420] Step 9:

[1421] The user submits feedback and the server modifies and regenerates the advertising content based on the feedback.

[1422] Input: User feedback.

[1423] Output: The modified and regenerated ad content.

[1424] How it works: The user sends feedback via their device, and the server then modifies and regenerates the ad content as needed.

[1425] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1426] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1427] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1428] [Fourth embodiment]

[1429] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1430] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1431] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1432] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1433] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1434] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1435] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1436] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1437] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1438] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1439] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1440] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1441] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1442] The present invention is embodied in a system including an image generating means, a language generating means, a voice generating means, an input means, and a display means, which supports corporate marketing and provides personalized information to consumers.

[1443] Server Processing

[1444] The server receives data provided by users and companies and creates content using the following three generation methods.

[1445] 1. Image generation method

[1446] The server generates visual content based on product information and target audience data provided by the user. The image generation means generates high-resolution images, for example, using a generative AI model.

[1447] Example: A server generates promotional images for "sportswear." The data includes the color, design, and background image of the clothing, and the server automatically generates an appropriate image based on this data.

[1448] 2. Language generation means

[1449] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in the language expression requested by the user.

[1450] Example: The server generates text describing the features of "new sportswear," such as "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[1451] 3. Voice Generation Method

[1452] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[1453] Example: Generate an audio file that reads the description of the generated sportswear in a professional voice.

[1454] Terminal handling

[1455] The terminal provides an interface for smooth communication between the user and the server.

[1456] 1. Data entry and customization

[1457] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[1458] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "unisex moisture-wicking, quick-drying sportswear, price, main target demographic").

[1459] 2. Product Labeling

[1460] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[1461] Example: Check the promotional sportswear images, introductory text, and audio files displayed on the device.

[1462] User operations

[1463] The user performs the following operations through the terminal.

[1464] 1. Input and Request Submission

[1465] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[1466] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[1467] 2. Feedback and correction requests

[1468] The user can check the generated content and make corrections or additions as needed.

[1469] Example: A user reviews the generated images and text and submits correction requests such as "make the image background a little brighter" or "make the text tone more casual."

[1470] In this way, the system of the present invention can efficiently generate and provide highly personalized marketing content that meets user requests.

[1471] The processing flow will be explained below.

[1472] Step 1:

[1473] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "latest running shoes," along with their features, target age group, and lifestyle.

[1474] Step 2:

[1475] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[1476] Step 3:

[1477] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[1478] Step 4:

[1479] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[1480] Step 5:

[1481] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme "latest running shoes."

[1482] Step 6:

[1483] The server generates text content using language generation tools. It uses a generative AI model to create product descriptions and reviews. For example, it generates a description such as, "These running shoes are lightweight and have excellent cushioning."

[1484] Step 7:

[1485] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[1486] Step 8:

[1487] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[1488] Step 9:

[1489] The device displays the result to the user. The device displays the result received from the server on the screen so that the user can review it. The screen displays an image of the promotional running shoes, introductory text, and an option to play the audio file.

[1490] Step 10:

[1491] Users can review the generated content and provide feedback and correction requests. Users can check the generated content, enter feedback such as "make the background of the image brighter" or "make the tone of the text more casual," and submit it.

[1492] Step 11:

[1493] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the image of the running shoes with a lighter background and adjusting the tone of the text.

[1494] Step 12:

[1495] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[1496] Step 13:

[1497] The device stores the approved content and prepares it for distribution. It stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[1498] Example 1

[1499] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1500] Conventional marketing support systems lack adequate means for efficiently generating highly personalized content that meets user requests. Furthermore, they lack the functionality to incorporate user feedback into the generated content, making it difficult to quickly and accurately revise the content. The present invention aims to solve these problems and more effectively support corporate marketing activities.

[1501] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1502] In this invention, the server includes an image generation means, a language generation means, and a voice generation means. This enables efficient generation of personalized marketing content in response to user requests. The server further includes a means for generating visual content using a generative AI model, a means for generating text using the generative AI model, a means for generating voice based on the generated text, a means for integrating the generated content and sending it to the user, and a means for receiving user feedback on the generated content, correcting it, and regenerating it. This enables content to be quickly and accurately corrected based on user feedback, thereby more effectively supporting corporate marketing activities.

[1503] "Image generation means" is a function for generating visual content (photographs, illustrations, etc.) based on information provided by the user.

[1504] "Language generation means" is a function for generating relevant text based on information and instructions provided by the user.

[1505] The "speech generation means" is a function for generating speech based on the generated text.

[1506] An "input means" is an interface through which a user inputs information into the system.

[1507] The "display means" is an interface for displaying the generated content (images, text, audio) to the user.

[1508] A "generative AI model" is an algorithm or platform that uses machine learning to generate content (images, text, audio, etc.).

[1509] A "prompt" is text containing commands or instructions that are input to a generative AI model to generate specific content.

[1510] The "data analysis means" is a function that analyzes input user information and converts it into a format suitable for the content to be generated.

[1511] The "feedback means" is a function that receives feedback from the user and modifies or regenerates the generated content.

[1512] "Integration means" is a function that combines images, text, and audio into a single piece of content.

[1513] "User feedback" refers to opinions and correction requests from users regarding the generated content.

[1514] The present invention is an advanced content generation system for effectively supporting corporate marketing activities. This system includes image generation means, language generation means, voice generation means, input means, and display means. The following describes in detail the use of specific hardware and software, as well as methods for data processing and data calculation.

[1515] Hardware and Software

[1516] The system is implemented using specific hardware and software as follows:

[1517] Server: A server with high-performance data processing capabilities is used to receive, analyze, generate content, and integrate data.

[1518] Terminal: The user operates the device using a PC, tablet, or other device. This terminal communicates with the server and provides an interface to the user.

[1519] Generative AI models: AI models such as DALL-E and MidJourney are used for image generation, GPT-4 for text generation, and Google Text-to-Speech and Amazon Polly for voice generation.

[1520] Image Generation Means

[1521] The server generates visual content based on product information and target audience data provided by the user. Specifically, it generates images using prompts from a generative AI model (e.g., DALL-E).

[1522] Example: "Generate a scene of a man and woman running in sportswear, moisture-wicking, with a bright background."

[1523] language generation means

[1524] The server has the functionality to generate text about products and services. It generates text using prompt sentences from a generative AI model (e.g., GPT-4).

[1525] Example: "Please explain the features and benefits of your new sportswear in 200 words or less."

[1526] Voice generation means

[1527] The server generates speech based on the generated text, and uses a speech generation AI (e.g., Google Text-to-Speech) to convert the generated text into a voice with a narration that sounds close to a human voice.

[1528] Example: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made with the latest moisture-wicking, quick-drying materials for maximum comfort during exercise.'"

[1529] Data processing and calculation

[1530] The server analyzes the received user data and converts it into a format suitable for the generative AI model. This includes classifying and filtering the data. It also modifies and regenerates the generated content based on feedback provided by the user. To do this, it inputs the prompt sentence into the generative AI model again to generate new content.

[1531] User operations

[1532] The user uses the terminal to input the necessary information through the interface and send a content generation request to the server. The generated content is displayed on the terminal, and the user can check it and send feedback if necessary. The terminal then sends the user's feedback to the server, and the server regenerates the content based on that feedback.

[1533] In this way, the system of the present invention can efficiently generate personalized marketing content in response to user requests and support the marketing activities of companies.

[1534] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1535] Step 1:

[1536] The user uses a terminal to input the necessary data, including product information (e.g., "sportswear"), target audience information (e.g., "runners in their 30s"), and other details required for content creation. The input information is then sent to the server.

[1537] Input: Product information, target audience information

[1538] Output: Data sent to the server

[1539] Step 2:

[1540] The terminal transmits the data entered by the user to the server, which receives the data and prepares it for analysis.

[1541] Input: Data from the user

[1542] Output: Data sent to the server

[1543] Step 3:

[1544] The server analyzes the received data, which may include categorizing and filtering the data, and converts it into a format suitable for the generative AI model.

[1545] Input: Data from the terminal

[1546] Output: Parsed data

[1547] Step 4:

[1548] The server activates the image generation means based on the analyzed data, sending specific prompts to the generative AI model (e.g., DALL-E) to generate visual content.

[1549] Input: Parsed data, prompt statement

[1550] Output: The generated image

[1551] Specifically, the prompt text includes, "Generate a scene of a man and woman running in sportswear, moisture-wicking and quick-drying, with a bright background."

[1552] Step 5:

[1553] The server then activates a language generation mechanism based on the analyzed data, sending specific prompts to a generative AI model (e.g., GPT-4) to generate text.

[1554] Input: Parsed data, prompt statement

[1555] Output: The generated text

[1556] Specific prompts include, "Please explain the features and benefits of your new sportswear in 200 words or less."

[1557] Step 6:

[1558] The server activates a voice generation means based on the generated text, sends the text to a voice generation AI (e.g., Google Text-to-Speech), and generates an audio file.

[1559] Input: Generated text

[1560] Output: Generated audio file

[1561] Specifically, the instructions include: "Generate the following text as an audio file in a professional male voice: 'This sportswear is made from the latest moisture-wicking, quick-drying material to maximize comfort during exercise.'"

[1562] Step 7:

[1563] The server integrates the generated images, text, and audio, and sends the integrated product to the terminal in a format that is easy for the user to view.

[1564] Input: Generated images, text, audio

[1565] Output: The integrated product

[1566] Step 8:

[1567] The server sends the integrated product to the terminal, which receives the product and displays it to the user.

[1568] Input: The integrated product

[1569] Output: The product sent to the terminal

[1570] Step 9:

[1571] The user uses the device to review the generated content and, if desired, enter feedback, which may include requests for image corrections or text adjustments.

[1572] Input: Generated content

[1573] Output: Feedback

[1574] Step 10:

[1575] The terminal sends feedback from the user to the server, which receives this feedback and prepares to modify and regenerate the content.

[1576] Input: User feedback

[1577] Output: Feedback sent to the server

[1578] Step 11:

[1579] The server modifies and regenerates the content based on the feedback, which may involve re-analyzing the data and adjusting the prompts.

[1580] Input: Feedback

[1581] Output: Modified and regenerated content

[1582] Through this process, the system can efficiently generate highly personalized marketing content that meets user requests and supports a company's marketing activities.

[1583] (Application example 1)

[1584] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1585] In today's online shopping environment, consumers cannot actually try on products, making it difficult to experience the feel and comfort of the products. Furthermore, it takes time and effort for consumers to select the right product for themselves. Therefore, there is a need for an easier and more intuitive way for consumers to try on products and understand information about them.

[1586] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1587] In this invention, the server includes an image generation means, a language generation means, a voice generation means, an input means for receiving user input, a display means for displaying the generated content to the user, a means for combining a user image with a product image, and a means for allowing the user to check the generated content, thereby enabling consumers to virtually try on products from the comfort of their own homes and check the product features visually and audibly.

[1588] The "image generation means" is a means for generating visual content based on data and product information provided by the user.

[1589] The "language generation means" is a means for generating text information about products and services.

[1590] The "voice generation means" is a means for generating voice based on the generated text, and creates narration that resembles a human voice.

[1591] The "input means for accepting user input" is an interface for the user to input necessary data.

[1592] The "display means for displaying the generated content to the user" is an interface for displaying the generated content such as images, text, and audio to the user.

[1593] The "means for combining a user image with a product image" refers to a means for combining an acquired user image with a product image, making it appear as if the user is trying on the product.

[1594] The "means for allowing the user to confirm the generated content" refers to a means for presenting the generated content to the user and receiving confirmation of the content and feedback therefrom.

[1595] In order to implement this invention, it is necessary to construct a system that includes an image generation means, a language generation means, a voice generation means, an input means for accepting user input, a display means for displaying the generated content to the user, a means for synthesizing the user image with a product image, and a means for allowing the user to confirm the generated content.

[1596] First, the server receives and analyzes data provided by the user. Specifically, it acquires the user's image data and product information and generates visual content based on this. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate images. This model is capable of generating high-resolution images. For example, a user can take a picture of themselves using a smartphone and send it to the server.

[1597] Next, the server uses language generation tools based on the product information to create a description of the product. A generative AI model (e.g., GPT-3) is used for language generation. For example, when the server generates text to describe the features of "moisture-wicking, quick-drying sportswear," it generates information such as, "This sportswear uses the latest moisture-wicking, quick-drying material to maximize comfort during exercise."

[1598] In addition, the server generates audio based on the generated text. The audio is generated using a speech synthesis AI model (e.g., WaveNet). This model allows for the creation of narration that sounds close to a human voice. Based on the generated description, an audio file introducing the product is generated.

[1599] The user's device displays these artifacts (images, text, and audio). Users can virtually try on items from the comfort of their own home using their smartphone, smart glasses, or head-mounted display. For example, a user can select a particular sportswear item, view its image and description, and listen to an audio description.

[1600] For example:

[1601] 1. The user enters product information for sportswear.

[1602] 2. The server receives image data and product information from the user and creates a synthetic image using a generative AI model.

[1603] 3. Automatically generate product descriptions using a language generation AI model.

[1604] 4. Convert the text generated by the speech synthesis AI model into speech.

[1605] 5. The device displays the generated content to the user, allowing them to virtually try on the clothes at home.

[1606] 6. An example of a prompt sentence is, "Please explain the characteristics of moisture-wicking, quick-drying sportswear."

[1607] Through these steps, consumers can virtually try on products and visually and audibly check the product's features from the comfort of their own home, helping them make product selections more easily and intuitively.

[1608] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1609] Step 1:

[1610] The user takes a picture of themselves using a smartphone or HMD and inputs detailed product information (e.g., the name, features, and size of the moisture-wicking, quick-drying sportswear). This data is sent to the server via the input means. The input data includes the user's image data and product information. The server analyzes this data and converts it into an appropriate format.

[1611] Step 2:

[1612] The server uses an image generation tool to generate visual content based on the image data and product information sent by the user. A generative AI model (e.g., DALL-E, StyleGAN) is used to generate a high-resolution image that synthesizes the user and the product. This image makes it appear as if the user is actually trying on the product. The input is the user image and product information, and the output is the synthesized image.

[1613] Step 3:

[1614] The server uses language generation tools to create product descriptions. It uses a generative AI model (e.g., GPT-3) to generate promotional text. For example, it uses a prompt such as "Please describe the features of moisture-wicking, quick-drying sportswear" to create a detailed product description. The input is product information, and the output is the product description.

[1615] Step 4:

[1616] The server uses a voice generation tool to create a narration based on the generated product description. It uses a speech synthesis AI model (e.g., WaveNet) to generate an audio file that sounds similar to a human voice. This is what the user hears when listening to the product description. The input is the product description, and the output is an audio file.

[1617] Step 5:

[1618] The server sends the generated content (synthetic images, product descriptions, and audio files) to the user's terminal. The terminal displays these contents so that the user can confirm them. Through the display means, the user can have a virtual try-on experience and visually and aurally confirm the generated content. The input is the generated content, and the output is the user's confirmation.

[1619] Step 6:

[1620] The user checks the provided content and inputs feedback as necessary. For example, they send requests for corrections such as "I want the background of the image changed" or "I want the tone of the description to be more casual." The server receives this and corrects and regenerates the content using the regeneration means. The input is the user feedback, and the output is the corrected content.

[1621] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1622] The present invention is embodied in a system including an image generation means, a language generation means, a voice generation means, an input means, a display means, a data analysis means, a feedback response means, and an emotion engine, which recognizes a user's emotions and personalizes content based on the emotions, thereby effectively supporting corporate marketing.

[1623] Server Processing

[1624] The server receives data provided by users and companies and creates content using the following generation methods:

[1625] 1. Image generation method

[1626] The server generates visual content based on user-provided product information and target audience data, for example by using generative AI models to generate high-resolution images.

[1627] Example: A server generates promotional images for the "latest smartphone." The data includes design details and background images, and the appropriate images are automatically generated based on this.

[1628] 2. Language generation means

[1629] The server has the functionality to generate text about products and services. Using a generative AI model, it creates text in linguistic expressions that correspond to the user's requests and emotions.

[1630] Example: The server generates text describing the features of a smartphone based on the user's emotions recognized by the emotion engine. The generated content might be something like, "This smartphone has an innovative design and a high-performance camera. It will enrich your daily life."

[1631] 3. Voice Generation Method

[1632] The server generates audio based on the generated text, using voice generation AI to create narration that sounds close to a human voice.

[1633] Example: Generate an audio file that reads the generated smartphone description in a professional voice.

[1634] Terminal handling

[1635] The terminal provides an interface for smooth communication between the user and the server.

[1636] 1. Data entry and customization

[1637] The terminal provides an interface where users can enter the necessary data, including product details and basic information about the target audience.

[1638] Example: A company's marketing staff uses a terminal to input information about a new product (e.g., "latest smartphone, price, target demographic").

[1639] 2. Product Labeling

[1640] The terminal has a function of displaying the products (images, text, and audio) sent from the server to the user for confirmation.

[1641] Example: Check the smartphone promotional images, introductory text, and audio files displayed on the device.

[1642] User operations

[1643] The user performs the following operations through the terminal.

[1644] 1. Input and Request Submission

[1645] The user inputs the necessary information through the terminal and sends a content generation request to the server.

[1646] Example: A marketer enters the necessary information to generate a promotional piece of content and clicks submit.

[1647] 2. Feedback and correction requests

[1648] The user can check the generated content and make corrections or additions as needed.

[1649] Example: A user reviews the generated images and text and provides feedback such as, "I wish the background of the images was a little brighter" or "I wish the tone of the text was more friendly."

[1650] Use of emotion engine

[1651] The emotion engine has the ability to recognize the user's emotions and reflect them in the generated content.

[1652] 1. Emotion analysis

[1653] The emotion engine analyzes emotions from user input and reactions, such as text input, facial expressions, and tone of voice.

[1654] Example: Determine whether a user has positive or negative emotions based on the comments and feedback they enter into their device.

[1655] 2. Adjust the tone of your content

[1656] The tone and style of the generated content is adjusted based on emotional data analyzed by the emotion engine.

[1657] Example: If the user is feeling positive, the tone of the text will change to be optimistic and uplifting, and if they are feeling negative, the tone will change to be calming and reassuring.

[1658] 3. Real-time emotional response

[1659] The emotion engine analyzes user emotions in real time and instantly modifies and regenerates content accordingly.

[1660] Example: If a user's emotions change (for example, they may initially show interest, but then become disappointed), detect that change and immediately change the way you present content or information.

[1661] In this way, the system of the present invention can provide highly personalized marketing content that takes into account the user's emotions, and effectively support a company's marketing activities.

[1662] The processing flow will be explained below.

[1663] Step 1:

[1664] The user uses the device to input information, specifically details about a new product and basic information about the target audience. For example, the user inputs the product name, such as "the latest smartphone," along with its features, target age group, and lifestyle.

[1665] Step 2:

[1666] The user submits a request to generate content. The entered information is confirmed and the generation request is sent to the server. The user enters a request such as "Please generate images, text, and audio to promote this product" and clicks the submit button.

[1667] Step 3:

[1668] The server receives the user's request. The server receives the data sent from the device and prepares it for analysis. It reads specific product information and target demographic data.

[1669] Step 4:

[1670] The server uses an emotion engine to analyze the user's emotions. The server analyzes the user's text input, facial expressions, tone of voice, etc. to recognize the emotion. For example, it can detect positive emotions from the user's input.

[1671] Step 5:

[1672] The server analyzes the data and converts it into a format suitable for each generation AI. The server uses data analysis methods to normalize information on product features and target demographics, and converts it into a format suitable for image generation AI, language generation AI, and voice generation AI.

[1673] Step 6:

[1674] The server generates visual content using an image generation means, for example, the server generates high-resolution images with the theme of "latest smartphones."

[1675] Step 7:

[1676] The server generates text content using language generation tools. It uses a generative AI model to create product introductions and review text. For example, if the emotion engine detects positive emotions, it generates an introduction such as, "This smartphone features an innovative design and a high-performance camera. It will enrich your daily life."

[1677] Step 8:

[1678] The server generates audio content using a speech generation means, and generates human-like speech based on the generated text. For example, it creates an audio file that reads the generated product description in a professional voice.

[1679] Step 9:

[1680] The server sends the generated content (images, text, audio) to the device, which then compiles these products and forwards them to the device for user review.

[1681] Step 10:

[1682] The device displays the product to the user. The device displays the product received from the server on the screen so that the user can review it. The screen displays an image of the promotional smartphone, introductory text, and an option to play the audio file.

[1683] Step 11:

[1684] Users can review the generated content and provide feedback and correction requests. Users can review the generated content and provide feedback such as "I would like the background of the images to be brighter" or "I would like the tone of the text to be more friendly."

[1685] Step 12:

[1686] The server receives the feedback and modifies and regenerates the content. Based on the user's feedback, the server makes any necessary modifications and regenerates the content, for example regenerating the smartphone image with a lighter background and adjusting the tone of the text.

[1687] Step 13:

[1688] The corrected content is sent back to the device for final confirmation. The user performs a final check and approves the content if there are no problems. They click the approve button to confirm that the content is OK.

[1689] Step 14:

[1690] The device stores the approved content and prepares it for distribution. It also stores the final visual, text, and audio files and distributes them to social media and advertising platforms. At this stage, the generated content is used in actual marketing activities.

[1691] Step 15:

[1692] The server monitors the user's reaction after distribution and evaluates it in real time using an emotion engine. For example, it analyzes how the user felt about the content and uses the feedback to generate the next content if necessary.

[1693] Example 2

[1694] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1695] Conventional marketing support systems have had difficulty generating highly personalized content that takes user emotions into account, and have been unable to respond to changes in user emotions in real time, resulting in reduced marketing effectiveness.

[1696] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1697] In this invention, the server includes means for receiving data provided by a user, image generation means for generating an image based on the received data, language generation means for generating text based on the received data, voice generation means for generating voice based on the generated text, input means for accepting user input, display means for displaying the generated content to the user, and an emotion engine for analyzing the user's emotions. This makes it possible to analyze the user's emotions in real time and generate and adjust images, text, and voice based on the analyzed emotions.

[1698] "Means for receiving data provided by users" refers to the interface and mechanism for receiving product information and target audience information sent by users and companies.

[1699] "Image generation means" refers to a device or software that creates visual content based on received data, specifically a mechanism that generates high-resolution images using a generative AI model.

[1700] "Language generation means" refers to a device or software that creates text based on received data and has the ability to generate descriptions and comments about products and services using generative AI models.

[1701] "Speech generation means" refers to a device or software that creates an audio file based on the generated text and has a mechanism for generating narration that resembles a human voice.

[1702] "Input means for accepting user input" refers to an interface for a user to input information and a mechanism for receiving data.

[1703] The "display means for displaying generated content to the user" refers to an interface and its display mechanism that allows the user to check generated content such as images, text, and audio.

[1704] An "emotion engine that analyzes user emotions" refers to a device or software that analyzes a user's input and reactions to infer their emotions, and has the ability to adjust the tone and style of the content it generates based on this.

[1705] "Data analysis means" refers to a device or software that analyzes input user information and converts it into a format suitable for the content to be generated.

[1706] "Means for generating images, text, and audio based on prompts using a generative AI model" refers to a generative AI model and its operating mechanism for inputting prompts to generate high-resolution images, appropriate text, natural-sounding audio, etc.

[1707] "Means for analyzing user emotions in real time and modifying and regenerating content in response to emotional changes" refers to a device or software that monitors a user's facial expressions, tone of voice, text input, etc. in real time and instantly adjusts the generated content in response to those emotional changes.

[1708] The present invention relates to a system that analyzes user emotions and generates highly personalized content based on those emotions. This system is composed of a server, a terminal, and a user, and each element has a specific function.

[1709] Server configuration and functions

[1710] The server contains the following main facilities:

[1711] 1. Means of receiving user-provided data

[1712] The server receives product information and target audience information provided by users and businesses through a specific interface, which is implemented using technologies such as HTTP requests and API calls.

[1713] 2. Image generation method

[1714] Based on the received data, the server uses a generative AI model to generate high-resolution images. For example, to generate images of the "latest smartphone" for product promotion, the following prompt sentence is used: "Generate a high-resolution promotional image of the latest smartphone."

[1715] 3. Language generation means

[1716] The server uses a generative AI model to generate text about products and services. It selects appropriate linguistic expressions based on the emotional data analyzed by the emotion engine. For example, when generating promotional text, the prompt is: "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera."

[1717] 4. Voice Generation Method

[1718] Based on the generated text, a voice generation AI is used to generate narration that sounds similar to a human voice, allowing users to create professional audio content. The voice generation prompt is "Please generate a natural voice based on the generated text."

[1719] 5. Emotion Engine

[1720] The emotion engine analyzes user input and responses to infer emotions, and this data is used to adjust the tone and style of the generated content.

[1721] Device configuration and functions

[1722] The terminal provides the following features:

[1723] 1. Data entry and customization

[1724] It provides an interface where users can input necessary information. For example, a company's marketing staff can input information about a new product and set specifications for the content to be generated.

[1725] 2. Product Labeling

[1726] It provides an interface that displays and allows users to check the generated content (images, text, audio) sent from the server. Users can check the generated content and enter feedback.

[1727] User operations

[1728] The user performs the following actions through the terminal:

[1729] 1. Input and Request Submission

[1730] A marketer inputs detailed information about a product or service and requests the server to generate content. For example, the marketer inputs basic product information and target information and clicks the submit button.

[1731] 2. Feedback and correction requests

[1732] Users can review the generated content and submit corrections or additions as needed. For example, they can use the device's feedback interface to submit requests such as "make the background brighter" or "make the text more friendly in tone."

[1733] Specific examples

[1734] For example, when generating promotional content for a new smartphone, a marketer might use prompts like this:

[1735] Image generation prompt: "Generate a high-resolution promotional image for the latest smartphone."

[1736] Text generation prompt: "Based on your emotions, create a text advertising the smartphone's innovative design and high-performance camera."

[1737] Speech generation prompt: "Generate natural-sounding speech based on the generated text."

[1738] By using this system, companies can quickly create highly personalized marketing content that responds to user emotions, thereby maximizing the effectiveness of their marketing activities.

[1739] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1740] Step 1:

[1741] Receiving user data

[1742] The server receives product information and target audience information provided by users and businesses. This data is obtained through HTTP requests and API calls.

[1743] Input: Product information (design details, target demographic, sales price, etc.)

[1744] Output: Received user data

[1745] Specific operation: The server accepts requests to the API endpoint and saves user data in the database.

[1746] Step 2:

[1747] Execution of image generation means

[1748] The server generates images based on the received data, using a generative AI model and prompt text to create high-resolution images.

[1749] Input: User data (product information)

[1750] Output: Generated image data

[1751] Specific operation: The generative AI model is given the prompt "Generate a high-resolution promotional image for the latest smartphone," and the image is generated and saved.

[1752] Step 3:

[1753] Execution of language generation methods

[1754] The server generates text about products and services by inputting prompt sentences into the generative AI model based on the received data and the analysis results of the emotion engine.

[1755] Input: User data (product information), sentiment analysis results

[1756] Output: The generated text

[1757] Specific operation: The generative AI model is given the prompt, "Based on emotions, please create text advertising the smartphone's innovative design and high-performance camera," and the generated text is saved in a database.

[1758] Step 4:

[1759] Execution of the speech generation means

[1760] The server generates speech based on the generated text, using a speech generation AI to input prompts and create narration.

[1761] Input: Generated text

[1762] Output: Generated audio file

[1763] Specific operation: Enter a prompt into the voice generation AI saying, "Please generate natural-sounding speech based on the generated text," and save the generated voice file.

[1764] Step 5:

[1765] Data Entry and Customization

[1766] The terminal provides an interface where the user can input the necessary information, such as product details and target audience information.

[1767] Input: Product information, target information

[1768] Output: User data sent

[1769] Specific operation: When the user enters the required information into the terminal interface and clicks the "Submit" button, the data is sent to the server.

[1770] Step 6:

[1771] Display of product

[1772] The terminal displays the products (images, text, and audio) sent from the server to the user for confirmation.

[1773] Input: Generated images, text, and audio files

[1774] Output: The product displayed to the user

[1775] Specific behavior: The device interface is updated to display the generated content to the user.

[1776] Step 7:

[1777] Input and Request Submission

[1778] The user inputs the necessary information through the terminal and sends a request for content generation to the server.

[1779] Input: Product information, target information

[1780] Output: Request to server

[1781] Specific operation: When the user enters information into the terminal and clicks the "Generate" button, the request is sent to the server.

[1782] Step 8:

[1783] Feedback and correction requests

[1784] The user checks the generated content and sends requests for corrections or additions as necessary.

[1785] Input: Feedback

[1786] Output: Submitting a fix request

[1787] Specific operation: The user uses the feedback interface on the device to enter correction requests or addition requests and clicks the submit button.

[1788] Step 9:

[1789] Sentiment analysis and content moderation

[1790] The server uses an emotion engine to analyze emotions based on user input and responses, and adjusts the tone and style of the generated content accordingly.

[1791] Input: User-entered data

[1792] Output: Reconciled content

[1793] Specific operation: The emotion engine analyzes the user's data and, based on the results, adjusts and inputs the prompt sentences into the generative AI model to regenerate the content.

[1794] Step 10:

[1795] Real-time emotional response

[1796] The server analyzes the user's emotions in real time and instantly modifies and regenerates content accordingly.

[1797] Input: Real-time user response data

[1798] Output: Instantly modified content

[1799] Specific operation: Monitors user reactions in real time, and if a change in emotion is detected, inputs new prompt sentences into the generative AI model to modify and regenerate content.

[1800] Through the above steps, highly personalized content based on the user's emotions can be generated quickly.

[1801] (Application example 2)

[1802] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1803] Conventional advertising content is uniform and difficult to fully reflect the diverse emotions and needs of users. As a result, ads that do not elicit a positive response from users are sometimes delivered, preventing the effectiveness of marketing activities from being maximized. Another issue is that modifying and regenerating content to reflect user feedback is time-consuming and laborious.

[1804] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1805] In this invention, the server includes emotion analysis means for analyzing user emotions, image generation means for generating visual advertising content, language generation means for generating text for the advertising content, and voice generation means for generating voice based on the generated text, thereby enabling personalized advertising content to be generated and delivered in real time according to the user's emotions.

[1806] "Emotion analysis means" refers to a device or program that analyzes the user's input information and behavioral data and recognizes the user's emotional state.

[1807] The "image generation means" is a device or program that generates visual advertising content based on specific conditions and input data.

[1808] The "language generation means" is a device or program that automatically generates text for advertising content based on specific conditions and input data.

[1809] The "voice generation means" is a device or program that generates natural voice based on the generated text data.

[1810] "Input means" refers to a device or interface for accepting user information or requests.

[1811] "Display means" refers to a device or interface for visually presenting and displaying the generated advertising content to the user.

[1812] "Data analysis means" refers to a device or program that analyzes input user information and emotion data and converts them into an appropriate format.

[1813] The "feedback receiving means" refers to a device or program for receiving feedback from users and correcting and regenerating the generated advertising content.

[1814] The "advertising content distribution means" refers to a device or program for distributing the generated advertising content to a user's device.

[1815] The system for implementing the present invention operates effectively by the mutual cooperation of the server, terminals, and users.

[1816] Server Processing

[1817] The server automatically generates advertising content using multiple generation means having the following functions:

[1818] Emotion analysis means

[1819] The server receives the user's input information and behavioral data and analyzes the user's emotional state using an emotion analysis tool. For this purpose, it uses a specialized emotion analysis library. For example, it analyzes input text and voice data to determine whether the user is in a positive, negative, excited, or other emotional state.

[1820] Image Generation Means

[1821] The server generates visual ad content based on data from the sentiment analysis method, for example using a generative AI model to generate high-resolution images based on prompts such as:

[1822] Example prompt: 'Create an image of a new smartphone that evokes excitement.'

[1823] language generation means

[1824] The server uses a language generation model to generate text for the advertising content, taking into account the sentiment analysis data and generating the text based on the following prompt:

[1825] Example prompt: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[1826] Voice generation means

[1827] The server generates narration audio based on the generated text, using an AI model specialized in voice generation.

[1828] Terminal handling

[1829] The terminal provides an interface to facilitate interaction between the user and the server.

[1830] Data Entry and Customization

[1831] Users input information about new products and target audiences through their devices, which then sends specific data to the server for generating advertising content.

[1832] Display of product

[1833] The content of the generated products (images, text, audio) sent from the server is displayed on the terminal, allowing the user to check the content.

[1834] User operations

[1835] The user performs the following operations through the terminal.

[1836] Input and Request Submission

[1837] The user enters the necessary information and sends a content generation request to the server.

[1838] Feedback and correction requests

[1839] The user checks the generated content and requests corrections or additions as necessary. The server regenerates the content based on that feedback.

[1840] Hardware and software used

[1841] Sentiment analysis method: EmotionDetector library

[1842] Image generation method: Generative AI model using Tensorflow

[1843] Language generation: transformers (GPT2)

[1844] Sound generation method: Tacotron2

[1845] Example prompt sentence:

[1846] Image Generation: 'Create an image of a new smartphone that evokes excitement.'

[1847] Language generation: 'Generate an advertisement text for a new smartphone that makes the user feel excited.'

[1848] The above system makes it possible to generate personalized advertising content in real time according to the user's emotions, supporting effective marketing activities that attract the user's attention.

[1849] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1850] Step 1:

[1851] The user inputs information about the new product into the terminal and sends a creation request.

[1852] Input: New product details (e.g. latest smartphone, price, target demographic).

[1853] Output: Generated request data.

[1854] Specific operation: The user uses the input interface of the terminal to input information about the new product, and then clicks the generation request button to send the data to the server.

[1855] Step 2:

[1856] The server analyzes the input information received from the user using a data analysis means and converts it into a required format.

[1857] Input: Generation request data (detailed information about new products, sentiment analysis data).

[1858] Output: The input data required to generate the advertising content.

[1859] Specific operation: The server uses data analysis means to analyze the input data and convert it into a standard format. It also uses emotion analysis means to determine the user's emotional state.

[1860] Step 3:

[1861] The server uses emotion analysis means to recognize the emotion of the user.

[1862] Input: User input information, emotion data.

[1863] Output: The user's emotional state.

[1864] Specific operation: The server analyzes the emotional state based on the received data using an emotion analysis means and recognizes emotions such as positive, negative, and excitement.

[1865] Step 4:

[1866] The server generates visual advertising content using an image generating means.

[1867] Input: Generate request data, user's emotional state.

[1868] Output: The generated ad image.

[1869] What it does: It uses a generative AI model to generate high-resolution advertising images based on a prompt (e.g., "Create an image of a new smartphone that evokes excitement").

[1870] Step 5:

[1871] The server generates text for the advertisement content using a language generation means.

[1872] Input: Generate request data, user's emotional state.

[1873] Output: The generated ad text.

[1874] How it works: Using the GPT2 model, we input an emotion-based prompt (e.g., "Generate an advertisement text for a new smartphone that makes the user feel excited") and generate advertisement text.

[1875] Step 6:

[1876] The server generates a voice based on the text generated by the voice generating means.

[1877] Input: The generated ad text.

[1878] Output: The generated audio file.

[1879] What it does: Uses the Tacotron2 model to convert the generated ad text into natural-sounding speech.

[1880] Step 7:

[1881] The server sends the generated advertising content (images, text, audio) to the terminal.

[1882] Input: Generated ad images, text and audio files.

[1883] Output: The ad content sent.

[1884] Specific operation: The server packages the generated content and sends it to the terminal.

[1885] Step 8:

[1886] The terminal displays the received advertisement content to the user.

[1887] Input: Submitted ad content.

[1888] Output: The ad content displayed to the user.

[1889] Specific operation: The terminal displays and plays the received advertising images, text, and audio to the user.

[1890] Step 9:

[1891] The user submits feedback and the server modifies and regenerates the advertising content based on the feedback.

[1892] Input: User feedback.

[1893] Output: The modified and regenerated ad content.

[1894] How it works: The user sends feedback via their device, and the server then modifies and regenerates the ad content as needed.

[1895] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1896] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1897] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1898] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1899] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1900] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1901] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1902] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1903] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1904] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1905] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1906] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1907] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1908] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1909] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1910] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1911] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1912] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1913] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1914] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1915] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1916] The following is further disclosed regarding the above embodiment.

[1917] (Claim 1)

[1918] image generating means;

[1919] a language generation means;

[1920] A voice generating means;

[1921] an input means for accepting user input;

[1922] A system including a display means for displaying generated content to a user.

[1923] (Claim 2)

[1924] 2. The system according to claim 1, further comprising a data analysis means for analyzing input user information and converting it into a format suitable for content to be generated.

[1925] (Claim 3)

[1926] 10. The system of claim 1, further comprising means for accepting feedback and modifying and regenerating the generated content.

[1927] "Example 1"

[1928] (Claim 1)

[1929] image generating means;

[1930] a language generation means;

[1931] A voice generating means;

[1932] an input means for accepting user input;

[1933] display means for displaying the generated content to a user;

[1934] a means for generating visual content using a generative AI model;

[1935] a means for generating text using a generative AI model;

[1936] means for generating speech based on the generated text;

[1937] means for aggregating and transmitting the generated content to the user;

[1938] A system that includes a means for accepting user feedback on generated content and for correcting and regenerating it.

[1939] (Claim 2)

[1940] 2. The system according to claim 1, further comprising a data analysis means for analyzing input user information and converting it into a format suitable for content to be generated.

[1941] (Claim 3)

[1942] 10. The system of claim 1, further comprising means for accepting feedback and modifying and regenerating the generated content.

[1943] "Application Example 1"

[1944] (Claim 1)

[1945] image generating means;

[1946] a language generation means;

[1947] A voice generating means;

[1948] an input means for accepting user input;

[1949] display means for displaying the generated content to a user;

[1950] A means for combining a user image and a product image;

[1951] A system including a means for allowing a user to review generated content.

[1952] (Claim 2)

[1953] 2. The system according to claim 1, further comprising a data analysis means for analyzing input user information and converting it into a format suitable for content to be generated.

[1954] (Claim 3)

[1955] 10. The system of claim 1, further comprising means for accepting feedback and modifying and regenerating the generated content.

[1956] "Example 2: Combining Emotion Engines"

[1957] (Claim 1)

[1958] means for receiving user-provided data;

[1959] an image generating means for generating an image based on the received data;

[1960] a language generation means for generating text based on the received data;

[1961] a speech generation means for generating speech based on the generated text;

[1962] an input means for accepting user input;

[1963] display means for displaying the generated content to a user;

[1964] A system including an emotion engine that analyzes user emotions.

[1965] (Claim 2)

[1966] The system of claim 1 further comprises a data analysis means for analyzing the input user information and converting it into a format suitable for the content to be generated, and a means for generating images, text, and audio based on prompt sentences using a generative AI model.

[1967] (Claim 3)

[1968] The system of claim 1 further comprising: means for adjusting the tone and style of the generated content based on the results of the user's emotion analysis; and means for analyzing the user's emotions in real time and modifying and regenerating the content in response to changes in emotion.

[1969] "Application example 2 when combining emotion engines"

[1970] (Claim 1)

[1971] emotion analysis means for analyzing the emotions of a user;

[1972] an image generating means for generating visual advertising content;

[1973] a language generation means for generating text for advertising content;

[1974] a speech generation means for generating speech based on the generated text;

[1975] an input means for accepting user input;

[1976] a display means for displaying the generated advertising content to a user;

[1977] A system including a means for delivering advertising content.

[1978] (Claim 2)

[1979] 2. The system according to claim 1, further comprising a data analysis means for analyzing input user information and emotion data and converting them into a format suitable for the advertising content to be generated.

[1980] (Claim 3)

[1981] 10. The system of claim 1, further comprising means for accepting feedback and modifying and regenerating the generated advertising content. [Explanation of symbols]

[1982] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. image generating means; a language generation means; a voice generating means; an input means for accepting user input; A system including a display means for displaying generated content to a user.

2. 2. The system according to claim 1, further comprising a data analysis means for analyzing input user information and converting it into a format suitable for the content to be generated.

3. The system of claim 1 further comprising means for accepting feedback, modifying and regenerating the generated content.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A