system
The system efficiently generates and publishes creative text by integrating user input, AI-driven text generation, and user review, overcoming the challenges of routinization and time constraints in routine tasks.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
In modern business operations, routine tasks such as email writing and blog article creation lead to routinization, decreasing information transmission effectiveness and making it difficult to attract recipient interest, while also hindering the creation of creative texts in a timely manner.
A system that includes an input receiving means for user requirements, an analysis means for a server to send requests to a generation AI, a response receiving means for displaying generated text, a correction means for user review, and a transmission means for final confirmation and sending, enabling efficient and creative text generation.
Enables users to easily generate, review, and publish high-quality creative text quickly, addressing the monotony and inefficiency of routine tasks.
Smart Images

Figure 2026064698000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the modern business environment, many operations are standardized. Especially for tasks like email writing and blog article creation, similar content is repeatedly used, leading to a tendency towards routinization. In such a situation, the effectiveness of information transmission decreases, and it becomes difficult to attract the interest of the recipients. Moreover, it is difficult to create creative texts in a short time during busy operations. To solve this problem, a system that can generate efficient and creative texts is needed.
Means for Solving the Problems
[0005] The present invention provides an input receiving means for a user to input the requirements of the text they wish to generate, and an analysis means for a server to analyze the input requirements and send a request to a generation AI. The generation AI has a text generation means for generating text based on the analyzed request, and the server includes a response receiving means for receiving the generated text and displaying it to the user. Furthermore, by providing a system that includes a correction means for the user to review and correct the displayed text, and a transmission means for the user to make a final confirmation of the corrected text and send it, the user can easily and efficiently generate, send, and publish creative text.
[0006] A "user" is a person who uses this system to input the requirements for a document, and then reviews, modifies, and submits the generated document.
[0007] "Input acceptance means" refers to an interface for a user to input the requirements of the text they wish to generate into the system, as well as the device or software that performs that processing.
[0008] A "server" is a computing device or system that receives and analyzes user requests, communicates with a generation AI, and generates and responds with text.
[0009] "Analysis means" refers to functions or devices that perform processing to analyze requirements received by the server from the user and create appropriate requests for the generating AI.
[0010] "Generative AI" is a type of artificial intelligence that refers to algorithms and their execution environments used to generate text based on input information.
[0011] "Text generation means" refers to the processes and functions that enable the generation AI to generate appropriate text based on a request.
[0012] A "response receiving means" refers to a function or device that receives generated text sent from a generation AI, formats it, and displays it to the user.
[0013] "Display means" refers to an interface and display device for visually presenting text received by the server from the AI to the user.
[0014] "Correction tools" refer to functions or interfaces that allow users to review generated text and modify its content as needed.
[0015] "Transmission means" refers to the functions or devices that allow the user to make a final check of the revised text and send it to the specified recipient.
[0016] The "recipient" is the final destination to which the generated text will be sent (for example, an email address or the public URL of a blog).
[0017] The above are the definitions of the important words. [Brief explanation of the drawing]
[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map where multiple emotions are mapped. [Figure 10] Shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0022] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention is a system for avoiding monotony in writing during routine tasks, and provides a method for efficiently generating creative text using generation AI. Specific embodiments of this system will be described below.
[0040] overview
[0041] This system consists of a series of steps: the user inputs the requirements for the text they want to generate, the system analyzes this input and sends a request to the generation AI, and the AI returns the response to the user. The user then modifies the provided text and submits / publishes the final text.
[0042] System Configuration
[0043] This system consists of the following main components:
[0044] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[0045] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives the request response, formats it, and displays it to the user. The server also sends the revised final text to the specified destination.
[0046] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[0047] Program processing
[0048] The program processing of this system is explained below in natural language.
[0049] 1. The user logs into the system via their terminal and opens a screen to enter the requirements for the document they want to generate. For example, the user might enter "I want to create a new product introduction email."
[0050] 2. The terminal formats the user's input into data packets and sends them to the server.
[0051] 3. The server receives the data packet and analyzes its contents. Based on the analysis, it generates an appropriate API request and sends it to the generation AI. For example, it sends a "POST request" to the generation AI's endpoint.
[0052] 4. The generation AI generates text based on the user's requirements and sends the results back to the server. For example, it generates specific text based on a template for a "new product introduction email."
[0053] 5. The server receives the response from the generating AI, formats its content, and sends it back to the user. The formatted response includes readability and formatting.
[0054] 6. The terminal displays the formatted text to the user. The user reviews the displayed content and makes corrections as needed.
[0055] 7. After making revisions, the user performs a final review and approves submission. At this time, the user clicks the "Confirm" button.
[0056] 8. The terminal sends the final, revised text to the server.
[0057] 9. The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0058] Specific example
[0059] Example 1: New product introduction email
[0060] 1. The user enters "Please create a new product introduction email."
[0061] 2. The server sends a request to the generation AI and receives the generated email content.
[0062] 3. The user reviews and corrects the displayed email content.
[0063] 4. The server sends the corrected email to the customer list.
[0064] Example 2: Creating a blog post
[0065] 1. The user enters "Please create a blog post about winter events."
[0066] 2. The server sends a request to the generation AI and receives the generated article content.
[0067] 3. The user reviews and edits the article.
[0068] 4. The server posts the corrected article to a specific URL on the blog.
[0069] These embodiments enable the system to easily generate creative and effective text and to quickly send and publish it. The above describes specific forms for carrying out the invention.
[0070] The following describes the processing flow.
[0071] Step 1:
[0072] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[0073] Step 2:
[0074] The terminal formats the user's input into data packets and sends them to the server. These data packets contain the user's requirements.
[0075] Step 3:
[0076] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and generates appropriate API requests.
[0077] Step 4:
[0078] The server sends a request to the generation AI. This request contains detailed requirements for the text to be generated.
[0079] Step 5:
[0080] The generation AI generates text using a predetermined algorithm based on a request received from the server.
[0081] Step 6:
[0082] The server receives the generated text sent back from the AI. The receiving method then formats this text into a format that can be displayed to the user.
[0083] Step 7:
[0084] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[0085] Step 8:
[0086] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0087] Step 9:
[0088] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0089] Step 10:
[0090] The server receives the final-approved text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint. This process allows users to quickly and efficiently generate and publish their desired text.
[0091] The above is a detailed explanation of the program's processing flow.
[0092] (Example 1)
[0093] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] Traditional document creation processes for routine tasks were prone to becoming monotonous and time-consuming. Furthermore, there was no system in place to efficiently generate creative documents based on specific user requirements. Therefore, there is a need for improved document quality and increased work efficiency.
[0095] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0096] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate; a packet generation means for the terminal to format the input requirements as data packets and transmit them; an analysis means for the server to analyze the received requirements and send a request to a generation AI; a text generation means for the generation AI to generate text based on the analyzed request; a response receiving means for the server to format the generated text and display it to the user; a correction means for the user to review and correct the displayed text; and a transmission means for the user to make a final confirmation of the corrected text and transmit it. This makes it possible to efficiently generate creative and high-quality text based on the requirements entered by the user and to quickly transmit and publish it.
[0097] A "user" is someone who accesses the system, inputs the requirements for text generation, and reviews and modifies the generated text.
[0098] A "terminal" is a device that provides an interface consisting of hardware and software for users to access a system and input requirements, verify and modify generated results.
[0099] A "server" is a core computer system that analyzes user requirements, sends requests to the generation AI, receives and formats the generated text, and sends it back to the user.
[0100] "Input acceptance means" refers to the interface and software functions used by the user to input the requirements for the text they wish to generate.
[0101] "Packet generation means" refers to the function by which a terminal formats user input into data packets and sends them to a server.
[0102] "Analysis means" refers to a function in which the server analyzes data packets received from the terminal and sends appropriate requests to the generating AI.
[0103] "Generative AI" refers to an algorithm and its execution environment for generating text based on user requirements.
[0104] A "request" is an API request sent from the server to the AI, which instructs the AI to generate text based on the user's requirements.
[0105] "Text generation means" refers to the function in which the generation AI generates text based on requests received from the server.
[0106] The "response receiving means" is a function in which the server formats the generated results received from the generating AI and sends them back to the user.
[0107] "Correction means" refers to the interface and software functions that allow the user to review the displayed generated results and make corrections as needed.
[0108] "Transmission method" refers to a function that allows the user to make a final check of the revised document and send it to the specified recipient.
[0109] This invention is a system for avoiding the monotony of repetitive writing in routine tasks and for efficiently generating creative text. The system consists of the following main components: a user terminal, a server, and a generation AI model. By combining these, it enables the generation of high-quality text based on user input requirements, and its rapid transmission and publication.
[0110] System Configuration
[0111] User's terminal
[0112] The user's device provides an interface for the user to access the system. This includes common computers, smartphones, tablets, etc. The device formats the user's input and sends it to the server as data packets.
[0113] server
[0114] The server has the core function of receiving requirements from the user, analyzing them, and sending requests to the generative AI model. The server also receives the generated text, formats it, and sends it back to the user. It also plays a role in sending the revised final text to the specified destination.
[0115] Generative AI Models
[0116] A generative AI model is an algorithm and its execution environment for generating text based on user requirements. The generative AI model generates appropriate text based on input prompts, referencing existing documents and knowledge bases.
[0117] Program processing
[0118] This system's program clearly defines the roles of the user, terminal, and server, enabling an efficient and creative text generation process. The specific processing steps are as follows, but the overall flow is shown here.
[0119] User actions
[0120] The user logs into the system using their own device. After logging in, an interface is displayed where the user can enter the requirements for the document they want to generate. For example, the user might enter a specific prompt such as "I want to create a new product introduction email."
[0121] Terminal operation
[0122] The terminal formats the user's input into a JSON data packet and sends it to the server using an HTTP request.
[0123] Server Operations
[0124] The server analyzes the received data packets and extracts the user's requirements. Based on the analysis results, it generates an appropriate API request for the generated AI model and sends the POST request.
[0125] Manipulation of Generative AI Models
[0126] The generation AI model generates text based on prompt messages received from the server. For example, in the case of a "new product introduction email," it generates an email body that includes appropriate product descriptions and promotional messages. The generated results are sent back to the server.
[0127] Server response
[0128] The server formats the text received from the AI generation model and sends it back to the user. The formatted text is then displayed on the user's device.
[0129] User modifications and final confirmation
[0130] The user reviews the displayed text and makes corrections as needed. Once corrections are complete, they click the "Confirm" button, and the device sends the final text to the server.
[0131] Server-based transmission and publication
[0132] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via an SMTP server, and in the case of a blog post, it sends a POST request to the specified API endpoint.
[0133] Specific example
[0134] Example 1: New product introduction email
[0135] The user enters "Please create a new product introduction email." The server sends a request to the AI generation model and receives the generated email content. The user reviews the displayed email content and makes corrections as needed. The server sends the corrected email to the customer list.
[0136] Example 2: Creating a blog post
[0137] The user enters "Please create a blog post about winter events." The server sends a request to the AI generation model and receives the generated article content. The user reviews the article and makes revisions as needed. The server posts the revised article to a specific URL on the blog.
[0138] In this way, this system enables users to easily generate creative and effective text and quickly send and publish it.
[0139] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0140] Step 1: User login and entry of requirements
[0141] The user logs into the system using their own device. The input is the user ID and password, and the output is the login success or failure status.
[0142] After logging in, the user accesses an interface to enter the requirements for the document they want to generate. The input consists of specific prompts, such as a text-based instruction like "I want to create a new product introduction email."
[0143] Step 2: Submitting Requirements
[0144] The terminal converts the prompt text entered by the user into a JSON data packet. The input is a text-based prompt text, and the output is a JSON data packet.
[0145] The terminal sends the formatted data packet to the server as an HTTP request. For example, it might be sent as a POST request.
[0146] Step 3: Requirements analysis and AI request generation
[0147] The server analyzes data packets received from the terminal. The input is data packets in JSON format, and the output is the analyzed requirements information.
[0148] The server generates an API request to the generated AI model based on the content of the received data packets. The input is the parsed requirements information, and the output is the API request. Specifically, it generates a POST request to the endpoint of the generated AI.
[0149] Step 4: Text Generation
[0150] The generative AI model generates text based on prompt messages received from the server. The input is an API request, and the output is the generated text.
[0151] For example, based on a prompt requesting a "new product introduction email," the system generates an email body containing appropriate product descriptions and promotional messages.
[0152] Step 5: Formatting and displaying the generated results
[0153] The server formats the text received from the generative AI model. The input is the generated text, and the output is the formatted text. Formatting includes paragraph breaks and formatting adjustments.
[0154] The server converts the formatted text into JSON format and sends it back to the terminal.
[0155] Step 6: User review and correction
[0156] The terminal displays generated text received from the server to the user. The input is data in JSON format, and the output is a display on the user interface.
[0157] The user reviews the displayed text and makes corrections as needed. The input is the generated text, and the output is the corrected text. The user enters the specific corrections in a text editor.
[0158] Step 7: Final confirmation and submission
[0159] The user reviews the revisions and clicks the "Confirm" button to approve the final submission. The input is the revised text, and the output is the final confirmation status.
[0160] The terminal sends the final text confirmed by the user to the server as a JSON data packet. The input is the revised text, and the output is the data packet.
[0161] Step 8: Send and publish
[0162] The server receives the final, confirmed message and sends it to the specified destination. The input is the final message, and the output is the transmission confirmation status.
[0163] For example, in the case of email, it is sent via an SMTP server, and in the case of a blog post, it is published as a POST request to the specified API endpoint.
[0164] (Application Example 1)
[0165] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0166] Traditional methods have limitations in terms of text generation speed and creativity, making it difficult to continuously generate large volumes of high-quality text. Furthermore, the time and effort required to properly edit and publish the generated text was also a problem. This issue is particularly serious for virtual stores, where rapid updates of new product descriptions and campaign information are essential.
[0167] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0168] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate, an analysis means for the server to analyze the input requirements and send a request to a generation AI, and a text generation means for the generation AI to generate text based on the analyzed request. This makes it possible to quickly generate creative text that meets the user's requirements and to publish the generated text on a designated website.
[0169] A "user" is an individual or group that uses a system to generate or edit text.
[0170] An "input acceptance mechanism" is an interface for users to input the requirements of the text they want to generate into the system.
[0171] A "server" is a central processing unit that receives and analyzes input data from users and sends requests to the generating AI.
[0172] "Analysis means" refers to a device or software that has the function of understanding the requirements entered by the user and converting them into an appropriate request.
[0173] "Generative AI" is an artificial intelligence system that generates text in a specified format based on user requirements.
[0174] "Text generation means" refers to the functions and execution environment for generating text using generation AI.
[0175] A "response receiving means" is a device or software that has the function of receiving a response from a generating AI, formatting it, and displaying it to the user.
[0176] A "correction mechanism" is an interface that allows users to review generated text and make corrections as needed.
[0177] "Transmission method" refers to a function that allows the user to make a final review of the revised text and send it to the specified recipient or platform.
[0178] "Publication means" refers to the process and execution environment for publishing the generated text on a designated website.
[0179] To implement this invention, it is necessary to construct a system in which a user, a server, and a generation AI model work in cooperation. This system performs a series of operations, including the user inputting the requirements for the text they want to generate, analyzing those requirements and sending a request to the generation AI, receiving the generated text and displaying it to the user, performing a final check and sending of the revised text, and publishing it on a designated website.
[0180] The server consists of the following main components:
[0181] 1. Input Reception Method: This is an interface for users to input the requirements of the text they want to generate into the system. This refers to the front-end interface of typical smartphone and PC applications.
[0182] 2. Analysis Method: This function analyzes the input requirements and sends a request to the generation AI. It captures the user's requirements as text data, formats it into an appropriate format, and sends it to the generation AI.
[0183] 3. Generative AI and Text Generation Means: The Generative AI is a system that generates text in a specified format based on user prompts. This Generative AI uses advanced text generation models such as OpenAI®'s GPT-3®.
[0184] 4. Response Reception Method: This function receives responses from the generation AI, formats them, and displays them to the user. It receives the generated text, converts it into a readable format, and sends it back to the user.
[0185] 5. Correction Method: This is an interface for the user to review the generated text and make corrections as needed. After the corrections are complete, a final confirmation is performed.
[0186] 6. Sending and Publishing Means: These are functions for sending and publishing the finalized document to a specified recipient and website. This includes email sending and publishing functions via web APIs.
[0187] Hardware and software to be used
[0188] This system often uses the following hardware and software:
[0189] Hardware: User terminals such as smartphones, tablets, and PCs, and servers.
[0190] Software: Generative AI (e.g., OpenAI GPT-3), HTTP libraries for API calls (e.g., Requests for Python), user interface development frameworks (e.g., React, Flutter®, etc.)
[0191] Example of a prompt
[0192] If a user wants to generate a description of a new product, they might enter a prompt like the following:
[0193] "Please write a description of the new product."
[0194] Specific example
[0195] For example, if a virtual store operator wants to generate a description for a new wireless earphone product, they would use the system in the following way:
[0196] 1. The user opens the smartphone app and enters "Please create a description for the new product."
[0197] 2. The server receives this requirement and uses an analysis tool to send a request to the generating AI.
[0198] 3. The AI generates a response such as, "The new wireless earphones feature the latest technology and offer clear sound quality and long battery life..."
[0199] 4. The server receives this generated text and displays it to the user.
[0200] 5. The user reviews and corrects the displayed text, then presses the "Confirm" button for final confirmation.
[0201] 6. The server receives the final, revised text and publishes it on the designated website.
[0202] Thus, by using this system, users can quickly generate high-quality documents and send or publish them to the appropriate recipients.
[0203] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0204] Step 1:
[0205] The user inputs the requirements for the text they want to generate through a smartphone app. This input is called a "prompt." For example, the user might input "Please create a description of the new product." The entered prompt is temporarily saved on the user's device.
[0206] Step 2:
[0207] The terminal formats the prompt text entered by the user into a data packet and sends it to the server. This data packet is often in JSON format. Specifically, the input requirements are formed into a JSON object and sent to the server as an HTTP POST request.
[0208] Step 3:
[0209] The server analyzes the data packets received from the terminal. Using the analysis tools, it converts the data from the prompt message into a format that is valid for the generation AI. For example, the user's prompt message, "Please create a description of the new product," is sent directly to the generation AI's API endpoint.
[0210] Step 4:
[0211] The server sends a request to the generative AI. Using the HTTP POST method, it sends a request containing a prompt to the API endpoint of the generative AI model (e.g., OpenAI GPT-3). The input is the prompt, and the output is the generated text.
[0212] Step 5:
[0213] The generative AI model generates text based on the received prompt. In this generation process, the AI model constructs appropriate text based on its learned knowledge. The output of the generative AI model is text in the specified format.
[0214] Step 6:
[0215] The server receives the response from the AI and formats the generated text. The generated text is formatted and converted into a user-friendly format. This formatting process includes grammatical checks and formatting adjustments.
[0216] Step 7:
[0217] The server returns the formatted text to the user's terminal. The formatted text is returned to the user's terminal as an HTTP response. The returned text is displayed on the user's terminal.
[0218] Step 8:
[0219] The user reviews the generated text and makes corrections as needed. If the user finds any errors, they can edit them directly within the app. The corrected text is temporarily saved on the user's device.
[0220] Step 9:
[0221] The user performs a final review and presses the "Confirm" button. Once the Confirm button is pressed, the revised text is determined to be the final text. This final text is then sent back to the server.
[0222] Step 10:
[0223] The server sends and publishes the finalized document to the specified recipient and website. For example, it may send it to a specified email address using the email sending function, or publish it to a website using a web API. At this time, the server selects the appropriate sending or publishing method.
[0224] The above outlines the program processing flow of this system. From user input to the final publication of the document, the process can be automated and efficient.
[0225] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0226] overview
[0227] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[0228] System Configuration
[0229] This system consists of the following main components.
[0230] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[0231] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives and formats the request response and displays it to the user. It also sends the revised final text to the specified destination.
[0232] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[0233] 4. Emotion Engine: This refers to an algorithm and its execution environment that recognizes the emotions of the user during input and adjusts the tone and content of the text based on those emotions.
[0234] Program processing
[0235] The program processing of this system is explained below in natural language.
[0236] 1. The user logs into the terminal and opens a screen to enter the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[0237] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[0238] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0239] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[0240] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends an appropriate API request to the AI.
[0241] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[0242] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[0243] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[0244] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0245] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0246] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0247] Specific example
[0248] Example 1: New product introduction email
[0249] 1. The user enters "Please create a new product introduction email."
[0250] 2. The emotion engine recognizes the user's feelings of joy.
[0251] 3. The AI generates text based on a joyful tone.
[0252] 4. The user reviews and corrects the email content displayed.
[0253] 5. The server sends the corrected email to the customer list.
[0254] Example 2: Creating a blog post
[0255] 1. The user enters "Please create a blog post about winter events."
[0256] 2. The emotion engine recognizes the user's emotion of surprise.
[0257] 3. The AI generates text based on a surprised tone.
[0258] 4. The user reviews and edits the article.
[0259] 5. The server posts the corrected article to a specific URL on the blog.
[0260] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions. The above describes a specific embodiment for carrying out the invention.
[0261] The following describes the processing flow.
[0262] Step 1:
[0263] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[0264] Step 2:
[0265] To recognize the user's emotions during input, the device uses its camera and microphone to collect the user's facial expressions and voice. This data is then sent to the emotion engine.
[0266] Step 3:
[0267] The emotion engine analyzes collected facial expressions, voice tone, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0268] Step 4:
[0269] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. The data packets include the user's requirements (e.g., "I want to create an email introducing a new product") and their emotional state (e.g., "joy").
[0270] Step 5:
[0271] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state and generates appropriate API requests.
[0272] Step 6:
[0273] The server sends a request to the generative AI. This request includes detailed requirements and emotional states of the text to be generated.
[0274] Step 7:
[0275] Based on the request received from the server by the generative AI, a text is generated using an algorithm. For example, if the user is in an emotional state of "joy", a text with a positive tone along with that emotion is generated.
[0276] Step 8:
[0277] The server receives the generated text returned from the generative AI. By the receiving means, this text is formatted into a form that can be displayed to the user and returned to the user.
[0278] Step 9:
[0279] The terminal displays the formatted text to the user. The user checks the text displayed through the terminal and modifies it if necessary.
[0280] Step 10:
[0281] The user makes a final check of the modified text and approves the transmission. The user clicks the "Confirm" button.
[0282] Step 11:
[0283] After the user's final check, the terminal formats the text again as a data packet and sends it to the server.
[0284] Step 12:
[0285] The server receives the text that has been finally confirmed and sends it to the specified destination. For example, in the case of an email, the email is sent via an SMTP server, and in the case of a blog post, it is posted to a specified API endpoint.
[0286] The above is a specific description of the processing steps specialized for the invention that combines the emotion engine.
[0287] (Example 2)
[0288] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".
[0289] In a conventional text generation system, simply inputting the requirements of the text that the user wants to generate does not necessarily result in the generated text conforming to the user's emotions and intentions. Therefore, the user has to revise the generated text many times, which is very time-consuming. Also, since the user's emotions are not considered, the generated text has a uniform tone and lacks flexibility according to individual situations. As a result, there is no guarantee that the generated text will be appropriately conveyed to the recipient.
[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means for analyzing the input requirements and emotional state and sending a request to the generation AI, a response receiving means for receiving the generated text and displaying it to the user, and a transmission means for the user to finally confirm and send the revised text. This makes it possible to provide a text generated considering the user's emotions.
[0291] The "user" refers to a person who logs in to the system, inputs the requirements of the text, and checks and revises the generated text.
[0292] A "device" refers to a device used by a user to access the system, acquire emotional data, and send it to the server. Specifically, this includes computers, smartphones, tablets, and other similar devices.
[0293] "Means of collecting emotional data" refers to the functions that a device uses to acquire the user's emotional state. Specifically, this includes functions that capture the user's facial expressions and voice using a camera and microphone.
[0294] A "server" refers to a central computing unit that includes an analysis process that analyzes input requirements and emotional states and sends requests to the generating AI.
[0295] "Analysis means" refers to the function in which the server analyzes the user's input requirements and emotional state, and determines what kind of text generation request to send to the generation AI.
[0296] "Generative AI" refers to an algorithm and its execution environment that generates text based on input requirements and analyzed emotional states.
[0297] "Text generation method" refers to a function provided by the generation AI, and the process of generating text based on analyzed data.
[0298] "Response receiving means" refers to the function that allows the server to receive text sent back from the generating AI and format it into a format that can be displayed to the user.
[0299] "Correction tools" refer to interfaces or tools that allow users to review generated text and make corrections as needed. Specifically, this includes text editors and similar tools.
[0300] "Transmission method" refers to the function used when a user sends a document they have finalized. This includes, for example, functions for sending data via email or a Web API.
[0301] System Overview
[0302] The present invention is a system in which a user inputs requirements for a desired text, and a generation AI creates a text based on those requirements. Furthermore, it combines an emotion engine for recognizing the user's emotion and adjusting the text based on that emotion. As a result, the tone and content of the text are adjusted to better match the user's intention.
[0303] System Configuration
[0304] This system consists of the following main components:
[0305] 1. User Terminal:
[0306] It provides an interface for the user to access the system, input, confirm, modify, and finally send the requirements for the text. General computers, smartphones, tablets, etc. can be used as the terminal.
[0307] 2. Server:
[0308] It is the core part that analyzes the input requirements and emotional state and sends requests to the generation AI. The server receives and formats the request responses and displays them to the user. It also sends the final modified text to the specified destination.
[0309] 3. Generation AI:
[0310] It is an algorithm for generating text based on the user's requirements and its execution environment. GPT-3, BERT, etc. can be used as the generation AI model.
[0311] 4. Emotion Engine:
[0312] It is an algorithm for recognizing the emotion at the time of the user's input and adjusting the tone and content of the text based on that emotion and its execution environment. Libraries such as DeepFace and OpenSmile can be used for emotion analysis.
[0313] Program processing
[0314] The program processing of this system is explained below in natural language.
[0315] 1. The user logs into the terminal and enters the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[0316] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[0317] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0318] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[0319] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends appropriate API requests to the AI.
[0320] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[0321] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[0322] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[0323] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0324] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0325] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0326] Specific example
[0327] Example 1: New product introduction email
[0328] User: "Please create a new product introduction email."
[0329] Emotion Engine: Recognizes the user's emotions of joy.
[0330] Generating AI: Generates text based on a joyful tone.
[0331] User: Review and correct the displayed email content.
[0332] Server: Send the corrected email to the customer list.
[0333] Example 2: Creating a blog post
[0334] User: "Please write a blog post about winter events."
[0335] Emotion Engine: Recognizes the user's emotions of surprise.
[0336] Generating AI: Generates text based on a tone of surprise.
[0337] User: Review and edit the article.
[0338] Server: Posts the corrected article to a specific URL on the blog.
[0339] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions.
[0340] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0341] Understood. Below, I will explain the process in detail, broken down into steps.
[0342] Step 1:
[0343] The user logs into their terminal and enters the requirements for the document they want to generate. Specifically, they access the system's login page using their terminal's browser or application, enter their user ID and password to authenticate, and then enter requirements such as "I want to create a new product introduction email." The data entered is requirement information in text format.
[0344] Step 2:
[0345] The device captures the user's facial expressions and voice using its built-in camera and microphone after user input. Specifically, it acquires data using a webcam or the smartphone's camera and microphone. The acquired data is sent to the emotion engine in real time. The input consists of image data and audio data, and the output is a stream of digital data for analysis.
[0346] Step 3:
[0347] The emotion engine analyzes the user's emotional state based on the received facial expressions and voice tone. Specifically, it uses emotion analysis libraries such as DeepFace and OpenSmile to analyze facial muscle movements and voice pitch to determine emotions such as "joy," "sadness," and "surprise." The input for this step is image data and voice data, and the output is emotional state data as a result of the analysis.
[0348] Step 4:
[0349] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. Specifically, it combines text data and emotional data into a JSON data packet and sends it to the server via an HTTP request. The input for this step is text data and emotional data, and the output is the data packet sent to the server.
[0350] Step 5:
[0351] The server analyzes the received data packets to understand what type of text to generate. Specifically, it uses natural language processing (NLP) algorithms to analyze the user's requirements and sends the appropriate API request to the generation AI. The input to this step is the data packets, and the output is the request data sent to the generation AI.
[0352] Step 6:
[0353] The generative AI generates text based on requirements and emotional states sent from the server. Specifically, it uses generative models such as GPT-3 and BERT to generate text with a positive tone that aligns with the emotional state of "joy." The input for this step is the request data, and the output is the text data of the generated text.
[0354] Step 7:
[0355] The server receives the text returned by the AI generator and formats it into a format that can be displayed to the user. Specifically, it converts it to HTML or JSON format and uses a filtering function to check that the generated text does not contain any inappropriate content. The input for this step is the text data of the text, and the output is the formatted text data.
[0356] Step 8:
[0357] The terminal displays the formatted text returned from the server to the user. Specifically, it uses JavaScript® and HTML elements to display the text on a web page. The input for this step is formatted text data, and the output is the text displayed to the user.
[0358] Step 9:
[0359] The user reviews the displayed text and makes corrections as needed. Specifically, they edit the text using the text editor included in the interface. The input for this step is the displayed text data, and the output is the corrected text data.
[0360] Step 10:
[0361] The user reviews the revised text and clicks the "Confirm" button to authorize submission. Specifically, clicking the confirmation button for submission sends the finalized text. The input for this step is the revised text data, and the output is the submission approval operation data.
[0362] Step 11:
[0363] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server. Specifically, it bundles the corrected text into a JSON data packet and sends it to the server via an HTTP request. The input for this step is the corrected text data, and the output is the data packet sent to the server.
[0364] Step 12:
[0365] The server receives the finalized text and sends it to the specified destination. Specifically, it uses an SMTP server for emails and posts to a specific API endpoint for blog posts. The input for this step is the finalized text data, and the output is confirmation data that the transmission was successful.
[0366] The above is a detailed explanation of the program processing of this system.
[0367] (Application Example 2)
[0368] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0369] Conventional text generation systems have the problem of being unable to generate text that reflects the user's emotions, making it difficult to produce text that aligns with the user's intentions and feelings. Furthermore, especially in customer service, it is necessary to respond in a way that is sensitive to the customer's emotions, but current systems are unable to do this adequately.
[0370] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and sends a request to the generation AI; a response receiving means that receives the generated text and displays it to the user; a transmission means that allows the user to make a final confirmation of the revised text and send it; an emotion recognition means that recognizes the user's emotions at the time of input; and an emotion adjustment means that adjusts the text based on the emotions recognized by the emotion recognition means. This enables the generation of text that reflects the user's emotions and appropriate customer service responses that correspond to the customer's emotions.
[0371] An "input acceptance means" is a means for a user to input the requirements for the text they want to generate.
[0372] "Analysis means" refers to a means of analyzing the input requirements and sending the analysis results as a request to the generating AI.
[0373] A "text generation means" is a means for generating text based on an analyzed request.
[0374] A "response receiving means" is a means of receiving text generated by a generation AI and displaying it to the user.
[0375] "Means of correction" refers to means by which the user can review the displayed text and correct it as needed.
[0376] "Means of transmission" refers to the means by which the user makes a final review of the revised text and then sends it.
[0377] "Emotion recognition means" refers to a means of recognizing the user's emotions at the time of input.
[0378] "Emotion adjustment means" are means for adjusting text based on emotions recognized by emotion recognition means.
[0379] System Configuration
[0380] This invention is a system that recognizes a user's emotions and generates and adjusts text based on those emotions. This system consists of the following main components.
[0381] 1. Input acceptance means
[0382] 2. Analysis method
[0383] 3. Sentence generation means
[0384] 4. Means of receiving replies
[0385] 5. Corrective measures
[0386] 6. Transmission method
[0387] 7. Emotion recognition means
[0388] 8. Emotional regulation tools
[0389] Hardware and software to be used
[0390] Hardware: Customer service robot, facial recognition camera, microphone, tablet display
[0391] Software: Emotion recognition engines (e.g., Google® Cloud Vision API, Microsoft® Azure® Face API), generative AI (e.g., OpenAI GPT series), speech recognition systems (e.g., Google Speech-to-Text API)
[0392] Program processing
[0393] 1. The user enters questions or requests to the customer service robot.
[0394] 2. The emotion recognition means analyzes the user's facial expressions and voice tone to determine their emotional state (e.g., joy, sadness, surprise).
[0395] 3. The analysis device formats the user's input and emotional state into data packets and sends them to the server.
[0396] 4. The server receives the data packet and analyzes its contents. It then sends the analysis results as a request to the generating AI.
[0397] 5. The generation AI generates text based on requirements and emotional state. For example, if the user is in a state of "joy," it will generate text with a positive tone that reflects that emotion.
[0398] 6. The response receiving device receives the generated text, formats it into a format that can be displayed to the user, and sends it back to the user.
[0399] 7. The user reviews the displayed text and makes corrections as needed.
[0400] 8. The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient.
[0401] Specific example
[0402] Example 1: Product Information
[0403] 1. The user asks the customer service robot, "Could you show me some of our new dresses?"
[0404] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that the user is excited.
[0405] 3. The AI generates text in a positive tone that matches the user's excitement, such as, "Here is our new dress. It's a very popular design and has been well-received by many customers."
[0406] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[0407] Example 2: Sales information announcement
[0408] 1. The user asks, "Tell me about this weekend's sale."
[0409] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that they are calm.
[0410] 3. The AI generates text in a calm and informative tone that matches the user's emotions, such as, "This weekend, we're having a sale with up to 50% off. Here are some of our recommended items."
[0411] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[0412] Example of a prompt
[0413] "Hello, welcome to the [store name] store. A customer is asking, 'Could you show me the new dresses?' The customer seems excited. Please respond in a positive tone."
[0414] This system will enable in-store customer service to be more emotionally resonant to the user, improving the customer experience.
[0415] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0416] Step 1:
[0417] The user inputs questions or requests to the customer service robot. For example, the user might input, "Could you show me some of your new dresses?" Input is done via voice or touch input on a tablet. Input data (voice data or text data) is obtained.
[0418] Step 2:
[0419] The emotion recognition system analyzes the user's facial expressions and voice tone to determine their emotional state. Image data of the face obtained from the facial recognition camera and voice data acquired from the microphone are input. Based on this input data, the emotion recognition engine (e.g., Google Cloud Vision API or Microsoft Azure Face API) performs data analysis and outputs the user's emotional state (e.g., joy, sadness, surprise).
[0420] Step 3:
[0421] The analysis device formats the user's input and emotional state into data packets and sends them to the server. The user's entered questions and request data, along with the emotional state data obtained from the emotion recognition device, are packaged to generate data packets for transmission to the server. These data packets are then sent to the server.
[0422] Step 4:
[0423] The server receives the data packet and analyzes its contents. The server analyzes the received data packet to extract the user's question and emotional state. The analysis reveals that the user requested to be shown new dresses and that the user's emotional state is "excited."
[0424] Step 5:
[0425] The server sends a request to the generating AI. Based on the extracted user question and emotional state, it generates a prompt. For example, it sends a request to the generating AI (e.g., OpenAI GPT series) with the prompt "The customer is asking 'Can you show me the new dresses?' The customer seems excited. Please respond in a positive tone." This prompt is input, and the generating AI generates and outputs an appropriate response.
[0426] Step 6:
[0427] The response receiving mechanism receives the generated text and formats it into a format that can be displayed to the user. It receives the response text generated by the generation AI and formats it into a format for display to the user. For example, it might receive the response text, "Here are our new dresses. They are a very popular design and have been well-received by many customers."
[0428] Step 7:
[0429] The user reviews the displayed text and makes corrections as needed. The formatted text is then displayed on the tablet screen. The user reviews the displayed text and makes corrections as needed using touch input or other methods. The corrected text data is then obtained.
[0430] Step 8:
[0431] The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient. The finalized text data is received and sent, for example, to the store's backend system or a specific email address. The submission is complete.
[0432] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0433] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0434] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0435] [Second Embodiment]
[0436] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0437] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0438] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0439] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0440] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0441] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0442] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0443] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0444] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0445] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0446] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0447] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0448] This invention is a system for avoiding monotony in writing during routine tasks, and provides a method for efficiently generating creative text using generation AI. Specific embodiments of this system will be described below.
[0449] overview
[0450] This system consists of a series of steps: the user inputs the requirements for the text they want to generate, the system analyzes this input and sends a request to the generation AI, and the AI returns the response to the user. The user then modifies the provided text and submits / publishes the final text.
[0451] System Configuration
[0452] This system consists of the following main components:
[0453] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[0454] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives the request response, formats it, and displays it to the user. The server also sends the revised final text to the specified destination.
[0455] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[0456] Program processing
[0457] The program processing of this system is explained below in natural language.
[0458] 1. The user logs into the system via their terminal and opens a screen to enter the requirements for the document they want to generate. For example, the user might enter "I want to create a new product introduction email."
[0459] 2. The terminal formats the user's input into data packets and sends them to the server.
[0460] 3. The server receives the data packet and analyzes its contents. Based on the analysis, it generates an appropriate API request and sends it to the generation AI. For example, it sends a "POST request" to the generation AI's endpoint.
[0461] 4. The generation AI generates text based on the user's requirements and sends the results back to the server. For example, it generates specific text based on a template for a "new product introduction email."
[0462] 5. The server receives the response from the generating AI, formats its content, and sends it back to the user. The formatted response includes readability and formatting.
[0463] 6. The terminal displays the formatted text to the user. The user reviews the displayed content and makes corrections as needed.
[0464] 7. After making revisions, the user performs a final review and approves submission. At this time, the user clicks the "Confirm" button.
[0465] 8. The terminal sends the final, revised text to the server.
[0466] 9. The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0467] Specific example
[0468] Example 1: New product introduction email
[0469] 1. The user enters "Please create a new product introduction email."
[0470] 2. The server sends a request to the generation AI and receives the generated email content.
[0471] 3. The user reviews and corrects the displayed email content.
[0472] 4. The server sends the corrected email to the customer list.
[0473] Example 2: Creating a blog post
[0474] 1. The user enters "Please create a blog post about winter events."
[0475] 2. The server sends a request to the generation AI and receives the generated article content.
[0476] 3. The user reviews and edits the article.
[0477] 4. The server posts the corrected article to a specific URL on the blog.
[0478] These embodiments enable the system to easily generate creative and effective text and to quickly send and publish it. The above describes specific forms for carrying out the invention.
[0479] The following describes the processing flow.
[0480] Step 1:
[0481] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[0482] Step 2:
[0483] The terminal formats the user's input into data packets and sends them to the server. These data packets contain the user's requirements.
[0484] Step 3:
[0485] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and generates appropriate API requests.
[0486] Step 4:
[0487] The server sends a request to the generation AI. This request contains detailed requirements for the text to be generated.
[0488] Step 5:
[0489] The generation AI generates text using a predetermined algorithm based on a request received from the server.
[0490] Step 6:
[0491] The server receives the generated text sent back from the AI. The receiving method then formats this text into a format that can be displayed to the user.
[0492] Step 7:
[0493] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[0494] Step 8:
[0495] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0496] Step 9:
[0497] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0498] Step 10:
[0499] The server receives the final-approved text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint. This process allows users to quickly and efficiently generate and publish their desired text.
[0500] The above is a detailed explanation of the program's processing flow.
[0501] (Example 1)
[0502] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0503] Traditional document creation processes for routine tasks were prone to becoming monotonous and time-consuming. Furthermore, there was no system in place to efficiently generate creative documents based on specific user requirements. Therefore, there is a need for improved document quality and increased work efficiency.
[0504] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0505] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate; a packet generation means for the terminal to format the input requirements as data packets and transmit them; an analysis means for the server to analyze the received requirements and send a request to a generation AI; a text generation means for the generation AI to generate text based on the analyzed request; a response receiving means for the server to format the generated text and display it to the user; a correction means for the user to review and correct the displayed text; and a transmission means for the user to make a final confirmation of the corrected text and transmit it. This makes it possible to efficiently generate creative and high-quality text based on the requirements entered by the user and to quickly transmit and publish it.
[0506] A "user" is someone who accesses the system, inputs the requirements for text generation, and reviews and modifies the generated text.
[0507] A "terminal" is a device that provides an interface consisting of hardware and software for users to access a system and input requirements, verify and modify generated results.
[0508] A "server" is a core computer system that analyzes user requirements, sends requests to the generation AI, receives and formats the generated text, and sends it back to the user.
[0509] "Input acceptance means" refers to the interface and software functions used by the user to input the requirements for the text they wish to generate.
[0510] "Packet generation means" refers to the function by which a terminal formats user input into data packets and sends them to a server.
[0511] "Analysis means" refers to a function in which the server analyzes data packets received from the terminal and sends appropriate requests to the generating AI.
[0512] "Generative AI" refers to an algorithm and its execution environment for generating text based on user requirements.
[0513] A "request" is an API request sent from the server to the AI, which instructs the AI to generate text based on the user's requirements.
[0514] "Text generation means" refers to the function in which the generation AI generates text based on requests received from the server.
[0515] The "response receiving means" is a function in which the server formats the generated results received from the generating AI and sends them back to the user.
[0516] "Correction means" refers to the interface and software functions that allow the user to review the displayed generated results and make corrections as needed.
[0517] "Transmission method" refers to a function that allows the user to make a final check of the revised document and send it to the specified recipient.
[0518] This invention is a system for avoiding the monotony of repetitive writing in routine tasks and for efficiently generating creative text. The system consists of the following main components: a user terminal, a server, and a generation AI model. By combining these, it enables the generation of high-quality text based on user input requirements, and its rapid transmission and publication.
[0519] System Configuration
[0520] User's terminal
[0521] The user's device provides an interface for the user to access the system. This includes common computers, smartphones, tablets, etc. The device formats the user's input and sends it to the server as data packets.
[0522] server
[0523] The server has the core function of receiving requirements from the user, analyzing them, and sending requests to the generative AI model. The server also receives the generated text, formats it, and sends it back to the user. It also plays a role in sending the revised final text to the specified destination.
[0524] Generative AI Models
[0525] A generative AI model is an algorithm and its execution environment for generating text based on user requirements. The generative AI model generates appropriate text based on input prompts, referencing existing documents and knowledge bases.
[0526] Program processing
[0527] This system's program clearly defines the roles of the user, terminal, and server, enabling an efficient and creative text generation process. The specific processing steps are as follows, but the overall flow is shown here.
[0528] User actions
[0529] The user logs into the system using their own device. After logging in, an interface is displayed where the user can enter the requirements for the document they want to generate. For example, the user might enter a specific prompt such as "I want to create a new product introduction email."
[0530] Terminal operation
[0531] The terminal formats the user's input into a JSON data packet and sends it to the server using an HTTP request.
[0532] Server Operations
[0533] The server analyzes the received data packets and extracts the user's requirements. Based on the analysis results, it generates an appropriate API request for the generated AI model and sends the POST request.
[0534] Manipulation of Generative AI Models
[0535] The generation AI model generates text based on prompt messages received from the server. For example, in the case of a "new product introduction email," it generates an email body that includes appropriate product descriptions and promotional messages. The generated results are sent back to the server.
[0536] Server response
[0537] The server formats the text received from the AI generation model and sends it back to the user. The formatted text is then displayed on the user's device.
[0538] User modifications and final confirmation
[0539] The user reviews the displayed text and makes corrections as needed. Once corrections are complete, they click the "Confirm" button, and the device sends the final text to the server.
[0540] Server-based transmission and publication
[0541] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via an SMTP server, and in the case of a blog post, it sends a POST request to the specified API endpoint.
[0542] Specific example
[0543] Example 1: New product introduction email
[0544] The user enters "Please create a new product introduction email." The server sends a request to the AI generation model and receives the generated email content. The user reviews the displayed email content and makes corrections as needed. The server sends the corrected email to the customer list.
[0545] Example 2: Creating a blog post
[0546] The user enters "Please create a blog post about winter events." The server sends a request to the AI generation model and receives the generated article content. The user reviews the article and makes revisions as needed. The server posts the revised article to a specific URL on the blog.
[0547] In this way, this system enables users to easily generate creative and effective text and quickly send and publish it.
[0548] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0549] Step 1: User login and entry of requirements
[0550] The user logs into the system using their own device. The input is the user ID and password, and the output is the login success or failure status.
[0551] After logging in, the user accesses an interface to enter the requirements for the document they want to generate. The input consists of specific prompts, such as a text-based instruction like "I want to create a new product introduction email."
[0552] Step 2: Submitting Requirements
[0553] The terminal converts the prompt text entered by the user into a JSON data packet. The input is a text-based prompt text, and the output is a JSON data packet.
[0554] The terminal sends the formatted data packet to the server as an HTTP request. For example, it might be sent as a POST request.
[0555] Step 3: Requirements analysis and AI request generation
[0556] The server analyzes data packets received from the terminal. The input is data packets in JSON format, and the output is the analyzed requirements information.
[0557] The server generates an API request to the generated AI model based on the content of the received data packets. The input is the parsed requirements information, and the output is the API request. Specifically, it generates a POST request to the endpoint of the generated AI.
[0558] Step 4: Text Generation
[0559] The generative AI model generates text based on prompt messages received from the server. The input is an API request, and the output is the generated text.
[0560] For example, based on a prompt requesting a "new product introduction email," the system generates an email body containing appropriate product descriptions and promotional messages.
[0561] Step 5: Formatting and displaying the generated results
[0562] The server formats the text received from the generative AI model. The input is the generated text, and the output is the formatted text. Formatting includes paragraph breaks and formatting adjustments.
[0563] The server converts the formatted text into JSON format and sends it back to the terminal.
[0564] Step 6: User review and correction
[0565] The terminal displays generated text received from the server to the user. The input is data in JSON format, and the output is a display on the user interface.
[0566] The user reviews the displayed text and makes corrections as needed. The input is the generated text, and the output is the corrected text. The user enters the specific corrections in a text editor.
[0567] Step 7: Final confirmation and submission
[0568] The user reviews the revisions and clicks the "Confirm" button to approve the final submission. The input is the revised text, and the output is the final confirmation status.
[0569] The terminal sends the final text confirmed by the user to the server as a JSON data packet. The input is the revised text, and the output is the data packet.
[0570] Step 8: Send and publish
[0571] The server receives the final, confirmed message and sends it to the specified destination. The input is the final message, and the output is the transmission confirmation status.
[0572] For example, in the case of email, it is sent via an SMTP server, and in the case of a blog post, it is published as a POST request to the specified API endpoint.
[0573] (Application Example 1)
[0574] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0575] Traditional methods have limitations in terms of text generation speed and creativity, making it difficult to continuously generate large volumes of high-quality text. Furthermore, the time and effort required to properly edit and publish the generated text was also a problem. This issue is particularly serious for virtual stores, where rapid updates of new product descriptions and campaign information are essential.
[0576] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0577] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate, an analysis means for the server to analyze the input requirements and send a request to a generation AI, and a text generation means for the generation AI to generate text based on the analyzed request. This makes it possible to quickly generate creative text that meets the user's requirements and to publish the generated text on a designated website.
[0578] A "user" is an individual or group that uses a system to generate or edit text.
[0579] An "input acceptance mechanism" is an interface for users to input the requirements of the text they want to generate into the system.
[0580] A "server" is a central processing unit that receives and analyzes input data from users and sends requests to the generating AI.
[0581] "Analysis means" refers to a device or software that has the function of understanding the requirements entered by the user and converting them into an appropriate request.
[0582] "Generative AI" is an artificial intelligence system that generates text in a specified format based on user requirements.
[0583] "Text generation means" refers to the functions and execution environment for generating text using generation AI.
[0584] A "response receiving means" is a device or software that has the function of receiving a response from a generating AI, formatting it, and displaying it to the user.
[0585] A "correction mechanism" is an interface that allows users to review generated text and make corrections as needed.
[0586] "Transmission method" refers to a function that allows the user to make a final review of the revised text and send it to the specified recipient or platform.
[0587] "Publication means" refers to the process and execution environment for publishing the generated text on a designated website.
[0588] To implement this invention, it is necessary to construct a system in which a user, a server, and a generation AI model work in cooperation. This system performs a series of operations, including the user inputting the requirements for the text they want to generate, analyzing those requirements and sending a request to the generation AI, receiving the generated text and displaying it to the user, performing a final check and sending of the revised text, and publishing it on a designated website.
[0589] The server consists of the following main components:
[0590] 1. Input Reception Method: This is an interface for users to input the requirements of the text they want to generate into the system. This refers to the front-end interface of typical smartphone and PC applications.
[0591] 2. Analysis Method: This function analyzes the input requirements and sends a request to the generation AI. It captures the user's requirements as text data, formats it into an appropriate format, and sends it to the generation AI.
[0592] 3. Generative AI and Text Generation Means: The Generative AI is a system that generates text in a specified format based on user prompts. This Generative AI uses advanced text generation models such as OpenAI's GPT-3.
[0593] 4. Response Reception Method: This function receives responses from the generation AI, formats them, and displays them to the user. It receives the generated text, converts it into a readable format, and sends it back to the user.
[0594] 5. Correction Method: This is an interface for the user to review the generated text and make corrections as needed. After the corrections are complete, a final confirmation is performed.
[0595] 6. Sending and Publishing Means: These are functions for sending and publishing the finalized document to a specified recipient and website. This includes email sending and publishing functions via web APIs.
[0596] Hardware and software to be used
[0597] This system often uses the following hardware and software:
[0598] Hardware: User terminals such as smartphones, tablets, and PCs, and servers.
[0599] Software: Generative AI (e.g., OpenAI GPT-3), HTTP libraries for API calls (e.g., Requests for Python), user interface development frameworks (e.g., React, Flutter, etc.)
[0600] Example of a prompt
[0601] If a user wants to generate a description of a new product, they might enter a prompt like the following:
[0602] "Please write a description of the new product."
[0603] Specific example
[0604] For example, if a virtual store operator wants to generate a description for a new wireless earphone product, they would use the system in the following way:
[0605] 1. The user opens the smartphone app and enters "Please create a description for the new product."
[0606] 2. The server receives this requirement and uses an analysis tool to send a request to the generating AI.
[0607] 3. The AI generates a response such as, "The new wireless earphones feature the latest technology and offer clear sound quality and long battery life..."
[0608] 4. The server receives this generated text and displays it to the user.
[0609] 5. The user reviews and corrects the displayed text, then presses the "Confirm" button for final confirmation.
[0610] 6. The server receives the final, revised text and publishes it on the designated website.
[0611] Thus, by using this system, users can quickly generate high-quality documents and send or publish them to the appropriate recipients.
[0612] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0613] Step 1:
[0614] The user inputs the requirements for the text they want to generate through a smartphone app. This input is called a "prompt." For example, the user might input "Please create a description of the new product." The entered prompt is temporarily saved on the user's device.
[0615] Step 2:
[0616] The terminal formats the prompt text entered by the user into a data packet and sends it to the server. This data packet is often in JSON format. Specifically, the input requirements are formed into a JSON object and sent to the server as an HTTP POST request.
[0617] Step 3:
[0618] The server analyzes the data packets received from the terminal. Using the analysis tools, it converts the data from the prompt message into a format that is valid for the generation AI. For example, the user's prompt message, "Please create a description of the new product," is sent directly to the generation AI's API endpoint.
[0619] Step 4:
[0620] The server sends a request to the generative AI. Using the HTTP POST method, it sends a request containing a prompt to the API endpoint of the generative AI model (e.g., OpenAI GPT-3). The input is the prompt, and the output is the generated text.
[0621] Step 5:
[0622] The generative AI model generates text based on the received prompt. In this generation process, the AI model constructs appropriate text based on its learned knowledge. The output of the generative AI model is text in the specified format.
[0623] Step 6:
[0624] The server receives the response from the AI and formats the generated text. The generated text is formatted and converted into a user-friendly format. This formatting process includes grammatical checks and formatting adjustments.
[0625] Step 7:
[0626] The server returns the formatted text to the user's terminal. The formatted text is returned to the user's terminal as an HTTP response. The returned text is displayed on the user's terminal.
[0627] Step 8:
[0628] The user reviews the generated text and makes corrections as needed. If the user finds any errors, they can edit them directly within the app. The corrected text is temporarily saved on the user's device.
[0629] Step 9:
[0630] The user performs a final review and presses the "Confirm" button. Once the Confirm button is pressed, the revised text is determined to be the final text. This final text is then sent back to the server.
[0631] Step 10:
[0632] The server sends and publishes the finalized document to the specified recipient and website. For example, it may send it to a specified email address using the email sending function, or publish it to a website using a web API. At this time, the server selects the appropriate sending or publishing method.
[0633] The above outlines the program processing flow of this system. From user input to the final publication of the document, the process can be automated and efficient.
[0634] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0635] overview
[0636] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[0637] System Configuration
[0638] This system consists of the following main components.
[0639] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[0640] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives and formats the request response and displays it to the user. It also sends the revised final text to the specified destination.
[0641] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[0642] 4. Emotion Engine: This refers to an algorithm and its execution environment that recognizes the emotions of the user during input and adjusts the tone and content of the text based on those emotions.
[0643] Program processing
[0644] The program processing of this system is explained below in natural language.
[0645] 1. The user logs into the terminal and opens a screen to enter the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[0646] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[0647] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0648] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[0649] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends an appropriate API request to the AI.
[0650] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[0651] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[0652] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[0653] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0654] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0655] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0656] Specific example
[0657] Example 1: New product introduction email
[0658] 1. The user enters "Please create a new product introduction email."
[0659] 2. The emotion engine recognizes the user's feelings of joy.
[0660] 3. The AI generates text based on a joyful tone.
[0661] 4. The user reviews and corrects the email content displayed.
[0662] 5. The server sends the corrected email to the customer list.
[0663] Example 2: Creating a blog post
[0664] 1. The user enters "Please create a blog post about winter events."
[0665] 2. The emotion engine recognizes the user's emotion of surprise.
[0666] 3. The AI generates text based on a surprised tone.
[0667] 4. The user reviews and edits the article.
[0668] 5. The server posts the corrected article to a specific URL on the blog.
[0669] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions. The above describes a specific embodiment for carrying out the invention.
[0670] The following describes the processing flow.
[0671] Step 1:
[0672] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[0673] Step 2:
[0674] To recognize the user's emotions during input, the device uses its camera and microphone to collect the user's facial expressions and voice. This data is then sent to the emotion engine.
[0675] Step 3:
[0676] The emotion engine analyzes collected facial expressions, voice tone, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0677] Step 4:
[0678] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. The data packets include the user's requirements (e.g., "I want to create an email introducing a new product") and their emotional state (e.g., "joy").
[0679] Step 5:
[0680] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state and generates appropriate API requests.
[0681] Step 6:
[0682] The server sends a request to the generation AI. This request includes detailed requirements for the text to be generated and its emotional state.
[0683] Step 7:
[0684] The generation AI uses an algorithm to generate text based on requests received from the server. For example, if the user is experiencing the emotion of "joy," it will generate text with a positive tone that aligns with that emotion.
[0685] Step 8:
[0686] The server receives the generated text sent back from the AI. The receiving device formats this text into a format that can be displayed to the user and sends it back to the user.
[0687] Step 9:
[0688] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[0689] Step 10:
[0690] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0691] Step 11:
[0692] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0693] Step 12:
[0694] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0695] The above is a detailed explanation of the processing steps specifically for inventions that combine an emotion engine.
[0696] (Example 2)
[0697] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0698] Conventional text generation systems often fail to accurately reflect the user's emotions and intentions, even when the user inputs the requirements for the text they want to generate. This necessitates numerous revisions, which is extremely time-consuming. Furthermore, because they fail to consider the user's emotions, the generated text tends to have a uniform tone, lacking flexibility to adapt to individual situations. Consequently, there is no guarantee that the generated text will effectively communicate with the recipient.
[0699] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and emotional state and sends a request to the generation AI, a response receiving means that receives the generated text and displays it to the user, and a transmission means that the user makes a final confirmation of the revised text and sends it. This makes it possible to provide text that is generated while taking the user's emotions into consideration.
[0700] A "user" refers to a person who logs into the system, enters the requirements for a document, and then reviews and modifies the generated document.
[0701] A "device" refers to a device used by a user to access the system, acquire emotional data, and send it to the server. Specifically, this includes computers, smartphones, tablets, and other similar devices.
[0702] "Means of collecting emotional data" refers to the functions that a device uses to acquire the user's emotional state. Specifically, this includes functions that capture the user's facial expressions and voice using a camera and microphone.
[0703] A "server" refers to a central computing unit that includes an analysis process that analyzes input requirements and emotional states and sends requests to the generating AI.
[0704] "Analysis means" refers to the function in which the server analyzes the user's input requirements and emotional state, and determines what kind of text generation request to send to the generation AI.
[0705] "Generative AI" refers to an algorithm and its execution environment that generates text based on input requirements and analyzed emotional states.
[0706] "Text generation method" refers to a function provided by the generation AI, and the process of generating text based on analyzed data.
[0707] "Response receiving means" refers to the function that allows the server to receive text sent back from the generating AI and format it into a format that can be displayed to the user.
[0708] "Correction tools" refer to interfaces or tools that allow users to review generated text and make corrections as needed. Specifically, this includes text editors and similar tools.
[0709] "Transmission method" refers to the function used when a user sends a document they have finalized. This includes, for example, functions for sending data via email or a Web API.
[0710] System Overview
[0711] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[0712] System Configuration
[0713] This system consists of the following main components:
[0714] 1. User's device:
[0715] It provides an interface for users to access the system and input, review, modify, and finally submit document requirements. The terminals can include standard computers, smartphones, and tablets.
[0716] 2. Server:
[0717] This is the core component that analyzes the input requirements and emotional state and sends a request to the generation AI. The server receives the request response, formats it, and displays it to the user. It also sends the revised final text to the specified destination.
[0718] 3. Generation AI:
[0719] This is an algorithm and execution environment for generating text based on user requirements. Generative AI models such as GPT-3 and BERT can be used.
[0720] 4. Emotional Engine:
[0721] This is an algorithm and its execution environment for recognizing the user's emotions during input and adjusting the tone and content of the text based on those emotions. Libraries such as DeepFace and OpenSmile can be used for emotion analysis.
[0722] Program processing
[0723] The program processing of this system is explained below in natural language.
[0724] 1. The user logs into the terminal and enters the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[0725] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[0726] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[0727] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[0728] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends appropriate API requests to the AI.
[0729] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[0730] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[0731] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[0732] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0733] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0734] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0735] Specific example
[0736] Example 1: New product introduction email
[0737] User: "Please create a new product introduction email."
[0738] Emotion Engine: Recognizes the user's emotions of joy.
[0739] Generating AI: Generates text based on a joyful tone.
[0740] User: Review and correct the displayed email content.
[0741] Server: Send the corrected email to the customer list.
[0742] Example 2: Creating a blog post
[0743] User: "Please write a blog post about winter events."
[0744] Emotion Engine: Recognizes the user's emotions of surprise.
[0745] Generating AI: Generates text based on a tone of surprise.
[0746] User: Review and edit the article.
[0747] Server: Posts the corrected article to a specific URL on the blog.
[0748] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions.
[0749] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0750] Understood. Below, I will explain the process in detail, broken down into steps.
[0751] Step 1:
[0752] The user logs into their terminal and enters the requirements for the document they want to generate. Specifically, they access the system's login page using their terminal's browser or application, enter their user ID and password to authenticate, and then enter requirements such as "I want to create a new product introduction email." The data entered is requirement information in text format.
[0753] Step 2:
[0754] The device captures the user's facial expressions and voice using its built-in camera and microphone after user input. Specifically, it acquires data using a webcam or the smartphone's camera and microphone. The acquired data is sent to the emotion engine in real time. The input consists of image data and audio data, and the output is a stream of digital data for analysis.
[0755] Step 3:
[0756] The emotion engine analyzes the user's emotional state based on the received facial expressions and voice tone. Specifically, it uses emotion analysis libraries such as DeepFace and OpenSmile to analyze facial muscle movements and voice pitch to determine emotions such as "joy," "sadness," and "surprise." The input for this step is image data and voice data, and the output is emotional state data as a result of the analysis.
[0757] Step 4:
[0758] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. Specifically, it combines text data and emotional data into a JSON data packet and sends it to the server via an HTTP request. The input for this step is text data and emotional data, and the output is the data packet sent to the server.
[0759] Step 5:
[0760] The server analyzes the received data packets to understand what type of text to generate. Specifically, it uses natural language processing (NLP) algorithms to analyze the user's requirements and sends the appropriate API request to the generation AI. The input to this step is the data packets, and the output is the request data sent to the generation AI.
[0761] Step 6:
[0762] The generative AI generates text based on requirements and emotional states sent from the server. Specifically, it uses generative models such as GPT-3 and BERT to generate text with a positive tone that aligns with the emotional state of "joy." The input for this step is the request data, and the output is the text data of the generated text.
[0763] Step 7:
[0764] The server receives the text returned by the AI generator and formats it into a format that can be displayed to the user. Specifically, it converts it to HTML or JSON format and uses a filtering function to check that the generated text does not contain any inappropriate content. The input for this step is the text data of the text, and the output is the formatted text data.
[0765] Step 8:
[0766] The terminal displays the formatted text returned from the server to the user. Specifically, it uses JavaScript and HTML elements to display the text on a web page. The input for this step is formatted text data, and the output is the text displayed to the user.
[0767] Step 9:
[0768] The user reviews the displayed text and makes corrections as needed. Specifically, they edit the text using the text editor included in the interface. The input for this step is the displayed text data, and the output is the corrected text data.
[0769] Step 10:
[0770] The user reviews the revised text and clicks the "Confirm" button to authorize submission. Specifically, clicking the confirmation button for submission sends the finalized text. The input for this step is the revised text data, and the output is the submission approval operation data.
[0771] Step 11:
[0772] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server. Specifically, it bundles the corrected text into a JSON data packet and sends it to the server via an HTTP request. The input for this step is the corrected text data, and the output is the data packet sent to the server.
[0773] Step 12:
[0774] The server receives the finalized text and sends it to the specified destination. Specifically, it uses an SMTP server for emails and posts to a specific API endpoint for blog posts. The input for this step is the finalized text data, and the output is confirmation data that the transmission was successful.
[0775] The above is a detailed explanation of the program processing of this system.
[0776] (Application Example 2)
[0777] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0778] Conventional text generation systems have the problem of being unable to generate text that reflects the user's emotions, making it difficult to produce text that aligns with the user's intentions and feelings. Furthermore, especially in customer service, it is necessary to respond in a way that is sensitive to the customer's emotions, but current systems are unable to do this adequately.
[0779] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and sends a request to the generation AI; a response receiving means that receives the generated text and displays it to the user; a transmission means that allows the user to make a final confirmation of the revised text and send it; an emotion recognition means that recognizes the user's emotions at the time of input; and an emotion adjustment means that adjusts the text based on the emotions recognized by the emotion recognition means. This enables the generation of text that reflects the user's emotions and appropriate customer service responses that correspond to the customer's emotions.
[0780] An "input acceptance means" is a means for a user to input the requirements for the text they want to generate.
[0781] "Analysis means" refers to a means of analyzing the input requirements and sending the analysis results as a request to the generating AI.
[0782] A "text generation means" is a means for generating text based on an analyzed request.
[0783] A "response receiving means" is a means of receiving text generated by a generation AI and displaying it to the user.
[0784] "Means of correction" refers to means by which the user can review the displayed text and correct it as needed.
[0785] "Means of transmission" refers to the means by which the user makes a final review of the revised text and then sends it.
[0786] "Emotion recognition means" refers to a means of recognizing the user's emotions at the time of input.
[0787] "Emotion adjustment means" are means for adjusting text based on emotions recognized by emotion recognition means.
[0788] System Configuration
[0789] This invention is a system that recognizes a user's emotions and generates and adjusts text based on those emotions. This system consists of the following main components.
[0790] 1. Input acceptance means
[0791] 2. Analysis method
[0792] 3. Sentence generation means
[0793] 4. Means of receiving replies
[0794] 5. Corrective measures
[0795] 6. Transmission method
[0796] 7. Emotion recognition means
[0797] 8. Emotional regulation tools
[0798] Hardware and software to be used
[0799] Hardware: Customer service robot, facial recognition camera, microphone, tablet display
[0800] Software: Emotion recognition engines (e.g., Google Cloud Vision API, Microsoft Azure Face API), generative AI (e.g., OpenAI GPT series), speech recognition systems (e.g., Google Speech-to-Text API)
[0801] Program processing
[0802] 1. The user enters questions or requests to the customer service robot.
[0803] 2. The emotion recognition means analyzes the user's facial expressions and voice tone to determine their emotional state (e.g., joy, sadness, surprise).
[0804] 3. The analysis device formats the user's input and emotional state into data packets and sends them to the server.
[0805] 4. The server receives the data packet and analyzes its contents. It then sends the analysis results as a request to the generating AI.
[0806] 5. The generation AI generates text based on requirements and emotional state. For example, if the user is in a state of "joy," it will generate text with a positive tone that reflects that emotion.
[0807] 6. The response receiving device receives the generated text, formats it into a format that can be displayed to the user, and sends it back to the user.
[0808] 7. The user reviews the displayed text and makes corrections as needed.
[0809] 8. The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient.
[0810] Specific example
[0811] Example 1: Product Information
[0812] 1. The user asks the customer service robot, "Could you show me some of our new dresses?"
[0813] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that the user is excited.
[0814] 3. The AI generates text in a positive tone that matches the user's excitement, such as, "Here is our new dress. It's a very popular design and has been well-received by many customers."
[0815] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[0816] Example 2: Sales information announcement
[0817] 1. The user asks, "Tell me about this weekend's sale."
[0818] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that they are calm.
[0819] 3. The AI generates text in a calm and informative tone that matches the user's emotions, such as, "This weekend, we're having a sale with up to 50% off. Here are some of our recommended items."
[0820] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[0821] Example of a prompt
[0822] "Hello, welcome to the [store name] store. A customer is asking, 'Could you show me the new dresses?' The customer seems excited. Please respond in a positive tone."
[0823] This system will enable in-store customer service to be more emotionally resonant to the user, improving the customer experience.
[0824] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0825] Step 1:
[0826] The user inputs questions or requests to the customer service robot. For example, the user might input, "Could you show me some of your new dresses?" Input is done via voice or touch input on a tablet. Input data (voice data or text data) is obtained.
[0827] Step 2:
[0828] The emotion recognition system analyzes the user's facial expressions and voice tone to determine their emotional state. Image data of the face obtained from the facial recognition camera and voice data acquired from the microphone are input. Based on this input data, the emotion recognition engine (e.g., Google Cloud Vision API or Microsoft Azure Face API) performs data analysis and outputs the user's emotional state (e.g., joy, sadness, surprise).
[0829] Step 3:
[0830] The analysis device formats the user's input and emotional state into data packets and sends them to the server. The user's entered questions and request data, along with the emotional state data obtained from the emotion recognition device, are packaged to generate data packets for transmission to the server. These data packets are then sent to the server.
[0831] Step 4:
[0832] The server receives the data packet and analyzes its contents. The server analyzes the received data packet to extract the user's question and emotional state. The analysis reveals that the user requested to be shown new dresses and that the user's emotional state is "excited."
[0833] Step 5:
[0834] The server sends a request to the generating AI. Based on the extracted user question and emotional state, it generates a prompt. For example, it sends a request to the generating AI (e.g., OpenAI GPT series) with the prompt "The customer is asking 'Can you show me the new dresses?' The customer seems excited. Please respond in a positive tone." This prompt is input, and the generating AI generates and outputs an appropriate response.
[0835] Step 6:
[0836] The response receiving mechanism receives the generated text and formats it into a format that can be displayed to the user. It receives the response text generated by the generation AI and formats it into a format for display to the user. For example, it might receive the response text, "Here are our new dresses. They are a very popular design and have been well-received by many customers."
[0837] Step 7:
[0838] The user reviews the displayed text and makes corrections as needed. The formatted text is then displayed on the tablet screen. The user reviews the displayed text and makes corrections as needed using touch input or other methods. The corrected text data is then obtained.
[0839] Step 8:
[0840] The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient. The finalized text data is received and sent, for example, to the store's backend system or a specific email address. The submission is complete.
[0841] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0842] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0843] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0844] [Third Embodiment]
[0845] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0846] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0847] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0848] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0849] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0850] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0851] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0852] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0853] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0854] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0855] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0856] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0857] This invention is a system for avoiding monotony in writing during routine tasks, and provides a method for efficiently generating creative text using generation AI. Specific embodiments of this system will be described below.
[0858] overview
[0859] This system consists of a series of steps: the user inputs the requirements for the text they want to generate, the system analyzes this input and sends a request to the generation AI, and the AI returns the response to the user. The user then modifies the provided text and submits / publishes the final text.
[0860] System Configuration
[0861] This system consists of the following main components:
[0862] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[0863] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives the request response, formats it, and displays it to the user. The server also sends the revised final text to the specified destination.
[0864] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[0865] Program processing
[0866] The program processing of this system is explained below in natural language.
[0867] 1. The user logs into the system via their terminal and opens a screen to enter the requirements for the document they want to generate. For example, the user might enter "I want to create a new product introduction email."
[0868] 2. The terminal formats the user's input into data packets and sends them to the server.
[0869] 3. The server receives the data packet and analyzes its contents. Based on the analysis, it generates an appropriate API request and sends it to the generation AI. For example, it sends a "POST request" to the generation AI's endpoint.
[0870] 4. The generation AI generates text based on the user's requirements and sends the results back to the server. For example, it generates specific text based on a template for a "new product introduction email."
[0871] 5. The server receives the response from the generating AI, formats its content, and sends it back to the user. The formatted response includes readability and formatting.
[0872] 6. The terminal displays the formatted text to the user. The user reviews the displayed content and makes corrections as needed.
[0873] 7. After making revisions, the user performs a final review and approves submission. At this time, the user clicks the "Confirm" button.
[0874] 8. The terminal sends the final, revised text to the server.
[0875] 9. The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[0876] Specific example
[0877] Example 1: New product introduction email
[0878] 1. The user enters "Please create a new product introduction email."
[0879] 2. The server sends a request to the generation AI and receives the generated email content.
[0880] 3. The user reviews and corrects the displayed email content.
[0881] 4. The server sends the corrected email to the customer list.
[0882] Example 2: Creating a blog post
[0883] 1. The user enters "Please create a blog post about winter events."
[0884] 2. The server sends a request to the generation AI and receives the generated article content.
[0885] 3. The user reviews and edits the article.
[0886] 4. The server posts the corrected article to a specific URL on the blog.
[0887] These embodiments enable the system to easily generate creative and effective text and to quickly send and publish it. The above describes specific forms for carrying out the invention.
[0888] The following describes the processing flow.
[0889] Step 1:
[0890] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[0891] Step 2:
[0892] The terminal formats the user's input into data packets and sends them to the server. These data packets contain the user's requirements.
[0893] Step 3:
[0894] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and generates appropriate API requests.
[0895] Step 4:
[0896] The server sends a request to the generation AI. This request contains detailed requirements for the text to be generated.
[0897] Step 5:
[0898] The generation AI generates text using a predetermined algorithm based on a request received from the server.
[0899] Step 6:
[0900] The server receives the generated text sent back from the AI. The receiving method then formats this text into a format that can be displayed to the user.
[0901] Step 7:
[0902] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[0903] Step 8:
[0904] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[0905] Step 9:
[0906] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[0907] Step 10:
[0908] The server receives the final-approved text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint. This process allows users to quickly and efficiently generate and publish their desired text.
[0909] The above is a detailed explanation of the program's processing flow.
[0910] (Example 1)
[0911] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0912] Traditional document creation processes for routine tasks were prone to becoming monotonous and time-consuming. Furthermore, there was no system in place to efficiently generate creative documents based on specific user requirements. Therefore, there is a need for improved document quality and increased work efficiency.
[0913] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0914] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate; a packet generation means for the terminal to format the input requirements as data packets and transmit them; an analysis means for the server to analyze the received requirements and send a request to a generation AI; a text generation means for the generation AI to generate text based on the analyzed request; a response receiving means for the server to format the generated text and display it to the user; a correction means for the user to review and correct the displayed text; and a transmission means for the user to make a final confirmation of the corrected text and transmit it. This makes it possible to efficiently generate creative and high-quality text based on the requirements entered by the user and to quickly transmit and publish it.
[0915] A "user" is someone who accesses the system, inputs the requirements for text generation, and reviews and modifies the generated text.
[0916] A "terminal" is a device that provides an interface consisting of hardware and software for users to access a system and input requirements, verify and modify generated results.
[0917] A "server" is a core computer system that analyzes user requirements, sends requests to the generation AI, receives and formats the generated text, and sends it back to the user.
[0918] "Input acceptance means" refers to the interface and software functions used by the user to input the requirements for the text they wish to generate.
[0919] "Packet generation means" refers to the function by which a terminal formats user input into data packets and sends them to a server.
[0920] "Analysis means" refers to a function in which the server analyzes data packets received from the terminal and sends appropriate requests to the generating AI.
[0921] "Generative AI" refers to an algorithm and its execution environment for generating text based on user requirements.
[0922] A "request" is an API request sent from the server to the AI, which instructs the AI to generate text based on the user's requirements.
[0923] "Text generation means" refers to the function in which the generation AI generates text based on requests received from the server.
[0924] The "response receiving means" is a function in which the server formats the generated results received from the generating AI and sends them back to the user.
[0925] "Correction means" refers to the interface and software functions that allow the user to review the displayed generated results and make corrections as needed.
[0926] "Transmission method" refers to a function that allows the user to make a final check of the revised document and send it to the specified recipient.
[0927] This invention is a system for avoiding the monotony of repetitive writing in routine tasks and for efficiently generating creative text. The system consists of the following main components: a user terminal, a server, and a generation AI model. By combining these, it enables the generation of high-quality text based on user input requirements, and its rapid transmission and publication.
[0928] System Configuration
[0929] User's terminal
[0930] The user's device provides an interface for the user to access the system. This includes common computers, smartphones, tablets, etc. The device formats the user's input and sends it to the server as data packets.
[0931] server
[0932] The server has the core function of receiving requirements from the user, analyzing them, and sending requests to the generative AI model. The server also receives the generated text, formats it, and sends it back to the user. It also plays a role in sending the revised final text to the specified destination.
[0933] Generative AI Models
[0934] A generative AI model is an algorithm and its execution environment for generating text based on user requirements. The generative AI model generates appropriate text based on input prompts, referencing existing documents and knowledge bases.
[0935] Program processing
[0936] This system's program clearly defines the roles of the user, terminal, and server, enabling an efficient and creative text generation process. The specific processing steps are as follows, but the overall flow is shown here.
[0937] User actions
[0938] The user logs into the system using their own device. After logging in, an interface is displayed where the user can enter the requirements for the document they want to generate. For example, the user might enter a specific prompt such as "I want to create a new product introduction email."
[0939] Terminal operation
[0940] The terminal formats the user's input into a JSON data packet and sends it to the server using an HTTP request.
[0941] Server Operations
[0942] The server analyzes the received data packets and extracts the user's requirements. Based on the analysis results, it generates an appropriate API request for the generated AI model and sends the POST request.
[0943] Manipulation of Generative AI Models
[0944] The generation AI model generates text based on prompt messages received from the server. For example, in the case of a "new product introduction email," it generates an email body that includes appropriate product descriptions and promotional messages. The generated results are sent back to the server.
[0945] Server response
[0946] The server formats the text received from the AI generation model and sends it back to the user. The formatted text is then displayed on the user's device.
[0947] User modifications and final confirmation
[0948] The user reviews the displayed text and makes corrections as needed. Once corrections are complete, they click the "Confirm" button, and the device sends the final text to the server.
[0949] Server-based transmission and publication
[0950] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via an SMTP server, and in the case of a blog post, it sends a POST request to the specified API endpoint.
[0951] Specific example
[0952] Example 1: New product introduction email
[0953] The user enters "Please create a new product introduction email." The server sends a request to the AI generation model and receives the generated email content. The user reviews the displayed email content and makes corrections as needed. The server sends the corrected email to the customer list.
[0954] Example 2: Creating a blog post
[0955] The user enters "Please create a blog post about winter events." The server sends a request to the AI generation model and receives the generated article content. The user reviews the article and makes revisions as needed. The server posts the revised article to a specific URL on the blog.
[0956] In this way, this system enables users to easily generate creative and effective text and quickly send and publish it.
[0957] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0958] Step 1: User login and entry of requirements
[0959] The user logs into the system using their own device. The input is the user ID and password, and the output is the login success or failure status.
[0960] After logging in, the user accesses an interface to enter the requirements for the document they want to generate. The input consists of specific prompts, such as a text-based instruction like "I want to create a new product introduction email."
[0961] Step 2: Submitting Requirements
[0962] The terminal converts the prompt text entered by the user into a JSON data packet. The input is a text-based prompt text, and the output is a JSON data packet.
[0963] The terminal sends the formatted data packet to the server as an HTTP request. For example, it might be sent as a POST request.
[0964] Step 3: Requirements analysis and AI request generation
[0965] The server analyzes data packets received from the terminal. The input is data packets in JSON format, and the output is the analyzed requirements information.
[0966] The server generates an API request to the generated AI model based on the content of the received data packets. The input is the parsed requirements information, and the output is the API request. Specifically, it generates a POST request to the endpoint of the generated AI.
[0967] Step 4: Text Generation
[0968] The generative AI model generates text based on prompt messages received from the server. The input is an API request, and the output is the generated text.
[0969] For example, based on a prompt requesting a "new product introduction email," the system generates an email body containing appropriate product descriptions and promotional messages.
[0970] Step 5: Formatting and displaying the generated results
[0971] The server formats the text received from the generative AI model. The input is the generated text, and the output is the formatted text. Formatting includes paragraph breaks and formatting adjustments.
[0972] The server converts the formatted text into JSON format and sends it back to the terminal.
[0973] Step 6: User review and correction
[0974] The terminal displays generated text received from the server to the user. The input is data in JSON format, and the output is a display on the user interface.
[0975] The user reviews the displayed text and makes corrections as needed. The input is the generated text, and the output is the corrected text. The user enters the specific corrections in a text editor.
[0976] Step 7: Final confirmation and submission
[0977] The user reviews the revisions and clicks the "Confirm" button to approve the final submission. The input is the revised text, and the output is the final confirmation status.
[0978] The terminal sends the final text confirmed by the user to the server as a JSON data packet. The input is the revised text, and the output is the data packet.
[0979] Step 8: Send and publish
[0980] The server receives the final, confirmed message and sends it to the specified destination. The input is the final message, and the output is the transmission confirmation status.
[0981] For example, in the case of email, it is sent via an SMTP server, and in the case of a blog post, it is published as a POST request to the specified API endpoint.
[0982] (Application Example 1)
[0983] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0984] Traditional methods have limitations in terms of text generation speed and creativity, making it difficult to continuously generate large volumes of high-quality text. Furthermore, the time and effort required to properly edit and publish the generated text was also a problem. This issue is particularly serious for virtual stores, where rapid updates of new product descriptions and campaign information are essential.
[0985] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0986] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate, an analysis means for the server to analyze the input requirements and send a request to a generation AI, and a text generation means for the generation AI to generate text based on the analyzed request. This makes it possible to quickly generate creative text that meets the user's requirements and to publish the generated text on a designated website.
[0987] A "user" is an individual or group that uses a system to generate or edit text.
[0988] An "input acceptance mechanism" is an interface for users to input the requirements of the text they want to generate into the system.
[0989] A "server" is a central processing unit that receives and analyzes input data from users and sends requests to the generating AI.
[0990] "Analysis means" refers to a device or software that has the function of understanding the requirements entered by the user and converting them into an appropriate request.
[0991] "Generative AI" is an artificial intelligence system that generates text in a specified format based on user requirements.
[0992] "Text generation means" refers to the functions and execution environment for generating text using generation AI.
[0993] A "response receiving means" is a device or software that has the function of receiving a response from a generating AI, formatting it, and displaying it to the user.
[0994] A "correction mechanism" is an interface that allows users to review generated text and make corrections as needed.
[0995] "Transmission method" refers to a function that allows the user to make a final review of the revised text and send it to the specified recipient or platform.
[0996] "Publication means" refers to the process and execution environment for publishing the generated text on a designated website.
[0997] To implement this invention, it is necessary to construct a system in which a user, a server, and a generation AI model work in cooperation. This system performs a series of operations, including the user inputting the requirements for the text they want to generate, analyzing those requirements and sending a request to the generation AI, receiving the generated text and displaying it to the user, performing a final check and sending of the revised text, and publishing it on a designated website.
[0998] The server consists of the following main components:
[0999] 1. Input Reception Method: This is an interface for users to input the requirements of the text they want to generate into the system. This refers to the front-end interface of typical smartphone and PC applications.
[1000] 2. Analysis Method: This function analyzes the input requirements and sends a request to the generation AI. It captures the user's requirements as text data, formats it into an appropriate format, and sends it to the generation AI.
[1001] 3. Generative AI and Text Generation Means: The Generative AI is a system that generates text in a specified format based on user prompts. This Generative AI uses advanced text generation models such as OpenAI's GPT-3.
[1002] 4. Response Reception Method: This function receives responses from the generation AI, formats them, and displays them to the user. It receives the generated text, converts it into a readable format, and sends it back to the user.
[1003] 5. Correction Method: This is an interface for the user to review the generated text and make corrections as needed. After the corrections are complete, a final confirmation is performed.
[1004] 6. Sending and Publishing Means: These are functions for sending and publishing the finalized document to a specified recipient and website. This includes email sending and publishing functions via web APIs.
[1005] Hardware and software to be used
[1006] This system often uses the following hardware and software:
[1007] Hardware: User terminals such as smartphones, tablets, and PCs, and servers.
[1008] Software: Generative AI (e.g., OpenAI GPT-3), HTTP libraries for API calls (e.g., Requests for Python), user interface development frameworks (e.g., React, Flutter, etc.)
[1009] Example of a prompt
[1010] If a user wants to generate a description of a new product, they might enter a prompt like the following:
[1011] "Please write a description of the new product."
[1012] Specific example
[1013] For example, if a virtual store operator wants to generate a description for a new wireless earphone product, they would use the system in the following way:
[1014] 1. The user opens the smartphone app and enters "Please create a description for the new product."
[1015] 2. The server receives this requirement and uses an analysis tool to send a request to the generating AI.
[1016] 3. The AI generates a response such as, "The new wireless earphones feature the latest technology and offer clear sound quality and long battery life..."
[1017] 4. The server receives this generated text and displays it to the user.
[1018] 5. The user reviews and corrects the displayed text, then presses the "Confirm" button for final confirmation.
[1019] 6. The server receives the final, revised text and publishes it on the designated website.
[1020] Thus, by using this system, users can quickly generate high-quality documents and send or publish them to the appropriate recipients.
[1021] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1022] Step 1:
[1023] The user inputs the requirements for the text they want to generate through a smartphone app. This input is called a "prompt." For example, the user might input "Please create a description of the new product." The entered prompt is temporarily saved on the user's device.
[1024] Step 2:
[1025] The terminal formats the prompt text entered by the user into a data packet and sends it to the server. This data packet is often in JSON format. Specifically, the input requirements are formed into a JSON object and sent to the server as an HTTP POST request.
[1026] Step 3:
[1027] The server analyzes the data packets received from the terminal. Using the analysis tools, it converts the data from the prompt message into a format that is valid for the generation AI. For example, the user's prompt message, "Please create a description of the new product," is sent directly to the generation AI's API endpoint.
[1028] Step 4:
[1029] The server sends a request to the generative AI. Using the HTTP POST method, it sends a request containing a prompt to the API endpoint of the generative AI model (e.g., OpenAI GPT-3). The input is the prompt, and the output is the generated text.
[1030] Step 5:
[1031] The generative AI model generates text based on the received prompt. In this generation process, the AI model constructs appropriate text based on its learned knowledge. The output of the generative AI model is text in the specified format.
[1032] Step 6:
[1033] The server receives the response from the AI and formats the generated text. The generated text is formatted and converted into a user-friendly format. This formatting process includes grammatical checks and formatting adjustments.
[1034] Step 7:
[1035] The server returns the formatted text to the user's terminal. The formatted text is returned to the user's terminal as an HTTP response. The returned text is displayed on the user's terminal.
[1036] Step 8:
[1037] The user reviews the generated text and makes corrections as needed. If the user finds any errors, they can edit them directly within the app. The corrected text is temporarily saved on the user's device.
[1038] Step 9:
[1039] The user performs a final review and presses the "Confirm" button. Once the Confirm button is pressed, the revised text is determined to be the final text. This final text is then sent back to the server.
[1040] Step 10:
[1041] The server sends and publishes the finalized document to the specified recipient and website. For example, it may send it to a specified email address using the email sending function, or publish it to a website using a web API. At this time, the server selects the appropriate sending or publishing method.
[1042] The above outlines the program processing flow of this system. From user input to the final publication of the document, the process can be automated and efficient.
[1043] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1044] overview
[1045] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[1046] System Configuration
[1047] This system consists of the following main components.
[1048] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[1049] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives and formats the request response and displays it to the user. It also sends the revised final text to the specified destination.
[1050] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[1051] 4. Emotion Engine: This refers to an algorithm and its execution environment that recognizes the emotions of the user during input and adjusts the tone and content of the text based on those emotions.
[1052] Program processing
[1053] The program processing of this system is explained below in natural language.
[1054] 1. The user logs into the terminal and opens a screen to enter the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[1055] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[1056] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1057] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[1058] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends an appropriate API request to the AI.
[1059] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[1060] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[1061] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[1062] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1063] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1064] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1065] Specific example
[1066] Example 1: New product introduction email
[1067] 1. The user enters "Please create a new product introduction email."
[1068] 2. The emotion engine recognizes the user's feelings of joy.
[1069] 3. The AI generates text based on a joyful tone.
[1070] 4. The user reviews and corrects the email content displayed.
[1071] 5. The server sends the corrected email to the customer list.
[1072] Example 2: Creating a blog post
[1073] 1. The user enters "Please create a blog post about winter events."
[1074] 2. The emotion engine recognizes the user's emotion of surprise.
[1075] 3. The AI generates text based on a surprised tone.
[1076] 4. The user reviews and edits the article.
[1077] 5. The server posts the corrected article to a specific URL on the blog.
[1078] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions. The above describes a specific embodiment for carrying out the invention.
[1079] The following describes the processing flow.
[1080] Step 1:
[1081] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[1082] Step 2:
[1083] To recognize the user's emotions during input, the device uses its camera and microphone to collect the user's facial expressions and voice. This data is then sent to the emotion engine.
[1084] Step 3:
[1085] The emotion engine analyzes collected facial expressions, voice tone, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1086] Step 4:
[1087] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. The data packets include the user's requirements (e.g., "I want to create an email introducing a new product") and their emotional state (e.g., "joy").
[1088] Step 5:
[1089] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state and generates appropriate API requests.
[1090] Step 6:
[1091] The server sends a request to the generation AI. This request includes detailed requirements for the text to be generated and its emotional state.
[1092] Step 7:
[1093] The generation AI uses an algorithm to generate text based on requests received from the server. For example, if the user is experiencing the emotion of "joy," it will generate text with a positive tone that aligns with that emotion.
[1094] Step 8:
[1095] The server receives the generated text sent back from the AI. The receiving device formats this text into a format that can be displayed to the user and sends it back to the user.
[1096] Step 9:
[1097] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[1098] Step 10:
[1099] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1100] Step 11:
[1101] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1102] Step 12:
[1103] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1104] The above is a detailed explanation of the processing steps specifically for inventions that combine an emotion engine.
[1105] (Example 2)
[1106] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1107] Conventional text generation systems often fail to accurately reflect the user's emotions and intentions, even when the user inputs the requirements for the text they want to generate. This necessitates numerous revisions, which is extremely time-consuming. Furthermore, because they fail to consider the user's emotions, the generated text tends to have a uniform tone, lacking flexibility to adapt to individual situations. Consequently, there is no guarantee that the generated text will effectively communicate with the recipient.
[1108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and emotional state and sends a request to the generation AI, a response receiving means that receives the generated text and displays it to the user, and a transmission means that the user makes a final confirmation of the revised text and sends it. This makes it possible to provide text that is generated while taking the user's emotions into consideration.
[1109] A "user" refers to a person who logs into the system, enters the requirements for a document, and then reviews and modifies the generated document.
[1110] A "device" refers to a device used by a user to access the system, acquire emotional data, and send it to the server. Specifically, this includes computers, smartphones, tablets, and other similar devices.
[1111] "Means of collecting emotional data" refers to the functions that a device uses to acquire the user's emotional state. Specifically, this includes functions that capture the user's facial expressions and voice using a camera and microphone.
[1112] A "server" refers to a central computing unit that includes an analysis process that analyzes input requirements and emotional states and sends requests to the generating AI.
[1113] "Analysis means" refers to the function in which the server analyzes the user's input requirements and emotional state, and determines what kind of text generation request to send to the generation AI.
[1114] "Generative AI" refers to an algorithm and its execution environment that generates text based on input requirements and analyzed emotional states.
[1115] "Text generation method" refers to a function provided by the generation AI, and the process of generating text based on analyzed data.
[1116] "Response receiving means" refers to the function that allows the server to receive text sent back from the generating AI and format it into a format that can be displayed to the user.
[1117] "Correction tools" refer to interfaces or tools that allow users to review generated text and make corrections as needed. Specifically, this includes text editors and similar tools.
[1118] "Transmission method" refers to the function used when a user sends a document they have finalized. This includes, for example, functions for sending data via email or a Web API.
[1119] System Overview
[1120] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[1121] System Configuration
[1122] This system consists of the following main components:
[1123] 1. User's device:
[1124] It provides an interface for users to access the system and input, review, modify, and finally submit document requirements. The terminals can include standard computers, smartphones, and tablets.
[1125] 2. Server:
[1126] This is the core component that analyzes the input requirements and emotional state and sends a request to the generation AI. The server receives the request response, formats it, and displays it to the user. It also sends the revised final text to the specified destination.
[1127] 3. Generation AI:
[1128] This is an algorithm and execution environment for generating text based on user requirements. Generative AI models such as GPT-3 and BERT can be used.
[1129] 4. Emotional Engine:
[1130] This is an algorithm and its execution environment for recognizing the user's emotions during input and adjusting the tone and content of the text based on those emotions. Libraries such as DeepFace and OpenSmile can be used for emotion analysis.
[1131] Program processing
[1132] The program processing of this system is explained below in natural language.
[1133] 1. The user logs into the terminal and enters the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[1134] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[1135] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1136] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[1137] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends appropriate API requests to the AI.
[1138] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[1139] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[1140] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[1141] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1142] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1143] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1144] Specific example
[1145] Example 1: New product introduction email
[1146] User: "Please create a new product introduction email."
[1147] Emotion Engine: Recognizes the user's emotions of joy.
[1148] Generating AI: Generates text based on a joyful tone.
[1149] User: Review and correct the displayed email content.
[1150] Server: Send the corrected email to the customer list.
[1151] Example 2: Creating a blog post
[1152] User: "Please write a blog post about winter events."
[1153] Emotion Engine: Recognizes the user's emotions of surprise.
[1154] Generating AI: Generates text based on a tone of surprise.
[1155] User: Review and edit the article.
[1156] Server: Posts the corrected article to a specific URL on the blog.
[1157] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions.
[1158] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1159] Understood. Below, I will explain the process in detail, broken down into steps.
[1160] Step 1:
[1161] The user logs into their terminal and enters the requirements for the document they want to generate. Specifically, they access the system's login page using their terminal's browser or application, enter their user ID and password to authenticate, and then enter requirements such as "I want to create a new product introduction email." The data entered is requirement information in text format.
[1162] Step 2:
[1163] The device captures the user's facial expressions and voice using its built-in camera and microphone after user input. Specifically, it acquires data using a webcam or the smartphone's camera and microphone. The acquired data is sent to the emotion engine in real time. The input consists of image data and audio data, and the output is a stream of digital data for analysis.
[1164] Step 3:
[1165] The emotion engine analyzes the user's emotional state based on the received facial expressions and voice tone. Specifically, it uses emotion analysis libraries such as DeepFace and OpenSmile to analyze facial muscle movements and voice pitch to determine emotions such as "joy," "sadness," and "surprise." The input for this step is image data and voice data, and the output is emotional state data as a result of the analysis.
[1166] Step 4:
[1167] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. Specifically, it combines text data and emotional data into a JSON data packet and sends it to the server via an HTTP request. The input for this step is text data and emotional data, and the output is the data packet sent to the server.
[1168] Step 5:
[1169] The server analyzes the received data packets to understand what type of text to generate. Specifically, it uses natural language processing (NLP) algorithms to analyze the user's requirements and sends the appropriate API request to the generation AI. The input to this step is the data packets, and the output is the request data sent to the generation AI.
[1170] Step 6:
[1171] The generative AI generates text based on requirements and emotional states sent from the server. Specifically, it uses generative models such as GPT-3 and BERT to generate text with a positive tone that aligns with the emotional state of "joy." The input for this step is the request data, and the output is the text data of the generated text.
[1172] Step 7:
[1173] The server receives the text returned by the AI generator and formats it into a format that can be displayed to the user. Specifically, it converts it to HTML or JSON format and uses a filtering function to check that the generated text does not contain any inappropriate content. The input for this step is the text data of the text, and the output is the formatted text data.
[1174] Step 8:
[1175] The terminal displays the formatted text returned from the server to the user. Specifically, it uses JavaScript and HTML elements to display the text on a web page. The input for this step is formatted text data, and the output is the text displayed to the user.
[1176] Step 9:
[1177] The user reviews the displayed text and makes corrections as needed. Specifically, they edit the text using the text editor included in the interface. The input for this step is the displayed text data, and the output is the corrected text data.
[1178] Step 10:
[1179] The user reviews the revised text and clicks the "Confirm" button to authorize submission. Specifically, clicking the confirmation button for submission sends the finalized text. The input for this step is the revised text data, and the output is the submission approval operation data.
[1180] Step 11:
[1181] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server. Specifically, it bundles the corrected text into a JSON data packet and sends it to the server via an HTTP request. The input for this step is the corrected text data, and the output is the data packet sent to the server.
[1182] Step 12:
[1183] The server receives the finalized text and sends it to the specified destination. Specifically, it uses an SMTP server for emails and posts to a specific API endpoint for blog posts. The input for this step is the finalized text data, and the output is confirmation data that the transmission was successful.
[1184] The above is a detailed explanation of the program processing of this system.
[1185] (Application Example 2)
[1186] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1187] Conventional text generation systems have the problem of being unable to generate text that reflects the user's emotions, making it difficult to produce text that aligns with the user's intentions and feelings. Furthermore, especially in customer service, it is necessary to respond in a way that is sensitive to the customer's emotions, but current systems are unable to do this adequately.
[1188] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and sends a request to the generation AI; a response receiving means that receives the generated text and displays it to the user; a transmission means that allows the user to make a final confirmation of the revised text and send it; an emotion recognition means that recognizes the user's emotions at the time of input; and an emotion adjustment means that adjusts the text based on the emotions recognized by the emotion recognition means. This enables the generation of text that reflects the user's emotions and appropriate customer service responses that correspond to the customer's emotions.
[1189] An "input acceptance means" is a means for a user to input the requirements for the text they want to generate.
[1190] "Analysis means" refers to a means of analyzing the input requirements and sending the analysis results as a request to the generating AI.
[1191] A "text generation means" is a means for generating text based on an analyzed request.
[1192] A "response receiving means" is a means of receiving text generated by a generation AI and displaying it to the user.
[1193] "Means of correction" refers to means by which the user can review the displayed text and correct it as needed.
[1194] "Means of transmission" refers to the means by which the user makes a final review of the revised text and then sends it.
[1195] "Emotion recognition means" refers to a means of recognizing the user's emotions at the time of input.
[1196] "Emotion adjustment means" are means for adjusting text based on emotions recognized by emotion recognition means.
[1197] System Configuration
[1198] This invention is a system that recognizes a user's emotions and generates and adjusts text based on those emotions. This system consists of the following main components.
[1199] 1. Input acceptance means
[1200] 2. Analysis method
[1201] 3. Sentence generation means
[1202] 4. Means of receiving replies
[1203] 5. Corrective measures
[1204] 6. Transmission method
[1205] 7. Emotion recognition means
[1206] 8. Emotional regulation tools
[1207] Hardware and software to be used
[1208] Hardware: Customer service robot, facial recognition camera, microphone, tablet display
[1209] Software: Emotion recognition engines (e.g., Google Cloud Vision API, Microsoft Azure Face API), generative AI (e.g., OpenAI GPT series), speech recognition systems (e.g., Google Speech-to-Text API)
[1210] Program processing
[1211] 1. The user enters questions or requests to the customer service robot.
[1212] 2. The emotion recognition means analyzes the user's facial expressions and voice tone to determine their emotional state (e.g., joy, sadness, surprise).
[1213] 3. The analysis device formats the user's input and emotional state into data packets and sends them to the server.
[1214] 4. The server receives the data packet and analyzes its contents. It then sends the analysis results as a request to the generating AI.
[1215] 5. The generation AI generates text based on requirements and emotional state. For example, if the user is in a state of "joy," it will generate text with a positive tone that reflects that emotion.
[1216] 6. The response receiving device receives the generated text, formats it into a format that can be displayed to the user, and sends it back to the user.
[1217] 7. The user reviews the displayed text and makes corrections as needed.
[1218] 8. The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient.
[1219] Specific example
[1220] Example 1: Product Information
[1221] 1. The user asks the customer service robot, "Could you show me some of our new dresses?"
[1222] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that the user is excited.
[1223] 3. The AI generates text in a positive tone that matches the user's excitement, such as, "Here is our new dress. It's a very popular design and has been well-received by many customers."
[1224] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[1225] Example 2: Sales information announcement
[1226] 1. The user asks, "Tell me about this weekend's sale."
[1227] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that they are calm.
[1228] 3. The AI generates text in a calm and informative tone that matches the user's emotions, such as, "This weekend, we're having a sale with up to 50% off. Here are some of our recommended items."
[1229] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[1230] Example of a prompt
[1231] "Hello, welcome to the [store name] store. A customer is asking, 'Could you show me the new dresses?' The customer seems excited. Please respond in a positive tone."
[1232] This system will enable in-store customer service to be more emotionally resonant to the user, improving the customer experience.
[1233] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1234] Step 1:
[1235] The user inputs questions or requests to the customer service robot. For example, the user might input, "Could you show me some of your new dresses?" Input is done via voice or touch input on a tablet. Input data (voice data or text data) is obtained.
[1236] Step 2:
[1237] The emotion recognition system analyzes the user's facial expressions and voice tone to determine their emotional state. Image data of the face obtained from the facial recognition camera and voice data acquired from the microphone are input. Based on this input data, the emotion recognition engine (e.g., Google Cloud Vision API or Microsoft Azure Face API) performs data analysis and outputs the user's emotional state (e.g., joy, sadness, surprise).
[1238] Step 3:
[1239] The analysis device formats the user's input and emotional state into data packets and sends them to the server. The user's entered questions and request data, along with the emotional state data obtained from the emotion recognition device, are packaged to generate data packets for transmission to the server. These data packets are then sent to the server.
[1240] Step 4:
[1241] The server receives the data packet and analyzes its contents. The server analyzes the received data packet to extract the user's question and emotional state. The analysis reveals that the user requested to be shown new dresses and that the user's emotional state is "excited."
[1242] Step 5:
[1243] The server sends a request to the generating AI. Based on the extracted user question and emotional state, it generates a prompt. For example, it sends a request to the generating AI (e.g., OpenAI GPT series) with the prompt "The customer is asking 'Can you show me the new dresses?' The customer seems excited. Please respond in a positive tone." This prompt is input, and the generating AI generates and outputs an appropriate response.
[1244] Step 6:
[1245] The response receiving mechanism receives the generated text and formats it into a format that can be displayed to the user. It receives the response text generated by the generation AI and formats it into a format for display to the user. For example, it might receive the response text, "Here are our new dresses. They are a very popular design and have been well-received by many customers."
[1246] Step 7:
[1247] The user reviews the displayed text and makes corrections as needed. The formatted text is then displayed on the tablet screen. The user reviews the displayed text and makes corrections as needed using touch input or other methods. The corrected text data is then obtained.
[1248] Step 8:
[1249] The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient. The finalized text data is received and sent, for example, to the store's backend system or a specific email address. The submission is complete.
[1250] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1251] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1252] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1253] [Fourth Embodiment]
[1254] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1255] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1256] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1257] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1258] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1259] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1260] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1261] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1262] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1263] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1264] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1265] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1266] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1267] This invention is a system for avoiding monotony in writing during routine tasks, and provides a method for efficiently generating creative text using generation AI. Specific embodiments of this system will be described below.
[1268] overview
[1269] This system consists of a series of steps: the user inputs the requirements for the text they want to generate, the system analyzes this input and sends a request to the generation AI, and the AI returns the response to the user. The user then modifies the provided text and submits / publishes the final text.
[1270] System Configuration
[1271] This system consists of the following main components:
[1272] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[1273] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives the request response, formats it, and displays it to the user. The server also sends the revised final text to the specified destination.
[1274] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[1275] Program processing
[1276] The program processing of this system is explained below in natural language.
[1277] 1. The user logs into the system via their terminal and opens a screen to enter the requirements for the document they want to generate. For example, the user might enter "I want to create a new product introduction email."
[1278] 2. The terminal formats the user's input into data packets and sends them to the server.
[1279] 3. The server receives the data packet and analyzes its contents. Based on the analysis, it generates an appropriate API request and sends it to the generation AI. For example, it sends a "POST request" to the generation AI's endpoint.
[1280] 4. The generation AI generates text based on the user's requirements and sends the results back to the server. For example, it generates specific text based on a template for a "new product introduction email."
[1281] 5. The server receives the response from the generating AI, formats its content, and sends it back to the user. The formatted response includes readability and formatting.
[1282] 6. The terminal displays the formatted text to the user. The user reviews the displayed content and makes corrections as needed.
[1283] 7. After making revisions, the user performs a final review and approves submission. At this time, the user clicks the "Confirm" button.
[1284] 8. The terminal sends the final, revised text to the server.
[1285] 9. The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1286] Specific example
[1287] Example 1: New product introduction email
[1288] 1. The user enters "Please create a new product introduction email."
[1289] 2. The server sends a request to the generation AI and receives the generated email content.
[1290] 3. The user reviews and corrects the displayed email content.
[1291] 4. The server sends the corrected email to the customer list.
[1292] Example 2: Creating a blog post
[1293] 1. The user enters "Please create a blog post about winter events."
[1294] 2. The server sends a request to the generation AI and receives the generated article content.
[1295] 3. The user reviews and edits the article.
[1296] 4. The server posts the corrected article to a specific URL on the blog.
[1297] These embodiments enable the system to easily generate creative and effective text and to quickly send and publish it. The above describes specific forms for carrying out the invention.
[1298] The following describes the processing flow.
[1299] Step 1:
[1300] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[1301] Step 2:
[1302] The terminal formats the user's input into data packets and sends them to the server. These data packets contain the user's requirements.
[1303] Step 3:
[1304] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and generates appropriate API requests.
[1305] Step 4:
[1306] The server sends a request to the generation AI. This request contains detailed requirements for the text to be generated.
[1307] Step 5:
[1308] The generation AI generates text using a predetermined algorithm based on a request received from the server.
[1309] Step 6:
[1310] The server receives the generated text sent back from the AI. The receiving method then formats this text into a format that can be displayed to the user.
[1311] Step 7:
[1312] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[1313] Step 8:
[1314] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1315] Step 9:
[1316] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1317] Step 10:
[1318] The server receives the final-approved text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint. This process allows users to quickly and efficiently generate and publish their desired text.
[1319] The above is a detailed explanation of the program's processing flow.
[1320] (Example 1)
[1321] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1322] Traditional document creation processes for routine tasks were prone to becoming monotonous and time-consuming. Furthermore, there was no system in place to efficiently generate creative documents based on specific user requirements. Therefore, there is a need for improved document quality and increased work efficiency.
[1323] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1324] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate; a packet generation means for the terminal to format the input requirements as data packets and transmit them; an analysis means for the server to analyze the received requirements and send a request to a generation AI; a text generation means for the generation AI to generate text based on the analyzed request; a response receiving means for the server to format the generated text and display it to the user; a correction means for the user to review and correct the displayed text; and a transmission means for the user to make a final confirmation of the corrected text and transmit it. This makes it possible to efficiently generate creative and high-quality text based on the requirements entered by the user and to quickly transmit and publish it.
[1325] A "user" is someone who accesses the system, inputs the requirements for text generation, and reviews and modifies the generated text.
[1326] A "terminal" is a device that provides an interface consisting of hardware and software for users to access a system and input requirements, verify and modify generated results.
[1327] A "server" is a core computer system that analyzes user requirements, sends requests to the generation AI, receives and formats the generated text, and sends it back to the user.
[1328] "Input acceptance means" refers to the interface and software functions used by the user to input the requirements for the text they wish to generate.
[1329] "Packet generation means" refers to the function by which a terminal formats user input into data packets and sends them to a server.
[1330] "Analysis means" refers to a function in which the server analyzes data packets received from the terminal and sends appropriate requests to the generating AI.
[1331] "Generative AI" refers to an algorithm and its execution environment for generating text based on user requirements.
[1332] A "request" is an API request sent from the server to the AI, which instructs the AI to generate text based on the user's requirements.
[1333] "Text generation means" refers to the function in which the generation AI generates text based on requests received from the server.
[1334] The "response receiving means" is a function in which the server formats the generated results received from the generating AI and sends them back to the user.
[1335] "Correction means" refers to the interface and software functions that allow the user to review the displayed generated results and make corrections as needed.
[1336] "Transmission method" refers to a function that allows the user to make a final check of the revised document and send it to the specified recipient.
[1337] This invention is a system for avoiding the monotony of repetitive writing in routine tasks and for efficiently generating creative text. The system consists of the following main components: a user terminal, a server, and a generation AI model. By combining these, it enables the generation of high-quality text based on user input requirements, and its rapid transmission and publication.
[1338] System Configuration
[1339] User's terminal
[1340] The user's device provides an interface for the user to access the system. This includes common computers, smartphones, tablets, etc. The device formats the user's input and sends it to the server as data packets.
[1341] server
[1342] The server has the core function of receiving requirements from the user, analyzing them, and sending requests to the generative AI model. The server also receives the generated text, formats it, and sends it back to the user. It also plays a role in sending the revised final text to the specified destination.
[1343] Generative AI Models
[1344] A generative AI model is an algorithm and its execution environment for generating text based on user requirements. The generative AI model generates appropriate text based on input prompts, referencing existing documents and knowledge bases.
[1345] Program processing
[1346] This system's program clearly defines the roles of the user, terminal, and server, enabling an efficient and creative text generation process. The specific processing steps are as follows, but the overall flow is shown here.
[1347] User actions
[1348] The user logs into the system using their own device. After logging in, an interface is displayed where the user can enter the requirements for the document they want to generate. For example, the user might enter a specific prompt such as "I want to create a new product introduction email."
[1349] Terminal operation
[1350] The terminal formats the user's input into a JSON data packet and sends it to the server using an HTTP request.
[1351] Server Operations
[1352] The server analyzes the received data packets and extracts the user's requirements. Based on the analysis results, it generates an appropriate API request for the generated AI model and sends the POST request.
[1353] Manipulation of Generative AI Models
[1354] The generation AI model generates text based on prompt messages received from the server. For example, in the case of a "new product introduction email," it generates an email body that includes appropriate product descriptions and promotional messages. The generated results are sent back to the server.
[1355] Server response
[1356] The server formats the text received from the AI generation model and sends it back to the user. The formatted text is then displayed on the user's device.
[1357] User modifications and final confirmation
[1358] The user reviews the displayed text and makes corrections as needed. Once corrections are complete, they click the "Confirm" button, and the device sends the final text to the server.
[1359] Server-based transmission and publication
[1360] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends it via an SMTP server, and in the case of a blog post, it sends a POST request to the specified API endpoint.
[1361] Specific example
[1362] Example 1: New product introduction email
[1363] The user enters "Please create a new product introduction email." The server sends a request to the AI generation model and receives the generated email content. The user reviews the displayed email content and makes corrections as needed. The server sends the corrected email to the customer list.
[1364] Example 2: Creating a blog post
[1365] The user enters "Please create a blog post about winter events." The server sends a request to the AI generation model and receives the generated article content. The user reviews the article and makes revisions as needed. The server posts the revised article to a specific URL on the blog.
[1366] In this way, this system enables users to easily generate creative and effective text and quickly send and publish it.
[1367] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1368] Step 1: User login and entry of requirements
[1369] The user logs into the system using their own device. The input is the user ID and password, and the output is the login success or failure status.
[1370] After logging in, the user accesses an interface to enter the requirements for the document they want to generate. The input consists of specific prompts, such as a text-based instruction like "I want to create a new product introduction email."
[1371] Step 2: Submitting Requirements
[1372] The terminal converts the prompt text entered by the user into a JSON data packet. The input is a text-based prompt text, and the output is a JSON data packet.
[1373] The terminal sends the formatted data packet to the server as an HTTP request. For example, it might be sent as a POST request.
[1374] Step 3: Requirements analysis and AI request generation
[1375] The server analyzes data packets received from the terminal. The input is data packets in JSON format, and the output is the analyzed requirements information.
[1376] The server generates an API request to the generated AI model based on the content of the received data packets. The input is the parsed requirements information, and the output is the API request. Specifically, it generates a POST request to the endpoint of the generated AI.
[1377] Step 4: Text Generation
[1378] The generative AI model generates text based on prompt messages received from the server. The input is an API request, and the output is the generated text.
[1379] For example, based on a prompt requesting a "new product introduction email," the system generates an email body containing appropriate product descriptions and promotional messages.
[1380] Step 5: Formatting and displaying the generated results
[1381] The server formats the text received from the generative AI model. The input is the generated text, and the output is the formatted text. Formatting includes paragraph breaks and formatting adjustments.
[1382] The server converts the formatted text into JSON format and sends it back to the terminal.
[1383] Step 6: User review and correction
[1384] The terminal displays generated text received from the server to the user. The input is data in JSON format, and the output is a display on the user interface.
[1385] The user reviews the displayed text and makes corrections as needed. The input is the generated text, and the output is the corrected text. The user enters the specific corrections in a text editor.
[1386] Step 7: Final confirmation and submission
[1387] The user reviews the revisions and clicks the "Confirm" button to approve the final submission. The input is the revised text, and the output is the final confirmation status.
[1388] The terminal sends the final text confirmed by the user to the server as a JSON data packet. The input is the revised text, and the output is the data packet.
[1389] Step 8: Send and publish
[1390] The server receives the final, confirmed message and sends it to the specified destination. The input is the final message, and the output is the transmission confirmation status.
[1391] For example, in the case of email, it is sent via an SMTP server, and in the case of a blog post, it is published as a POST request to the specified API endpoint.
[1392] (Application Example 1)
[1393] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1394] Traditional methods have limitations in terms of text generation speed and creativity, making it difficult to continuously generate large volumes of high-quality text. Furthermore, the time and effort required to properly edit and publish the generated text was also a problem. This issue is particularly serious for virtual stores, where rapid updates of new product descriptions and campaign information are essential.
[1395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1396] In this invention, the server includes an input receiving means for the user to input the requirements for the text they wish to generate, an analysis means for the server to analyze the input requirements and send a request to a generation AI, and a text generation means for the generation AI to generate text based on the analyzed request. This makes it possible to quickly generate creative text that meets the user's requirements and to publish the generated text on a designated website.
[1397] A "user" is an individual or group that uses a system to generate or edit text.
[1398] An "input acceptance mechanism" is an interface for users to input the requirements of the text they want to generate into the system.
[1399] A "server" is a central processing unit that receives and analyzes input data from users and sends requests to the generating AI.
[1400] "Analysis means" refers to a device or software that has the function of understanding the requirements entered by the user and converting them into an appropriate request.
[1401] "Generative AI" is an artificial intelligence system that generates text in a specified format based on user requirements.
[1402] "Text generation means" refers to the functions and execution environment for generating text using generation AI.
[1403] A "response receiving means" is a device or software that has the function of receiving a response from a generating AI, formatting it, and displaying it to the user.
[1404] A "correction mechanism" is an interface that allows users to review generated text and make corrections as needed.
[1405] "Transmission method" refers to a function that allows the user to make a final review of the revised text and send it to the specified recipient or platform.
[1406] "Publication means" refers to the process and execution environment for publishing the generated text on a designated website.
[1407] To implement this invention, it is necessary to construct a system in which a user, a server, and a generation AI model work in cooperation. This system performs a series of operations, including the user inputting the requirements for the text they want to generate, analyzing those requirements and sending a request to the generation AI, receiving the generated text and displaying it to the user, performing a final check and sending of the revised text, and publishing it on a designated website.
[1408] The server consists of the following main components:
[1409] 1. Input Reception Method: This is an interface for users to input the requirements of the text they want to generate into the system. This refers to the front-end interface of typical smartphone and PC applications.
[1410] 2. Analysis Method: This function analyzes the input requirements and sends a request to the generation AI. It captures the user's requirements as text data, formats it into an appropriate format, and sends it to the generation AI.
[1411] 3. Generative AI and Text Generation Means: The Generative AI is a system that generates text in a specified format based on user prompts. This Generative AI uses advanced text generation models such as OpenAI's GPT-3.
[1412] 4. Response Reception Method: This function receives responses from the generation AI, formats them, and displays them to the user. It receives the generated text, converts it into a readable format, and sends it back to the user.
[1413] 5. Correction Method: This is an interface for the user to review the generated text and make corrections as needed. After the corrections are complete, a final confirmation is performed.
[1414] 6. Sending and Publishing Means: These are functions for sending and publishing the finalized document to a specified recipient and website. This includes email sending and publishing functions via web APIs.
[1415] Hardware and software to be used
[1416] This system often uses the following hardware and software:
[1417] Hardware: User terminals such as smartphones, tablets, and PCs, and servers.
[1418] Software: Generative AI (e.g., OpenAI GPT-3), HTTP libraries for API calls (e.g., Requests for Python), user interface development frameworks (e.g., React, Flutter, etc.)
[1419] Example of a prompt
[1420] If a user wants to generate a description of a new product, they might enter a prompt like the following:
[1421] "Please write a description of the new product."
[1422] Specific example
[1423] For example, if a virtual store operator wants to generate a description for a new wireless earphone product, they would use the system in the following way:
[1424] 1. The user opens the smartphone app and enters "Please create a description for the new product."
[1425] 2. The server receives this requirement and uses an analysis tool to send a request to the generating AI.
[1426] 3. The AI generates a response such as, "The new wireless earphones feature the latest technology and offer clear sound quality and long battery life..."
[1427] 4. The server receives this generated text and displays it to the user.
[1428] 5. The user reviews and corrects the displayed text, then presses the "Confirm" button for final confirmation.
[1429] 6. The server receives the final, revised text and publishes it on the designated website.
[1430] Thus, by using this system, users can quickly generate high-quality documents and send or publish them to the appropriate recipients.
[1431] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1432] Step 1:
[1433] The user inputs the requirements for the text they want to generate through a smartphone app. This input is called a "prompt." For example, the user might input "Please create a description of the new product." The entered prompt is temporarily saved on the user's device.
[1434] Step 2:
[1435] The terminal formats the prompt text entered by the user into a data packet and sends it to the server. This data packet is often in JSON format. Specifically, the input requirements are formed into a JSON object and sent to the server as an HTTP POST request.
[1436] Step 3:
[1437] The server analyzes the data packets received from the terminal. Using the analysis tools, it converts the data from the prompt message into a format that is valid for the generation AI. For example, the user's prompt message, "Please create a description of the new product," is sent directly to the generation AI's API endpoint.
[1438] Step 4:
[1439] The server sends a request to the generative AI. Using the HTTP POST method, it sends a request containing a prompt to the API endpoint of the generative AI model (e.g., OpenAI GPT-3). The input is the prompt, and the output is the generated text.
[1440] Step 5:
[1441] The generative AI model generates text based on the received prompt. In this generation process, the AI model constructs appropriate text based on its learned knowledge. The output of the generative AI model is text in the specified format.
[1442] Step 6:
[1443] The server receives the response from the AI and formats the generated text. The generated text is formatted and converted into a user-friendly format. This formatting process includes grammatical checks and formatting adjustments.
[1444] Step 7:
[1445] The server returns the formatted text to the user's terminal. The formatted text is returned to the user's terminal as an HTTP response. The returned text is displayed on the user's terminal.
[1446] Step 8:
[1447] The user reviews the generated text and makes corrections as needed. If the user finds any errors, they can edit them directly within the app. The corrected text is temporarily saved on the user's device.
[1448] Step 9:
[1449] The user performs a final review and presses the "Confirm" button. Once the Confirm button is pressed, the revised text is determined to be the final text. This final text is then sent back to the server.
[1450] Step 10:
[1451] The server sends and publishes the finalized document to the specified recipient and website. For example, it may send it to a specified email address using the email sending function, or publish it to a website using a web API. At this time, the server selects the appropriate sending or publishing method.
[1452] The above outlines the program processing flow of this system. From user input to the final publication of the document, the process can be automated and efficient.
[1453] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1454] overview
[1455] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[1456] System Configuration
[1457] This system consists of the following main components.
[1458] 1. User's terminal: Provides an interface for the user to access the system, input, review, modify, and finally submit document requirements. The terminal can be a standard computer, smartphone, tablet, etc.
[1459] 2. Server: This is the core component that analyzes the input requirements and sends requests to the generation AI. The server receives and formats the request response and displays it to the user. It also sends the revised final text to the specified destination.
[1460] 3. Generative AI: Refers to an algorithm that generates text based on user requirements, and its execution environment.
[1461] 4. Emotion Engine: This refers to an algorithm and its execution environment that recognizes the emotions of the user during input and adjusts the tone and content of the text based on those emotions.
[1462] Program processing
[1463] The program processing of this system is explained below in natural language.
[1464] 1. The user logs into the terminal and opens a screen to enter the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[1465] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[1466] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1467] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[1468] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends an appropriate API request to the AI.
[1469] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[1470] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[1471] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[1472] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1473] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1474] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1475] Specific example
[1476] Example 1: New product introduction email
[1477] 1. The user enters "Please create a new product introduction email."
[1478] 2. The emotion engine recognizes the user's feelings of joy.
[1479] 3. The AI generates text based on a joyful tone.
[1480] 4. The user reviews and corrects the email content displayed.
[1481] 5. The server sends the corrected email to the customer list.
[1482] Example 2: Creating a blog post
[1483] 1. The user enters "Please create a blog post about winter events."
[1484] 2. The emotion engine recognizes the user's emotion of surprise.
[1485] 3. The AI generates text based on a surprised tone.
[1486] 4. The user reviews and edits the article.
[1487] 5. The server posts the corrected article to a specific URL on the blog.
[1488] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions. The above describes a specific embodiment for carrying out the invention.
[1489] The following describes the processing flow.
[1490] Step 1:
[1491] The user logs into their device and opens a screen where they can enter the requirements for the document they want to generate. On this screen, the user enters specific requirements, such as "I want to create an email introducing a new product."
[1492] Step 2:
[1493] To recognize the user's emotions during input, the device uses its camera and microphone to collect the user's facial expressions and voice. This data is then sent to the emotion engine.
[1494] Step 3:
[1495] The emotion engine analyzes collected facial expressions, voice tone, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1496] Step 4:
[1497] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. The data packets include the user's requirements (e.g., "I want to create an email introducing a new product") and their emotional state (e.g., "joy").
[1498] Step 5:
[1499] The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state and generates appropriate API requests.
[1500] Step 6:
[1501] The server sends a request to the generation AI. This request includes detailed requirements for the text to be generated and its emotional state.
[1502] Step 7:
[1503] The generation AI uses an algorithm to generate text based on requests received from the server. For example, if the user is experiencing the emotion of "joy," it will generate text with a positive tone that aligns with that emotion.
[1504] Step 8:
[1505] The server receives the generated text sent back from the AI. The receiving device formats this text into a format that can be displayed to the user and sends it back to the user.
[1506] Step 9:
[1507] The device displays formatted text to the user. The user reviews the text displayed on the device and makes corrections as needed.
[1508] Step 10:
[1509] The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1510] Step 11:
[1511] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1512] Step 12:
[1513] The server receives the finalized text and sends it to the specified recipient. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1514] The above is a detailed explanation of the processing steps specifically for inventions that combine an emotion engine.
[1515] (Example 2)
[1516] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1517] Conventional text generation systems often fail to accurately reflect the user's emotions and intentions, even when the user inputs the requirements for the text they want to generate. This necessitates numerous revisions, which is extremely time-consuming. Furthermore, because they fail to consider the user's emotions, the generated text tends to have a uniform tone, lacking flexibility to adapt to individual situations. Consequently, there is no guarantee that the generated text will effectively communicate with the recipient.
[1518] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and emotional state and sends a request to the generation AI, a response receiving means that receives the generated text and displays it to the user, and a transmission means that the user makes a final confirmation of the revised text and sends it. This makes it possible to provide text that is generated while taking the user's emotions into consideration.
[1519] A "user" refers to a person who logs into the system, enters the requirements for a document, and then reviews and modifies the generated document.
[1520] A "device" refers to a device used by a user to access the system, acquire emotional data, and send it to the server. Specifically, this includes computers, smartphones, tablets, and other similar devices.
[1521] "Means of collecting emotional data" refers to the functions that a device uses to acquire the user's emotional state. Specifically, this includes functions that capture the user's facial expressions and voice using a camera and microphone.
[1522] A "server" refers to a central computing unit that includes an analysis process that analyzes input requirements and emotional states and sends requests to the generating AI.
[1523] "Analysis means" refers to the function in which the server analyzes the user's input requirements and emotional state, and determines what kind of text generation request to send to the generation AI.
[1524] "Generative AI" refers to an algorithm and its execution environment that generates text based on input requirements and analyzed emotional states.
[1525] "Text generation method" refers to a function provided by the generation AI, and the process of generating text based on analyzed data.
[1526] "Response receiving means" refers to the function that allows the server to receive text sent back from the generating AI and format it into a format that can be displayed to the user.
[1527] "Correction tools" refer to interfaces or tools that allow users to review generated text and make corrections as needed. Specifically, this includes text editors and similar tools.
[1528] "Transmission method" refers to the function used when a user sends a document they have finalized. This includes, for example, functions for sending data via email or a Web API.
[1529] System Overview
[1530] This invention is a system in which a user inputs the requirements for the text they wish to generate, and a generation AI creates the text based on those requirements. Furthermore, it incorporates an emotion engine that recognizes the user's emotions and adjusts the text based on those emotions. As a result, the tone and content of the text are adjusted to better match the user's intentions.
[1531] System Configuration
[1532] This system consists of the following main components:
[1533] 1. User's device:
[1534] It provides an interface for users to access the system and input, review, modify, and finally submit document requirements. The terminals can include standard computers, smartphones, and tablets.
[1535] 2. Server:
[1536] This is the core component that analyzes the input requirements and emotional state and sends a request to the generation AI. The server receives the request response, formats it, and displays it to the user. It also sends the revised final text to the specified destination.
[1537] 3. Generation AI:
[1538] This is an algorithm and execution environment for generating text based on user requirements. Generative AI models such as GPT-3 and BERT can be used.
[1539] 4. Emotional Engine:
[1540] This is an algorithm and its execution environment for recognizing the user's emotions during input and adjusting the tone and content of the text based on those emotions. Libraries such as DeepFace and OpenSmile can be used for emotion analysis.
[1541] Program processing
[1542] The program processing of this system is explained below in natural language.
[1543] 1. The user logs into the terminal and enters the requirements for the document they want to generate. For example, they might enter, "I want to create a new product introduction email."
[1544] 2. The device uses sensors and microphones to analyze the user's facial expressions and voice during input, and transmits the user's emotional state to the emotion engine.
[1545] 3. The emotion engine analyzes the user's facial expressions, voice tone, input content, etc., to determine the user's emotional state (e.g., joy, sadness, surprise, etc.).
[1546] 4. The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server.
[1547] 5. The server receives data packets and analyzes their contents. Using the analysis tools, it understands the user's requirements and emotional state, and sends appropriate API requests to the AI.
[1548] 6. The generating AI uses an algorithm to generate text based on the requirements and emotional state sent from the server. For example, if the user is in an emotional state of "joy," it will generate text with a positive tone that aligns with that emotion.
[1549] 7. The server receives the generated text returned from the generation AI. The receiving means formats this text into a format that can be displayed to the user and sends it back to the user.
[1550] 8. The terminal displays the formatted text to the user. The user reviews the text displayed on the terminal and makes corrections as needed.
[1551] 9. The user reviews the revised text and approves submission. The user clicks the "Confirm" button.
[1552] 10. The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server.
[1553] 11. The server receives the finalized text and sends it to the specified destination. For example, in the case of email, it sends the email via the SMTP server, and in the case of a blog post, it posts it to the specified API endpoint.
[1554] Specific example
[1555] Example 1: New product introduction email
[1556] User: "Please create a new product introduction email."
[1557] Emotion Engine: Recognizes the user's emotions of joy.
[1558] Generating AI: Generates text based on a joyful tone.
[1559] User: Review and correct the displayed email content.
[1560] Server: Send the corrected email to the customer list.
[1561] Example 2: Creating a blog post
[1562] User: "Please write a blog post about winter events."
[1563] Emotion Engine: Recognizes the user's emotions of surprise.
[1564] Generating AI: Generates text based on a tone of surprise.
[1565] User: Review and edit the article.
[1566] Server: Posts the corrected article to a specific URL on the blog.
[1567] In this way, the system can recognize the user's emotions and adjust the generated text based on those emotions, making it possible to generate and send more effective text that better reflects the user's intentions.
[1568] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1569] Understood. Below, I will explain the process in detail, broken down into steps.
[1570] Step 1:
[1571] The user logs into their terminal and enters the requirements for the document they want to generate. Specifically, they access the system's login page using their terminal's browser or application, enter their user ID and password to authenticate, and then enter requirements such as "I want to create a new product introduction email." The data entered is requirement information in text format.
[1572] Step 2:
[1573] The device captures the user's facial expressions and voice using its built-in camera and microphone after user input. Specifically, it acquires data using a webcam or the smartphone's camera and microphone. The acquired data is sent to the emotion engine in real time. The input consists of image data and audio data, and the output is a stream of digital data for analysis.
[1574] Step 3:
[1575] The emotion engine analyzes the user's emotional state based on the received facial expressions and voice tone. Specifically, it uses emotion analysis libraries such as DeepFace and OpenSmile to analyze facial muscle movements and voice pitch to determine emotions such as "joy," "sadness," and "surprise." The input for this step is image data and voice data, and the output is emotional state data as a result of the analysis.
[1576] Step 4:
[1577] The terminal formats the user's input and the emotional state recognized by the emotion engine into data packets and sends them to the server. Specifically, it combines text data and emotional data into a JSON data packet and sends it to the server via an HTTP request. The input for this step is text data and emotional data, and the output is the data packet sent to the server.
[1578] Step 5:
[1579] The server analyzes the received data packets to understand what type of text to generate. Specifically, it uses natural language processing (NLP) algorithms to analyze the user's requirements and sends the appropriate API request to the generation AI. The input to this step is the data packets, and the output is the request data sent to the generation AI.
[1580] Step 6:
[1581] The generative AI generates text based on requirements and emotional states sent from the server. Specifically, it uses generative models such as GPT-3 and BERT to generate text with a positive tone that aligns with the emotional state of "joy." The input for this step is the request data, and the output is the text data of the generated text.
[1582] Step 7:
[1583] The server receives the text returned by the AI generator and formats it into a format that can be displayed to the user. Specifically, it converts it to HTML or JSON format and uses a filtering function to check that the generated text does not contain any inappropriate content. The input for this step is the text data of the text, and the output is the formatted text data.
[1584] Step 8:
[1585] The terminal displays the formatted text returned from the server to the user. Specifically, it uses JavaScript and HTML elements to display the text on a web page. The input for this step is formatted text data, and the output is the text displayed to the user.
[1586] Step 9:
[1587] The user reviews the displayed text and makes corrections as needed. Specifically, they edit the text using the text editor included in the interface. The input for this step is the displayed text data, and the output is the corrected text data.
[1588] Step 10:
[1589] The user reviews the revised text and clicks the "Confirm" button to authorize submission. Specifically, clicking the confirmation button for submission sends the finalized text. The input for this step is the revised text data, and the output is the submission approval operation data.
[1590] Step 11:
[1591] The terminal reformats the text after the user's final confirmation into a data packet and sends it to the server. Specifically, it bundles the corrected text into a JSON data packet and sends it to the server via an HTTP request. The input for this step is the corrected text data, and the output is the data packet sent to the server.
[1592] Step 12:
[1593] The server receives the finalized text and sends it to the specified destination. Specifically, it uses an SMTP server for emails and posts to a specific API endpoint for blog posts. The input for this step is the finalized text data, and the output is confirmation data that the transmission was successful.
[1594] The above is a detailed explanation of the program processing of this system.
[1595] (Application Example 2)
[1596] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1597] Conventional text generation systems have the problem of being unable to generate text that reflects the user's emotions, making it difficult to produce text that aligns with the user's intentions and feelings. Furthermore, especially in customer service, it is necessary to respond in a way that is sensitive to the customer's emotions, but current systems are unable to do this adequately.
[1598] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an analysis means that analyzes the input requirements and sends a request to the generation AI; a response receiving means that receives the generated text and displays it to the user; a transmission means that allows the user to make a final confirmation of the revised text and send it; an emotion recognition means that recognizes the user's emotions at the time of input; and an emotion adjustment means that adjusts the text based on the emotions recognized by the emotion recognition means. This enables the generation of text that reflects the user's emotions and appropriate customer service responses that correspond to the customer's emotions.
[1599] An "input acceptance means" is a means for a user to input the requirements for the text they want to generate.
[1600] "Analysis means" refers to a means of analyzing the input requirements and sending the analysis results as a request to the generating AI.
[1601] A "text generation means" is a means for generating text based on an analyzed request.
[1602] A "response receiving means" is a means of receiving text generated by a generation AI and displaying it to the user.
[1603] "Means of correction" refers to means by which the user can review the displayed text and correct it as needed.
[1604] "Means of transmission" refers to the means by which the user makes a final review of the revised text and then sends it.
[1605] "Emotion recognition means" refers to a means of recognizing the user's emotions at the time of input.
[1606] "Emotion adjustment means" are means for adjusting text based on emotions recognized by emotion recognition means.
[1607] System Configuration
[1608] This invention is a system that recognizes a user's emotions and generates and adjusts text based on those emotions. This system consists of the following main components.
[1609] 1. Input acceptance means
[1610] 2. Analysis method
[1611] 3. Sentence generation means
[1612] 4. Means of receiving replies
[1613] 5. Corrective measures
[1614] 6. Transmission method
[1615] 7. Emotion recognition means
[1616] 8. Emotional regulation tools
[1617] Hardware and software to be used
[1618] Hardware: Customer service robot, facial recognition camera, microphone, tablet display
[1619] Software: Emotion recognition engines (e.g., Google Cloud Vision API, Microsoft Azure Face API), generative AI (e.g., OpenAI GPT series), speech recognition systems (e.g., Google Speech-to-Text API)
[1620] Program processing
[1621] 1. The user enters questions or requests to the customer service robot.
[1622] 2. The emotion recognition means analyzes the user's facial expressions and voice tone to determine their emotional state (e.g., joy, sadness, surprise).
[1623] 3. The analysis device formats the user's input and emotional state into data packets and sends them to the server.
[1624] 4. The server receives the data packet and analyzes its contents. It then sends the analysis results as a request to the generating AI.
[1625] 5. The generation AI generates text based on requirements and emotional state. For example, if the user is in a state of "joy," it will generate text with a positive tone that reflects that emotion.
[1626] 6. The response receiving device receives the generated text, formats it into a format that can be displayed to the user, and sends it back to the user.
[1627] 7. The user reviews the displayed text and makes corrections as needed.
[1628] 8. The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient.
[1629] Specific example
[1630] Example 1: Product Information
[1631] 1. The user asks the customer service robot, "Could you show me some of our new dresses?"
[1632] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that the user is excited.
[1633] 3. The AI generates text in a positive tone that matches the user's excitement, such as, "Here is our new dress. It's a very popular design and has been well-received by many customers."
[1634] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[1635] Example 2: Sales information announcement
[1636] 1. The user asks, "Tell me about this weekend's sale."
[1637] 2. The emotion recognition system analyzes the user's facial expressions and voice tone and determines that they are calm.
[1638] 3. The AI generates text in a calm and informative tone that matches the user's emotions, such as, "This weekend, we're having a sale with up to 50% off. Here are some of our recommended items."
[1639] 4. The robot displays the generated text on a tablet screen and reads it aloud.
[1640] Example of a prompt
[1641] "Hello, welcome to the [store name] store. A customer is asking, 'Could you show me the new dresses?' The customer seems excited. Please respond in a positive tone."
[1642] This system will enable in-store customer service to be more emotionally resonant to the user, improving the customer experience.
[1643] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1644] Step 1:
[1645] The user inputs questions or requests to the customer service robot. For example, the user might input, "Could you show me some of your new dresses?" Input is done via voice or touch input on a tablet. Input data (voice data or text data) is obtained.
[1646] Step 2:
[1647] The emotion recognition system analyzes the user's facial expressions and voice tone to determine their emotional state. Image data of the face obtained from the facial recognition camera and voice data acquired from the microphone are input. Based on this input data, the emotion recognition engine (e.g., Google Cloud Vision API or Microsoft Azure Face API) performs data analysis and outputs the user's emotional state (e.g., joy, sadness, surprise).
[1648] Step 3:
[1649] The analysis device formats the user's input and emotional state into data packets and sends them to the server. The user's entered questions and request data, along with the emotional state data obtained from the emotion recognition device, are packaged to generate data packets for transmission to the server. These data packets are then sent to the server.
[1650] Step 4:
[1651] The server receives the data packet and analyzes its contents. The server analyzes the received data packet to extract the user's question and emotional state. The analysis reveals that the user requested to be shown new dresses and that the user's emotional state is "excited."
[1652] Step 5:
[1653] The server sends a request to the generating AI. Based on the extracted user question and emotional state, it generates a prompt. For example, it sends a request to the generating AI (e.g., OpenAI GPT series) with the prompt "The customer is asking 'Can you show me the new dresses?' The customer seems excited. Please respond in a positive tone." This prompt is input, and the generating AI generates and outputs an appropriate response.
[1654] Step 6:
[1655] The response receiving mechanism receives the generated text and formats it into a format that can be displayed to the user. It receives the response text generated by the generation AI and formats it into a format for display to the user. For example, it might receive the response text, "Here are our new dresses. They are a very popular design and have been well-received by many customers."
[1656] Step 7:
[1657] The user reviews the displayed text and makes corrections as needed. The formatted text is then displayed on the tablet screen. The user reviews the displayed text and makes corrections as needed using touch input or other methods. The corrected text data is then obtained.
[1658] Step 8:
[1659] The user reviews the revised text and submits it. The submission method then sends the finalized text to the specified recipient. The finalized text data is received and sent, for example, to the store's backend system or a specific email address. The submission is complete.
[1660] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1661] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1662] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1663] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1664] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1665] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1666] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1667] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1668] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1669] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1670] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1671] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1672] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1673] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1674] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1675] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1676] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1677] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1678] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1679] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1680] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1681] The following is further disclosed regarding the embodiments described above.
[1682] (Claim 1)
[1683] An input receiving method for the user to input the requirements of the text they want to generate,
[1684] An analysis means that analyzes the input requirements on the server and sends a request to the generating AI,
[1685] A text generation means that generates text based on a request analyzed by a generation AI,
[1686] A server receives the generated text and a response receiving means to display it to the user,
[1687] A means for the user to review and correct the displayed text,
[1688] A means of sending the revised text after the user has made a final check,
[1689] A system that includes this.
[1690] (Claim 2)
[1691] The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
[1692] (Claim 3)
[1693] The system according to claim 1, comprising a function for sending a generated document to a specified recipient.
[1694] "Example 1"
[1695] (Claim 1)
[1696] An input receiving method for the user to input the requirements of the text they want to generate,
[1697] A packet generation means that formats the input requirements of the terminal into a data packet and transmits it,
[1698] An analysis means that analyzes the requirements received by the server and sends a request to the generating AI,
[1699] A text generation means that generates text based on a request analyzed by a generation AI,
[1700] A server formats the generated text and provides a response receiving mechanism to display it to the user.
[1701] A means for the user to review and correct the displayed text,
[1702] A means of sending the revised text after the user has made a final check,
[1703] A system that includes this.
[1704] (Claim 2)
[1705] The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
[1706] (Claim 3)
[1707] The system according to claim 1, comprising a function for sending a generated document to a specified recipient.
[1708] "Application Example 1"
[1709] (Claim 1)
[1710] An input receiving method for the user to input the requirements of the text they want to generate,
[1711] An analysis means that analyzes the input requirements on the server and sends a request to the generating AI,
[1712] A text generation means that generates text based on a request analyzed by a generation AI,
[1713] A server receives the generated text and a response receiving means to display it to the user,
[1714] A means for the user to review and correct the displayed text,
[1715] A means of sending the revised text after the user has made a final check,
[1716] A means of publishing the generated text to a specified website,
[1717] A system that includes this.
[1718] (Claim 2)
[1719] The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
[1720] (Claim 3)
[1721] The system according to claim 1, comprising the function of sending the generated text to a specified recipient and publishing it on a website.
[1722] "Example 2 of combining an emotion engine"
[1723] (Claim 1)
[1724] An input receiving method for the user to input the requirements of the text they want to generate,
[1725] A means for collecting emotional data in which the terminal acquires the user's emotional state and sends it to a server,
[1726] The server analyzes the input requirements and emotional state and sends a request to the generating AI;
[1727] A text generation means that generates text based on a request analyzed by a generation AI,
[1728] A server receives the generated text and a response receiving means to display it to the user,
[1729] A means for the user to review and correct the displayed text,
[1730] A means of sending the revised text after the user has made a final check,
[1731] A system that includes this.
[1732] (Claim 2)
[1733] The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
[1734] (Claim 3)
[1735] The system according to claim 1, comprising a function for sending a generated document to a specified recipient.
[1736] "Application example 2 when combining with an emotional engine"
[1737] (Claim 1)
[1738] An input receiving method for the user to input the requirements of the text they want to generate,
[1739] An analysis means that analyzes the input requirements on the server and sends a request to the generating AI,
[1740] A text generation means that generates text based on a request analyzed by a generation AI,
[1741] A server receives the generated text and a response receiving means to display it to the user,
[1742] A means for the user to review and correct the displayed text,
[1743] A means of sending the revised text after the user has made a final check,
[1744] An emotion recognition means that recognizes the user's emotions at the time of input,
[1745] An emotion adjustment means that adjusts the text based on the emotion recognized by the emotion recognition means,
[1746] A system that includes this.
[1747] (Claim 2)
[1748] The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
[1749] (Claim 3)
[1750] The system according to claim 1, comprising a function for sending a generated document to a specified recipient. [Explanation of Symbols]
[1751] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input receiving method for the user to input the requirements of the text they want to generate, An analysis means that analyzes the input requirements on the server and sends a request to the generating AI, A text generation means that generates text based on a request analyzed by a generation AI, A server receives the generated text and a response receiving means to display it to the user, A means for the user to review and correct the displayed text, A means of sending the revised text after the user has made a final check, A system that includes this.
2. The system according to claim 1, further comprising display means for displaying generated text when the user reviews and corrects it.
3. The system according to claim 1, further comprising a function for sending a generated document to a specified recipient.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A