System
The system addresses communication challenges by allowing users to input keywords and settings, generating natural conversational sentences through a server, and displaying them, enhancing conversational skills and relationship building.
Patent Information
- Application Number
- JP2024133490
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Users who are not good at communication face challenges in creating conversational flow, leading to stalled conversations and isolation, with existing technologies lacking effective methods to support such users and generate natural-sounding conversational sentences efficiently.
A system that includes a user terminal for inputting keywords and settings, a server for analyzing and generating natural conversational sentences using generative AI, and a display for presenting the generated sentences, allowing users to improve their conversational skills and build relationships.
Enables users to generate and use natural-sounding conversational sentences to improve their communication abilities, reducing feelings of loneliness and facilitating relationship building.
Smart Images

Figure 2026030507000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Users who are not good at communication have difficulty creating a flow in the conversation and conversations tend to stall. This can make it difficult to build and maintain relationships. Existing technologies do not provide adequate methods to fully support such users, which can lead to users feeling isolated. In addition, there is a lack of a low-cost, highly efficient way to generate natural conversational sentences, making it difficult to accommodate a wide range of users. [Means for solving the problem]
[0005] The present invention provides a system including a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a user terminal to a server via a communications network, a means for analyzing the received data and generating natural-sounding conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the user terminal. This allows users to automatically generate natural-sounding conversational sentences based on specified keywords and settings and use them as reference for actual conversations. Furthermore, by including a means for the user terminal to receive the generated conversational sentence data and display it to the user, and a means for generating prompts for the generation AI, more accurate conversational sentence generation is achieved. This improves the user's conversational ability and makes it easier to build and maintain relationships.
[0006] "Generated conversation data" refers to natural conversation data created by the generation AI based on keywords and settings specified by the user.
[0007] A "communications network" is an infrastructure for transmitting and receiving data between terminals and servers.
[0008] A "user terminal" is a device that a user directly operates to input and display data.
[0009] A "server" is a computer system that receives and analyzes data sent from a user terminal.
[0010] "Keywords" are words or phrases that indicate specific topics or conditions that users specify for the AI generator.
[0011] "Settings" are additional conditions or parameters that the user specifies to influence the response of the generated AI.
[0012] "Generative AI" is an artificial intelligence technology that generates natural-sounding conversational sentences based on specified keywords and settings.
[0013] A "prompt" is a specific instruction or input information given to the generation AI to generate a conversational sentence.
[0014] "Parsing" is the process by which the server understands the data received from the user and generates an appropriate response.
[0015] "Natural conversational text" is text created by generative AI that has the same flow and language as human conversation.
[0016] The "display means" is a function for visually presenting the generated conversation sentence data on the screen of the user terminal.
[0017] The "transmitting means" is a function for sending data from a user terminal to a server or from a server to a user terminal. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention comprises a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the user terminal. This allows the user to automatically generate natural conversational sentences based on specified keywords and settings and use them as a reference for actual conversations.
[0040] Program processing flow
[0041] 1. User Input
[0042] The user inputs specific keywords and settings into the terminal, for example, the keyword "dinner conversation with friends" and the setting "casual atmosphere."
[0043] 2. Data Transmission
[0044] The terminal sends the input keywords and setting data to the server via a communication network.
[0045] 3. Data Receipt and Analysis
[0046] The server receives the data sent from the device, analyzes the received keywords and settings, and generates the appropriate prompts.
[0047] 4. Conversation generation using generative AI
[0048] The server uses a generative AI (e.g., GPT-4) to generate natural-sounding conversational sentences based on the generated prompts. The generative AI takes into account the input keywords and settings to create natural-sounding conversational sentences that fit the context.
[0049] 5. Sending conversation text
[0050] The server transmits the generated conversation sentence data to the terminal again via the communication network.
[0051] 6. User Visibility
[0052] The device displays the conversation data received from the server, and the user can use the displayed conversation data as a reference for actual conversations.
[0053] Specific examples
[0054] 1. User Input
[0055] The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0056] 2. Data Transmission
[0057] The terminal sends the entered data to the server. Communication is carried out using the HTTPS protocol.
[0058] 3. Data Receipt and Analysis
[0059] The server receives the data, analyzes the keywords and settings, and generates a prompt based on the analysis results, which it prepares to send to the AI.
[0060] 4. Conversation generation using generative AI
[0061] The server sends the following prompt to the AI: "Generate a casual conversation about a meal with a friend." Based on this instruction, the AI generates the following conversation:
[0062] A: What do you want to eat today? Pizza?
[0063] B: That's great! But I hear the pasta here is delicious too.
[0064] A: Right, then let's order both pasta and pizza and share.
[0065] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0066] 5. Sending conversation text
[0067] The server transmits the generated conversation sentence to the terminal via a communication network.
[0068] 6. User Visibility
[0069] The device receives the conversation data and displays it to the user, who can use it as a reference when dining with real friends.
[0070] This invention allows users to improve their conversational skills and obtain natural conversational sentences that can be used as reference in actual communication situations. This allows even users who are not good at communication to enjoy smooth conversations, which greatly contributes to reducing feelings of loneliness and building relationships.
[0071] The processing flow will be explained below.
[0072] Step 1:
[0073] The user inputs a specific keyword and setting into the terminal. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. After inputting, the user presses the data transmission button.
[0074] Step 2:
[0075] The device organizes the input keywords and settings data and sends it to the server. Specifically, it converts the keywords and settings into JSON format and sends it to the server's API endpoint via an HTTPS request.
[0076] Step 3:
[0077] The server receives the data sent from the device. Specifically, the server's API endpoint processes the HTTPS request and converts the received data into an internal format for analysis.
[0078] Step 4:
[0079] The server analyzes the received keywords and settings, and generates prompts for the AI based on the keywords and settings. These prompts reflect the context and mood of the conversation the user desires.
[0080] Step 5:
[0081] The server sends a prompt to the generating AI (e.g., GPT-4), which contains information based on user-specified keywords and preferences.
[0082] Step 6:
[0083] Generative AI generates natural-sounding conversations based on prompts, taking into account the context, flow of the dialogue, and word usage to create conversations that fit the specified conditions.
[0084] Step 7:
[0085] The server receives the conversation data generated by the AI, checks the quality of the data, and formats it before providing it to the user.
[0086] Step 8:
[0087] The server sends the generated conversation data to the device. Again, the data is sent in JSON format as an HTTPS response.
[0088] Step 9:
[0089] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[0090] Step 10:
[0091] The terminal displays the formatted conversation sentences to the user, who can refer to the generated natural conversation sentences and use them in actual conversations.
[0092] This allows users to improve their conversational skills and communicate more smoothly.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] Conventional conversation support systems have struggled to automatically generate natural-sounding conversational sentences for users to refer to. Furthermore, generating custom conversational sentences based on user-specified keywords and conversational tones requires advanced expertise, making them difficult for average users to use. Furthermore, there was a lack of a way to quickly generate natural-sounding conversational sentences and provide them to users, making real-time use difficult.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes means for a user to input specific keywords and settings to a terminal, means for transmitting the input keywords and settings to the server via a communication network, means for the server to analyze the received keywords and settings and generate prompts for a generative AI model, means for generating natural conversational sentences based on the prompts using the generative AI model, means for transmitting the generated conversational sentence data from the server again via the communication network to the terminal, and means for displaying the generated conversational sentence data received by the terminal to the user. This enables the user to easily generate and refer to natural conversational sentences in real time according to the set keywords and atmosphere.
[0098] "User" refers to an individual or group that uses the system and is the entity that inputs keywords and settings via a terminal.
[0099] "Terminal" refers to a device operated by a user, including devices such as smartphones, personal computers, and tablets.
[0100] "Keywords" refer to specific words or phrases that users input to generate conversation text, and they play a role in determining the theme and content of the generated conversation text.
[0101] "Settings" refers to additional conditions or attributes that a user specifies when generating a conversation, including details such as the tone and atmosphere of the conversation.
[0102] "Communications network" refers to an infrastructure that enables the transmission and reception of data, including the Internet and local area networks.
[0103] "Server" refers to a central processing unit that processes data received from the terminal and sends prompts to the generating AI.
[0104] A "prompt" refers to an instruction entered into a generative AI model, which serves as a guide for the generated conversational text.
[0105] "Generative AI models" refer to artificial intelligence algorithms that generate natural-sounding conversational sentences based on input prompts, including, for example, large-scale language models.
[0106] "Conversational text" refers to part or all of a sentence that is generated in a form that can be referenced by the user, and refers to text that has a natural dialogue format.
[0107] "Data" refers to keywords, settings, generated dialogue, and related information in general.
[0108] "JSON format" refers to a standard data representation format for structuring data in a way that is easy for humans to read and machines to analyze.
[0109] This invention is a system that allows a user to input specific keywords and settings and generates natural conversational sentences based on the input. The system is implemented mainly using the following hardware and software.
[0110] System configuration and operation
[0111] Hardware
[0112] 1. User device: A device such as a smartphone, personal computer, or tablet. These devices provide an interface for users to enter keywords and settings.
[0113] 2. Server: A high-performance server or cloud environment. The server is used to receive and analyze data and operate the generative AI model.
[0114] software
[0115] 1. Generative AI models: Large-scale language models such as GPT-4 from OpenAI, which generate highly natural-sounding conversational sentences.
[0116] 2. Communication protocol: The HTTPS protocol is used for communication between the terminal and the server to ensure data security and reliability.
[0117] 3. Data analysis tools: Use analysis tools such as Python's json module to receive and analyze data.
[0118] Data processing and calculation
[0119] The user inputs specific keywords, such as "dinner conversation with friends," and settings, such as "casual atmosphere," into the terminal. The specific data processing and calculations performed at each step are described in detail below.
[0120] 1. User Input: The user enters keywords and settings into the terminal application. The input is done in a text box, and when the submit button is pressed, the data is converted into JSON format.
[0121] 2. Data transmission: The device uses the HTTPS protocol to send JSON format data to the server. The data is encrypted before being sent.
[0122] 3. Receiving and parsing data: The server parses the received JSON data and extracts keywords and settings. This parsing is done using the Python json module.
[0123] 4. Prompt generation: Based on the extracted keywords and settings, the server generates a prompt to send to the generative AI model. An example of a prompt is "Generate a casual dinner conversation with a friend."
[0124] 5. Conversational Sentence Generation by Generative AI: The server sends prompts to the generative AI model, which then generates natural-sounding conversational sentences using hundreds of billions of parameters to provide context-appropriate answers.
[0125] 6. Sending the conversation: The generated conversation is converted back to JSON format and sent to the device via HTTPS.
[0126] 7. Display to user: The device parses the received JSON data and displays the generated conversation text to the user. The user can refer to this conversation text and use it in actual conversations.
[0127] Specific examples
[0128] The specific flow of system usage is shown below.
[0129] 1. User Input: The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0130] 2. Prompt generation: The server sends the generative AI model a prompt like this: "Generate a casual dinner conversation with a friend."
[0131] 3. Conversation generation results: The generative AI model generates the following conversation:
[0132] A: What do you want to eat today? Pizza?
[0133] B: That's great! But I hear the pasta here is delicious too.
[0134] A: Right, then let's order both pasta and pizza and share.
[0135] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0136] This system allows users to improve their conversational skills, easily generate natural conversational sentences in real time, and use them in real-life communications.
[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0138] Step 1:
[0139] The user opens a smartphone application or web browser, inputs specific keywords and settings, such as "dinner conversation with friends" and "casual atmosphere" in the text boxes, and then presses the submit button. This input data is passed to the next step.
[0140] Step 2:
[0141] The device converts the keywords and settings entered by the user into JSON format data. For example, the following JSON data is generated:
[0142] json
[0143] {
[0144] "keyword": "dinner conversation with friends",
[0145] "setting": "casual atmosphere"
[0146] }
[0147] This JSON data is sent to the server using the HTTPS protocol, and the device confirms that the data is sent encrypted.
[0148] Step 3:
[0149] The server receives the JSON data sent from the device. It parses the data using Python's json module and extracts keywords and settings. The analysis results are passed to the next step. From the input JSON data, "dinner conversation with friends" and "casual atmosphere" are extracted.
[0150] Step 4:
[0151] Based on the extracted keywords and settings, the server generates a prompt to be sent to the generative AI model. Specifically, it generates the prompt "Generate a casual dinner conversation with a friend." This prompt is passed to the next step.
[0152] Step 5:
[0153] The server sends the prompt to a generative AI model (e.g., GPT-4) and requests it to generate a conversation. The generative AI model analyzes the prompt and generates natural-sounding conversational sentences that meet the required conditions. For example, the following conversational sentences are generated:
[0154] A: What do you want to eat today? Pizza?
[0155] B: That's great! But I hear the pasta here is delicious too.
[0156] A: Right, then let's order both pasta and pizza and share.
[0157] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0158] The generated conversation sentence is passed to the next step.
[0159] Step 6:
[0160] The server converts the generated conversation text back into JSON format and sends it to the device. At this time, the server checks the integrity of the generated data and performs error handling if necessary. The generated conversation text is converted into JSON data as follows:
[0161] json
[0162] {
[0163] "conversation": [
[0164] "A: What do you want to eat today? Pizza?"
[0165] "B: That's great! But I heard the pasta here is also delicious."
[0166] "A: Right, then let's order both pasta and pizza and share."
[0167] "B: That might be a good idea. By the way, have you traveled anywhere recently?"
[0168] ]
[0169] }
[0170] This JSON data is sent to the device.
[0171] Step 7:
[0172] The device parses the JSON data received from the server and displays it on the application's UI. The user can refer to the natural conversational text displayed on the screen and use it in actual conversations. The device displays the generated conversational text to the user through the user interface, allowing the user to perform additional operations such as scrolling and saving.
[0173] By following the steps above, users can easily and quickly generate natural conversational sentences based on specified conditions and use them in actual conversations.
[0174] (Application example 1)
[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0176] Customer service in modern brick-and-mortar stores requires store clerks to respond quickly and appropriately to customer questions and requests. However, this depends on the experience and knowledge of the store clerk, which can lead to inconsistencies in individual responses. Furthermore, responding immediately in a busy environment can be difficult, which can lead to a decrease in customer satisfaction. A solution to this issue is needed, enabling store clerks to consistently provide high-quality service to customers.
[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0178] In this invention, the server includes means for receiving customer questions and requests by voice input, means for converting the voice input data into text data, and means for generating prompts based on the text data, thereby enabling store staff to provide quick and appropriate responses to customer questions and requests.
[0179] "Conversational data" is text information of natural conversations created by generative AI.
[0180] A "communications network" is an information pathway such as the Internet or a local area network that allows data to be sent and received.
[0181] A "terminal" is a user device such as a smartphone, smart glasses, or head-mounted display operated by a user or store clerk.
[0182] A "server" is a computer system that processes data and generates conversational sentences using AI.
[0183] "Keywords" are specific words or phrases entered by the user, and serve as basic information for generating conversational text.
[0184] "Settings" are parameters related to the tone and topic of the conversation specified by the user.
[0185] "Generative AI" is an artificial intelligence system that generates natural-sounding conversational sentences based on input keywords and settings.
[0186] "Voice input" is a method in which a user or store clerk speaks to a terminal and the voice data is acquired.
[0187] "Text data" refers to data obtained by converting information input by voice into a character string format.
[0188] A "prompt" is a text instruction that serves as the basis for the generation AI to generate conversational sentences.
[0189] This invention is a system including a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the terminal. Also, by adding a means for generating prompts based on keywords and settings input by a user and a means for converting voice-input data into text data, this is a specific embodiment for supporting communication between store clerks and customers.
[0190] Hardware and software used
[0191] Hardware
[0192] Devices: smartphones, smart glasses, head-mounted displays, etc.
[0193] Server: A high-performance cloud server, such as an AWS EC2 server.
[0194] software
[0195] Speech recognition engine: Google Speech-to-Text API
[0196] Natural language analysis engines: SpaCy, NLTK
[0197] Generative AI model: OpenAI GPT-4
[0198] Communication protocol: HTTPS
[0199] What the system does
[0200] First, the store clerk (the user) uses the voice input function of the terminal to input the customer's question or request by voice. For example, a customer might ask, "Do you have this shirt in my size?" This voice input is converted into text data on the terminal and sent to the server using the HTTPS protocol.
[0201] The server analyzes the received text data using a natural language analysis engine (SpaCy or NLTK) and generates an appropriate prompt. The generated prompt is passed to a generative AI model (OpenAI GPT-4), which then generates natural-sounding conversational sentences based on the prompt. For example, in response to the prompt, "You have requested stock information about the size of this shirt. Please generate an appropriate response," the generative AI generates the following conversational sentences:
[0202] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0203] The generated conversation sentences are sent from the server to the terminal again and displayed in the field of view of the store clerk. The store clerk can refer to the displayed conversation sentences and provide an appropriate response to the customer.
[0204] Specific examples
[0205] 1. Example prompt:
[0206] When a customer asks, "Do you have this shirt in my size?", generate an appropriate response.
[0207] 2. Example of generated dialogue:
[0208] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0209] This invention enables store staff to provide quick and appropriate responses to customer questions and requests, which is expected to improve the quality of customer service in physical stores.
[0210] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0211] Step 1:
[0212] A user uses the smart glasses to input a customer question or request by voice. For example, a question such as "Do you have this shirt in my size?" The voice data is the input data.
[0213] Step 2:
[0214] The device uses a speech recognition engine (Google Speech-to-Text API) to convert the voice input data into text data. The voice data is converted into text data.
[0215] Step 3:
[0216] The terminal sends the converted text data to the server via a communication protocol (HTTPS). The text data sent is the input data, and the text data received by the server is the output data.
[0217] Step 4:
[0218] The server receives the text data and analyzes it using a natural language analysis engine (SpaCy, NLTK). The text data is the input data, and the analysis results are the output data.
[0219] Step 5:
[0220] Based on the analysis results, the server generates prompts to send to the generative AI model (GPT-4). The analysis results are the input data, and the generated prompts are the output data.
[0221] Step 6:
[0222] The server sends prompts to the generative AI model (GPT-4) to generate natural-sounding conversational sentences. The generated prompts are the input data, and the generated conversational sentences are the output data.
[0223] Step 7:
[0224] The server sends the generated conversation text to the terminal via HTTPS. The generated conversation text is the input data, and the conversation text received by the terminal is the output data.
[0225] Step 8:
[0226] The conversational text received by the terminal is displayed on the smart glasses. The received conversational text is input data, and the conversational text displayed on the smart glasses is output data.
[0227] Step 9:
[0228] The user refers to the displayed conversation sentence and provides an appropriate response to the customer. The displayed conversation sentence is input data, and the user's response is output data.
[0229] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0230] The present invention provides a system including a means for displaying generated conversational text data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational text using a generation AI, a means for transmitting the generated conversational text data to the user terminal, and an emotion engine for recognizing the user's emotions.
[0231] Program processing flow
[0232] 1. User Input
[0233] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. Furthermore, the user's facial expressions and tone of voice are monitored by the device's emotion engine.
[0234] 2. Data Transmission
[0235] The device sends the input keywords and settings, as well as the user's emotional data analyzed by the emotion engine, to the server via a communication network.
[0236] 3. Data Receipt and Analysis
[0237] The server receives the data sent from the device, analyzes the received keywords, settings, and emotional data, and generates prompts appropriate to the user's emotions.
[0238] 4. Conversation generation using generative AI
[0239] The server sends emotion-based prompts to the AI, which contain information corresponding to the user's emotional state. The AI then uses these prompts to generate natural-sounding conversational sentences.
[0240] 5. Emotional Adaptation Check
[0241] The server checks whether the generated dialogue is appropriate to the user's emotions. The emotion engine performs emotion analysis again, and if necessary, modifies the prompt and generates it again.
[0242] 6. Sending conversation text
[0243] The server transmits the generated conversation data to the terminal, and the data is transmitted again via the communication network.
[0244] 7. User Visibility
[0245] The device receives the generated conversation data and displays it to the user, who can then use it in actual conversations.
[0246] Specific examples
[0247] 1. User Input
[0248] When a user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere," the emotion engine detects that the user is slightly nervous.
[0249] 2. Data Transmission
[0250] The device sends data, including the entered keywords, settings, and tension generated by the emotion engine, to the server using the HTTPS protocol.
[0251] 3. Data Receipt and Analysis
[0252] The server receives the data and analyzes it for keywords, settings, and tension. As a result of the analysis, it generates prompts that help the user relax.
[0253] 4. Conversation generation using generative AI
[0254] The server sends the following prompt to the AI: "Generate a conversation scenario for a user who is nervous about having a casual meal with a friend." Based on this instruction, the AI generates the following dialogue:
[0255] A: Where do you want to eat today?
[0256] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[0257] A: Great, let's do that! How have you been lately?
[0258] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[0259] 5. Emotional Adaptation Check
[0260] The server checks whether the generated dialogue relieves the user's tension. The emotion engine analyzes the user's emotions again, and if necessary, modifies and regenerates the prompt.
[0261] 6. Sending conversation text
[0262] The server sends the confirmed conversation data to the terminal as an HTTPS response.
[0263] 7. User Visibility
[0264] The terminal receives the conversation sentence data and displays it to the user, who can then refer to the generated natural conversation sentences and have a relaxed conversation.
[0265] The present invention allows users to obtain natural conversational sentences that are adapted to their emotional state, enabling smooth communication. This allows even users who are not good at communication to enjoy conversations with ease, and also contributes to building interpersonal relationships.
[0266] The processing flow will be explained below.
[0267] Step 1:
[0268] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. At this time, the device analyzes the user's facial expression data and tone of voice using an emotion engine to extract emotional data.
[0269] Step 2:
[0270] The device organizes the input keywords and settings, as well as the user's emotional data extracted by the emotion engine, and sends them to the server. Specifically, it converts the keywords, settings, and emotional data into JSON format and sends it as an HTTPS request to the server's API endpoint.
[0271] Step 3:
[0272] The server receives data sent from the device, which is then processed by the API endpoint and converted into an internal format for analysis.
[0273] Step 4:
[0274] The server analyzes the received keywords, preferences, and emotional data. Based on the analysis, it generates prompts that take the user's emotions into account. For example, if the user is nervous, the prompt may include instructions such as "generate relaxing conversation content."
[0275] Step 5:
[0276] The server sends a generated prompt to the generation AI (e.g., GPT-4), which includes the user's emotional state and the specified keywords and preferences.
[0277] Step 6:
[0278] The AI generates natural-sounding dialogue based on prompts, taking into account context, emotions, and the flow of the conversation, and tailors the dialogue to match the user's emotions.
[0279] Step 7:
[0280] The server receives the conversation data generated by the generation AI, checks the quality of the received conversation data, and again determines whether it is appropriate for the user's emotional state.
[0281] Step 8:
[0282] The server uses the emotion engine to check whether the generated dialogue is appropriate to the user's emotions. If it is not appropriate, it modifies the prompt and sends it to the generation AI again.
[0283] Step 9:
[0284] The server sends the confirmed conversation data to the terminal, which is then sent in JSON format as an HTTPS response.
[0285] Step 10:
[0286] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[0287] Step 11:
[0288] The device displays the formatted conversational text to the user, who can then refer to the natural-sounding conversational text and use it in actual conversations. The emotion engine ensures that the conversational text takes into account the user's current emotional state, allowing the user to continue the conversation with peace of mind.
[0289] This allows the user to obtain natural conversational sentences that are adapted to his or her own emotional state, making actual communication smoother.
[0290] Example 2
[0291] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0292] Conventional systems have the problem that when a user inputs specific keywords or settings to generate conversational sentences, they are unable to generate conversational sentences that correspond to the user's emotional state. As a result, it is difficult to obtain conversational content that matches the user's emotional state, such as when they are tense or relaxed, making it difficult to build natural communication.
[0293] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0294] In this invention, the server includes means for a user to input keywords and settings into a terminal and collect emotional data, means for transmitting the keywords, settings, and emotional data from the user terminal to the server via a communication network, means for analyzing the received data and generating prompts based on the user's emotional state, means for generating natural conversational sentences from the prompts using a generative AI model, means for checking whether the generated conversational sentence data is appropriate for the user's emotions and for regenerating the data as necessary, and means for transmitting the generated conversational sentence data to the user terminal. This makes it possible to generate natural conversational sentences that are adapted to the user's emotional state and support smooth communication.
[0295] A "user" is a person who operates a device to input keywords and settings and obtain conversational text based on their emotional state.
[0296] A "terminal" is a device operated by a user, which collects input keywords, settings, and emotional data and transmits them to a server.
[0297] "Emotional data" is data analyzed by the emotion engine based on the user's facial expressions, tone of voice, etc.
[0298] A "communication network" is a network infrastructure for data communication between a user terminal and a server.
[0299] A "server" is a computer system that receives data sent from a user terminal, analyzes it, and generates conversational sentences.
[0300] A "generative AI model" is an artificial intelligence that uses pre-trained algorithms to generate natural-sounding conversational sentences from generated prompts.
[0301] A "prompt" is an instruction or question-type sentence that tells the generative AI model to generate natural conversational sentences.
[0302] "Conversational text" is a natural conversational text generated by a generative AI model and provided to the user.
[0303] The present invention is a system that generates natural conversational sentences that correspond to the user's emotional state using a user terminal, a server, a generative AI model, and an emotion engine.
[0304] The system includes a user terminal, a communication network, a server, a generative AI model, and an emotion engine. The system operates as follows.
[0305] Hardware and software used
[0306] User devices include smartphones, tablets, and other devices. These devices include an interface for entering keywords and settings using a keyboard or touchscreen. They also include an emotion engine for analyzing facial expressions and tone of voice. The emotion engine may use, for example, facial recognition software or voice analysis software.
[0307] The communication network is the infrastructure that connects the user terminal and the server. Specifically, the Internet or mobile data communication is often used. Data is transmitted and received securely using the HTTPS protocol.
[0308] The server is a computer system that receives data sent from the user's device, analyzes it, and generates conversational sentences. The server is equipped with a generative AI model, such as GPT-4, that has the ability to generate advanced natural language.
[0309] Data processing and calculation
[0310] The user enters keywords and settings into the device to collect emotional data. This data is then sent to a server via a communications network. The server analyzes the received data and generates prompts based on the user's emotional state. These prompts are then input into a generative AI model to generate natural-sounding conversational sentences.
[0311] The generated conversation sentences are checked by the server to see if they are appropriate for the user's emotions, and are regenerated if necessary. Finally, the generated conversation sentence data is sent to the user's terminal and displayed to the user.
[0312] Specific examples
[0313] A user opens a smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere." The emotion engine analyzes the user's facial expressions and tone of voice and detects that the user is "slightly nervous." This data is sent to the server, which then analyzes the keywords, settings, and emotion data.
[0314] The analysis results in prompts that help users relax. Based on these prompts, the generative AI model generates natural-sounding conversational sentences like the following:
[0315] A: Where do you want to eat today?
[0316] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[0317] A: Great, let's do that! How have you been lately?
[0318] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[0319] The generated conversation sentences are then checked again by the emotion engine, and after corrections are made as necessary, they are sent to the user's device. The user can refer to these conversation sentences and proceed with the actual conversation in a relaxed manner.
[0320] As described above, the present invention generates natural conversational sentences that are adapted to the emotional state of the user, thereby supporting smooth communication.
[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0322] Step 1:
[0323] The user inputs keywords and settings into the device to collect emotional data. Specifically, the user inputs keywords and settings using the device's keyboard or touch screen. At this time, the device's camera and microphone analyze the user's facial expressions and tone of voice using an emotion engine to obtain emotional data.
[0324] Input: Keywords, settings, user facial expression and tone of voice
[0325] Output: Emotion data, keywords and settings
[0326] Step 2:
[0327] The device then transmits the collected keywords, settings, and emotion data to a server, where the data is transmitted securely over a communications network using the HTTPS protocol.
[0328] Input: Keywords, settings, emotion data
[0329] Output: Data sent to the server
[0330] Step 3:
[0331] The server receives the data sent from the device and analyzes keywords, preferences, and emotional data. Specifically, it uses an analysis algorithm on the received data to generate prompts based on the user's emotional state.
[0332] Input: Keywords, settings, emotion data
[0333] Output: Generated prompt
[0334] Step 4:
[0335] The server inputs the generated prompts to the generative AI model to generate natural conversational sentences. In this process, the generative AI model generates dialogue-style sentences based on the prompts.
[0336] Input: Generated prompt
[0337] Output: Natural conversational sentences
[0338] Step 5:
[0339] The server checks whether the generated dialogue is appropriate for the user's emotional state, again using the emotion engine to analyze the user's emotional changes in response to the dialogue, modifying the prompts as necessary, and re-inputting them into the generative AI model.
[0340] Input: Generated conversation sentences, user reaction data
[0341] Output: Modified prompt (if needed)
[0342] Step 6:
[0343] The server then sends the final generated conversation data to the device, again securely using the HTTPS protocol.
[0344] Input: Final generated dialogue
[0345] Output: Conversation data sent to the device
[0346] Step 7:
[0347] The device displays the received conversation data to the user, who can then refer to the displayed conversation data and use it in actual conversations. Specifically, the generated conversation data is displayed on the device screen.
[0348] Input: Received conversation data
[0349] Output: The dialogue displayed to the user
[0350] (Application example 2)
[0351] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0352] Conventional conversation generation systems have had difficulty generating natural conversations that reflect the user's emotions and situation. In particular, when ordering food delivery, appropriate conversations are required that reflect the user's stress or relaxation state, but these systems have not been able to meet this need. This often results in an unsmooth ordering process and a poor user experience.
[0353] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for displaying the generated conversation sentence data, means for transmitting data including keywords and settings from the user terminal to the server via a communication network, means for analyzing the received data and generating natural conversation sentences using a generation AI, means for transmitting the generated conversation sentence data to the user terminal, and means for generating natural ordering conversation sentences based on the user's emotion data in the case of ordering food delivery. This makes it possible to generate natural conversation sentences that correspond to the user's emotions and situation.
[0354] The "means for displaying the generated conversation sentence data" is a function or device for displaying the generated conversation sentence data on the screen of a user terminal or the like.
[0355] The "means for transmitting data including keywords and settings from a user terminal to a server via a communications network" refers to a communications means and protocol for transmitting keywords and settings entered by a user to a server.
[0356] The "means for analyzing received data and generating natural conversational sentences using a generation AI" refers to an algorithm and processing device for analyzing received data and generating natural conversational sentences using the analysis results and a generation AI.
[0357] The "means for transmitting the generated conversation sentence data to the user terminal" refers to a communication means and protocol for transmitting the generated conversation sentence data from the server to the user terminal.
[0358] The "means for generating natural ordering conversations based on the user's emotional data when ordering food delivery" refers to an algorithm and processing device that takes into account the user's emotional data when ordering food delivery and generates natural ordering conversations based on that data.
[0359] This invention relates to a system that facilitates the food delivery ordering process by generating natural ordering conversations based on the user's emotions. Specific embodiments for implementing this system are described below.
[0360] System Overview
[0361] This system generates natural conversational sentences and provides them to users when they order food delivery using their smartphones. The system consists of the following main components:
[0362] Hardware / Software Components
[0363] 1. Smartphone
[0364] Emotion engine: An engine that recognizes the user's emotions, for example by using a camera and microphone to analyze the user's facial expressions and tone of voice.
[0365] Communication module: A module to support HTTPS communication.
[0366] 2. Server
[0367] Data analysis engine: Analyzes keywords, settings, and sentiment data entered by users.
[0368] Generative AI models (e.g., GPT-3): Generate prompts based on user sentiment and generate natural-sounding conversations.
[0369] Communication module: A module for receiving data from the user terminal and transmitting the generated conversation data to the user terminal.
[0370] Data processing and calculation
[0371] User Input
[0372] When a user opens the smartphone application and enters keywords related to the order (e.g., "Italian restaurant") and a preference (e.g., "relaxed atmosphere"), the emotion engine collects the user's emotional data. For example, if the user is relaxed, that data is collected.
[0373] Data transmission and analysis
[0374] The smartphone's communication module sends keywords, settings, and emotional data to the server using the HTTPS protocol, where the server's data analysis engine analyzes the received data and generates prompts for the generative AI model.
[0375] Conversation generation using generative AI
[0376] The generative AI model generates natural-sounding sentences based on prompts such as:
[0377] "Generate an Italian restaurant ordering scenario for a relaxed user."
[0378] Here are some examples of dialogue generated by the generative AI:
[0379] A: Hello, I'd like some suggestions for Italian restaurants.
[0380] B: Of course! I recommend the Margherita pizza. It's a very relaxing dish.
[0381] Sending and displaying conversation text
[0382] The server then sends the generated conversational text to the smartphone, which receives and displays it on the user's device. The user can refer to the natural conversational text displayed and place their order smoothly.
[0383] This will enable the system to provide natural-sounding conversations that reflect the user's emotions and situation, which is expected to make the food delivery ordering process go more smoothly.
[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0385] Step 1:
[0386] A user launches the smartphone application and enters keywords related to their order (e.g., "Italian restaurant") and preferences (e.g., "relaxed atmosphere"). At this time, the smartphone's emotion engine analyzes the user's facial expressions and tone of voice in real time to collect emotional data (whether they are relaxed or not).
[0387] Input data: Keywords, settings, and sentiment data
[0388] Output data: User input dataset
[0389] Step 2:
[0390] The terminal's communication module compiles the collected keywords, preferences, and emotion data and transmits them to a server using the HTTPS protocol.
[0391] Input data: User input dataset
[0392] Output data: JSON formatted data packet to the server
[0393] Step 3:
[0394] The server receives data packets from the user terminal, and the analysis engine analyzes them. The analysis engine extracts keywords, preferences, and emotion data from the data packets, and then creates prompts for the generative AI model.
[0395] Input data: JSON format data packet to the server
[0396] Output data: prompt statement
[0397] Step 4:
[0398] The server's generative AI model generates natural-sounding conversational sentences based on the generated prompt, for example, "Generate an ordering scenario at an Italian restaurant for a relaxed user."
[0399] Input data: Prompt statement
[0400] Output data: Generated conversation
[0401] Step 5:
[0402] The server then rechecks the generated dialogue to verify its suitability for the user's emotional data, and if necessary performs additional analysis to refine it and create the complete dialogue.
[0403] Input data: Generated conversation
[0404] Output: Verified and possibly corrected dialogue
[0405] Step 6:
[0406] The server then sends the final generated conversation text to the user's device using the HTTPS protocol, allowing the user to obtain the conversation text to smoothly complete the ordering process.
[0407] Input data: Verified and corrected conversation text
[0408] Output data: Conversation data packets to the user terminal
[0409] Step 7:
[0410] The terminal receives the conversation data sent from the server and displays it on the smartphone screen. The user can then place an order based on the natural conversation displayed.
[0411] Input data: Conversation data packets to the user terminal
[0412] Output data: Conversation displayed on smartphone screen
[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0416] [Second embodiment]
[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0427] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0429] The present invention comprises a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the user terminal. This allows the user to automatically generate natural conversational sentences based on specified keywords and settings and use them as a reference for actual conversations.
[0430] Program processing flow
[0431] 1. User Input
[0432] The user inputs specific keywords and settings into the terminal, for example, the keyword "dinner conversation with friends" and the setting "casual atmosphere."
[0433] 2. Data Transmission
[0434] The terminal sends the input keywords and setting data to the server via a communication network.
[0435] 3. Data Receipt and Analysis
[0436] The server receives the data sent from the device, analyzes the received keywords and settings, and generates the appropriate prompts.
[0437] 4. Conversation generation using generative AI
[0438] The server uses a generative AI (e.g., GPT-4) to generate natural-sounding conversational sentences based on the generated prompts. The generative AI takes into account the input keywords and settings to create natural-sounding conversational sentences that fit the context.
[0439] 5. Sending conversation text
[0440] The server transmits the generated conversation sentence data to the terminal again via the communication network.
[0441] 6. User Visibility
[0442] The device displays the conversation data received from the server, and the user can use the displayed conversation data as a reference for actual conversations.
[0443] Specific examples
[0444] 1. User Input
[0445] The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0446] 2. Data Transmission
[0447] The terminal sends the entered data to the server. Communication is carried out using the HTTPS protocol.
[0448] 3. Data Receipt and Analysis
[0449] The server receives the data, analyzes the keywords and settings, and generates a prompt based on the analysis results, which it prepares to send to the AI.
[0450] 4. Conversation generation using generative AI
[0451] The server sends the following prompt to the AI: "Generate a casual conversation about a meal with a friend." Based on this instruction, the AI generates the following conversation:
[0452] A: What do you want to eat today? Pizza?
[0453] B: That's great! But I hear the pasta here is delicious too.
[0454] A: Right, then let's order both pasta and pizza and share.
[0455] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0456] 5. Sending conversation text
[0457] The server transmits the generated conversation sentence to the terminal via a communication network.
[0458] 6. User Visibility
[0459] The device receives the conversation data and displays it to the user, who can use it as a reference when dining with real friends.
[0460] This invention allows users to improve their conversational skills and obtain natural conversational sentences that can be used as reference in actual communication situations. This allows even users who are not good at communication to enjoy smooth conversations, which greatly contributes to reducing feelings of loneliness and building relationships.
[0461] The processing flow will be explained below.
[0462] Step 1:
[0463] The user inputs a specific keyword and setting into the terminal. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. After inputting, the user presses the data transmission button.
[0464] Step 2:
[0465] The device organizes the input keywords and settings data and sends it to the server. Specifically, it converts the keywords and settings into JSON format and sends it to the server's API endpoint via an HTTPS request.
[0466] Step 3:
[0467] The server receives the data sent from the device. Specifically, the server's API endpoint processes the HTTPS request and converts the received data into an internal format for analysis.
[0468] Step 4:
[0469] The server analyzes the received keywords and settings, and generates prompts for the AI based on the keywords and settings. These prompts reflect the context and mood of the conversation the user desires.
[0470] Step 5:
[0471] The server sends a prompt to the generating AI (e.g., GPT-4), which contains information based on user-specified keywords and preferences.
[0472] Step 6:
[0473] Generative AI generates natural-sounding conversations based on prompts, taking into account the context, flow of the dialogue, and word usage to create conversations that fit the specified conditions.
[0474] Step 7:
[0475] The server receives the conversation data generated by the AI, checks the quality of the data, and formats it before providing it to the user.
[0476] Step 8:
[0477] The server sends the generated conversation data to the device. Again, the data is sent in JSON format as an HTTPS response.
[0478] Step 9:
[0479] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[0480] Step 10:
[0481] The terminal displays the formatted conversation sentences to the user, who can refer to the generated natural conversation sentences and use them in actual conversations.
[0482] This allows users to improve their conversational skills and communicate more smoothly.
[0483] Example 1
[0484] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0485] Conventional conversation support systems have struggled to automatically generate natural-sounding conversational sentences for users to refer to. Furthermore, generating custom conversational sentences based on user-specified keywords and conversational tones requires advanced expertise, making them difficult for average users to use. Furthermore, there was a lack of a way to quickly generate natural-sounding conversational sentences and provide them to users, making real-time use difficult.
[0486] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0487] In this invention, the server includes means for a user to input specific keywords and settings to a terminal, means for transmitting the input keywords and settings to the server via a communication network, means for the server to analyze the received keywords and settings and generate prompts for a generative AI model, means for generating natural conversational sentences based on the prompts using the generative AI model, means for transmitting the generated conversational sentence data from the server again via the communication network to the terminal, and means for displaying the generated conversational sentence data received by the terminal to the user. This enables the user to easily generate and refer to natural conversational sentences in real time according to the set keywords and atmosphere.
[0488] "User" refers to an individual or group that uses the system and is the entity that inputs keywords and settings via a terminal.
[0489] "Terminal" refers to a device operated by a user, including devices such as smartphones, personal computers, and tablets.
[0490] "Keywords" refer to specific words or phrases that users input to generate conversation text, and they play a role in determining the theme and content of the generated conversation text.
[0491] "Settings" refers to additional conditions or attributes that a user specifies when generating a conversation, including details such as the tone and atmosphere of the conversation.
[0492] "Communications network" refers to an infrastructure that enables the transmission and reception of data, including the Internet and local area networks.
[0493] "Server" refers to a central processing unit that processes data received from the terminal and sends prompts to the generating AI.
[0494] A "prompt" refers to an instruction entered into a generative AI model, which serves as a guide for the generated conversational text.
[0495] "Generative AI models" refer to artificial intelligence algorithms that generate natural-sounding conversational sentences based on input prompts, including, for example, large-scale language models.
[0496] "Conversational text" refers to part or all of a sentence that is generated in a form that can be referenced by the user, and refers to text that has a natural dialogue format.
[0497] "Data" refers to keywords, settings, generated dialogue, and related information in general.
[0498] "JSON format" refers to a standard data representation format for structuring data in a way that is easy for humans to read and machines to analyze.
[0499] This invention is a system that allows a user to input specific keywords and settings and generates natural conversational sentences based on the input. The system is implemented mainly using the following hardware and software.
[0500] System configuration and operation
[0501] Hardware
[0502] 1. User device: A device such as a smartphone, personal computer, or tablet. These devices provide an interface for users to enter keywords and settings.
[0503] 2. Server: A high-performance server or cloud environment. The server is used to receive and analyze data and operate the generative AI model.
[0504] software
[0505] 1. Generative AI models: Large-scale language models such as GPT-4 from OpenAI, which generate highly natural-sounding conversational sentences.
[0506] 2. Communication protocol: The HTTPS protocol is used for communication between the terminal and the server to ensure data security and reliability.
[0507] 3. Data analysis tools: Use analysis tools such as Python's json module to receive and analyze data.
[0508] Data processing and calculation
[0509] The user inputs specific keywords, such as "dinner conversation with friends," and settings, such as "casual atmosphere," into the terminal. The specific data processing and calculations performed at each step are described in detail below.
[0510] 1. User Input: The user enters keywords and settings into the terminal application. The input is done in a text box, and when the submit button is pressed, the data is converted into JSON format.
[0511] 2. Data transmission: The device uses the HTTPS protocol to send JSON format data to the server. The data is encrypted before being sent.
[0512] 3. Receiving and parsing data: The server parses the received JSON data and extracts keywords and settings. This parsing is done using the Python json module.
[0513] 4. Prompt generation: Based on the extracted keywords and settings, the server generates a prompt to send to the generative AI model. An example of a prompt is "Generate a casual dinner conversation with a friend."
[0514] 5. Conversational Sentence Generation by Generative AI: The server sends prompts to the generative AI model, which then generates natural-sounding conversational sentences using hundreds of billions of parameters to provide context-appropriate answers.
[0515] 6. Sending the conversation: The generated conversation is converted back to JSON format and sent to the device via HTTPS.
[0516] 7. Display to user: The device parses the received JSON data and displays the generated conversation text to the user. The user can refer to this conversation text and use it in actual conversations.
[0517] Specific examples
[0518] The specific flow of system usage is shown below.
[0519] 1. User Input: The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0520] 2. Prompt generation: The server sends the generative AI model a prompt like this: "Generate a casual dinner conversation with a friend."
[0521] 3. Conversation generation results: The generative AI model generates the following conversation:
[0522] A: What do you want to eat today? Pizza?
[0523] B: That's great! But I hear the pasta here is delicious too.
[0524] A: Right, then let's order both pasta and pizza and share.
[0525] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0526] This system allows users to improve their conversational skills, easily generate natural conversational sentences in real time, and use them in real-life communications.
[0527] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0528] Step 1:
[0529] The user opens a smartphone application or web browser, inputs specific keywords and settings, such as "dinner conversation with friends" and "casual atmosphere" in the text boxes, and then presses the submit button. This input data is passed to the next step.
[0530] Step 2:
[0531] The device converts the keywords and settings entered by the user into JSON format data. For example, the following JSON data is generated:
[0532] json
[0533] {
[0534] "keyword": "dinner conversation with friends",
[0535] "setting": "casual atmosphere"
[0536] }
[0537] This JSON data is sent to the server using the HTTPS protocol, and the device confirms that the data is sent encrypted.
[0538] Step 3:
[0539] The server receives the JSON data sent from the device. It parses the data using Python's json module and extracts keywords and settings. The analysis results are passed to the next step. From the input JSON data, "dinner conversation with friends" and "casual atmosphere" are extracted.
[0540] Step 4:
[0541] Based on the extracted keywords and settings, the server generates a prompt to be sent to the generative AI model. Specifically, it generates the prompt "Generate a casual dinner conversation with a friend." This prompt is passed to the next step.
[0542] Step 5:
[0543] The server sends the prompt to a generative AI model (e.g., GPT-4) and requests it to generate a conversation. The generative AI model analyzes the prompt and generates natural-sounding conversational sentences that meet the required conditions. For example, the following conversational sentences are generated:
[0544] A: What do you want to eat today? Pizza?
[0545] B: That's great! But I hear the pasta here is delicious too.
[0546] A: Right, then let's order both pasta and pizza and share.
[0547] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0548] The generated conversation sentence is passed to the next step.
[0549] Step 6:
[0550] The server converts the generated conversation text back into JSON format and sends it to the device. At this time, the server checks the integrity of the generated data and performs error handling if necessary. The generated conversation text is converted into JSON data as follows:
[0551] json
[0552] {
[0553] "conversation": [
[0554] "A: What do you want to eat today? Pizza?"
[0555] "B: That's great! But I heard the pasta here is also delicious."
[0556] "A: Right, then let's order both pasta and pizza and share."
[0557] "B: That might be a good idea. By the way, have you traveled anywhere recently?"
[0558] ]
[0559] }
[0560] This JSON data is sent to the device.
[0561] Step 7:
[0562] The device parses the JSON data received from the server and displays it on the application's UI. The user can refer to the natural conversational text displayed on the screen and use it in actual conversations. The device displays the generated conversational text to the user through the user interface, allowing the user to perform additional operations such as scrolling and saving.
[0563] By following the steps above, users can easily and quickly generate natural conversational sentences based on specified conditions and use them in actual conversations.
[0564] (Application example 1)
[0565] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0566] Customer service in modern brick-and-mortar stores requires store clerks to respond quickly and appropriately to customer questions and requests. However, this depends on the experience and knowledge of the store clerk, which can lead to inconsistencies in individual responses. Furthermore, responding immediately in a busy environment can be difficult, which can lead to a decrease in customer satisfaction. A solution to this issue is needed, enabling store clerks to consistently provide high-quality service to customers.
[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0568] In this invention, the server includes means for receiving customer questions and requests by voice input, means for converting the voice input data into text data, and means for generating prompts based on the text data, thereby enabling store staff to provide quick and appropriate responses to customer questions and requests.
[0569] "Conversational data" is text information of natural conversations created by generative AI.
[0570] A "communications network" is an information pathway such as the Internet or a local area network that allows data to be sent and received.
[0571] A "terminal" is a user device such as a smartphone, smart glasses, or head-mounted display operated by a user or store clerk.
[0572] A "server" is a computer system that processes data and generates conversational sentences using AI.
[0573] "Keywords" are specific words or phrases entered by the user, and serve as basic information for generating conversational text.
[0574] "Settings" are parameters related to the tone and topic of the conversation specified by the user.
[0575] "Generative AI" is an artificial intelligence system that generates natural-sounding conversational sentences based on input keywords and settings.
[0576] "Voice input" is a method in which a user or store clerk speaks to a terminal and the voice data is acquired.
[0577] "Text data" refers to data obtained by converting information input by voice into a character string format.
[0578] A "prompt" is a text instruction that serves as the basis for the generation AI to generate conversational sentences.
[0579] This invention is a system including a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the terminal. Also, by adding a means for generating prompts based on keywords and settings input by a user and a means for converting voice-input data into text data, this is a specific embodiment for supporting communication between store clerks and customers.
[0580] Hardware and software used
[0581] Hardware
[0582] Devices: smartphones, smart glasses, head-mounted displays, etc.
[0583] Server: A high-performance cloud server, such as an AWS EC2 server.
[0584] software
[0585] Speech recognition engine: Google Speech-to-Text API
[0586] Natural language analysis engines: SpaCy, NLTK
[0587] Generative AI model: OpenAI GPT-4
[0588] Communication protocol: HTTPS
[0589] What the system does
[0590] First, the store clerk (the user) uses the voice input function of the terminal to input the customer's question or request by voice. For example, a customer might ask, "Do you have this shirt in my size?" This voice input is converted into text data on the terminal and sent to the server using the HTTPS protocol.
[0591] The server analyzes the received text data using a natural language analysis engine (SpaCy or NLTK) and generates an appropriate prompt. The generated prompt is passed to a generative AI model (OpenAI GPT-4), which then generates natural-sounding conversational sentences based on the prompt. For example, in response to the prompt, "You have requested stock information about the size of this shirt. Please generate an appropriate response," the generative AI generates the following conversational sentences:
[0592] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0593] The generated conversation sentences are sent from the server to the terminal again and displayed in the field of view of the store clerk. The store clerk can refer to the displayed conversation sentences and provide an appropriate response to the customer.
[0594] Specific examples
[0595] 1. Example prompt:
[0596] When a customer asks, "Do you have this shirt in my size?", generate an appropriate response.
[0597] 2. Example of generated dialogue:
[0598] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0599] This invention enables store staff to provide quick and appropriate responses to customer questions and requests, which is expected to improve the quality of customer service in physical stores.
[0600] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0601] Step 1:
[0602] A user uses the smart glasses to input a customer question or request by voice. For example, a question such as "Do you have this shirt in my size?" The voice data is the input data.
[0603] Step 2:
[0604] The device uses a speech recognition engine (Google Speech-to-Text API) to convert the voice input data into text data. The voice data is converted into text data.
[0605] Step 3:
[0606] The terminal sends the converted text data to the server via a communication protocol (HTTPS). The text data sent is the input data, and the text data received by the server is the output data.
[0607] Step 4:
[0608] The server receives the text data and analyzes it using a natural language analysis engine (SpaCy, NLTK). The text data is the input data, and the analysis results are the output data.
[0609] Step 5:
[0610] Based on the analysis results, the server generates prompts to send to the generative AI model (GPT-4). The analysis results are the input data, and the generated prompts are the output data.
[0611] Step 6:
[0612] The server sends prompts to the generative AI model (GPT-4) to generate natural-sounding conversational sentences. The generated prompts are the input data, and the generated conversational sentences are the output data.
[0613] Step 7:
[0614] The server sends the generated conversation text to the terminal via HTTPS. The generated conversation text is the input data, and the conversation text received by the terminal is the output data.
[0615] Step 8:
[0616] The conversational text received by the terminal is displayed on the smart glasses. The received conversational text is input data, and the conversational text displayed on the smart glasses is output data.
[0617] Step 9:
[0618] The user refers to the displayed conversation sentence and provides an appropriate response to the customer. The displayed conversation sentence is input data, and the user's response is output data.
[0619] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0620] The present invention provides a system including a means for displaying generated conversational text data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational text using a generation AI, a means for transmitting the generated conversational text data to the user terminal, and an emotion engine for recognizing the user's emotions.
[0621] Program processing flow
[0622] 1. User Input
[0623] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. Furthermore, the user's facial expressions and tone of voice are monitored by the device's emotion engine.
[0624] 2. Data Transmission
[0625] The device sends the input keywords and settings, as well as the user's emotional data analyzed by the emotion engine, to the server via a communication network.
[0626] 3. Data Receipt and Analysis
[0627] The server receives the data sent from the device, analyzes the received keywords, settings, and emotional data, and generates prompts appropriate to the user's emotions.
[0628] 4. Conversation generation using generative AI
[0629] The server sends emotion-based prompts to the AI, which contain information corresponding to the user's emotional state. The AI then uses these prompts to generate natural-sounding conversational sentences.
[0630] 5. Emotional Adaptation Check
[0631] The server checks whether the generated dialogue is appropriate to the user's emotions. The emotion engine performs emotion analysis again, and if necessary, modifies the prompt and generates it again.
[0632] 6. Sending conversation text
[0633] The server transmits the generated conversation data to the terminal, and the data is transmitted again via the communication network.
[0634] 7. User Visibility
[0635] The device receives the generated conversation data and displays it to the user, who can then use it in actual conversations.
[0636] Specific examples
[0637] 1. User Input
[0638] When a user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere," the emotion engine detects that the user is slightly nervous.
[0639] 2. Data Transmission
[0640] The device sends data, including the entered keywords, settings, and tension generated by the emotion engine, to the server using the HTTPS protocol.
[0641] 3. Data Receipt and Analysis
[0642] The server receives the data and analyzes it for keywords, settings, and tension. As a result of the analysis, it generates prompts that help the user relax.
[0643] 4. Conversation generation using generative AI
[0644] The server sends the following prompt to the AI: "Generate a conversation scenario for a user who is nervous about having a casual meal with a friend." Based on this instruction, the AI generates the following dialogue:
[0645] A: Where do you want to eat today?
[0646] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[0647] A: Great, let's do that! How have you been lately?
[0648] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[0649] 5. Emotional Adaptation Check
[0650] The server checks whether the generated dialogue relieves the user's tension. The emotion engine analyzes the user's emotions again, and if necessary, modifies and regenerates the prompt.
[0651] 6. Sending conversation text
[0652] The server sends the confirmed conversation data to the terminal as an HTTPS response.
[0653] 7. User Visibility
[0654] The terminal receives the conversation sentence data and displays it to the user, who can then refer to the generated natural conversation sentences and have a relaxed conversation.
[0655] The present invention allows users to obtain natural conversational sentences that are adapted to their emotional state, enabling smooth communication. This allows even users who are not good at communication to enjoy conversations with ease, and also contributes to building interpersonal relationships.
[0656] The processing flow will be explained below.
[0657] Step 1:
[0658] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. At this time, the device analyzes the user's facial expression data and tone of voice using an emotion engine to extract emotional data.
[0659] Step 2:
[0660] The device organizes the input keywords and settings, as well as the user's emotional data extracted by the emotion engine, and sends them to the server. Specifically, it converts the keywords, settings, and emotional data into JSON format and sends it as an HTTPS request to the server's API endpoint.
[0661] Step 3:
[0662] The server receives data sent from the device, which is then processed by the API endpoint and converted into an internal format for analysis.
[0663] Step 4:
[0664] The server analyzes the received keywords, preferences, and emotional data. Based on the analysis, it generates prompts that take the user's emotions into account. For example, if the user is nervous, the prompt may include instructions such as "generate relaxing conversation content."
[0665] Step 5:
[0666] The server sends a generated prompt to the generation AI (e.g., GPT-4), which includes the user's emotional state and the specified keywords and preferences.
[0667] Step 6:
[0668] The AI generates natural-sounding dialogue based on prompts, taking into account context, emotions, and the flow of the conversation, and tailors the dialogue to match the user's emotions.
[0669] Step 7:
[0670] The server receives the conversation data generated by the generation AI, checks the quality of the received conversation data, and again determines whether it is appropriate for the user's emotional state.
[0671] Step 8:
[0672] The server uses the emotion engine to check whether the generated dialogue is appropriate to the user's emotions. If it is not appropriate, it modifies the prompt and sends it to the generation AI again.
[0673] Step 9:
[0674] The server sends the confirmed conversation data to the terminal, which is then sent in JSON format as an HTTPS response.
[0675] Step 10:
[0676] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[0677] Step 11:
[0678] The device displays the formatted conversational text to the user, who can then refer to the natural-sounding conversational text and use it in actual conversations. The emotion engine ensures that the conversational text takes into account the user's current emotional state, allowing the user to continue the conversation with peace of mind.
[0679] This allows the user to obtain natural conversational sentences that are adapted to his or her own emotional state, making actual communication smoother.
[0680] Example 2
[0681] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0682] Conventional systems have the problem that when a user inputs specific keywords or settings to generate conversational sentences, they are unable to generate conversational sentences that correspond to the user's emotional state. As a result, it is difficult to obtain conversational content that matches the user's emotional state, such as when they are tense or relaxed, making it difficult to build natural communication.
[0683] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0684] In this invention, the server includes means for a user to input keywords and settings into a terminal and collect emotional data, means for transmitting the keywords, settings, and emotional data from the user terminal to the server via a communication network, means for analyzing the received data and generating prompts based on the user's emotional state, means for generating natural conversational sentences from the prompts using a generative AI model, means for checking whether the generated conversational sentence data is appropriate for the user's emotions and for regenerating the data as necessary, and means for transmitting the generated conversational sentence data to the user terminal. This makes it possible to generate natural conversational sentences that are adapted to the user's emotional state and support smooth communication.
[0685] A "user" is a person who operates a device to input keywords and settings and obtain conversational text based on their emotional state.
[0686] A "terminal" is a device operated by a user, which collects input keywords, settings, and emotional data and transmits them to a server.
[0687] "Emotional data" is data analyzed by the emotion engine based on the user's facial expressions, tone of voice, etc.
[0688] A "communication network" is a network infrastructure for data communication between a user terminal and a server.
[0689] A "server" is a computer system that receives data sent from a user terminal, analyzes it, and generates conversational sentences.
[0690] A "generative AI model" is an artificial intelligence that uses pre-trained algorithms to generate natural-sounding conversational sentences from generated prompts.
[0691] A "prompt" is an instruction or question-type sentence that tells the generative AI model to generate natural conversational sentences.
[0692] "Conversational text" is a natural conversational text generated by a generative AI model and provided to the user.
[0693] The present invention is a system that generates natural conversational sentences that correspond to the user's emotional state using a user terminal, a server, a generative AI model, and an emotion engine.
[0694] The system includes a user terminal, a communication network, a server, a generative AI model, and an emotion engine. The system operates as follows.
[0695] Hardware and software used
[0696] User devices include smartphones, tablets, and other devices. These devices include an interface for entering keywords and settings using a keyboard or touchscreen. They also include an emotion engine for analyzing facial expressions and tone of voice. The emotion engine may use, for example, facial recognition software or voice analysis software.
[0697] The communication network is the infrastructure that connects the user terminal and the server. Specifically, the Internet or mobile data communication is often used. Data is transmitted and received securely using the HTTPS protocol.
[0698] The server is a computer system that receives data sent from the user's device, analyzes it, and generates conversational sentences. The server is equipped with a generative AI model, such as GPT-4, that has the ability to generate advanced natural language.
[0699] Data processing and calculation
[0700] The user enters keywords and settings into the device to collect emotional data. This data is then sent to a server via a communications network. The server analyzes the received data and generates prompts based on the user's emotional state. These prompts are then input into a generative AI model to generate natural-sounding conversational sentences.
[0701] The generated conversation sentences are checked by the server to see if they are appropriate for the user's emotions, and are regenerated if necessary. Finally, the generated conversation sentence data is sent to the user's terminal and displayed to the user.
[0702] Specific examples
[0703] A user opens a smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere." The emotion engine analyzes the user's facial expressions and tone of voice and detects that the user is "slightly nervous." This data is sent to the server, which then analyzes the keywords, settings, and emotion data.
[0704] The analysis results in prompts that help users relax. Based on these prompts, the generative AI model generates natural-sounding conversational sentences like the following:
[0705] A: Where do you want to eat today?
[0706] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[0707] A: Great, let's do that! How have you been lately?
[0708] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[0709] The generated conversation sentences are then checked again by the emotion engine, and after corrections are made as necessary, they are sent to the user's device. The user can refer to these conversation sentences and proceed with the actual conversation in a relaxed manner.
[0710] As described above, the present invention generates natural conversational sentences that are adapted to the emotional state of the user, thereby supporting smooth communication.
[0711] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0712] Step 1:
[0713] The user inputs keywords and settings into the device to collect emotional data. Specifically, the user inputs keywords and settings using the device's keyboard or touch screen. At this time, the device's camera and microphone analyze the user's facial expressions and tone of voice using an emotion engine to obtain emotional data.
[0714] Input: Keywords, settings, user facial expression and tone of voice
[0715] Output: Emotion data, keywords and settings
[0716] Step 2:
[0717] The device then transmits the collected keywords, settings, and emotion data to a server, where the data is transmitted securely over a communications network using the HTTPS protocol.
[0718] Input: Keywords, settings, emotion data
[0719] Output: Data sent to the server
[0720] Step 3:
[0721] The server receives the data sent from the device and analyzes keywords, preferences, and emotional data. Specifically, it uses an analysis algorithm on the received data to generate prompts based on the user's emotional state.
[0722] Input: Keywords, settings, emotion data
[0723] Output: Generated prompt
[0724] Step 4:
[0725] The server inputs the generated prompts to the generative AI model to generate natural conversational sentences. In this process, the generative AI model generates dialogue-style sentences based on the prompts.
[0726] Input: Generated prompt
[0727] Output: Natural conversational sentences
[0728] Step 5:
[0729] The server checks whether the generated dialogue is appropriate for the user's emotional state, again using the emotion engine to analyze the user's emotional changes in response to the dialogue, modifying the prompts as necessary, and re-inputting them into the generative AI model.
[0730] Input: Generated conversation sentences, user reaction data
[0731] Output: Modified prompt (if needed)
[0732] Step 6:
[0733] The server then sends the final generated conversation data to the device, again securely using the HTTPS protocol.
[0734] Input: Final generated dialogue
[0735] Output: Conversation data sent to the device
[0736] Step 7:
[0737] The device displays the received conversation data to the user, who can then refer to the displayed conversation data and use it in actual conversations. Specifically, the generated conversation data is displayed on the device screen.
[0738] Input: Received conversation data
[0739] Output: The dialogue displayed to the user
[0740] (Application example 2)
[0741] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0742] Conventional conversation generation systems have had difficulty generating natural conversations that reflect the user's emotions and situation. In particular, when ordering food delivery, appropriate conversations are required that reflect the user's stress or relaxation state, but these systems have not been able to meet this need. This often results in an unsmooth ordering process and a poor user experience.
[0743] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for displaying the generated conversation sentence data, means for transmitting data including keywords and settings from the user terminal to the server via a communication network, means for analyzing the received data and generating natural conversation sentences using a generation AI, means for transmitting the generated conversation sentence data to the user terminal, and means for generating natural ordering conversation sentences based on the user's emotion data in the case of ordering food delivery. This makes it possible to generate natural conversation sentences that correspond to the user's emotions and situation.
[0744] The "means for displaying the generated conversation sentence data" is a function or device for displaying the generated conversation sentence data on the screen of a user terminal or the like.
[0745] The "means for transmitting data including keywords and settings from a user terminal to a server via a communications network" refers to a communications means and protocol for transmitting keywords and settings entered by a user to a server.
[0746] The "means for analyzing received data and generating natural conversational sentences using a generation AI" refers to an algorithm and processing device for analyzing received data and generating natural conversational sentences using the analysis results and a generation AI.
[0747] The "means for transmitting the generated conversation sentence data to the user terminal" refers to a communication means and protocol for transmitting the generated conversation sentence data from the server to the user terminal.
[0748] The "means for generating natural ordering conversations based on the user's emotional data when ordering food delivery" refers to an algorithm and processing device that takes into account the user's emotional data when ordering food delivery and generates natural ordering conversations based on that data.
[0749] This invention relates to a system that facilitates the food delivery ordering process by generating natural ordering conversations based on the user's emotions. Specific embodiments for implementing this system are described below.
[0750] System Overview
[0751] This system generates natural conversational sentences and provides them to users when they order food delivery using their smartphones. The system consists of the following main components:
[0752] Hardware / Software Components
[0753] 1. Smartphone
[0754] Emotion engine: An engine that recognizes the user's emotions, for example by using a camera and microphone to analyze the user's facial expressions and tone of voice.
[0755] Communication module: A module to support HTTPS communication.
[0756] 2. Server
[0757] Data analysis engine: Analyzes keywords, settings, and sentiment data entered by users.
[0758] Generative AI models (e.g., GPT-3): Generate prompts based on user sentiment and generate natural-sounding conversations.
[0759] Communication module: A module for receiving data from the user terminal and transmitting the generated conversation data to the user terminal.
[0760] Data processing and calculation
[0761] User Input
[0762] When a user opens the smartphone application and enters keywords related to the order (e.g., "Italian restaurant") and a preference (e.g., "relaxed atmosphere"), the emotion engine collects the user's emotional data. For example, if the user is relaxed, that data is collected.
[0763] Data transmission and analysis
[0764] The smartphone's communication module sends keywords, settings, and emotional data to the server using the HTTPS protocol, where the server's data analysis engine analyzes the received data and generates prompts for the generative AI model.
[0765] Conversation generation using generative AI
[0766] The generative AI model generates natural-sounding sentences based on prompts such as:
[0767] "Generate an Italian restaurant ordering scenario for a relaxed user."
[0768] Here are some examples of dialogue generated by the generative AI:
[0769] A: Hello, I'd like some suggestions for Italian restaurants.
[0770] B: Of course! I recommend the Margherita pizza. It's a very relaxing dish.
[0771] Sending and displaying conversation text
[0772] The server then sends the generated conversational text to the smartphone, which receives and displays it on the user's device. The user can refer to the natural conversational text displayed and place their order smoothly.
[0773] This will enable the system to provide natural-sounding conversations that reflect the user's emotions and situation, which is expected to make the food delivery ordering process go more smoothly.
[0774] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0775] Step 1:
[0776] A user launches the smartphone application and enters keywords related to their order (e.g., "Italian restaurant") and preferences (e.g., "relaxed atmosphere"). At this time, the smartphone's emotion engine analyzes the user's facial expressions and tone of voice in real time to collect emotional data (whether they are relaxed or not).
[0777] Input data: Keywords, settings, and sentiment data
[0778] Output data: User input dataset
[0779] Step 2:
[0780] The terminal's communication module compiles the collected keywords, preferences, and emotion data and transmits them to a server using the HTTPS protocol.
[0781] Input data: User input dataset
[0782] Output data: JSON formatted data packet to the server
[0783] Step 3:
[0784] The server receives data packets from the user terminal, and the analysis engine analyzes them. The analysis engine extracts keywords, preferences, and emotion data from the data packets, and then creates prompts for the generative AI model.
[0785] Input data: JSON format data packet to the server
[0786] Output data: prompt statement
[0787] Step 4:
[0788] The server's generative AI model generates natural-sounding conversational sentences based on the generated prompt, for example, "Generate an ordering scenario at an Italian restaurant for a relaxed user."
[0789] Input data: Prompt statement
[0790] Output data: Generated conversation
[0791] Step 5:
[0792] The server then rechecks the generated dialogue to verify its suitability for the user's emotional data, and if necessary performs additional analysis to refine it and create the complete dialogue.
[0793] Input data: Generated conversation
[0794] Output: Verified and possibly corrected dialogue
[0795] Step 6:
[0796] The server then sends the final generated conversation text to the user's device using the HTTPS protocol, allowing the user to obtain the conversation text to smoothly complete the ordering process.
[0797] Input data: Verified and corrected conversation text
[0798] Output data: Conversation data packets to the user terminal
[0799] Step 7:
[0800] The terminal receives the conversation data sent from the server and displays it on the smartphone screen. The user can then place an order based on the natural conversation displayed.
[0801] Input data: Conversation data packets to the user terminal
[0802] Output data: Conversation displayed on smartphone screen
[0803] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0804] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0805] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0806] [Third embodiment]
[0807] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0808] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0809] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0810] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0811] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0812] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0813] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0814] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0815] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0816] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0817] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0818] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0819] The present invention comprises a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the user terminal. This allows the user to automatically generate natural conversational sentences based on specified keywords and settings and use them as a reference for actual conversations.
[0820] Program processing flow
[0821] 1. User Input
[0822] The user inputs specific keywords and settings into the terminal, for example, the keyword "dinner conversation with friends" and the setting "casual atmosphere."
[0823] 2. Data Transmission
[0824] The terminal sends the input keywords and setting data to the server via a communication network.
[0825] 3. Data Receipt and Analysis
[0826] The server receives the data sent from the device, analyzes the received keywords and settings, and generates the appropriate prompts.
[0827] 4. Conversation generation using generative AI
[0828] The server uses a generative AI (e.g., GPT-4) to generate natural-sounding conversational sentences based on the generated prompts. The generative AI takes into account the input keywords and settings to create natural-sounding conversational sentences that fit the context.
[0829] 5. Sending conversation text
[0830] The server transmits the generated conversation sentence data to the terminal again via the communication network.
[0831] 6. User Visibility
[0832] The device displays the conversation data received from the server, and the user can use the displayed conversation data as a reference for actual conversations.
[0833] Specific examples
[0834] 1. User Input
[0835] The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0836] 2. Data Transmission
[0837] The terminal sends the entered data to the server. Communication is carried out using the HTTPS protocol.
[0838] 3. Data Receipt and Analysis
[0839] The server receives the data, analyzes the keywords and settings, and generates a prompt based on the analysis results, which it prepares to send to the AI.
[0840] 4. Conversation generation using generative AI
[0841] The server sends the following prompt to the AI: "Generate a casual conversation about a meal with a friend." Based on this instruction, the AI generates the following conversation:
[0842] A: What do you want to eat today? Pizza?
[0843] B: That's great! But I hear the pasta here is delicious too.
[0844] A: Right, then let's order both pasta and pizza and share.
[0845] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0846] 5. Sending conversation text
[0847] The server transmits the generated conversation sentence to the terminal via a communication network.
[0848] 6. User Visibility
[0849] The device receives the conversation data and displays it to the user, who can use it as a reference when dining with real friends.
[0850] This invention allows users to improve their conversational skills and obtain natural conversational sentences that can be used as reference in actual communication situations. This allows even users who are not good at communication to enjoy smooth conversations, which greatly contributes to reducing feelings of loneliness and building relationships.
[0851] The processing flow will be explained below.
[0852] Step 1:
[0853] The user inputs a specific keyword and setting into the terminal. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. After inputting, the user presses the data transmission button.
[0854] Step 2:
[0855] The device organizes the input keywords and settings data and sends it to the server. Specifically, it converts the keywords and settings into JSON format and sends it to the server's API endpoint via an HTTPS request.
[0856] Step 3:
[0857] The server receives the data sent from the device. Specifically, the server's API endpoint processes the HTTPS request and converts the received data into an internal format for analysis.
[0858] Step 4:
[0859] The server analyzes the received keywords and settings, and generates prompts for the AI based on the keywords and settings. These prompts reflect the context and mood of the conversation the user desires.
[0860] Step 5:
[0861] The server sends a prompt to the generating AI (e.g., GPT-4), which contains information based on user-specified keywords and preferences.
[0862] Step 6:
[0863] Generative AI generates natural-sounding conversations based on prompts, taking into account the context, flow of the dialogue, and word usage to create conversations that fit the specified conditions.
[0864] Step 7:
[0865] The server receives the conversation data generated by the AI, checks the quality of the data, and formats it before providing it to the user.
[0866] Step 8:
[0867] The server sends the generated conversation data to the device. Again, the data is sent in JSON format as an HTTPS response.
[0868] Step 9:
[0869] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[0870] Step 10:
[0871] The terminal displays the formatted conversation sentences to the user, who can refer to the generated natural conversation sentences and use them in actual conversations.
[0872] This allows users to improve their conversational skills and communicate more smoothly.
[0873] Example 1
[0874] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0875] Conventional conversation support systems have struggled to automatically generate natural-sounding conversational sentences for users to refer to. Furthermore, generating custom conversational sentences based on user-specified keywords and conversational tones requires advanced expertise, making them difficult for average users to use. Furthermore, there was a lack of a way to quickly generate natural-sounding conversational sentences and provide them to users, making real-time use difficult.
[0876] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0877] In this invention, the server includes means for a user to input specific keywords and settings to a terminal, means for transmitting the input keywords and settings to the server via a communication network, means for the server to analyze the received keywords and settings and generate prompts for a generative AI model, means for generating natural conversational sentences based on the prompts using the generative AI model, means for transmitting the generated conversational sentence data from the server again via the communication network to the terminal, and means for displaying the generated conversational sentence data received by the terminal to the user. This enables the user to easily generate and refer to natural conversational sentences in real time according to the set keywords and atmosphere.
[0878] "User" refers to an individual or group that uses the system and is the entity that inputs keywords and settings via a terminal.
[0879] "Terminal" refers to a device operated by a user, including devices such as smartphones, personal computers, and tablets.
[0880] "Keywords" refer to specific words or phrases that users input to generate conversation text, and they play a role in determining the theme and content of the generated conversation text.
[0881] "Settings" refers to additional conditions or attributes that a user specifies when generating a conversation, including details such as the tone and atmosphere of the conversation.
[0882] "Communications network" refers to an infrastructure that enables the transmission and reception of data, including the Internet and local area networks.
[0883] "Server" refers to a central processing unit that processes data received from the terminal and sends prompts to the generating AI.
[0884] A "prompt" refers to an instruction entered into a generative AI model, which serves as a guide for the generated conversational text.
[0885] "Generative AI models" refer to artificial intelligence algorithms that generate natural-sounding conversational sentences based on input prompts, including, for example, large-scale language models.
[0886] "Conversational text" refers to part or all of a sentence that is generated in a form that can be referenced by the user, and refers to text that has a natural dialogue format.
[0887] "Data" refers to keywords, settings, generated dialogue, and related information in general.
[0888] "JSON format" refers to a standard data representation format for structuring data in a way that is easy for humans to read and machines to analyze.
[0889] This invention is a system that allows a user to input specific keywords and settings and generates natural conversational sentences based on the input. The system is implemented mainly using the following hardware and software.
[0890] System configuration and operation
[0891] Hardware
[0892] 1. User device: A device such as a smartphone, personal computer, or tablet. These devices provide an interface for users to enter keywords and settings.
[0893] 2. Server: A high-performance server or cloud environment. The server is used to receive and analyze data and operate the generative AI model.
[0894] software
[0895] 1. Generative AI models: Large-scale language models such as GPT-4 from OpenAI, which generate highly natural-sounding conversational sentences.
[0896] 2. Communication protocol: The HTTPS protocol is used for communication between the terminal and the server to ensure data security and reliability.
[0897] 3. Data analysis tools: Use analysis tools such as Python's json module to receive and analyze data.
[0898] Data processing and calculation
[0899] The user inputs specific keywords, such as "dinner conversation with friends," and settings, such as "casual atmosphere," into the terminal. The specific data processing and calculations performed at each step are described in detail below.
[0900] 1. User Input: The user enters keywords and settings into the terminal application. The input is done in a text box, and when the submit button is pressed, the data is converted into JSON format.
[0901] 2. Data transmission: The device uses the HTTPS protocol to send JSON format data to the server. The data is encrypted before being sent.
[0902] 3. Receiving and parsing data: The server parses the received JSON data and extracts keywords and settings. This parsing is done using the Python json module.
[0903] 4. Prompt generation: Based on the extracted keywords and settings, the server generates a prompt to send to the generative AI model. An example of a prompt is "Generate a casual dinner conversation with a friend."
[0904] 5. Conversational Sentence Generation by Generative AI: The server sends prompts to the generative AI model, which then generates natural-sounding conversational sentences using hundreds of billions of parameters to provide context-appropriate answers.
[0905] 6. Sending the conversation: The generated conversation is converted back to JSON format and sent to the device via HTTPS.
[0906] 7. Display to user: The device parses the received JSON data and displays the generated conversation text to the user. The user can refer to this conversation text and use it in actual conversations.
[0907] Specific examples
[0908] The specific flow of system usage is shown below.
[0909] 1. User Input: The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[0910] 2. Prompt generation: The server sends the generative AI model a prompt like this: "Generate a casual dinner conversation with a friend."
[0911] 3. Conversation generation results: The generative AI model generates the following conversation:
[0912] A: What do you want to eat today? Pizza?
[0913] B: That's great! But I hear the pasta here is delicious too.
[0914] A: Right, then let's order both pasta and pizza and share.
[0915] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0916] This system allows users to improve their conversational skills, easily generate natural conversational sentences in real time, and use them in real-life communications.
[0917] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0918] Step 1:
[0919] The user opens a smartphone application or web browser, inputs specific keywords and settings, such as "dinner conversation with friends" and "casual atmosphere" in the text boxes, and then presses the submit button. This input data is passed to the next step.
[0920] Step 2:
[0921] The device converts the keywords and settings entered by the user into JSON format data. For example, the following JSON data is generated:
[0922] json
[0923] {
[0924] "keyword": "dinner conversation with friends",
[0925] "setting": "casual atmosphere"
[0926] }
[0927] This JSON data is sent to the server using the HTTPS protocol, and the device confirms that the data is sent encrypted.
[0928] Step 3:
[0929] The server receives the JSON data sent from the device. It parses the data using Python's json module and extracts keywords and settings. The analysis results are passed to the next step. From the input JSON data, "dinner conversation with friends" and "casual atmosphere" are extracted.
[0930] Step 4:
[0931] Based on the extracted keywords and settings, the server generates a prompt to be sent to the generative AI model. Specifically, it generates the prompt "Generate a casual dinner conversation with a friend." This prompt is passed to the next step.
[0932] Step 5:
[0933] The server sends the prompt to a generative AI model (e.g., GPT-4) and requests it to generate a conversation. The generative AI model analyzes the prompt and generates natural-sounding conversational sentences that meet the required conditions. For example, the following conversational sentences are generated:
[0934] A: What do you want to eat today? Pizza?
[0935] B: That's great! But I hear the pasta here is delicious too.
[0936] A: Right, then let's order both pasta and pizza and share.
[0937] B: That might be a good idea. By the way, have you traveled anywhere recently?
[0938] The generated conversation sentence is passed to the next step.
[0939] Step 6:
[0940] The server converts the generated conversation text back into JSON format and sends it to the device. At this time, the server checks the integrity of the generated data and performs error handling if necessary. The generated conversation text is converted into JSON data as follows:
[0941] json
[0942] {
[0943] "conversation": [
[0944] "A: What do you want to eat today? Pizza?"
[0945] "B: That's great! But I heard the pasta here is also delicious."
[0946] "A: Right, then let's order both pasta and pizza and share."
[0947] "B: That might be a good idea. By the way, have you traveled anywhere recently?"
[0948] ]
[0949] }
[0950] This JSON data is sent to the device.
[0951] Step 7:
[0952] The device parses the JSON data received from the server and displays it on the application's UI. The user can refer to the natural conversational text displayed on the screen and use it in actual conversations. The device displays the generated conversational text to the user through the user interface, allowing the user to perform additional operations such as scrolling and saving.
[0953] By following the steps above, users can easily and quickly generate natural conversational sentences based on specified conditions and use them in actual conversations.
[0954] (Application example 1)
[0955] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0956] Customer service in modern brick-and-mortar stores requires store clerks to respond quickly and appropriately to customer questions and requests. However, this depends on the experience and knowledge of the store clerk, which can lead to inconsistencies in individual responses. Furthermore, responding immediately in a busy environment can be difficult, which can lead to a decrease in customer satisfaction. A solution to this issue is needed, enabling store clerks to consistently provide high-quality service to customers.
[0957] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0958] In this invention, the server includes means for receiving customer questions and requests by voice input, means for converting the voice input data into text data, and means for generating prompts based on the text data, thereby enabling store staff to provide quick and appropriate responses to customer questions and requests.
[0959] "Conversational data" is text information of natural conversations created by generative AI.
[0960] A "communications network" is an information pathway such as the Internet or a local area network that allows data to be sent and received.
[0961] A "terminal" is a user device such as a smartphone, smart glasses, or head-mounted display operated by a user or store clerk.
[0962] A "server" is a computer system that processes data and generates conversational sentences using AI.
[0963] "Keywords" are specific words or phrases entered by the user, and serve as basic information for generating conversational text.
[0964] "Settings" are parameters related to the tone and topic of the conversation specified by the user.
[0965] "Generative AI" is an artificial intelligence system that generates natural-sounding conversational sentences based on input keywords and settings.
[0966] "Voice input" is a method in which a user or store clerk speaks to a terminal and the voice data is acquired.
[0967] "Text data" refers to data obtained by converting information input by voice into a character string format.
[0968] A "prompt" is a text instruction that serves as the basis for the generation AI to generate conversational sentences.
[0969] This invention is a system including a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the terminal. Also, by adding a means for generating prompts based on keywords and settings input by a user and a means for converting voice-input data into text data, this is a specific embodiment for supporting communication between store clerks and customers.
[0970] Hardware and software used
[0971] Hardware
[0972] Devices: smartphones, smart glasses, head-mounted displays, etc.
[0973] Server: A high-performance cloud server, such as an AWS EC2 server.
[0974] software
[0975] Speech recognition engine: Google Speech-to-Text API
[0976] Natural language analysis engines: SpaCy, NLTK
[0977] Generative AI model: OpenAI GPT-4
[0978] Communication protocol: HTTPS
[0979] What the system does
[0980] First, the store clerk (the user) uses the voice input function of the terminal to input the customer's question or request by voice. For example, a customer might ask, "Do you have this shirt in my size?" This voice input is converted into text data on the terminal and sent to the server using the HTTPS protocol.
[0981] The server analyzes the received text data using a natural language analysis engine (SpaCy or NLTK) and generates an appropriate prompt. The generated prompt is passed to a generative AI model (OpenAI GPT-4), which then generates natural-sounding conversational sentences based on the prompt. For example, in response to the prompt, "You have requested stock information about the size of this shirt. Please generate an appropriate response," the generative AI generates the following conversational sentences:
[0982] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0983] The generated conversation sentences are sent from the server to the terminal again and displayed in the field of view of the store clerk. The store clerk can refer to the displayed conversation sentences and provide an appropriate response to the customer.
[0984] Specific examples
[0985] 1. Example prompt:
[0986] When a customer asks, "Do you have this shirt in my size?", generate an appropriate response.
[0987] 2. Example of generated dialogue:
[0988] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[0989] This invention enables store staff to provide quick and appropriate responses to customer questions and requests, which is expected to improve the quality of customer service in physical stores.
[0990] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0991] Step 1:
[0992] A user uses the smart glasses to input a customer question or request by voice. For example, a question such as "Do you have this shirt in my size?" The voice data is the input data.
[0993] Step 2:
[0994] The device uses a speech recognition engine (Google Speech-to-Text API) to convert the voice input data into text data. The voice data is converted into text data.
[0995] Step 3:
[0996] The terminal sends the converted text data to the server via a communication protocol (HTTPS). The text data sent is the input data, and the text data received by the server is the output data.
[0997] Step 4:
[0998] The server receives the text data and analyzes it using a natural language analysis engine (SpaCy, NLTK). The text data is the input data, and the analysis results are the output data.
[0999] Step 5:
[1000] Based on the analysis results, the server generates prompts to send to the generative AI model (GPT-4). The analysis results are the input data, and the generated prompts are the output data.
[1001] Step 6:
[1002] The server sends prompts to the generative AI model (GPT-4) to generate natural-sounding conversational sentences. The generated prompts are the input data, and the generated conversational sentences are the output data.
[1003] Step 7:
[1004] The server sends the generated conversation text to the terminal via HTTPS. The generated conversation text is the input data, and the conversation text received by the terminal is the output data.
[1005] Step 8:
[1006] The conversational text received by the terminal is displayed on the smart glasses. The received conversational text is input data, and the conversational text displayed on the smart glasses is output data.
[1007] Step 9:
[1008] The user refers to the displayed conversation sentence and provides an appropriate response to the customer. The displayed conversation sentence is input data, and the user's response is output data.
[1009] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1010] The present invention provides a system including a means for displaying generated conversational text data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational text using a generation AI, a means for transmitting the generated conversational text data to the user terminal, and an emotion engine for recognizing the user's emotions.
[1011] Program processing flow
[1012] 1. User Input
[1013] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. Furthermore, the user's facial expressions and tone of voice are monitored by the device's emotion engine.
[1014] 2. Data Transmission
[1015] The device sends the input keywords and settings, as well as the user's emotional data analyzed by the emotion engine, to the server via a communication network.
[1016] 3. Data Receipt and Analysis
[1017] The server receives the data sent from the device, analyzes the received keywords, settings, and emotional data, and generates prompts appropriate to the user's emotions.
[1018] 4. Conversation generation using generative AI
[1019] The server sends emotion-based prompts to the AI, which contain information corresponding to the user's emotional state. The AI then uses these prompts to generate natural-sounding conversational sentences.
[1020] 5. Emotional Adaptation Check
[1021] The server checks whether the generated dialogue is appropriate to the user's emotions. The emotion engine performs emotion analysis again, and if necessary, modifies the prompt and generates it again.
[1022] 6. Sending conversation text
[1023] The server transmits the generated conversation data to the terminal, and the data is transmitted again via the communication network.
[1024] 7. User Visibility
[1025] The device receives the generated conversation data and displays it to the user, who can then use it in actual conversations.
[1026] Specific examples
[1027] 1. User Input
[1028] When a user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere," the emotion engine detects that the user is slightly nervous.
[1029] 2. Data Transmission
[1030] The device sends data, including the entered keywords, settings, and tension generated by the emotion engine, to the server using the HTTPS protocol.
[1031] 3. Data Receipt and Analysis
[1032] The server receives the data and analyzes it for keywords, settings, and tension. As a result of the analysis, it generates prompts that help the user relax.
[1033] 4. Conversation generation using generative AI
[1034] The server sends the following prompt to the AI: "Generate a conversation scenario for a user who is nervous about having a casual meal with a friend." Based on this instruction, the AI generates the following dialogue:
[1035] A: Where do you want to eat today?
[1036] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[1037] A: Great, let's do that! How have you been lately?
[1038] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[1039] 5. Emotional Adaptation Check
[1040] The server checks whether the generated dialogue relieves the user's tension. The emotion engine analyzes the user's emotions again, and if necessary, modifies and regenerates the prompt.
[1041] 6. Sending conversation text
[1042] The server sends the confirmed conversation data to the terminal as an HTTPS response.
[1043] 7. User Visibility
[1044] The terminal receives the conversation sentence data and displays it to the user, who can then refer to the generated natural conversation sentences and have a relaxed conversation.
[1045] The present invention allows users to obtain natural conversational sentences that are adapted to their emotional state, enabling smooth communication. This allows even users who are not good at communication to enjoy conversations with ease, and also contributes to building interpersonal relationships.
[1046] The processing flow will be explained below.
[1047] Step 1:
[1048] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. At this time, the device analyzes the user's facial expression data and tone of voice using an emotion engine to extract emotional data.
[1049] Step 2:
[1050] The device organizes the input keywords and settings, as well as the user's emotional data extracted by the emotion engine, and sends them to the server. Specifically, it converts the keywords, settings, and emotional data into JSON format and sends it as an HTTPS request to the server's API endpoint.
[1051] Step 3:
[1052] The server receives data sent from the device, which is then processed by the API endpoint and converted into an internal format for analysis.
[1053] Step 4:
[1054] The server analyzes the received keywords, preferences, and emotional data. Based on the analysis, it generates prompts that take the user's emotions into account. For example, if the user is nervous, the prompt may include instructions such as "generate relaxing conversation content."
[1055] Step 5:
[1056] The server sends a generated prompt to the generation AI (e.g., GPT-4), which includes the user's emotional state and the specified keywords and preferences.
[1057] Step 6:
[1058] The AI generates natural-sounding dialogue based on prompts, taking into account context, emotions, and the flow of the conversation, and tailors the dialogue to match the user's emotions.
[1059] Step 7:
[1060] The server receives the conversation data generated by the generation AI, checks the quality of the received conversation data, and again determines whether it is appropriate for the user's emotional state.
[1061] Step 8:
[1062] The server uses the emotion engine to check whether the generated dialogue is appropriate to the user's emotions. If it is not appropriate, it modifies the prompt and sends it to the generation AI again.
[1063] Step 9:
[1064] The server sends the confirmed conversation data to the terminal, which is then sent in JSON format as an HTTPS response.
[1065] Step 10:
[1066] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[1067] Step 11:
[1068] The device displays the formatted conversational text to the user, who can then refer to the natural-sounding conversational text and use it in actual conversations. The emotion engine ensures that the conversational text takes into account the user's current emotional state, allowing the user to continue the conversation with peace of mind.
[1069] This allows the user to obtain natural conversational sentences that are adapted to his or her own emotional state, making actual communication smoother.
[1070] Example 2
[1071] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1072] Conventional systems have the problem that when a user inputs specific keywords or settings to generate conversational sentences, they are unable to generate conversational sentences that correspond to the user's emotional state. As a result, it is difficult to obtain conversational content that matches the user's emotional state, such as when they are tense or relaxed, making it difficult to build natural communication.
[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1074] In this invention, the server includes means for a user to input keywords and settings into a terminal and collect emotional data, means for transmitting the keywords, settings, and emotional data from the user terminal to the server via a communication network, means for analyzing the received data and generating prompts based on the user's emotional state, means for generating natural conversational sentences from the prompts using a generative AI model, means for checking whether the generated conversational sentence data is appropriate for the user's emotions and for regenerating the data as necessary, and means for transmitting the generated conversational sentence data to the user terminal. This makes it possible to generate natural conversational sentences that are adapted to the user's emotional state and support smooth communication.
[1075] A "user" is a person who operates a device to input keywords and settings and obtain conversational text based on their emotional state.
[1076] A "terminal" is a device operated by a user, which collects input keywords, settings, and emotional data and transmits them to a server.
[1077] "Emotional data" is data analyzed by the emotion engine based on the user's facial expressions, tone of voice, etc.
[1078] A "communication network" is a network infrastructure for data communication between a user terminal and a server.
[1079] A "server" is a computer system that receives data sent from a user terminal, analyzes it, and generates conversational sentences.
[1080] A "generative AI model" is an artificial intelligence that uses pre-trained algorithms to generate natural-sounding conversational sentences from generated prompts.
[1081] A "prompt" is an instruction or question-type sentence that tells the generative AI model to generate natural conversational sentences.
[1082] "Conversational text" is a natural conversational text generated by a generative AI model and provided to the user.
[1083] The present invention is a system that generates natural conversational sentences that correspond to the user's emotional state using a user terminal, a server, a generative AI model, and an emotion engine.
[1084] The system includes a user terminal, a communication network, a server, a generative AI model, and an emotion engine. The system operates as follows.
[1085] Hardware and software used
[1086] User devices include smartphones, tablets, and other devices. These devices include an interface for entering keywords and settings using a keyboard or touchscreen. They also include an emotion engine for analyzing facial expressions and tone of voice. The emotion engine may use, for example, facial recognition software or voice analysis software.
[1087] The communication network is the infrastructure that connects the user terminal and the server. Specifically, the Internet or mobile data communication is often used. Data is transmitted and received securely using the HTTPS protocol.
[1088] The server is a computer system that receives data sent from the user's device, analyzes it, and generates conversational sentences. The server is equipped with a generative AI model, such as GPT-4, that has the ability to generate advanced natural language.
[1089] Data processing and calculation
[1090] The user enters keywords and settings into the device to collect emotional data. This data is then sent to a server via a communications network. The server analyzes the received data and generates prompts based on the user's emotional state. These prompts are then input into a generative AI model to generate natural-sounding conversational sentences.
[1091] The generated conversation sentences are checked by the server to see if they are appropriate for the user's emotions, and are regenerated if necessary. Finally, the generated conversation sentence data is sent to the user's terminal and displayed to the user.
[1092] Specific examples
[1093] A user opens a smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere." The emotion engine analyzes the user's facial expressions and tone of voice and detects that the user is "slightly nervous." This data is sent to the server, which then analyzes the keywords, settings, and emotion data.
[1094] The analysis results in prompts that help users relax. Based on these prompts, the generative AI model generates natural-sounding conversational sentences like the following:
[1095] A: Where do you want to eat today?
[1096] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[1097] A: Great, let's do that! How have you been lately?
[1098] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[1099] The generated conversation sentences are then checked again by the emotion engine, and after corrections are made as necessary, they are sent to the user's device. The user can refer to these conversation sentences and proceed with the actual conversation in a relaxed manner.
[1100] As described above, the present invention generates natural conversational sentences that are adapted to the emotional state of the user, thereby supporting smooth communication.
[1101] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1102] Step 1:
[1103] The user inputs keywords and settings into the device to collect emotional data. Specifically, the user inputs keywords and settings using the device's keyboard or touch screen. At this time, the device's camera and microphone analyze the user's facial expressions and tone of voice using an emotion engine to obtain emotional data.
[1104] Input: Keywords, settings, user facial expression and tone of voice
[1105] Output: Emotion data, keywords and settings
[1106] Step 2:
[1107] The device then transmits the collected keywords, settings, and emotion data to a server, where the data is transmitted securely over a communications network using the HTTPS protocol.
[1108] Input: Keywords, settings, emotion data
[1109] Output: Data sent to the server
[1110] Step 3:
[1111] The server receives the data sent from the device and analyzes keywords, preferences, and emotional data. Specifically, it uses an analysis algorithm on the received data to generate prompts based on the user's emotional state.
[1112] Input: Keywords, settings, emotion data
[1113] Output: Generated prompt
[1114] Step 4:
[1115] The server inputs the generated prompts to the generative AI model to generate natural conversational sentences. In this process, the generative AI model generates dialogue-style sentences based on the prompts.
[1116] Input: Generated prompt
[1117] Output: Natural conversational sentences
[1118] Step 5:
[1119] The server checks whether the generated dialogue is appropriate for the user's emotional state, again using the emotion engine to analyze the user's emotional changes in response to the dialogue, modifying the prompts as necessary, and re-inputting them into the generative AI model.
[1120] Input: Generated conversation sentences, user reaction data
[1121] Output: Modified prompt (if needed)
[1122] Step 6:
[1123] The server then sends the final generated conversation data to the device, again securely using the HTTPS protocol.
[1124] Input: Final generated dialogue
[1125] Output: Conversation data sent to the device
[1126] Step 7:
[1127] The device displays the received conversation data to the user, who can then refer to the displayed conversation data and use it in actual conversations. Specifically, the generated conversation data is displayed on the device screen.
[1128] Input: Received conversation data
[1129] Output: The dialogue displayed to the user
[1130] (Application example 2)
[1131] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1132] Conventional conversation generation systems have had difficulty generating natural conversations that reflect the user's emotions and situation. In particular, when ordering food delivery, appropriate conversations are required that reflect the user's stress or relaxation state, but these systems have not been able to meet this need. This often results in an unsmooth ordering process and a poor user experience.
[1133] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for displaying the generated conversation sentence data, means for transmitting data including keywords and settings from the user terminal to the server via a communication network, means for analyzing the received data and generating natural conversation sentences using a generation AI, means for transmitting the generated conversation sentence data to the user terminal, and means for generating natural ordering conversation sentences based on the user's emotion data in the case of ordering food delivery. This makes it possible to generate natural conversation sentences that correspond to the user's emotions and situation.
[1134] The "means for displaying the generated conversation sentence data" is a function or device for displaying the generated conversation sentence data on the screen of a user terminal or the like.
[1135] The "means for transmitting data including keywords and settings from a user terminal to a server via a communications network" refers to a communications means and protocol for transmitting keywords and settings entered by a user to a server.
[1136] The "means for analyzing received data and generating natural conversational sentences using a generation AI" refers to an algorithm and processing device for analyzing received data and generating natural conversational sentences using the analysis results and a generation AI.
[1137] The "means for transmitting the generated conversation sentence data to the user terminal" refers to a communication means and protocol for transmitting the generated conversation sentence data from the server to the user terminal.
[1138] The "means for generating natural ordering conversations based on the user's emotional data when ordering food delivery" refers to an algorithm and processing device that takes into account the user's emotional data when ordering food delivery and generates natural ordering conversations based on that data.
[1139] This invention relates to a system that facilitates the food delivery ordering process by generating natural ordering conversations based on the user's emotions. Specific embodiments for implementing this system are described below.
[1140] System Overview
[1141] This system generates natural conversational sentences and provides them to users when they order food delivery using their smartphones. The system consists of the following main components:
[1142] Hardware / Software Components
[1143] 1. Smartphone
[1144] Emotion engine: An engine that recognizes the user's emotions, for example by using a camera and microphone to analyze the user's facial expressions and tone of voice.
[1145] Communication module: A module to support HTTPS communication.
[1146] 2. Server
[1147] Data analysis engine: Analyzes keywords, settings, and sentiment data entered by users.
[1148] Generative AI models (e.g., GPT-3): Generate prompts based on user sentiment and generate natural-sounding conversations.
[1149] Communication module: A module for receiving data from the user terminal and transmitting the generated conversation data to the user terminal.
[1150] Data processing and calculation
[1151] User Input
[1152] When a user opens the smartphone application and enters keywords related to the order (e.g., "Italian restaurant") and a preference (e.g., "relaxed atmosphere"), the emotion engine collects the user's emotional data. For example, if the user is relaxed, that data is collected.
[1153] Data transmission and analysis
[1154] The smartphone's communication module sends keywords, settings, and emotional data to the server using the HTTPS protocol, where the server's data analysis engine analyzes the received data and generates prompts for the generative AI model.
[1155] Conversation generation using generative AI
[1156] The generative AI model generates natural-sounding sentences based on prompts such as:
[1157] "Generate an Italian restaurant ordering scenario for a relaxed user."
[1158] Here are some examples of dialogue generated by the generative AI:
[1159] A: Hello, I'd like some suggestions for Italian restaurants.
[1160] B: Of course! I recommend the Margherita pizza. It's a very relaxing dish.
[1161] Sending and displaying conversation text
[1162] The server then sends the generated conversational text to the smartphone, which receives and displays it on the user's device. The user can refer to the natural conversational text displayed and place their order smoothly.
[1163] This will enable the system to provide natural-sounding conversations that reflect the user's emotions and situation, which is expected to make the food delivery ordering process go more smoothly.
[1164] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1165] Step 1:
[1166] A user launches the smartphone application and enters keywords related to their order (e.g., "Italian restaurant") and preferences (e.g., "relaxed atmosphere"). At this time, the smartphone's emotion engine analyzes the user's facial expressions and tone of voice in real time to collect emotional data (whether they are relaxed or not).
[1167] Input data: Keywords, settings, and sentiment data
[1168] Output data: User input dataset
[1169] Step 2:
[1170] The terminal's communication module compiles the collected keywords, preferences, and emotion data and transmits them to a server using the HTTPS protocol.
[1171] Input data: User input dataset
[1172] Output data: JSON formatted data packet to the server
[1173] Step 3:
[1174] The server receives data packets from the user terminal, and the analysis engine analyzes them. The analysis engine extracts keywords, preferences, and emotion data from the data packets, and then creates prompts for the generative AI model.
[1175] Input data: JSON format data packet to the server
[1176] Output data: prompt statement
[1177] Step 4:
[1178] The server's generative AI model generates natural-sounding conversational sentences based on the generated prompt, for example, "Generate an ordering scenario at an Italian restaurant for a relaxed user."
[1179] Input data: Prompt statement
[1180] Output data: Generated conversation
[1181] Step 5:
[1182] The server then rechecks the generated dialogue to verify its suitability for the user's emotional data, and if necessary performs additional analysis to refine it and create the complete dialogue.
[1183] Input data: Generated conversation
[1184] Output: Verified and possibly corrected dialogue
[1185] Step 6:
[1186] The server then sends the final generated conversation text to the user's device using the HTTPS protocol, allowing the user to obtain the conversation text to smoothly complete the ordering process.
[1187] Input data: Verified and corrected conversation text
[1188] Output data: Conversation data packets to the user terminal
[1189] Step 7:
[1190] The terminal receives the conversation data sent from the server and displays it on the smartphone screen. The user can then place an order based on the natural conversation displayed.
[1191] Input data: Conversation data packets to the user terminal
[1192] Output data: Conversation displayed on smartphone screen
[1193] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1194] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1195] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1196] [Fourth embodiment]
[1197] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1198] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1199] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1200] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1201] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1202] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1203] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1204] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1205] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1206] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1207] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1208] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1209] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] The present invention comprises a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the user terminal. This allows the user to automatically generate natural conversational sentences based on specified keywords and settings and use them as a reference for actual conversations.
[1211] Program processing flow
[1212] 1. User Input
[1213] The user inputs specific keywords and settings into the terminal, for example, the keyword "dinner conversation with friends" and the setting "casual atmosphere."
[1214] 2. Data Transmission
[1215] The terminal sends the input keywords and setting data to the server via a communication network.
[1216] 3. Data Receipt and Analysis
[1217] The server receives the data sent from the device, analyzes the received keywords and settings, and generates the appropriate prompts.
[1218] 4. Conversation generation using generative AI
[1219] The server uses a generative AI (e.g., GPT-4) to generate natural-sounding conversational sentences based on the generated prompts. The generative AI takes into account the input keywords and settings to create natural-sounding conversational sentences that fit the context.
[1220] 5. Sending conversation text
[1221] The server transmits the generated conversation sentence data to the terminal again via the communication network.
[1222] 6. User Visibility
[1223] The device displays the conversation data received from the server, and the user can use the displayed conversation data as a reference for actual conversations.
[1224] Specific examples
[1225] 1. User Input
[1226] The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[1227] 2. Data Transmission
[1228] The terminal sends the entered data to the server. Communication is carried out using the HTTPS protocol.
[1229] 3. Data Receipt and Analysis
[1230] The server receives the data, analyzes the keywords and settings, and generates a prompt based on the analysis results, which it prepares to send to the AI.
[1231] 4. Conversation generation using generative AI
[1232] The server sends the following prompt to the AI: "Generate a casual conversation about a meal with a friend." Based on this instruction, the AI generates the following conversation:
[1233] A: What do you want to eat today? Pizza?
[1234] B: That's great! But I hear the pasta here is delicious too.
[1235] A: Right, then let's order both pasta and pizza and share.
[1236] B: That might be a good idea. By the way, have you traveled anywhere recently?
[1237] 5. Sending conversation text
[1238] The server transmits the generated conversation sentence to the terminal via a communication network.
[1239] 6. User Visibility
[1240] The device receives the conversation data and displays it to the user, who can use it as a reference when dining with real friends.
[1241] This invention allows users to improve their conversational skills and obtain natural conversational sentences that can be used as reference in actual communication situations. This allows even users who are not good at communication to enjoy smooth conversations, which greatly contributes to reducing feelings of loneliness and building relationships.
[1242] The processing flow will be explained below.
[1243] Step 1:
[1244] The user inputs a specific keyword and setting into the terminal. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. After inputting, the user presses the data transmission button.
[1245] Step 2:
[1246] The device organizes the input keywords and settings data and sends it to the server. Specifically, it converts the keywords and settings into JSON format and sends it to the server's API endpoint via an HTTPS request.
[1247] Step 3:
[1248] The server receives the data sent from the device. Specifically, the server's API endpoint processes the HTTPS request and converts the received data into an internal format for analysis.
[1249] Step 4:
[1250] The server analyzes the received keywords and settings, and generates prompts for the AI based on the keywords and settings. These prompts reflect the context and mood of the conversation the user desires.
[1251] Step 5:
[1252] The server sends a prompt to the generating AI (e.g., GPT-4), which contains information based on user-specified keywords and preferences.
[1253] Step 6:
[1254] Generative AI generates natural-sounding conversations based on prompts, taking into account the context, flow of the dialogue, and word usage to create conversations that fit the specified conditions.
[1255] Step 7:
[1256] The server receives the conversation data generated by the AI, checks the quality of the data, and formats it before providing it to the user.
[1257] Step 8:
[1258] The server sends the generated conversation data to the device. Again, the data is sent in JSON format as an HTTPS response.
[1259] Step 9:
[1260] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[1261] Step 10:
[1262] The terminal displays the formatted conversation sentences to the user, who can refer to the generated natural conversation sentences and use them in actual conversations.
[1263] This allows users to improve their conversational skills and communicate more smoothly.
[1264] Example 1
[1265] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1266] Conventional conversation support systems have struggled to automatically generate natural-sounding conversational sentences for users to refer to. Furthermore, generating custom conversational sentences based on user-specified keywords and conversational tones requires advanced expertise, making them difficult for average users to use. Furthermore, there was a lack of a way to quickly generate natural-sounding conversational sentences and provide them to users, making real-time use difficult.
[1267] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1268] In this invention, the server includes means for a user to input specific keywords and settings to a terminal, means for transmitting the input keywords and settings to the server via a communication network, means for the server to analyze the received keywords and settings and generate prompts for a generative AI model, means for generating natural conversational sentences based on the prompts using the generative AI model, means for transmitting the generated conversational sentence data from the server again via the communication network to the terminal, and means for displaying the generated conversational sentence data received by the terminal to the user. This enables the user to easily generate and refer to natural conversational sentences in real time according to the set keywords and atmosphere.
[1269] "User" refers to an individual or group that uses the system and is the entity that inputs keywords and settings via a terminal.
[1270] "Terminal" refers to a device operated by a user, including devices such as smartphones, personal computers, and tablets.
[1271] "Keywords" refer to specific words or phrases that users input to generate conversation text, and they play a role in determining the theme and content of the generated conversation text.
[1272] "Settings" refers to additional conditions or attributes that a user specifies when generating a conversation, including details such as the tone and atmosphere of the conversation.
[1273] "Communications network" refers to an infrastructure that enables the transmission and reception of data, including the Internet and local area networks.
[1274] "Server" refers to a central processing unit that processes data received from the terminal and sends prompts to the generating AI.
[1275] A "prompt" refers to an instruction entered into a generative AI model, which serves as a guide for the generated conversational text.
[1276] "Generative AI models" refer to artificial intelligence algorithms that generate natural-sounding conversational sentences based on input prompts, including, for example, large-scale language models.
[1277] "Conversational text" refers to part or all of a sentence that is generated in a form that can be referenced by the user, and refers to text that has a natural dialogue format.
[1278] "Data" refers to keywords, settings, generated dialogue, and related information in general.
[1279] "JSON format" refers to a standard data representation format for structuring data in a way that is easy for humans to read and machines to analyze.
[1280] This invention is a system that allows a user to input specific keywords and settings and generates natural conversational sentences based on the input. The system is implemented mainly using the following hardware and software.
[1281] System configuration and operation
[1282] Hardware
[1283] 1. User device: A device such as a smartphone, personal computer, or tablet. These devices provide an interface for users to enter keywords and settings.
[1284] 2. Server: A high-performance server or cloud environment. The server is used to receive and analyze data and operate the generative AI model.
[1285] software
[1286] 1. Generative AI models: Large-scale language models such as GPT-4 from OpenAI, which generate highly natural-sounding conversational sentences.
[1287] 2. Communication protocol: The HTTPS protocol is used for communication between the terminal and the server to ensure data security and reliability.
[1288] 3. Data analysis tools: Use analysis tools such as Python's json module to receive and analyze data.
[1289] Data processing and calculation
[1290] The user inputs specific keywords, such as "dinner conversation with friends," and settings, such as "casual atmosphere," into the terminal. The specific data processing and calculations performed at each step are described in detail below.
[1291] 1. User Input: The user enters keywords and settings into the terminal application. The input is done in a text box, and when the submit button is pressed, the data is converted into JSON format.
[1292] 2. Data transmission: The device uses the HTTPS protocol to send JSON format data to the server. The data is encrypted before being sent.
[1293] 3. Receiving and parsing data: The server parses the received JSON data and extracts keywords and settings. This parsing is done using the Python json module.
[1294] 4. Prompt generation: Based on the extracted keywords and settings, the server generates a prompt to send to the generative AI model. An example of a prompt is "Generate a casual dinner conversation with a friend."
[1295] 5. Conversational Sentence Generation by Generative AI: The server sends prompts to the generative AI model, which then generates natural-sounding conversational sentences using hundreds of billions of parameters to provide context-appropriate answers.
[1296] 6. Sending the conversation: The generated conversation is converted back to JSON format and sent to the device via HTTPS.
[1297] 7. Display to user: The device parses the received JSON data and displays the generated conversation text to the user. The user can refer to this conversation text and use it in actual conversations.
[1298] Specific examples
[1299] The specific flow of system usage is shown below.
[1300] 1. User Input: The user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere."
[1301] 2. Prompt generation: The server sends the generative AI model a prompt like this: "Generate a casual dinner conversation with a friend."
[1302] 3. Conversation generation results: The generative AI model generates the following conversation:
[1303] A: What do you want to eat today? Pizza?
[1304] B: That's great! But I hear the pasta here is delicious too.
[1305] A: Right, then let's order both pasta and pizza and share.
[1306] B: That might be a good idea. By the way, have you traveled anywhere recently?
[1307] This system allows users to improve their conversational skills, easily generate natural conversational sentences in real time, and use them in real-life communications.
[1308] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1309] Step 1:
[1310] The user opens a smartphone application or web browser, inputs specific keywords and settings, such as "dinner conversation with friends" and "casual atmosphere" in the text boxes, and then presses the submit button. This input data is passed to the next step.
[1311] Step 2:
[1312] The device converts the keywords and settings entered by the user into JSON format data. For example, the following JSON data is generated:
[1313] json
[1314] {
[1315] "keyword": "dinner conversation with friends",
[1316] "setting": "casual atmosphere"
[1317] }
[1318] This JSON data is sent to the server using the HTTPS protocol, and the device confirms that the data is sent encrypted.
[1319] Step 3:
[1320] The server receives the JSON data sent from the device. It parses the data using Python's json module and extracts keywords and settings. The analysis results are passed to the next step. From the input JSON data, "dinner conversation with friends" and "casual atmosphere" are extracted.
[1321] Step 4:
[1322] Based on the extracted keywords and settings, the server generates a prompt to be sent to the generative AI model. Specifically, it generates the prompt "Generate a casual dinner conversation with a friend." This prompt is passed to the next step.
[1323] Step 5:
[1324] The server sends the prompt to a generative AI model (e.g., GPT-4) and requests it to generate a conversation. The generative AI model analyzes the prompt and generates natural-sounding conversational sentences that meet the required conditions. For example, the following conversational sentences are generated:
[1325] A: What do you want to eat today? Pizza?
[1326] B: That's great! But I hear the pasta here is delicious too.
[1327] A: Right, then let's order both pasta and pizza and share.
[1328] B: That might be a good idea. By the way, have you traveled anywhere recently?
[1329] The generated conversation sentence is passed to the next step.
[1330] Step 6:
[1331] The server converts the generated conversation text back into JSON format and sends it to the device. At this time, the server checks the integrity of the generated data and performs error handling if necessary. The generated conversation text is converted into JSON data as follows:
[1332] json
[1333] {
[1334] "conversation": [
[1335] "A: What do you want to eat today? Pizza?"
[1336] "B: That's great! But I heard the pasta here is also delicious."
[1337] "A: Right, then let's order both pasta and pizza and share."
[1338] "B: That might be a good idea. By the way, have you traveled anywhere recently?"
[1339] ]
[1340] }
[1341] This JSON data is sent to the device.
[1342] Step 7:
[1343] The device parses the JSON data received from the server and displays it on the application's UI. The user can refer to the natural conversational text displayed on the screen and use it in actual conversations. The device displays the generated conversational text to the user through the user interface, allowing the user to perform additional operations such as scrolling and saving.
[1344] By following the steps above, users can easily and quickly generate natural conversational sentences based on specified conditions and use them in actual conversations.
[1345] (Application example 1)
[1346] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1347] Customer service in modern brick-and-mortar stores requires store clerks to respond quickly and appropriately to customer questions and requests. However, this depends on the experience and knowledge of the store clerk, which can lead to inconsistencies in individual responses. Furthermore, responding immediately in a busy environment can be difficult, which can lead to a decrease in customer satisfaction. A solution to this issue is needed, enabling store clerks to consistently provide high-quality service to customers.
[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1349] In this invention, the server includes means for receiving customer questions and requests by voice input, means for converting the voice input data into text data, and means for generating prompts based on the text data, thereby enabling store staff to provide quick and appropriate responses to customer questions and requests.
[1350] "Conversational data" is text information of natural conversations created by generative AI.
[1351] A "communications network" is an information pathway such as the Internet or a local area network that allows data to be sent and received.
[1352] A "terminal" is a user device such as a smartphone, smart glasses, or head-mounted display operated by a user or store clerk.
[1353] A "server" is a computer system that processes data and generates conversational sentences using AI.
[1354] "Keywords" are specific words or phrases entered by the user, and serve as basic information for generating conversational text.
[1355] "Settings" are parameters related to the tone and topic of the conversation specified by the user.
[1356] "Generative AI" is an artificial intelligence system that generates natural-sounding conversational sentences based on input keywords and settings.
[1357] "Voice input" is a method in which a user or store clerk speaks to a terminal and the voice data is acquired.
[1358] "Text data" refers to data obtained by converting information input by voice into a character string format.
[1359] A "prompt" is a text instruction that serves as the basis for the generation AI to generate conversational sentences.
[1360] This invention is a system including a means for displaying generated conversational sentence data, a means for transmitting data including keywords and settings from a terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational sentences using a generation AI, and a means for transmitting the generated conversational sentence data to the terminal. Also, by adding a means for generating prompts based on keywords and settings input by a user and a means for converting voice-input data into text data, this is a specific embodiment for supporting communication between store clerks and customers.
[1361] Hardware and software used
[1362] Hardware
[1363] Devices: smartphones, smart glasses, head-mounted displays, etc.
[1364] Server: A high-performance cloud server, such as an AWS EC2 server.
[1365] software
[1366] Speech recognition engine: Google Speech-to-Text API
[1367] Natural language analysis engines: SpaCy, NLTK
[1368] Generative AI model: OpenAI GPT-4
[1369] Communication protocol: HTTPS
[1370] What the system does
[1371] First, the store clerk (the user) uses the voice input function of the terminal to input the customer's question or request by voice. For example, a customer might ask, "Do you have this shirt in my size?" This voice input is converted into text data on the terminal and sent to the server using the HTTPS protocol.
[1372] The server analyzes the received text data using a natural language analysis engine (SpaCy or NLTK) and generates an appropriate prompt. The generated prompt is passed to a generative AI model (OpenAI GPT-4), which then generates natural-sounding conversational sentences based on the prompt. For example, in response to the prompt, "You have requested stock information about the size of this shirt. Please generate an appropriate response," the generative AI generates the following conversational sentences:
[1373] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[1374] The generated conversation sentences are sent from the server to the terminal again and displayed in the field of view of the store clerk. The store clerk can refer to the displayed conversation sentences and provide an appropriate response to the customer.
[1375] Specific examples
[1376] 1. Example prompt:
[1377] When a customer asks, "Do you have this shirt in my size?", generate an appropriate response.
[1378] 2. Example of generated dialogue:
[1379] Salesperson: We currently have this shirt in sizes S, M, and L. Which size are you looking for?
[1380] This invention enables store staff to provide quick and appropriate responses to customer questions and requests, which is expected to improve the quality of customer service in physical stores.
[1381] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1382] Step 1:
[1383] A user uses the smart glasses to input a customer question or request by voice. For example, a question such as "Do you have this shirt in my size?" The voice data is the input data.
[1384] Step 2:
[1385] The device uses a speech recognition engine (Google Speech-to-Text API) to convert the voice input data into text data. The voice data is converted into text data.
[1386] Step 3:
[1387] The terminal sends the converted text data to the server via a communication protocol (HTTPS). The text data sent is the input data, and the text data received by the server is the output data.
[1388] Step 4:
[1389] The server receives the text data and analyzes it using a natural language analysis engine (SpaCy, NLTK). The text data is the input data, and the analysis results are the output data.
[1390] Step 5:
[1391] Based on the analysis results, the server generates prompts to send to the generative AI model (GPT-4). The analysis results are the input data, and the generated prompts are the output data.
[1392] Step 6:
[1393] The server sends prompts to the generative AI model (GPT-4) to generate natural-sounding conversational sentences. The generated prompts are the input data, and the generated conversational sentences are the output data.
[1394] Step 7:
[1395] The server sends the generated conversation text to the terminal via HTTPS. The generated conversation text is the input data, and the conversation text received by the terminal is the output data.
[1396] Step 8:
[1397] The conversational text received by the terminal is displayed on the smart glasses. The received conversational text is input data, and the conversational text displayed on the smart glasses is output data.
[1398] Step 9:
[1399] The user refers to the displayed conversation sentence and provides an appropriate response to the customer. The displayed conversation sentence is input data, and the user's response is output data.
[1400] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1401] The present invention provides a system including a means for displaying generated conversational text data, a means for transmitting data including keywords and settings from a user terminal to a server via a communication network, a means for analyzing the received data and generating natural conversational text using a generation AI, a means for transmitting the generated conversational text data to the user terminal, and an emotion engine for recognizing the user's emotions.
[1402] Program processing flow
[1403] 1. User Input
[1404] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. Furthermore, the user's facial expressions and tone of voice are monitored by the device's emotion engine.
[1405] 2. Data Transmission
[1406] The device sends the input keywords and settings, as well as the user's emotional data analyzed by the emotion engine, to the server via a communication network.
[1407] 3. Data Receipt and Analysis
[1408] The server receives the data sent from the device, analyzes the received keywords, settings, and emotional data, and generates prompts appropriate to the user's emotions.
[1409] 4. Conversation generation using generative AI
[1410] The server sends emotion-based prompts to the AI, which contain information corresponding to the user's emotional state. The AI then uses these prompts to generate natural-sounding conversational sentences.
[1411] 5. Emotional Adaptation Check
[1412] The server checks whether the generated dialogue is appropriate to the user's emotions. The emotion engine performs emotion analysis again, and if necessary, modifies the prompt and generates it again.
[1413] 6. Sending conversation text
[1414] The server transmits the generated conversation data to the terminal, and the data is transmitted again via the communication network.
[1415] 7. User Visibility
[1416] The device receives the generated conversation data and displays it to the user, who can then use it in actual conversations.
[1417] Specific examples
[1418] 1. User Input
[1419] When a user opens the smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere," the emotion engine detects that the user is slightly nervous.
[1420] 2. Data Transmission
[1421] The device sends data, including the entered keywords, settings, and tension generated by the emotion engine, to the server using the HTTPS protocol.
[1422] 3. Data Receipt and Analysis
[1423] The server receives the data and analyzes it for keywords, settings, and tension. As a result of the analysis, it generates prompts that help the user relax.
[1424] 4. Conversation generation using generative AI
[1425] The server sends the following prompt to the AI: "Generate a conversation scenario for a user who is nervous about having a casual meal with a friend." Based on this instruction, the AI generates the following dialogue:
[1426] A: Where do you want to eat today?
[1427] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[1428] A: Great, let's do that! How have you been lately?
[1429] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[1430] 5. Emotional Adaptation Check
[1431] The server checks whether the generated dialogue relieves the user's tension. The emotion engine analyzes the user's emotions again, and if necessary, modifies and regenerates the prompt.
[1432] 6. Sending conversation text
[1433] The server sends the confirmed conversation data to the terminal as an HTTPS response.
[1434] 7. User Visibility
[1435] The terminal receives the conversation sentence data and displays it to the user, who can then refer to the generated natural conversation sentences and have a relaxed conversation.
[1436] The present invention allows users to obtain natural conversational sentences that are adapted to their emotional state, enabling smooth communication. This allows even users who are not good at communication to enjoy conversations with ease, and also contributes to building interpersonal relationships.
[1437] The processing flow will be explained below.
[1438] Step 1:
[1439] The user inputs specific keywords and settings into the device. For example, the keyword "dinner conversation with friends" and the setting "casual atmosphere" are input. At this time, the device analyzes the user's facial expression data and tone of voice using an emotion engine to extract emotional data.
[1440] Step 2:
[1441] The device organizes the input keywords and settings, as well as the user's emotional data extracted by the emotion engine, and sends them to the server. Specifically, it converts the keywords, settings, and emotional data into JSON format and sends it as an HTTPS request to the server's API endpoint.
[1442] Step 3:
[1443] The server receives data sent from the device, which is then processed by the API endpoint and converted into an internal format for analysis.
[1444] Step 4:
[1445] The server analyzes the received keywords, preferences, and emotional data. Based on the analysis, it generates prompts that take the user's emotions into account. For example, if the user is nervous, the prompt may include instructions such as "generate relaxing conversation content."
[1446] Step 5:
[1447] The server sends a generated prompt to the generation AI (e.g., GPT-4), which includes the user's emotional state and the specified keywords and preferences.
[1448] Step 6:
[1449] The AI generates natural-sounding dialogue based on prompts, taking into account context, emotions, and the flow of the conversation, and tailors the dialogue to match the user's emotions.
[1450] Step 7:
[1451] The server receives the conversation data generated by the generation AI, checks the quality of the received conversation data, and again determines whether it is appropriate for the user's emotional state.
[1452] Step 8:
[1453] The server uses the emotion engine to check whether the generated dialogue is appropriate to the user's emotions. If it is not appropriate, it modifies the prompt and sends it to the generation AI again.
[1454] Step 9:
[1455] The server sends the confirmed conversation data to the terminal, which is then sent in JSON format as an HTTPS response.
[1456] Step 10:
[1457] The device parses the conversation data received from the server and prepares it for display to the user. Specifically, it parses the JSON formatted data and formats it so that it can be displayed appropriately on the user interface.
[1458] Step 11:
[1459] The device displays the formatted conversational text to the user, who can then refer to the natural-sounding conversational text and use it in actual conversations. The emotion engine ensures that the conversational text takes into account the user's current emotional state, allowing the user to continue the conversation with peace of mind.
[1460] This allows the user to obtain natural conversational sentences that are adapted to his or her own emotional state, making actual communication smoother.
[1461] Example 2
[1462] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1463] Conventional systems have the problem that when a user inputs specific keywords or settings to generate conversational sentences, they are unable to generate conversational sentences that correspond to the user's emotional state. As a result, it is difficult to obtain conversational content that matches the user's emotional state, such as when they are tense or relaxed, making it difficult to build natural communication.
[1464] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1465] In this invention, the server includes means for a user to input keywords and settings into a terminal and collect emotional data, means for transmitting the keywords, settings, and emotional data from the user terminal to the server via a communication network, means for analyzing the received data and generating prompts based on the user's emotional state, means for generating natural conversational sentences from the prompts using a generative AI model, means for checking whether the generated conversational sentence data is appropriate for the user's emotions and for regenerating the data as necessary, and means for transmitting the generated conversational sentence data to the user terminal. This makes it possible to generate natural conversational sentences that are adapted to the user's emotional state and support smooth communication.
[1466] A "user" is a person who operates a device to input keywords and settings and obtain conversational text based on their emotional state.
[1467] A "terminal" is a device operated by a user, which collects input keywords, settings, and emotional data and transmits them to a server.
[1468] "Emotional data" is data analyzed by the emotion engine based on the user's facial expressions, tone of voice, etc.
[1469] A "communication network" is a network infrastructure for data communication between a user terminal and a server.
[1470] A "server" is a computer system that receives data sent from a user terminal, analyzes it, and generates conversational sentences.
[1471] A "generative AI model" is an artificial intelligence that uses pre-trained algorithms to generate natural-sounding conversational sentences from generated prompts.
[1472] A "prompt" is an instruction or question-type sentence that tells the generative AI model to generate natural conversational sentences.
[1473] "Conversational text" is a natural conversational text generated by a generative AI model and provided to the user.
[1474] The present invention is a system that generates natural conversational sentences that correspond to the user's emotional state using a user terminal, a server, a generative AI model, and an emotion engine.
[1475] The system includes a user terminal, a communication network, a server, a generative AI model, and an emotion engine. The system operates as follows.
[1476] Hardware and software used
[1477] User devices include smartphones, tablets, and other devices. These devices include an interface for entering keywords and settings using a keyboard or touchscreen. They also include an emotion engine for analyzing facial expressions and tone of voice. The emotion engine may use, for example, facial recognition software or voice analysis software.
[1478] The communication network is the infrastructure that connects the user terminal and the server. Specifically, the Internet or mobile data communication is often used. Data is transmitted and received securely using the HTTPS protocol.
[1479] The server is a computer system that receives data sent from the user's device, analyzes it, and generates conversational sentences. The server is equipped with a generative AI model, such as GPT-4, that has the ability to generate advanced natural language.
[1480] Data processing and calculation
[1481] The user enters keywords and settings into the device to collect emotional data. This data is then sent to a server via a communications network. The server analyzes the received data and generates prompts based on the user's emotional state. These prompts are then input into a generative AI model to generate natural-sounding conversational sentences.
[1482] The generated conversation sentences are checked by the server to see if they are appropriate for the user's emotions, and are regenerated if necessary. Finally, the generated conversation sentence data is sent to the user's terminal and displayed to the user.
[1483] Specific examples
[1484] A user opens a smartphone application and enters the keywords "dinner conversation with friends" and the setting "casual atmosphere." The emotion engine analyzes the user's facial expressions and tone of voice and detects that the user is "slightly nervous." This data is sent to the server, which then analyzes the keywords, settings, and emotion data.
[1485] The analysis results in prompts that help users relax. Based on these prompts, the generative AI model generates natural-sounding conversational sentences like the following:
[1486] A: Where do you want to eat today?
[1487] B: Shall we go to our usual cafe? I like the relaxed atmosphere.
[1488] A: Great, let's do that! How have you been lately?
[1489] B: Not much has changed, but I've recently picked up a new hobby and have been meeting up with friends from high school.
[1490] The generated conversation sentences are then checked again by the emotion engine, and after corrections are made as necessary, they are sent to the user's device. The user can refer to these conversation sentences and proceed with the actual conversation in a relaxed manner.
[1491] As described above, the present invention generates natural conversational sentences that are adapted to the emotional state of the user, thereby supporting smooth communication.
[1492] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1493] Step 1:
[1494] The user inputs keywords and settings into the device to collect emotional data. Specifically, the user inputs keywords and settings using the device's keyboard or touch screen. At this time, the device's camera and microphone analyze the user's facial expressions and tone of voice using an emotion engine to obtain emotional data.
[1495] Input: Keywords, settings, user facial expression and tone of voice
[1496] Output: Emotion data, keywords and settings
[1497] Step 2:
[1498] The device then transmits the collected keywords, settings, and emotion data to a server, where the data is transmitted securely over a communications network using the HTTPS protocol.
[1499] Input: Keywords, settings, emotion data
[1500] Output: Data sent to the server
[1501] Step 3:
[1502] The server receives the data sent from the device and analyzes keywords, preferences, and emotional data. Specifically, it uses an analysis algorithm on the received data to generate prompts based on the user's emotional state.
[1503] Input: Keywords, settings, emotion data
[1504] Output: Generated prompt
[1505] Step 4:
[1506] The server inputs the generated prompts to the generative AI model to generate natural conversational sentences. In this process, the generative AI model generates dialogue-style sentences based on the prompts.
[1507] Input: Generated prompt
[1508] Output: Natural conversational sentences
[1509] Step 5:
[1510] The server checks whether the generated dialogue is appropriate for the user's emotional state, again using the emotion engine to analyze the user's emotional changes in response to the dialogue, modifying the prompts as necessary, and re-inputting them into the generative AI model.
[1511] Input: Generated conversation sentences, user reaction data
[1512] Output: Modified prompt (if needed)
[1513] Step 6:
[1514] The server then sends the final generated conversation data to the device, again securely using the HTTPS protocol.
[1515] Input: Final generated dialogue
[1516] Output: Conversation data sent to the device
[1517] Step 7:
[1518] The device displays the received conversation data to the user, who can then refer to the displayed conversation data and use it in actual conversations. Specifically, the generated conversation data is displayed on the device screen.
[1519] Input: Received conversation data
[1520] Output: The dialogue displayed to the user
[1521] (Application example 2)
[1522] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1523] Conventional conversation generation systems have had difficulty generating natural conversations that reflect the user's emotions and situation. In particular, when ordering food delivery, appropriate conversations are required that reflect the user's stress or relaxation state, but these systems have not been able to meet this need. This often results in an unsmooth ordering process and a poor user experience.
[1524] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for displaying the generated conversation sentence data, means for transmitting data including keywords and settings from the user terminal to the server via a communication network, means for analyzing the received data and generating natural conversation sentences using a generation AI, means for transmitting the generated conversation sentence data to the user terminal, and means for generating natural ordering conversation sentences based on the user's emotion data in the case of ordering food delivery. This makes it possible to generate natural conversation sentences that correspond to the user's emotions and situation.
[1525] The "means for displaying the generated conversation sentence data" is a function or device for displaying the generated conversation sentence data on the screen of a user terminal or the like.
[1526] The "means for transmitting data including keywords and settings from a user terminal to a server via a communications network" refers to a communications means and protocol for transmitting keywords and settings entered by a user to a server.
[1527] The "means for analyzing received data and generating natural conversational sentences using a generation AI" refers to an algorithm and processing device for analyzing received data and generating natural conversational sentences using the analysis results and a generation AI.
[1528] The "means for transmitting the generated conversation sentence data to the user terminal" refers to a communication means and protocol for transmitting the generated conversation sentence data from the server to the user terminal.
[1529] The "means for generating natural ordering conversations based on the user's emotional data when ordering food delivery" refers to an algorithm and processing device that takes into account the user's emotional data when ordering food delivery and generates natural ordering conversations based on that data.
[1530] This invention relates to a system that facilitates the food delivery ordering process by generating natural ordering conversations based on the user's emotions. Specific embodiments for implementing this system are described below.
[1531] System Overview
[1532] This system generates natural conversational sentences and provides them to users when they order food delivery using their smartphones. The system consists of the following main components:
[1533] Hardware / Software Components
[1534] 1. Smartphone
[1535] Emotion engine: An engine that recognizes the user's emotions, for example by using a camera and microphone to analyze the user's facial expressions and tone of voice.
[1536] Communication module: A module to support HTTPS communication.
[1537] 2. Server
[1538] Data analysis engine: Analyzes keywords, settings, and sentiment data entered by users.
[1539] Generative AI models (e.g., GPT-3): Generate prompts based on user sentiment and generate natural-sounding conversations.
[1540] Communication module: A module for receiving data from the user terminal and transmitting the generated conversation data to the user terminal.
[1541] Data processing and calculation
[1542] User Input
[1543] When a user opens the smartphone application and enters keywords related to the order (e.g., "Italian restaurant") and a preference (e.g., "relaxed atmosphere"), the emotion engine collects the user's emotional data. For example, if the user is relaxed, that data is collected.
[1544] Data transmission and analysis
[1545] The smartphone's communication module sends keywords, settings, and emotional data to the server using the HTTPS protocol, where the server's data analysis engine analyzes the received data and generates prompts for the generative AI model.
[1546] Conversation generation using generative AI
[1547] The generative AI model generates natural-sounding sentences based on prompts such as:
[1548] "Generate an Italian restaurant ordering scenario for a relaxed user."
[1549] Here are some examples of dialogue generated by the generative AI:
[1550] A: Hello, I'd like some suggestions for Italian restaurants.
[1551] B: Of course! I recommend the Margherita pizza. It's a very relaxing dish.
[1552] Sending and displaying conversation text
[1553] The server then sends the generated conversational text to the smartphone, which receives and displays it on the user's device. The user can refer to the natural conversational text displayed and place their order smoothly.
[1554] This will enable the system to provide natural-sounding conversations that reflect the user's emotions and situation, which is expected to make the food delivery ordering process go more smoothly.
[1555] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1556] Step 1:
[1557] A user launches the smartphone application and enters keywords related to their order (e.g., "Italian restaurant") and preferences (e.g., "relaxed atmosphere"). At this time, the smartphone's emotion engine analyzes the user's facial expressions and tone of voice in real time to collect emotional data (whether they are relaxed or not).
[1558] Input data: Keywords, settings, and sentiment data
[1559] Output data: User input dataset
[1560] Step 2:
[1561] The terminal's communication module compiles the collected keywords, preferences, and emotion data and transmits them to a server using the HTTPS protocol.
[1562] Input data: User input dataset
[1563] Output data: JSON formatted data packet to the server
[1564] Step 3:
[1565] The server receives data packets from the user terminal, and the analysis engine analyzes them. The analysis engine extracts keywords, preferences, and emotion data from the data packets, and then creates prompts for the generative AI model.
[1566] Input data: JSON format data packet to the server
[1567] Output data: prompt statement
[1568] Step 4:
[1569] The server's generative AI model generates natural-sounding conversational sentences based on the generated prompt, for example, "Generate an ordering scenario at an Italian restaurant for a relaxed user."
[1570] Input data: Prompt statement
[1571] Output data: Generated conversation
[1572] Step 5:
[1573] The server then rechecks the generated dialogue to verify its suitability for the user's emotional data, and if necessary performs additional analysis to refine it and create the complete dialogue.
[1574] Input data: Generated conversation
[1575] Output: Verified and possibly corrected dialogue
[1576] Step 6:
[1577] The server then sends the final generated conversation text to the user's device using the HTTPS protocol, allowing the user to obtain the conversation text to smoothly complete the ordering process.
[1578] Input data: Verified and corrected conversation text
[1579] Output data: Conversation data packets to the user terminal
[1580] Step 7:
[1581] The terminal receives the conversation data sent from the server and displays it on the smartphone screen. The user can then place an order based on the natural conversation displayed.
[1582] Input data: Conversation data packets to the user terminal
[1583] Output data: Conversation displayed on smartphone screen
[1584] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1585] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1586] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1587] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1588] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1589] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1590] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1591] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1592] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1593] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1594] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1595] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1596] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1597] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1598] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1599] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1600] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1601] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1602] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1603] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1604] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1605] The following is further disclosed regarding the above embodiment.
[1606] (Claim 1)
[1607] a means for displaying the generated conversation sentence data;
[1608] means for transmitting data including keywords and settings from a user terminal to a server via a communication network;
[1609] A means for analyzing the received data and generating natural conversational sentences using a generation AI;
[1610] means for transmitting the generated conversation sentence data to a user terminal;
[1611] A system including:
[1612] (Claim 2)
[1613] 2. The system of claim 1, wherein the user terminal further comprises means for receiving the generated conversation sentence data and displaying it to the user.
[1614] (Claim 3)
[1615] 10. The system of claim 1, further comprising: means for generating a prompt for the generating AI based on keywords and settings input by the user.
[1616] "Example 1"
[1617] (Claim 1)
[1618] a means for the user to input specific keywords and settings into the device;
[1619] means for transmitting the input keywords and settings to a server via a communication network;
[1620] a means for analyzing the received keywords and settings by the server and generating prompts for the generative AI model;
[1621] a means for generating natural-sounding conversational sentences based on prompts using a generative AI model;
[1622] means for transmitting the generated conversation sentence data from the server to the terminal again via a communication network;
[1623] means for displaying the generated conversation data received by the terminal to a user;
[1624] A system including:
[1625] (Claim 2)
[1626] 2. The system of claim 1, wherein the user terminal further comprises means for receiving the generated conversation sentence data and displaying it to the user.
[1627] (Claim 3)
[1628] 10. The system of claim 1, further comprising means for generating prompts for the generative AI model based on keywords and settings input by a user.
[1629] "Application Example 1"
[1630] (Claim 1)
[1631] a means for displaying the generated conversation sentence data;
[1632] means for transmitting data including keywords and settings from the terminal to the server via a communication network;
[1633] A means for analyzing the received data and generating natural conversational sentences using a generation AI;
[1634] means for transmitting the generated conversation sentence data to a terminal;
[1635] A means of capturing customer questions and requests through voice input;
[1636] A means for converting voice input data into text data;
[1637] means for generating a prompt based on the text data;
[1638] A system including:
[1639] (Claim 2)
[1640] 2. The system of claim 1, wherein the terminal further comprises means for receiving the generated conversation sentence data and displaying it to the user.
[1641] (Claim 3)
[1642] 10. The system of claim 1, further comprising: means for generating a prompt for the generating AI based on keywords and settings input by the user.
[1643] "Example 2: Combining Emotion Engines"
[1644] (Claim 1)
[1645] A means for users to input keywords and settings into the device to collect emotional data;
[1646] means for transmitting keywords, settings, and emotion data from a user terminal to a server via a communication network;
[1647] means for analyzing the received data and generating prompts based on the user's emotional state;
[1648] a means for generating natural-sounding conversational sentences from the prompts using a generative AI model;
[1649] a means for checking whether the generated conversation data is appropriate for the user's emotions and regenerating the data as necessary;
[1650] means for transmitting the generated conversation sentence data to a user terminal;
[1651] A system including:
[1652] (Claim 2)
[1653] 2. The system of claim 1, wherein the user terminal further comprises means for receiving the generated conversation sentence data and displaying it to the user.
[1654] (Claim 3)
[1655] 10. The system of claim 1, further comprising: means for generating a prompt for the generation AI based on keywords and settings and emotion data input by the user.
[1656] "Application example 2 when combining emotion engines"
[1657] (Claim 1)
[1658] a means for displaying the generated conversation sentence data;
[1659] means for transmitting data including keywords and settings from a user terminal to a server via a communication network;
[1660] A means for analyzing the received data and generating natural conversational sentences using a generation AI;
[1661] means for transmitting the generated conversation sentence data to a user terminal;
[1662] A means for generating natural ordering conversation sentences based on user emotion data when ordering food delivery;
[1663] A system including:
[1664] (Claim 2)
[1665] 2. The system of claim 1, wherein the user terminal further comprises means for receiving the generated conversation sentence data and displaying it to the user.
[1666] (Claim 3)
[1667] 10. The system of claim 1, further comprising: means for generating a prompt for the generating AI based on keywords and settings input by the user. [Explanation of symbols]
[1668] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for displaying the generated conversation sentence data; means for transmitting data including keywords and settings from a user terminal to a server via a communication network; A means for analyzing the received data and generating natural conversational sentences using a generation AI; means for transmitting the generated conversation sentence data to a user terminal; A system including:
2. 2. The system of claim 1, wherein the user terminal further comprises means for receiving said generated conversation sentence data and displaying it to the user.
3. 10. The system of claim 1, further comprising: means for generating a prompt for the generating AI based on keywords and settings input by the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A