System
A system using a natural language processing engine and generative AI model addresses the inefficiencies of conventional customer support by offering immediate and accurate responses, improving user experience and reducing physical store burdens.
Patent Information
- Application Number
- JP2024123954
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional customer support systems face challenges in providing quick and clear solutions to middle-aged and elderly users, young people seeking immediate responses, business people, and students, due to long waiting times, limited business hours, and inefficient communication, leading to a poor customer experience.
A system that utilizes a natural language processing engine to analyze user questions and generates responses using a generative AI model, enabling 24/7 support and reducing the burden on physical store staff.
The system provides efficient, immediate, and accurate responses to user inquiries, enhancing customer experience by eliminating waiting times and reducing the need for in-person visits, thus optimizing customer service operations.
Smart Images

Figure 2026022437000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With conventional customer support systems, it was difficult for middle-aged and elderly people who are not familiar with technology, young people who want a quick response, business people, and students to get a clear and quick solution when they had a problem. Furthermore, there were issues that caused inconvenience to users, such as long waiting times at the support center and limited business hours. This resulted in a poor customer experience. [Means for solving the problem]
[0005] This invention solves the above-mentioned problems with a system that includes a means for receiving product- or service-related questions from users, a means for analyzing the questions using a natural language processing engine, a means for generating responses using a generative AI model based on the analysis results, and a means for returning the generated responses to the users. As a result, users can reduce wait times and efficiently resolve problems through chat-based support available 24 hours a day, 365 days a year. It also reduces the burden on counter staff at physical stores and enables efficient customer service.
[0006] "Questions about products and services" are inquiries about doubts or problems that users have regarding products or services they use.
[0007] A "user" is someone who uses a product or service and has a question or problem regarding its use.
[0008] "Means for receiving" refers to the function or process by which the system receives a question entered by a user.
[0009] A "natural language processing engine" is software or algorithms that analyze user input and understand its meaning and intent.
[0010] "Means for analyzing" refers to the process of analyzing the received question using a natural language processing engine and extracting meaning.
[0011] A "generative AI model" is an artificial intelligence algorithm that generates appropriate responses to user questions.
[0012] "Means for generating responses" refers to a function that utilizes a generative AI model to create an appropriate response based on the analyzed question.
[0013] A "return mechanism" is a communication mechanism or process for returning the generated response to the user.
[0014] "Sentiment analysis" is the process of identifying emotional nuances and tones in a user's question and generating an appropriate response based on that.
[0015] The "means for generating responses in an interactive manner" is a process for creating responses in a manner that follows the natural flow of conversation with the user. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system that receives questions about products and services from users, analyzes them using a natural language processing engine, and generates appropriate responses using a generative AI model. Specific program processing and examples are described below.
[0038] System Configuration
[0039] server
[0040] The server receives a question from the user, analyzes it using a natural language processing engine, and then generates an appropriate response using a generative AI model based on the analysis results and sends it back to the user. The server's main functions are:
[0041] Receiving questions:
[0042] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0043] Natural language analysis:
[0044] The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[0045] Response generation:
[0046] The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for delivery back to the user.
[0047] Sending a response:
[0048] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0049] Terminal
[0050] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0051] Capturing input:
[0052] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0053] Sending an API request:
[0054] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0055] Receiving and displaying the response:
[0056] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0057] Specific examples
[0058] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0059] 1. User:
[0060] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0061] 2. Terminal:
[0062] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0063] 3. Server:
[0064] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[0065] 4. Server:
[0066] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0067] 5. Server:
[0068] The generated response is sent back to the terminal in JSON format.
[0069] 6. Terminal:
[0070] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0071] In this way, users can quickly and accurately obtain information to resolve their issues. This system improves the customer experience by eliminating waiting times and trips to physical stores. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[0072] The processing flow will be explained below.
[0073] Step 1:
[0074] User: The user uses a device such as a smartphone or PC to enter a question into an input form. Example: "My battery is draining quickly. What should I do?"
[0075] Step 2:
[0076] Terminal: When the user presses the "Submit" button, the terminal captures this input and creates an API request, formats the question in JSON format, and sends it as an HTTP POST request to the server's API endpoint.
[0077] Step 3:
[0078] Server: The server receives the API request sent from the device, parses the JSON data included in the request body, and extracts the user's question.
[0079] Step 4:
[0080] Server: The server sends the extracted questions to a natural language processing engine, which analyzes the questions to understand their main content and sentiment, and obtains the analysis results.
[0081] Step 5:
[0082] Server: Based on the results of natural language analysis, the server inputs the analysis results into a generative AI model and generates a response. Example: "We recommend that you check which apps are using a lot of battery and delete any unnecessary apps."
[0083] Step 6:
[0084] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[0085] Step 7:
[0086] Device: The device receives the response sent back from the server. It analyzes the response content and displays it on the screen in a user-friendly format. Example: "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0087] Step 8:
[0088] User: The user checks the response displayed on the device and follows the instructions to take appropriate measures to resolve the problem.
[0089] Through the above steps, the system responds promptly and appropriately to the user's questions and provides information for resolving the problem.
[0090] Example 1
[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0092] Modern consumers expect quick responses to general product and service-related questions and problem-solving, necessitating an efficient and effective support system for end users. However, traditional inquiry systems have limitations in accurately analyzing user intent and emotions and providing appropriate responses, which can lead to a decline in customer satisfaction. In addition, individual responses require significant resources, creating cost challenges.
[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0094] In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for generating a response using a generative AI model based on the analysis result, means for a terminal to capture the received question, convert it into JSON format, and send it to the server, means for the server to receive the converted data and generate a response based on the analysis result, and means for returning the generated response to the user and for the terminal to display the response to the user, thereby enabling the user to quickly and accurately obtain information for solving their problem.
[0095] The "means for receiving questions" is a function for receiving questions about products or services sent by users from the terminal.
[0096] A "natural language processing engine" is software that analyzes questions received from users and understands their intentions and emotions.
[0097] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on the results analyzed by a natural language processing engine.
[0098] The "means for generating a response" is a function that uses a generative AI model to create a response based on the analysis results.
[0099] The "means for capturing questions" is a function for recognizing questions entered by users into the terminal and capturing them as data.
[0100] The "means for converting into JSON format" is a function for converting the captured user question into JSON format in order to send the data to the server.
[0101] An "API endpoint" is a specific URL or URI on a server that a device accesses to send a question or inquiry to the server.
[0102] The "means for receiving data" is a function that allows the server to receive data sent from the terminal.
[0103] A "terminal" is a device that a user uses to enter and display questions and responses.
[0104] A "user" is someone who enters a question about a product or service.
[0105] This invention is a system that receives questions about products and services from users, analyzes them with a natural language processing engine, generates appropriate responses using a generative AI model, and returns them to the users. Specific program processing and examples are described below.
[0106] System Configuration
[0107] server
[0108] The server receives a question from the user, analyzes the question using a natural language processing engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are:
[0109] Receiving questions:
[0110] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0111] Natural language analysis:
[0112] The server analyzes the received question using a natural language processing engine (e.g., Google NLP or SpaCy) to understand the user's intent and emotions. The results of this analysis are used to generate a response.
[0113] Response generation:
[0114] The server uses a generative AI model (e.g., OpenAI GPT-4) based on the parsed question to generate an appropriate response.
[0115] Sending a response:
[0116] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0117] Terminal
[0118] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0119] Capturing input:
[0120] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0121] Sending an API request:
[0122] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0123] Receiving and displaying the response:
[0124] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0125] Specific examples
[0126] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0127] 1. User:
[0128] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0129] 2. Terminal:
[0130] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0131] 3. Server:
[0132] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[0133] 4. Server:
[0134] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0135] 5. Server:
[0136] The generated response is sent back to the terminal in JSON format.
[0137] 6. Terminal:
[0138] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0139] This system allows users to quickly and accurately obtain information to resolve their issues, eliminating waiting times and the hassle of visiting a physical store, improving the customer experience. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[0140] Examples of prompt statements
[0141] Here are some examples of prompts:
[0142] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the submit button.
[0143] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0144] Step 1:
[0145] User:
[0146] The user enters a question into the input form on the device and presses the send button. For example, the user might enter, "My battery is draining quickly. What should I do?"
[0147] Input: The question the user entered into the input form.
[0148] Output: None.
[0149] Specific operation: The user enters their question or problem directly into the input form on the device.
[0150] Step 2:
[0151] Device:
[0152] The device captures the questions entered by the user, converts them into JSON format, and sends them as an HTTP POST request to the server's API endpoint.
[0153] Input: Questions entered by the user in the input form.
[0154] Output: The converted questions in JSON format.
[0155] Specific operation: The terminal captures the entered question, converts it into the appropriate format (JSON), and sends it to the server.
[0156] Step 3:
[0157] server:
[0158] The server receives the HTTP POST request sent from the terminal.
[0159] Input: Question sent from the terminal in JSON format.
[0160] Output: The received question data.
[0161] Specific operation: The server receives the question sent from the terminal and prepares for analysis.
[0162] Step 4:
[0163] server:
[0164] The server analyzes the received question using a natural language processing engine, specifically by performing grammatical analysis, keyword extraction, and intent recognition.
[0165] Input: Received question data.
[0166] Output: Analysis results (including user intent and sentiment).
[0167] Specific operation: The natural language processing engine analyzes the question and extracts important keywords and the user's intent.
[0168] Step 5:
[0169] server:
[0170] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[0171] Input: Analysis results of the natural language processing engine.
[0172] Output: The response generated by the generative AI model.
[0173] Specific operation: Based on the analysis results, the generative AI model generates an appropriate response sentence.
[0174] Step 6:
[0175] server:
[0176] The server converts the generated response into JSON format and sends it to the terminal as an HTTP response.
[0177] Input: The generated response.
[0178] Output: The response formatted as JSON.
[0179] Specific operation: The server converts the generated response into an appropriate format and sends it to the terminal.
[0180] Step 7:
[0181] Device:
[0182] The terminal receives the response sent by the server and displays it to the user.
[0183] Input: The response sent by the server in JSON format.
[0184] Output: The response that is displayed to the user.
[0185] Specific operation: The terminal analyzes the received response and displays it in a format that is easy for the user to understand.
[0186] This series of steps allows users to quickly and accurately find a solution to their question.
[0187] (Application example 1)
[0188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0189] In modern brick-and-mortar stores, customers need to be able to quickly and accurately resolve their product and service-related questions. However, traditional brick-and-mortar stores have a limited number of customer service representatives, making it difficult to respond to all customers in real time. This has led to problems such as a decline in customer satisfaction and an increased burden on store staff. It is also difficult to guarantee the accuracy and consistency of the information provided to customers. To solve these problems, there is a need to provide effective technology.
[0190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0191] In this invention, the server includes a means for receiving questions about products or services from users, a means for analyzing the questions using a natural language processing engine, and a means for generating responses using a generative AI model based on the analysis results, thereby enabling the server to provide quick and accurate responses to customer questions.
[0192] "Products and services" means goods or activities offered for commercial purposes that are purchased or used by consumers.
[0193] "User" means any person or entity that uses the System or Services.
[0194] "Means for receiving" refers to the device or process used to obtain information or data from a user.
[0195] A "natural language processing engine" refers to a computer program or algorithm that analyzes human language and understands its meaning and structure.
[0196] "Means for analysis" refers to the methods and techniques used to break down received data and extract meaning and significance.
[0197] A "generative AI model" refers to an artificial intelligence algorithm or program that generates appropriate responses or suggestions based on training data.
[0198] "Means for generating a response" refers to the methods and techniques for generating an answer to a question based on the analysis results.
[0199] "Returning means" refers to the method or process for sending the generated response to the user.
[0200] "Means for generating API requests" refers to the code and protocols used to send requests to a server through an application interface.
[0201] "User terminal" refers to an electronic device used by a user, such as a smartphone or tablet.
[0202] "Means for displaying" refers to a display device and its control program for visually presenting information to a user.
[0203] A "server" refers to a central processing unit that receives requests from clients via a network and returns responses.
[0204] This invention provides a system for quickly and accurately responding to questions about products and services when a user asks a question in a physical store. This system includes a server and a user terminal (e.g., a smartphone) and operates in the following steps.
[0205] System Configuration
[0206] server
[0207] The server receives a question from the user and analyzes it using a natural language processing engine. It then generates an appropriate response using a generative AI model based on the analysis results. The generated response is then sent back to the user's device. The server's main functions are as follows:
[0208] Receiving a question: The server receives a question sent from the user's terminal. The question is sent as an HTTP POST request.
[0209] Natural language analysis: The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[0210] Response Generation: The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for sending back to the user.
[0211] Sending a response: The server sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0212] Terminal
[0213] The user terminal captures input from the user, sends it to the server, and displays the responses sent back from the server. The terminal's main functions are:
[0214] Input capture: When a user enters a question into an input form and presses the submit button, the device captures this input.
[0215] Send API request: Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0216] Receive and display the response: Receive the response sent back from the server and display it on the screen in a user-friendly format.
[0217] Specific examples
[0218] For example, if a user types the question "What material is this shirt made of?", the question is processed as follows:
[0219] 1. User: The user enters "What material is this shirt made of?" into the input form on the device and presses the send button.
[0220] 2. Terminal: The terminal captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0221] 3. Server: The server receives the query and analyzes it using a natural language processing engine. The keywords "shirt" and "material" are identified as the analysis results.
[0222] 4. Server: Then, using a generative AI model, it generates a response such as "This shirt is made of 100% cotton."
[0223] 5. Server: Returns the generated response to the user terminal.
[0224] 6. Terminal: The terminal receives the response from the server and displays "This shirt is made of 100% cotton" on the user screen.
[0225] Prompt Sentence Examples
[0226] For example, if a user asks, "Does this shampoo contain any allergens?", the generative AI model might generate a prompt like this:
[0227] Question analysis:
[0228] Product: Shampoo
[0229] Question: Are there any ingredients that can cause allergies?
[0230] Produces response: This shampoo is allergen-free. However, if you would like to know more about the ingredients, please see the back of the package.
[0231] As described above, this system allows users to receive quick and accurate responses to questions about products and services in physical stores.
[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0233] Step 1:
[0234] The user enters a question into the input form on the terminal and presses the send button, which causes the terminal to capture the user's question.
[0235] Input: User question (e.g., "What material is this shirt made of?")
[0236] Output: Captured question information
[0237] Specific operation: The user enters a question into the input field on the device's UI and presses the "Send" button.
[0238] Step 2:
[0239] The device converts the captured questions into JSON format and sends it as an HTTP POST request to the server's API endpoint.
[0240] Input: Captured question information
[0241] Output: HTTP POST request to the server
[0242] Specific operation: The terminal encodes the question into JSON format and sends an HTTP POST request to the server using the requests library.
[0243] Step 3:
[0244] The server receives a question from a user, which is then passed to a natural language processing engine for analysis.
[0245] Input: Question sent as an HTTP POST request
[0246] Output: Parsing request passed to the natural language processing engine
[0247] Specific operation: The server retrieves the question data received and uses the requests library to send an analysis request to the natural language processing engine.
[0248] Step 4:
[0249] The server receives the results analyzed by the natural language processing engine, which include the user's intent and keywords.
[0250] Input: Parsing request to the natural language processing engine
[0251] Output: Analysis results (e.g., keywords "shirt" and "material")
[0252] Specific behavior: Receives analysis results returned from a natural language processing engine.
[0253] Step 5:
[0254] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[0255] Input: Analysis results
[0256] Output: The generated response (e.g., "This shirt is made of 100% cotton")
[0257] Specific operation: The analysis results are input into the generative AI model as a prompt sentence to obtain an appropriate response.
[0258] Step 6:
[0259] The server returns the generated response to the user terminal.
[0260] Input: Generated response
[0261] Output: HTTP response to the user's device
[0262] Specific operation: The generated response is encoded in JSON format and sent to the user's terminal as an HTTP response.
[0263] Step 7:
[0264] The terminal receives the response sent back from the server and displays it on the user screen.
[0265] Input: HTTP response from the server
[0266] Output: The displayed response (e.g., "This shirt is made of 100% cotton")
[0267] Specific behavior: Displaying the received response in the UI (e.g., displaying it in a text box or alert).
[0268] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0269] This invention is a system that receives questions about products or services from users, analyzes the emotions contained in the questions using a natural language processing engine and an emotion engine, and generates a response using a generative AI model based on the analysis results. Specific program processing and examples are described below.
[0270] System Configuration
[0271] server
[0272] The server receives questions from users, analyzes them using a natural language processing engine and an emotion engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are as follows:
[0273] Receiving questions:
[0274] The server receives a question sent from the user's device, which is sent as an HTTP POST request.
[0275] Natural language analysis:
[0276] The server analyzes the received question using a natural language processing engine to understand the user's intent and emotions, and the analysis results are used to generate a response.
[0277] Emotion analysis:
[0278] The emotion engine recognizes the user's emotions in response to questions analyzed by the natural language processing engine. The results of the emotion analysis are used to generate more detailed responses.
[0279] Response generation:
[0280] The server uses a generative AI model based on the results of natural language analysis and sentiment analysis to generate an appropriate response, such as "We recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0281] Sending a response:
[0282] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response.
[0283] Terminal
[0284] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0285] Capturing input:
[0286] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0287] Sending an API request:
[0288] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0289] Receiving and displaying the response:
[0290] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0291] Specific examples
[0292] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0293] 1. User:
[0294] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0295] 2. Terminal:
[0296] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0297] 3. Server:
[0298] The server receives the question and analyzes it using a natural language processing engine and an emotion engine. The natural language processing engine identifies the problem of "low battery," and the emotion engine analyzes the user's emotion.
[0299] 4. Server:
[0300] The generative AI model is then used to generate a response such as "We recommend you review the apps that are using a lot of battery power and delete any unnecessary apps." This response is adjusted based on the user's emotions.
[0301] 5. Server:
[0302] The generated response is sent back to the terminal in JSON format.
[0303] 6. Terminal:
[0304] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0305] In this way, users can quickly and accurately obtain information to solve their problems. This system improves the customer experience by eliminating waiting times and the hassle of visiting a physical store. It also reduces the burden on counter staff at physical stores, enabling more efficient support. The addition of an emotion engine enables more personalized responses, further improving user satisfaction.
[0306] The processing flow will be explained below.
[0307] Step 1:
[0308] User: The user uses a device such as a smartphone or PC to enter a question into an input form. For example, they might enter, "My battery is draining quickly. What should I do?"
[0309] Step 2:
[0310] Terminal: When the user presses the "Submit" button, the terminal captures this input, converts the captured question into JSON format, and sends it as an HTTP POST request to the server's API endpoint. The content of the request is JSON data containing the question text.
[0311] Step 3:
[0312] Server: The server receives the API request sent from the device, parses the body of the received request, and extracts the user's question.
[0313] Step 4:
[0314] Server: The server sends the extracted question to a natural language processing engine, which analyzes the intent and content of the question. As a result of the analysis, the subject of the question and related keywords are identified.
[0315] Step 5:
[0316] Server: Based on the parsed question, the emotion engine is used to recognize the user's emotions, for example, emotions such as "anxiety" or "dissatisfaction."
[0317] Step 6:
[0318] Server: Based on the analysis results of the emotion engine, the analysis results (question content and emotion) are input into the generative AI model and a response is generated. The generated response takes into account the user's emotions, and might be something like, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0319] Step 7:
[0320] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[0321] Step 8:
[0322] Device: The device receives the response sent back from the server. It parses the received JSON data and displays it on the screen in a user-friendly format. The displayed message is, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0323] Step 9:
[0324] User: The user checks the response displayed on the device and takes the suggested action. For example, the user opens the smartphone settings, checks which apps are using a lot of battery power, and resolves the battery issue by deleting unnecessary apps.
[0325] In this way, a series of steps is completed: the user's question is captured on the device, sent to the server, analyzed by the natural language processing engine and emotion engine, an appropriate response is generated by the generative AI model, and the response is sent back to the device. Through specific actions at each step, the user can quickly and effectively obtain information to solve their problem.
[0326] Example 2
[0327] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0328] Conventional systems have had difficulty generating appropriate responses to user questions. This is due to the lack of accurate analysis of the question content or the provision of personalized responses that reflect the user's emotions. This has made it difficult to improve user satisfaction and provide efficient support.
[0329] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0330] In this invention, the server includes means for capturing input from a user, converting it into JSON format, and transmitting it, means for using a natural language processing engine to analyze the question received from the user, means for using an emotion engine to analyze emotions based on the analysis results obtained by the natural language processing, means for generating a response using a generative AI model based on the results of the natural language analysis and the emotion analysis, and means for converting the generated response into JSON format and returning it to the user. This makes it possible to accurately analyze the content of the user's question and quickly provide a personalized response that reflects the user's emotions.
[0331] "Means for capturing user input, converting it into JSON format, and sending it" refers to a method for obtaining information entered by a user into an input form, converting it into the JSON format commonly used for data exchange, and sending it to a server.
[0332] The "means using a natural language processing engine" is a software module that analyzes questions and text data received from users to understand their meaning and intent.
[0333] The "means for using an emotion engine" is a software module for identifying and analyzing a user's emotion from text data analyzed by a natural language processing engine.
[0334] "Means for generating responses using a generative AI model" refers to a method for automatically generating appropriate responses using AI technology based on the results of natural language analysis and sentiment analysis.
[0335] "Means for converting the generated response into JSON format and returning it to the user" refers to a method for converting the response generated by the generative AI model into JSON format and returning it to the user.
[0336] "Means for inputting a prompt sentence and generating a response in an interactive format" refers to a method for inputting a prompt sentence containing a question or instruction to an AI model and generating a natural, interactive response based on the prompt.
[0337] This invention is a system that receives questions about products and services from users, analyzes the content and sentiment contained in the questions, and generates responses using a generative AI model based on the analysis results. The specific hardware and software configurations and processing of this system are described below.
[0338] System Configuration
[0339] server
[0340] The server receives queries from users, analyzes their content, and generates and sends back appropriate responses. It has the following main functions:
[0341] 1. Receiving Questions:
[0342] The server receives questions sent from the user's device as HTTP POST requests using common web server technologies (e.g., Node.js and the Express framework).
[0343] 2. Natural language analysis:
[0344] The received question is analyzed using a natural language processing engine, using the Google Cloud Natural Language API to extract the user's intent and keywords.
[0345] 3. Emotion analysis:
[0346] The system uses an emotion engine to recognize the user's emotions in response to questions analyzed by a natural language processing engine. An example of an emotion engine used is the IBM Watson Tone Analyzer.
[0347] 4. Response Generation:
[0348] Based on the results of natural language analysis and sentiment analysis, an appropriate response is generated using a generative AI model such as OpenAI GPT-4.
[0349] 5. Sending the response:
[0350] The generated response is converted to JSON format and sent back to the terminal as an HTTP response.
[0351] Terminal
[0352] The terminal is responsible for capturing input from the user, sending it to the server, and displaying the responses sent back from the server. It has the following main functions:
[0353] 1. Capturing input:
[0354] When a user enters a question into the input form and presses the submit button, the terminal captures this input. The form, which runs on a web browser, is implemented using JavaScript.
[0355] 2. Sending API requests:
[0356] The captured input is converted to JSON format and sent as an HTTP POST request to the server's API endpoint. A common method for sending requests is to use the JavaScript "axios" library.
[0357] 3. Receiving and displaying responses:
[0358] It receives the response sent back from the server and displays it on the screen in a user-friendly format. You can use "Vue.js" or "React" for front-end processing.
[0359] Specific examples
[0360] For example, if a user types the question "My battery is draining quickly. What should I do?", here's what happens:
[0361] 1. User:
[0362] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0363] 2. Terminal:
[0364] The device captures the question, converts it into JSON format, and sends it to the server's API endpoint.
[0365] 3. Server:
[0366] The server receives the question, uses a natural language processing engine to identify the problem of "low battery," and uses an emotion engine to analyze the user's emotions, such as "dissatisfaction" or "confusion."
[0367] 4. Server:
[0368] The generative AI model generates a response such as "Check which apps are using a lot of battery power and recommend deleting any unnecessary apps." This response is adjusted based on the user's emotions.
[0369] 5. Server:
[0370] The generated response is sent back to the terminal in JSON format.
[0371] 6. Terminal:
[0372] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0373] Prompt Sentence Examples
[0374] User Question: "My battery is draining quickly. What should I do?"
[0375] Prompt for generative AI model: "A user complains that their battery is draining too quickly. What would be a good solution?"
[0376] This system allows users to quickly and accurately obtain information to solve their problems, and by adding sentiment analysis, it enables more personalized responses, improving user satisfaction.
[0377] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0378] Step 1: User enters question
[0379] Input: The user inputs a question into the input form on the terminal.
[0380] Specific actions: The user opens a browser on their smartphone or computer, accesses the support page, types in "My battery is draining quickly. What should I do?", and presses the send button.
[0381] Output: The user sends the question entered in the input form to the terminal.
[0382] Step 2: The device captures the question and sends it to the server
[0383] Input: The question the user entered into the input form.
[0384] Specific operation: The device uses JavaScript to retrieve questions from the input form, convert them to JSON format, and then sends the JSON data as an HTTP POST request to the server's API endpoint, using the "axios" library, for example.
[0385] Output: The device sends the question data in JSON format to the server.
[0386] Step 3: The server receives the query
[0387] Input: Question data in JSON format sent from the terminal.
[0388] How it works: The server sets up an endpoint that listens for HTTP POST requests, and when a request arrives, it extracts the question data in JSON format from the request body. This process is done using Node.js and Express, for example.
[0389] Output: The server receives the question data in JSON format and stores it as parseable text data.
[0390] Step 4: The server performs natural language analysis
[0391] Input: Text data of the received question.
[0392] How it works: The server calls the Google Cloud Natural Language API and sends the question data. The natural language processing engine analyzes the text and extracts the user's intent and keywords. For example, the keyword "low battery" is identified.
[0393] Output: The server receives the analysis results from the natural language processing engine and obtains data including the user's intent and important keywords.
[0394] Step 5: The server performs sentiment analysis
[0395] Input: Analysis results obtained from the natural language processing engine.
[0396] How it works: The server calls the IBM Watson Tone Analyzer API and sends the analysis results. The emotion engine identifies emotions in the text and detects emotions such as "frustrated" or "confused."
[0397] Output: The server receives the emotion analysis results and obtains data that indicates the user's emotional state.
[0398] Step 6: Server Generates Response
[0399] Input: Results of natural language analysis and sentiment analysis.
[0400] Specific operation: The server calls the OpenAI GPT-4 API and sends a prompt message. The prompt message is in the form of "The user complains that the battery is draining quickly. Please tell us what to do." Based on this prompt, the generative AI model generates a response such as "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0401] Output: The server receives the response text from the generative AI model and stores it.
[0402] Step 7: The server converts the response into JSON format and sends it back to the device.
[0403] Input: The generated response text.
[0404] Specific operation: The server formats the response text into JSON format and sends the formatted JSON data to the terminal as an HTTP response.
[0405] Output: The server sends the generated response back to the device in JSON format.
[0406] Step 8: The device receives the response from the server and displays it to the user
[0407] Input: JSON formatted response data returned from the server.
[0408] Specific operation: The device receives the HTTP response and parses the JSON data. The parsed response is inserted into an HTML element and displayed on the user's screen. For example, the response might say, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0409] Output: The terminal displays the parsed response to the user.
[0410] (Application example 2)
[0411] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0412] In recent years, in online shopping and virtual stores, users expect their questions and problems to be resolved quickly and accurately. However, conventional systems have difficulty providing appropriate responses to user questions, particularly in terms of personalized responses that reflect the user's emotions. Furthermore, voice interfaces are inadequate, making it difficult to achieve natural interactions, especially through wearable devices such as smart glasses. This has resulted in the challenge of not fully improving the customer experience.
[0413] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for analyzing the emotion of the question using an emotion analysis engine, means for generating a response using a generative AI model based on the analysis result, means for returning the generated response to the user, means for capturing the user's question as voice and converting it into text using voice recognition technology, and means for notifying the user of the generated response visually or audibly. This enables real-time analysis and response to the user's question, and further enables generation of a personalized response according to the emotion, thereby increasing user satisfaction.
[0414] The "means for receiving questions about products and services from users" is a function that allows users to input questions about products and services into the system and transmit them to the server.
[0415] "Means for analysis using a natural language processing engine" refers to a function that analyzes text data received from a user using machine learning and statistical methods to understand the intent and content of the question.
[0416] The "means for analyzing emotions using an emotion analysis engine" is a function for analyzing the emotions contained in the user's question text and determining whether the emotion corresponds to positive, negative, neutral, or the like.
[0417] "Means for generating responses using a generative AI model" refers to a function for using an artificial intelligence model that automatically generates appropriate responses based on the results of natural language processing and sentiment analysis.
[0418] The "means for returning to the user" is a function for sending the generated response to the user and displaying or reproducing it on the screen or through audio.
[0419] "Means for converting to text using speech recognition technology" is a function for using a speech recognition algorithm to convert a question input by a user into text data.
[0420] "Visual or audio notification means" refers to a function for visually displaying or audibly playing the generated response to the user.
[0421] In order to put the present invention into practice, it is necessary to build a system in which the elements of the user, terminal, and server function in cooperation with each other.
[0422] Program Overview
[0423] The server has the following features:
[0424] 1. Receiving a question: This is a function to receive a question sent from the user's terminal. This question is received as voice or text.
[0425] 2. Natural Language Analysis: This function analyzes the received question using a natural language processing engine to understand the user's intent. Natural language processing technologies such as Google Cloud Natural Language API are used for the analysis.
[0426] 3. Sentiment analysis: This function analyzes the emotions contained in questions using a sentiment analysis engine. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[0427] 4. Response Generation: This function generates an appropriate response using a generative AI model (such as OpenAI GPT-3) based on the analysis results.
[0428] 5. Sending the response: This function converts the generated response into JSON format and sends it to the user's device.
[0429] The terminal has the following features:
[0430] 1. Input capture: When a user types a question by voice, the smart glasses' microphone is used to capture the voice and convert it into text using voice recognition technology (such as Google Cloud Speech-to-Text).
[0431] 2. Send API request: The converted text is converted into JSON format and sent as an HTTP POST request to the server's API endpoint.
[0432] 3. Receiving and displaying response: The response sent back from the server is received and displayed on the smart glasses display for the user to visually confirm, or output as an audio output.
[0433] Specific examples
[0434] For example, if a user uses smart glasses and asks, "I'm not sleeping well these days. How can I improve it?", the following steps will occur:
[0435] 1. Input capture: Audio is captured using the microphone on the smart glasses and converted to text using Google Cloud Speech-to-Text.
[0436] 2. Natural Language Analysis: We use the Google Cloud Natural Language API to analyze the question "Poor quality of sleep."
[0437] 3. Sentiment Analysis: Analyze emotions using IBM Watson Tone Analyzer to identify user frustrations and concerns.
[0438] 4. Response Generation: Using Open AI GPT-3, we generate a response like, "First, I recommend reducing your screen time before bed and trying yoga or meditation as a way to relax."
[0439] 5. Send and display response: The generated response is sent back to the terminal in JSON format and displayed on the smart glasses or read aloud.
[0440] Prompt Sentence Examples
[0441] If the user's question is "I'm not sleeping well lately. What can I do to improve it?", the prompt might look like this:
[0442] User's query: I've been having trouble sleeping lately. How can I improve it?
[0443] Sentiment: -0.3
[0444] Emotions: {"anger": 0.1, "sadness": 0.5, "joy": 0.2}
[0445] Response:
[0446] As described above, the system works in cooperation with the server and the terminal, generating appropriate responses to user questions in real time. This system improves the user's customer experience and provides efficient support.
[0447] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0448] Step 1:
[0449] The user uses the smart glasses to input a question by voice, and the user's voice is captured by the microphone of the smart glasses. The input voice data is collected.
[0450] Step 2:
[0451] The device converts the captured audio into text using the Google Cloud Speech-to-Text API. During this conversion process, the audio data is converted into text data, and the user's question is obtained in text format.
[0452] Step 3:
[0453] The terminal converts the converted text data into JSON format and sends it to the server's API endpoint as an HTTP POST request. The input information is text data, and a JSON format request is generated as the output.
[0454] Step 4:
[0455] The server analyzes the question received from the device using a natural language processing engine (for example, Google Cloud Natural Language API). During this analysis, data calculations are performed to understand the intent and content of the question. The analysis results are output as data indicating the intent.
[0456] Step 5:
[0457] The server analyzes emotions using a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) based on the analysis results of the natural language processing engine. The input is data containing intent, and the output is a sentiment score. During this process, data calculations are performed to identify sentiment categories (positive, negative, neutral, etc.).
[0458] Step 6:
[0459] The server combines the results of natural language processing and sentiment analysis to generate a response using a generative AI model (e.g., OpenAI GPT-3). The input is data containing intent and sentiment scores, and the output is an appropriate response text. The generated response is then processed based on the prompt.
[0460] Step 7:
[0461] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response. The input is the generated response text and the output is the JSON formatted response.
[0462] Step 8:
[0463] The device parses the response received from the server and displays or reads it out loud to the user in a user-friendly format. The input is the response data in JSON format, and the output is visual or audio feedback. During this process, the data is reformatted to make it easier for the user to understand.
[0464] These steps result in a system that generates real-time responses to user questions and provides feedback in an appropriate format.
[0465] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0466] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0467] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0468] [Second embodiment]
[0469] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0470] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0471] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0472] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0473] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0474] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0475] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0476] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0477] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0478] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0479] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0480] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0481] This invention is a system that receives questions about products and services from users, analyzes them using a natural language processing engine, and generates appropriate responses using a generative AI model. Specific program processing and examples are described below.
[0482] System Configuration
[0483] server
[0484] The server receives a question from the user, analyzes it using a natural language processing engine, and then generates an appropriate response using a generative AI model based on the analysis results and sends it back to the user. The server's main functions are:
[0485] Receiving questions:
[0486] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0487] Natural language analysis:
[0488] The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[0489] Response generation:
[0490] The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for delivery back to the user.
[0491] Sending a response:
[0492] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0493] Terminal
[0494] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0495] Capturing input:
[0496] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0497] Sending an API request:
[0498] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0499] Receiving and displaying the response:
[0500] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0501] Specific examples
[0502] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0503] 1. User:
[0504] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0505] 2. Terminal:
[0506] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0507] 3. Server:
[0508] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[0509] 4. Server:
[0510] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0511] 5. Server:
[0512] The generated response is sent back to the terminal in JSON format.
[0513] 6. Terminal:
[0514] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0515] In this way, users can quickly and accurately obtain information to resolve their issues. This system improves the customer experience by eliminating waiting times and trips to physical stores. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[0516] The processing flow will be explained below.
[0517] Step 1:
[0518] User: The user uses a device such as a smartphone or PC to enter a question into an input form. Example: "My battery is draining quickly. What should I do?"
[0519] Step 2:
[0520] Terminal: When the user presses the "Submit" button, the terminal captures this input and creates an API request, formats the question in JSON format, and sends it as an HTTP POST request to the server's API endpoint.
[0521] Step 3:
[0522] Server: The server receives the API request sent from the device, parses the JSON data included in the request body, and extracts the user's question.
[0523] Step 4:
[0524] Server: The server sends the extracted questions to a natural language processing engine, which analyzes the questions to understand their main content and sentiment, and obtains the analysis results.
[0525] Step 5:
[0526] Server: Based on the results of natural language analysis, the server inputs the analysis results into a generative AI model and generates a response. Example: "We recommend that you check which apps are using a lot of battery and delete any unnecessary apps."
[0527] Step 6:
[0528] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[0529] Step 7:
[0530] Device: The device receives the response sent back from the server. It analyzes the response content and displays it on the screen in a user-friendly format. Example: "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0531] Step 8:
[0532] User: The user checks the response displayed on the device and follows the instructions to take appropriate measures to resolve the problem.
[0533] Through the above steps, the system responds promptly and appropriately to the user's questions and provides information for resolving the problem.
[0534] Example 1
[0535] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0536] Modern consumers expect quick responses to general product and service-related questions and problem-solving, necessitating an efficient and effective support system for end users. However, traditional inquiry systems have limitations in accurately analyzing user intent and emotions and providing appropriate responses, which can lead to a decline in customer satisfaction. In addition, individual responses require significant resources, creating cost challenges.
[0537] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0538] In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for generating a response using a generative AI model based on the analysis result, means for a terminal to capture the received question, convert it into JSON format, and send it to the server, means for the server to receive the converted data and generate a response based on the analysis result, and means for returning the generated response to the user and for the terminal to display the response to the user, thereby enabling the user to quickly and accurately obtain information for solving their problem.
[0539] The "means for receiving questions" is a function for receiving questions about products or services sent by users from the terminal.
[0540] A "natural language processing engine" is software that analyzes questions received from users and understands their intentions and emotions.
[0541] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on the results analyzed by a natural language processing engine.
[0542] The "means for generating a response" is a function that uses a generative AI model to create a response based on the analysis results.
[0543] The "means for capturing questions" is a function for recognizing questions entered by users into the terminal and capturing them as data.
[0544] The "means for converting into JSON format" is a function for converting the captured user question into JSON format in order to send the data to the server.
[0545] An "API endpoint" is a specific URL or URI on a server that a device accesses to send a question or inquiry to the server.
[0546] The "means for receiving data" is a function that allows the server to receive data sent from the terminal.
[0547] A "terminal" is a device that a user uses to enter and display questions and responses.
[0548] A "user" is someone who enters a question about a product or service.
[0549] This invention is a system that receives questions about products and services from users, analyzes them with a natural language processing engine, generates appropriate responses using a generative AI model, and returns them to the users. Specific program processing and examples are described below.
[0550] System Configuration
[0551] server
[0552] The server receives a question from the user, analyzes the question using a natural language processing engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are:
[0553] Receiving questions:
[0554] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0555] Natural language analysis:
[0556] The server analyzes the received question using a natural language processing engine (e.g., Google NLP or SpaCy) to understand the user's intent and emotions. The results of this analysis are used to generate a response.
[0557] Response generation:
[0558] The server uses a generative AI model (e.g., OpenAI GPT-4) based on the parsed question to generate an appropriate response.
[0559] Sending a response:
[0560] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0561] Terminal
[0562] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0563] Capturing input:
[0564] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0565] Sending an API request:
[0566] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0567] Receiving and displaying the response:
[0568] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0569] Specific examples
[0570] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0571] 1. User:
[0572] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0573] 2. Terminal:
[0574] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0575] 3. Server:
[0576] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[0577] 4. Server:
[0578] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0579] 5. Server:
[0580] The generated response is sent back to the terminal in JSON format.
[0581] 6. Terminal:
[0582] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0583] This system allows users to quickly and accurately obtain information to resolve their issues, eliminating waiting times and the hassle of visiting a physical store, improving the customer experience. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[0584] Examples of prompt statements
[0585] Here are some examples of prompts:
[0586] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the submit button.
[0587] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0588] Step 1:
[0589] User:
[0590] The user enters a question into the input form on the device and presses the send button. For example, the user might enter, "My battery is draining quickly. What should I do?"
[0591] Input: The question the user entered into the input form.
[0592] Output: None.
[0593] Specific operation: The user enters their question or problem directly into the input form on the device.
[0594] Step 2:
[0595] Device:
[0596] The device captures the questions entered by the user, converts them into JSON format, and sends them as an HTTP POST request to the server's API endpoint.
[0597] Input: Questions entered by the user in the input form.
[0598] Output: The converted questions in JSON format.
[0599] Specific operation: The terminal captures the entered question, converts it into the appropriate format (JSON), and sends it to the server.
[0600] Step 3:
[0601] server:
[0602] The server receives the HTTP POST request sent from the terminal.
[0603] Input: Question sent from the terminal in JSON format.
[0604] Output: The received question data.
[0605] Specific operation: The server receives the question sent from the terminal and prepares for analysis.
[0606] Step 4:
[0607] server:
[0608] The server analyzes the received question using a natural language processing engine, specifically by performing grammatical analysis, keyword extraction, and intent recognition.
[0609] Input: Received question data.
[0610] Output: Analysis results (including user intent and sentiment).
[0611] Specific operation: The natural language processing engine analyzes the question and extracts important keywords and the user's intent.
[0612] Step 5:
[0613] server:
[0614] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[0615] Input: Analysis results of the natural language processing engine.
[0616] Output: The response generated by the generative AI model.
[0617] Specific operation: Based on the analysis results, the generative AI model generates an appropriate response sentence.
[0618] Step 6:
[0619] server:
[0620] The server converts the generated response into JSON format and sends it to the terminal as an HTTP response.
[0621] Input: The generated response.
[0622] Output: The response formatted as JSON.
[0623] Specific operation: The server converts the generated response into an appropriate format and sends it to the terminal.
[0624] Step 7:
[0625] Device:
[0626] The terminal receives the response sent by the server and displays it to the user.
[0627] Input: The response sent by the server in JSON format.
[0628] Output: The response that is displayed to the user.
[0629] Specific operation: The terminal analyzes the received response and displays it in a format that is easy for the user to understand.
[0630] This series of steps allows users to quickly and accurately find a solution to their question.
[0631] (Application example 1)
[0632] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0633] In modern brick-and-mortar stores, customers need to be able to quickly and accurately resolve their product and service-related questions. However, traditional brick-and-mortar stores have a limited number of customer service representatives, making it difficult to respond to all customers in real time. This has led to problems such as a decline in customer satisfaction and an increased burden on store staff. It is also difficult to guarantee the accuracy and consistency of the information provided to customers. To solve these problems, there is a need to provide effective technology.
[0634] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0635] In this invention, the server includes a means for receiving questions about products or services from users, a means for analyzing the questions using a natural language processing engine, and a means for generating responses using a generative AI model based on the analysis results, thereby enabling the server to provide quick and accurate responses to customer questions.
[0636] "Products and services" means goods or activities offered for commercial purposes that are purchased or used by consumers.
[0637] "User" means any person or entity that uses the System or Services.
[0638] "Means for receiving" refers to the device or process used to obtain information or data from a user.
[0639] A "natural language processing engine" refers to a computer program or algorithm that analyzes human language and understands its meaning and structure.
[0640] "Means for analysis" refers to the methods and techniques used to break down received data and extract meaning and significance.
[0641] A "generative AI model" refers to an artificial intelligence algorithm or program that generates appropriate responses or suggestions based on training data.
[0642] "Means for generating a response" refers to the methods and techniques for generating an answer to a question based on the analysis results.
[0643] "Returning means" refers to the method or process for sending the generated response to the user.
[0644] "Means for generating API requests" refers to the code and protocols used to send requests to a server through an application interface.
[0645] "User terminal" refers to an electronic device used by a user, such as a smartphone or tablet.
[0646] "Means for displaying" refers to a display device and its control program for visually presenting information to a user.
[0647] A "server" refers to a central processing unit that receives requests from clients via a network and returns responses.
[0648] This invention provides a system for quickly and accurately responding to questions about products and services when a user asks a question in a physical store. This system includes a server and a user terminal (e.g., a smartphone) and operates in the following steps.
[0649] System Configuration
[0650] server
[0651] The server receives a question from the user and analyzes it using a natural language processing engine. It then generates an appropriate response using a generative AI model based on the analysis results. The generated response is then sent back to the user's device. The server's main functions are as follows:
[0652] Receiving a question: The server receives a question sent from the user's terminal. The question is sent as an HTTP POST request.
[0653] Natural language analysis: The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[0654] Response Generation: The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for sending back to the user.
[0655] Sending a response: The server sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0656] Terminal
[0657] The user terminal captures input from the user, sends it to the server, and displays the responses sent back from the server. The terminal's main functions are:
[0658] Input capture: When a user enters a question into an input form and presses the submit button, the device captures this input.
[0659] Send API request: Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0660] Receive and display the response: Receive the response sent back from the server and display it on the screen in a user-friendly format.
[0661] Specific examples
[0662] For example, if a user types the question "What material is this shirt made of?", the question is processed as follows:
[0663] 1. User: The user enters "What material is this shirt made of?" into the input form on the device and presses the send button.
[0664] 2. Terminal: The terminal captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0665] 3. Server: The server receives the query and analyzes it using a natural language processing engine. The keywords "shirt" and "material" are identified as the analysis results.
[0666] 4. Server: Then, using a generative AI model, it generates a response such as "This shirt is made of 100% cotton."
[0667] 5. Server: Returns the generated response to the user terminal.
[0668] 6. Terminal: The terminal receives the response from the server and displays "This shirt is made of 100% cotton" on the user screen.
[0669] Prompt Sentence Examples
[0670] For example, if a user asks, "Does this shampoo contain any allergens?", the generative AI model might generate a prompt like this:
[0671] Question analysis:
[0672] Product: Shampoo
[0673] Question: Are there any ingredients that can cause allergies?
[0674] Produces response: This shampoo is allergen-free. However, if you would like to know more about the ingredients, please see the back of the package.
[0675] As described above, this system allows users to receive quick and accurate responses to questions about products and services in physical stores.
[0676] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0677] Step 1:
[0678] The user enters a question into the input form on the terminal and presses the send button, which causes the terminal to capture the user's question.
[0679] Input: User question (e.g., "What material is this shirt made of?")
[0680] Output: Captured question information
[0681] Specific operation: The user enters a question into the input field on the device's UI and presses the "Send" button.
[0682] Step 2:
[0683] The device converts the captured questions into JSON format and sends it as an HTTP POST request to the server's API endpoint.
[0684] Input: Captured question information
[0685] Output: HTTP POST request to the server
[0686] Specific operation: The terminal encodes the question into JSON format and sends an HTTP POST request to the server using the requests library.
[0687] Step 3:
[0688] The server receives a question from a user, which is then passed to a natural language processing engine for analysis.
[0689] Input: Question sent as an HTTP POST request
[0690] Output: Parsing request passed to the natural language processing engine
[0691] Specific operation: The server retrieves the question data received and uses the requests library to send an analysis request to the natural language processing engine.
[0692] Step 4:
[0693] The server receives the results analyzed by the natural language processing engine, which include the user's intent and keywords.
[0694] Input: Parsing request to the natural language processing engine
[0695] Output: Analysis results (e.g., keywords "shirt" and "material")
[0696] Specific behavior: Receives analysis results returned from a natural language processing engine.
[0697] Step 5:
[0698] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[0699] Input: Analysis results
[0700] Output: The generated response (e.g., "This shirt is made of 100% cotton")
[0701] Specific operation: The analysis results are input into the generative AI model as a prompt sentence to obtain an appropriate response.
[0702] Step 6:
[0703] The server returns the generated response to the user terminal.
[0704] Input: Generated response
[0705] Output: HTTP response to the user's device
[0706] Specific operation: The generated response is encoded in JSON format and sent to the user's terminal as an HTTP response.
[0707] Step 7:
[0708] The terminal receives the response sent back from the server and displays it on the user screen.
[0709] Input: HTTP response from the server
[0710] Output: The displayed response (e.g., "This shirt is made of 100% cotton")
[0711] Specific behavior: Displaying the received response in the UI (e.g., displaying it in a text box or alert).
[0712] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0713] This invention is a system that receives questions about products or services from users, analyzes the emotions contained in the questions using a natural language processing engine and an emotion engine, and generates a response using a generative AI model based on the analysis results. Specific program processing and examples are described below.
[0714] System Configuration
[0715] server
[0716] The server receives questions from users, analyzes them using a natural language processing engine and an emotion engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are as follows:
[0717] Receiving questions:
[0718] The server receives a question sent from the user's device, which is sent as an HTTP POST request.
[0719] Natural language analysis:
[0720] The server analyzes the received question using a natural language processing engine to understand the user's intent and emotions, and the analysis results are used to generate a response.
[0721] Emotion analysis:
[0722] The emotion engine recognizes the user's emotions in response to questions analyzed by the natural language processing engine. The results of the emotion analysis are used to generate more detailed responses.
[0723] Response generation:
[0724] The server uses a generative AI model based on the results of natural language analysis and sentiment analysis to generate an appropriate response, such as "We recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0725] Sending a response:
[0726] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response.
[0727] Terminal
[0728] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0729] Capturing input:
[0730] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0731] Sending an API request:
[0732] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0733] Receiving and displaying the response:
[0734] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0735] Specific examples
[0736] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0737] 1. User:
[0738] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0739] 2. Terminal:
[0740] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0741] 3. Server:
[0742] The server receives the question and analyzes it using a natural language processing engine and an emotion engine. The natural language processing engine identifies the problem of "low battery," and the emotion engine analyzes the user's emotion.
[0743] 4. Server:
[0744] The generative AI model is then used to generate a response such as "We recommend you review the apps that are using a lot of battery power and delete any unnecessary apps." This response is adjusted based on the user's emotions.
[0745] 5. Server:
[0746] The generated response is sent back to the terminal in JSON format.
[0747] 6. Terminal:
[0748] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0749] In this way, users can quickly and accurately obtain information to solve their problems. This system improves the customer experience by eliminating waiting times and the hassle of visiting a physical store. It also reduces the burden on counter staff at physical stores, enabling more efficient support. The addition of an emotion engine enables more personalized responses, further improving user satisfaction.
[0750] The processing flow will be explained below.
[0751] Step 1:
[0752] User: The user uses a device such as a smartphone or PC to enter a question into an input form. For example, they might enter, "My battery is draining quickly. What should I do?"
[0753] Step 2:
[0754] Terminal: When the user presses the "Submit" button, the terminal captures this input, converts the captured question into JSON format, and sends it as an HTTP POST request to the server's API endpoint. The content of the request is JSON data containing the question text.
[0755] Step 3:
[0756] Server: The server receives the API request sent from the device, parses the body of the received request, and extracts the user's question.
[0757] Step 4:
[0758] Server: The server sends the extracted question to a natural language processing engine, which analyzes the intent and content of the question. As a result of the analysis, the subject of the question and related keywords are identified.
[0759] Step 5:
[0760] Server: Based on the parsed question, the emotion engine is used to recognize the user's emotions, for example, emotions such as "anxiety" or "dissatisfaction."
[0761] Step 6:
[0762] Server: Based on the analysis results of the emotion engine, the analysis results (question content and emotion) are input into the generative AI model and a response is generated. The generated response takes into account the user's emotions, and might be something like, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0763] Step 7:
[0764] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[0765] Step 8:
[0766] Device: The device receives the response sent back from the server. It parses the received JSON data and displays it on the screen in a user-friendly format. The displayed message is, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[0767] Step 9:
[0768] User: The user checks the response displayed on the device and takes the suggested action. For example, the user opens the smartphone settings, checks which apps are using a lot of battery power, and resolves the battery issue by deleting unnecessary apps.
[0769] In this way, a series of steps is completed: the user's question is captured on the device, sent to the server, analyzed by the natural language processing engine and emotion engine, an appropriate response is generated by the generative AI model, and the response is sent back to the device. Through specific actions at each step, the user can quickly and effectively obtain information to solve their problem.
[0770] Example 2
[0771] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0772] Conventional systems have had difficulty generating appropriate responses to user questions. This is due to the lack of accurate analysis of the question content or the provision of personalized responses that reflect the user's emotions. This has made it difficult to improve user satisfaction and provide efficient support.
[0773] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0774] In this invention, the server includes means for capturing input from a user, converting it into JSON format, and transmitting it, means for using a natural language processing engine to analyze the question received from the user, means for using an emotion engine to analyze emotions based on the analysis results obtained by the natural language processing, means for generating a response using a generative AI model based on the results of the natural language analysis and the emotion analysis, and means for converting the generated response into JSON format and returning it to the user. This makes it possible to accurately analyze the content of the user's question and quickly provide a personalized response that reflects the user's emotions.
[0775] "Means for capturing user input, converting it into JSON format, and sending it" refers to a method for obtaining information entered by a user into an input form, converting it into the JSON format commonly used for data exchange, and sending it to a server.
[0776] The "means using a natural language processing engine" is a software module that analyzes questions and text data received from users to understand their meaning and intent.
[0777] The "means for using an emotion engine" is a software module for identifying and analyzing a user's emotion from text data analyzed by a natural language processing engine.
[0778] "Means for generating responses using a generative AI model" refers to a method for automatically generating appropriate responses using AI technology based on the results of natural language analysis and sentiment analysis.
[0779] "Means for converting the generated response into JSON format and returning it to the user" refers to a method for converting the response generated by the generative AI model into JSON format and returning it to the user.
[0780] "Means for inputting a prompt sentence and generating a response in an interactive format" refers to a method for inputting a prompt sentence containing a question or instruction to an AI model and generating a natural, interactive response based on the prompt.
[0781] This invention is a system that receives questions about products and services from users, analyzes the content and sentiment contained in the questions, and generates responses using a generative AI model based on the analysis results. The specific hardware and software configurations and processing of this system are described below.
[0782] System Configuration
[0783] server
[0784] The server receives queries from users, analyzes their content, and generates and sends back appropriate responses. It has the following main functions:
[0785] 1. Receiving Questions:
[0786] The server receives questions sent from the user's device as HTTP POST requests using common web server technologies (e.g., Node.js and the Express framework).
[0787] 2. Natural language analysis:
[0788] The received question is analyzed using a natural language processing engine, using the Google Cloud Natural Language API to extract the user's intent and keywords.
[0789] 3. Emotion analysis:
[0790] The system uses an emotion engine to recognize the user's emotions in response to questions analyzed by a natural language processing engine. An example of an emotion engine used is the IBM Watson Tone Analyzer.
[0791] 4. Response Generation:
[0792] Based on the results of natural language analysis and sentiment analysis, an appropriate response is generated using a generative AI model such as OpenAI GPT-4.
[0793] 5. Sending the response:
[0794] The generated response is converted to JSON format and sent back to the terminal as an HTTP response.
[0795] Terminal
[0796] The terminal is responsible for capturing input from the user, sending it to the server, and displaying the responses sent back from the server. It has the following main functions:
[0797] 1. Capturing input:
[0798] When a user enters a question into the input form and presses the submit button, the terminal captures this input. The form, which runs on a web browser, is implemented using JavaScript.
[0799] 2. Sending API requests:
[0800] The captured input is converted to JSON format and sent as an HTTP POST request to the server's API endpoint. A common method for sending requests is to use the JavaScript "axios" library.
[0801] 3. Receiving and displaying responses:
[0802] It receives the response sent back from the server and displays it on the screen in a user-friendly format. You can use "Vue.js" or "React" for front-end processing.
[0803] Specific examples
[0804] For example, if a user types the question "My battery is draining quickly. What should I do?", here's what happens:
[0805] 1. User:
[0806] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0807] 2. Terminal:
[0808] The device captures the question, converts it into JSON format, and sends it to the server's API endpoint.
[0809] 3. Server:
[0810] The server receives the question, uses a natural language processing engine to identify the problem of "low battery," and uses an emotion engine to analyze the user's emotions, such as "dissatisfaction" or "confusion."
[0811] 4. Server:
[0812] The generative AI model generates a response such as "Check which apps are using a lot of battery power and recommend deleting any unnecessary apps." This response is adjusted based on the user's emotions.
[0813] 5. Server:
[0814] The generated response is sent back to the terminal in JSON format.
[0815] 6. Terminal:
[0816] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0817] Prompt Sentence Examples
[0818] User Question: "My battery is draining quickly. What should I do?"
[0819] Prompt for generative AI model: "A user complains that their battery is draining too quickly. What would be a good solution?"
[0820] This system allows users to quickly and accurately obtain information to solve their problems, and by adding sentiment analysis, it enables more personalized responses, improving user satisfaction.
[0821] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0822] Step 1: User enters question
[0823] Input: The user inputs a question into the input form on the terminal.
[0824] Specific actions: The user opens a browser on their smartphone or computer, accesses the support page, types in "My battery is draining quickly. What should I do?", and presses the send button.
[0825] Output: The user sends the question entered in the input form to the terminal.
[0826] Step 2: The device captures the question and sends it to the server
[0827] Input: The question the user entered into the input form.
[0828] Specific operation: The device uses JavaScript to retrieve questions from the input form, convert them to JSON format, and then sends the JSON data as an HTTP POST request to the server's API endpoint, using the "axios" library, for example.
[0829] Output: The device sends the question data in JSON format to the server.
[0830] Step 3: The server receives the query
[0831] Input: Question data in JSON format sent from the terminal.
[0832] How it works: The server sets up an endpoint that listens for HTTP POST requests, and when a request arrives, it extracts the question data in JSON format from the request body. This process is done using Node.js and Express, for example.
[0833] Output: The server receives the question data in JSON format and stores it as parseable text data.
[0834] Step 4: The server performs natural language analysis
[0835] Input: Text data of the received question.
[0836] How it works: The server calls the Google Cloud Natural Language API and sends the question data. The natural language processing engine analyzes the text and extracts the user's intent and keywords. For example, the keyword "low battery" is identified.
[0837] Output: The server receives the analysis results from the natural language processing engine and obtains data including the user's intent and important keywords.
[0838] Step 5: The server performs sentiment analysis
[0839] Input: Analysis results obtained from the natural language processing engine.
[0840] How it works: The server calls the IBM Watson Tone Analyzer API and sends the analysis results. The emotion engine identifies emotions in the text and detects emotions such as "frustrated" or "confused."
[0841] Output: The server receives the emotion analysis results and obtains data that indicates the user's emotional state.
[0842] Step 6: Server Generates Response
[0843] Input: Results of natural language analysis and sentiment analysis.
[0844] Specific operation: The server calls the OpenAI GPT-4 API and sends a prompt message. The prompt message is in the form of "The user complains that the battery is draining quickly. Please tell us what to do." Based on this prompt, the generative AI model generates a response such as "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0845] Output: The server receives the response text from the generative AI model and stores it.
[0846] Step 7: The server converts the response into JSON format and sends it back to the device.
[0847] Input: The generated response text.
[0848] Specific operation: The server formats the response text into JSON format and sends the formatted JSON data to the terminal as an HTTP response.
[0849] Output: The server sends the generated response back to the device in JSON format.
[0850] Step 8: The device receives the response from the server and displays it to the user
[0851] Input: JSON formatted response data returned from the server.
[0852] Specific operation: The device receives the HTTP response and parses the JSON data. The parsed response is inserted into an HTML element and displayed on the user's screen. For example, the response might say, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0853] Output: The terminal displays the parsed response to the user.
[0854] (Application example 2)
[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0856] In recent years, in online shopping and virtual stores, users expect their questions and problems to be resolved quickly and accurately. However, conventional systems have difficulty providing appropriate responses to user questions, particularly in terms of personalized responses that reflect the user's emotions. Furthermore, voice interfaces are inadequate, making it difficult to achieve natural interactions, especially through wearable devices such as smart glasses. This has resulted in the challenge of not fully improving the customer experience.
[0857] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for analyzing the emotion of the question using an emotion analysis engine, means for generating a response using a generative AI model based on the analysis result, means for returning the generated response to the user, means for capturing the user's question as voice and converting it into text using voice recognition technology, and means for notifying the user of the generated response visually or audibly. This enables real-time analysis and response to the user's question, and further enables generation of a personalized response according to the emotion, thereby increasing user satisfaction.
[0858] The "means for receiving questions about products and services from users" is a function that allows users to input questions about products and services into the system and transmit them to the server.
[0859] "Means for analysis using a natural language processing engine" refers to a function that analyzes text data received from a user using machine learning and statistical methods to understand the intent and content of the question.
[0860] The "means for analyzing emotions using an emotion analysis engine" is a function for analyzing the emotions contained in the user's question text and determining whether the emotion corresponds to positive, negative, neutral, or the like.
[0861] "Means for generating responses using a generative AI model" refers to a function for using an artificial intelligence model that automatically generates appropriate responses based on the results of natural language processing and sentiment analysis.
[0862] The "means for returning to the user" is a function for sending the generated response to the user and displaying or reproducing it on the screen or through audio.
[0863] "Means for converting to text using speech recognition technology" is a function for using a speech recognition algorithm to convert a question input by a user into text data.
[0864] "Visual or audio notification means" refers to a function for visually displaying or audibly playing the generated response to the user.
[0865] In order to put the present invention into practice, it is necessary to build a system in which the elements of the user, terminal, and server function in cooperation with each other.
[0866] Program Overview
[0867] The server has the following features:
[0868] 1. Receiving a question: This is a function to receive a question sent from the user's terminal. This question is received as voice or text.
[0869] 2. Natural Language Analysis: This function analyzes the received question using a natural language processing engine to understand the user's intent. Natural language processing technologies such as Google Cloud Natural Language API are used for the analysis.
[0870] 3. Sentiment analysis: This function analyzes the emotions contained in questions using a sentiment analysis engine. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[0871] 4. Response Generation: This function generates an appropriate response using a generative AI model (such as OpenAI GPT-3) based on the analysis results.
[0872] 5. Sending the response: This function converts the generated response into JSON format and sends it to the user's device.
[0873] The terminal has the following features:
[0874] 1. Input capture: When a user types a question by voice, the smart glasses' microphone is used to capture the voice and convert it into text using voice recognition technology (such as Google Cloud Speech-to-Text).
[0875] 2. Send API request: The converted text is converted into JSON format and sent as an HTTP POST request to the server's API endpoint.
[0876] 3. Receiving and displaying response: The response sent back from the server is received and displayed on the smart glasses display for the user to visually confirm, or output as an audio output.
[0877] Specific examples
[0878] For example, if a user uses smart glasses and asks, "I'm not sleeping well these days. How can I improve it?", the following steps will occur:
[0879] 1. Input capture: Audio is captured using the microphone on the smart glasses and converted to text using Google Cloud Speech-to-Text.
[0880] 2. Natural Language Analysis: We use the Google Cloud Natural Language API to analyze the question "Poor quality of sleep."
[0881] 3. Sentiment Analysis: Analyze emotions using IBM Watson Tone Analyzer to identify user frustrations and concerns.
[0882] 4. Response Generation: Using Open AI GPT-3, we generate a response like, "First, I recommend reducing your screen time before bed and trying yoga or meditation as a way to relax."
[0883] 5. Send and display response: The generated response is sent back to the terminal in JSON format and displayed on the smart glasses or read aloud.
[0884] Prompt Sentence Examples
[0885] If the user's question is "I'm not sleeping well lately. What can I do to improve it?", the prompt might look like this:
[0886] User's query: I've been having trouble sleeping lately. How can I improve it?
[0887] Sentiment: -0.3
[0888] Emotions: {"anger": 0.1, "sadness": 0.5, "joy": 0.2}
[0889] Response:
[0890] As described above, the system works in cooperation with the server and the terminal, generating appropriate responses to user questions in real time. This system improves the user's customer experience and provides efficient support.
[0891] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0892] Step 1:
[0893] The user uses the smart glasses to input a question by voice, and the user's voice is captured by the microphone of the smart glasses. The input voice data is collected.
[0894] Step 2:
[0895] The device converts the captured audio into text using the Google Cloud Speech-to-Text API. During this conversion process, the audio data is converted into text data, and the user's question is obtained in text format.
[0896] Step 3:
[0897] The terminal converts the converted text data into JSON format and sends it to the server's API endpoint as an HTTP POST request. The input information is text data, and a JSON format request is generated as the output.
[0898] Step 4:
[0899] The server analyzes the question received from the device using a natural language processing engine (for example, Google Cloud Natural Language API). During this analysis, data calculations are performed to understand the intent and content of the question. The analysis results are output as data indicating the intent.
[0900] Step 5:
[0901] The server analyzes emotions using a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) based on the analysis results of the natural language processing engine. The input is data containing intent, and the output is a sentiment score. During this process, data calculations are performed to identify sentiment categories (positive, negative, neutral, etc.).
[0902] Step 6:
[0903] The server combines the results of natural language processing and sentiment analysis to generate a response using a generative AI model (e.g., OpenAI GPT-3). The input is data containing intent and sentiment scores, and the output is an appropriate response text. The generated response is then processed based on the prompt.
[0904] Step 7:
[0905] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response. The input is the generated response text and the output is the JSON formatted response.
[0906] Step 8:
[0907] The device parses the response received from the server and displays or reads it out loud to the user in a user-friendly format. The input is the response data in JSON format, and the output is visual or audio feedback. During this process, the data is reformatted to make it easier for the user to understand.
[0908] These steps result in a system that generates real-time responses to user questions and provides feedback in an appropriate format.
[0909] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0910] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0911] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0912] [Third embodiment]
[0913] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0914] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0915] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0916] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0917] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0918] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0919] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0920] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0921] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0922] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0923] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0924] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0925] This invention is a system that receives questions about products and services from users, analyzes them using a natural language processing engine, and generates appropriate responses using a generative AI model. Specific program processing and examples are described below.
[0926] System Configuration
[0927] server
[0928] The server receives a question from the user, analyzes it using a natural language processing engine, and then generates an appropriate response using a generative AI model based on the analysis results and sends it back to the user. The server's main functions are:
[0929] Receiving questions:
[0930] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0931] Natural language analysis:
[0932] The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[0933] Response generation:
[0934] The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for delivery back to the user.
[0935] Sending a response:
[0936] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[0937] Terminal
[0938] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[0939] Capturing input:
[0940] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[0941] Sending an API request:
[0942] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[0943] Receiving and displaying the response:
[0944] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[0945] Specific examples
[0946] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[0947] 1. User:
[0948] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[0949] 2. Terminal:
[0950] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[0951] 3. Server:
[0952] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[0953] 4. Server:
[0954] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0955] 5. Server:
[0956] The generated response is sent back to the terminal in JSON format.
[0957] 6. Terminal:
[0958] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[0959] In this way, users can quickly and accurately obtain information to resolve their issues. This system improves the customer experience by eliminating waiting times and trips to physical stores. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[0960] The processing flow will be explained below.
[0961] Step 1:
[0962] User: The user uses a device such as a smartphone or PC to enter a question into an input form. Example: "My battery is draining quickly. What should I do?"
[0963] Step 2:
[0964] Terminal: When the user presses the "Submit" button, the terminal captures this input and creates an API request, formats the question in JSON format, and sends it as an HTTP POST request to the server's API endpoint.
[0965] Step 3:
[0966] Server: The server receives the API request sent from the device, parses the JSON data included in the request body, and extracts the user's question.
[0967] Step 4:
[0968] Server: The server sends the extracted questions to a natural language processing engine, which analyzes the questions to understand their main content and sentiment, and obtains the analysis results.
[0969] Step 5:
[0970] Server: Based on the results of natural language analysis, the server inputs the analysis results into a generative AI model and generates a response. Example: "We recommend that you check which apps are using a lot of battery and delete any unnecessary apps."
[0971] Step 6:
[0972] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[0973] Step 7:
[0974] Device: The device receives the response sent back from the server. It analyzes the response content and displays it on the screen in a user-friendly format. Example: "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[0975] Step 8:
[0976] User: The user checks the response displayed on the device and follows the instructions to take appropriate measures to resolve the problem.
[0977] Through the above steps, the system responds promptly and appropriately to the user's questions and provides information for resolving the problem.
[0978] Example 1
[0979] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0980] Modern consumers expect quick responses to general product and service-related questions and problem-solving, necessitating an efficient and effective support system for end users. However, traditional inquiry systems have limitations in accurately analyzing user intent and emotions and providing appropriate responses, which can lead to a decline in customer satisfaction. In addition, individual responses require significant resources, creating cost challenges.
[0981] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0982] In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for generating a response using a generative AI model based on the analysis result, means for a terminal to capture the received question, convert it into JSON format, and send it to the server, means for the server to receive the converted data and generate a response based on the analysis result, and means for returning the generated response to the user and for the terminal to display the response to the user, thereby enabling the user to quickly and accurately obtain information for solving their problem.
[0983] The "means for receiving questions" is a function for receiving questions about products or services sent by users from the terminal.
[0984] A "natural language processing engine" is software that analyzes questions received from users and understands their intentions and emotions.
[0985] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on the results analyzed by a natural language processing engine.
[0986] The "means for generating a response" is a function that uses a generative AI model to create a response based on the analysis results.
[0987] The "means for capturing questions" is a function for recognizing questions entered by users into the terminal and capturing them as data.
[0988] The "means for converting into JSON format" is a function for converting the captured user question into JSON format in order to send the data to the server.
[0989] An "API endpoint" is a specific URL or URI on a server that a device accesses to send a question or inquiry to the server.
[0990] The "means for receiving data" is a function that allows the server to receive data sent from the terminal.
[0991] A "terminal" is a device that a user uses to enter and display questions and responses.
[0992] A "user" is someone who enters a question about a product or service.
[0993] This invention is a system that receives questions about products and services from users, analyzes them with a natural language processing engine, generates appropriate responses using a generative AI model, and returns them to the users. Specific program processing and examples are described below.
[0994] System Configuration
[0995] server
[0996] The server receives a question from the user, analyzes the question using a natural language processing engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are:
[0997] Receiving questions:
[0998] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[0999] Natural language analysis:
[1000] The server analyzes the received question using a natural language processing engine (e.g., Google NLP or SpaCy) to understand the user's intent and emotions. The results of this analysis are used to generate a response.
[1001] Response generation:
[1002] The server uses a generative AI model (e.g., OpenAI GPT-4) based on the parsed question to generate an appropriate response.
[1003] Sending a response:
[1004] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[1005] Terminal
[1006] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[1007] Capturing input:
[1008] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[1009] Sending an API request:
[1010] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1011] Receiving and displaying the response:
[1012] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[1013] Specific examples
[1014] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[1015] 1. User:
[1016] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1017] 2. Terminal:
[1018] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1019] 3. Server:
[1020] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[1021] 4. Server:
[1022] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1023] 5. Server:
[1024] The generated response is sent back to the terminal in JSON format.
[1025] 6. Terminal:
[1026] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1027] This system allows users to quickly and accurately obtain information to resolve their issues, eliminating waiting times and the hassle of visiting a physical store, improving the customer experience. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[1028] Examples of prompt statements
[1029] Here are some examples of prompts:
[1030] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the submit button.
[1031] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1032] Step 1:
[1033] User:
[1034] The user enters a question into the input form on the device and presses the send button. For example, the user might enter, "My battery is draining quickly. What should I do?"
[1035] Input: The question the user entered into the input form.
[1036] Output: None.
[1037] Specific operation: The user enters their question or problem directly into the input form on the device.
[1038] Step 2:
[1039] Device:
[1040] The device captures the questions entered by the user, converts them into JSON format, and sends them as an HTTP POST request to the server's API endpoint.
[1041] Input: Questions entered by the user in the input form.
[1042] Output: The converted questions in JSON format.
[1043] Specific operation: The terminal captures the entered question, converts it into the appropriate format (JSON), and sends it to the server.
[1044] Step 3:
[1045] server:
[1046] The server receives the HTTP POST request sent from the terminal.
[1047] Input: Question sent from the terminal in JSON format.
[1048] Output: The received question data.
[1049] Specific operation: The server receives the question sent from the terminal and prepares for analysis.
[1050] Step 4:
[1051] server:
[1052] The server analyzes the received question using a natural language processing engine, specifically by performing grammatical analysis, keyword extraction, and intent recognition.
[1053] Input: Received question data.
[1054] Output: Analysis results (including user intent and sentiment).
[1055] Specific operation: The natural language processing engine analyzes the question and extracts important keywords and the user's intent.
[1056] Step 5:
[1057] server:
[1058] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[1059] Input: Analysis results of the natural language processing engine.
[1060] Output: The response generated by the generative AI model.
[1061] Specific operation: Based on the analysis results, the generative AI model generates an appropriate response sentence.
[1062] Step 6:
[1063] server:
[1064] The server converts the generated response into JSON format and sends it to the terminal as an HTTP response.
[1065] Input: The generated response.
[1066] Output: The response formatted as JSON.
[1067] Specific operation: The server converts the generated response into an appropriate format and sends it to the terminal.
[1068] Step 7:
[1069] Device:
[1070] The terminal receives the response sent by the server and displays it to the user.
[1071] Input: The response sent by the server in JSON format.
[1072] Output: The response that is displayed to the user.
[1073] Specific operation: The terminal analyzes the received response and displays it in a format that is easy for the user to understand.
[1074] This series of steps allows users to quickly and accurately find a solution to their question.
[1075] (Application example 1)
[1076] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1077] In modern brick-and-mortar stores, customers need to be able to quickly and accurately resolve their product and service-related questions. However, traditional brick-and-mortar stores have a limited number of customer service representatives, making it difficult to respond to all customers in real time. This has led to problems such as a decline in customer satisfaction and an increased burden on store staff. It is also difficult to guarantee the accuracy and consistency of the information provided to customers. To solve these problems, there is a need to provide effective technology.
[1078] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1079] In this invention, the server includes a means for receiving questions about products or services from users, a means for analyzing the questions using a natural language processing engine, and a means for generating responses using a generative AI model based on the analysis results, thereby enabling the server to provide quick and accurate responses to customer questions.
[1080] "Products and services" means goods or activities offered for commercial purposes that are purchased or used by consumers.
[1081] "User" means any person or entity that uses the System or Services.
[1082] "Means for receiving" refers to the device or process used to obtain information or data from a user.
[1083] A "natural language processing engine" refers to a computer program or algorithm that analyzes human language and understands its meaning and structure.
[1084] "Means for analysis" refers to the methods and techniques used to break down received data and extract meaning and significance.
[1085] A "generative AI model" refers to an artificial intelligence algorithm or program that generates appropriate responses or suggestions based on training data.
[1086] "Means for generating a response" refers to the methods and techniques for generating an answer to a question based on the analysis results.
[1087] "Returning means" refers to the method or process for sending the generated response to the user.
[1088] "Means for generating API requests" refers to the code and protocols used to send requests to a server through an application interface.
[1089] "User terminal" refers to an electronic device used by a user, such as a smartphone or tablet.
[1090] "Means for displaying" refers to a display device and its control program for visually presenting information to a user.
[1091] A "server" refers to a central processing unit that receives requests from clients via a network and returns responses.
[1092] This invention provides a system for quickly and accurately responding to questions about products and services when a user asks a question in a physical store. This system includes a server and a user terminal (e.g., a smartphone) and operates in the following steps.
[1093] System Configuration
[1094] server
[1095] The server receives a question from the user and analyzes it using a natural language processing engine. It then generates an appropriate response using a generative AI model based on the analysis results. The generated response is then sent back to the user's device. The server's main functions are as follows:
[1096] Receiving a question: The server receives a question sent from the user's terminal. The question is sent as an HTTP POST request.
[1097] Natural language analysis: The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[1098] Response Generation: The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for sending back to the user.
[1099] Sending a response: The server sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[1100] Terminal
[1101] The user terminal captures input from the user, sends it to the server, and displays the responses sent back from the server. The terminal's main functions are:
[1102] Input capture: When a user enters a question into an input form and presses the submit button, the device captures this input.
[1103] Send API request: Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1104] Receive and display the response: Receive the response sent back from the server and display it on the screen in a user-friendly format.
[1105] Specific examples
[1106] For example, if a user types the question "What material is this shirt made of?", the question is processed as follows:
[1107] 1. User: The user enters "What material is this shirt made of?" into the input form on the device and presses the send button.
[1108] 2. Terminal: The terminal captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1109] 3. Server: The server receives the query and analyzes it using a natural language processing engine. The keywords "shirt" and "material" are identified as the analysis results.
[1110] 4. Server: Then, using a generative AI model, it generates a response such as "This shirt is made of 100% cotton."
[1111] 5. Server: Returns the generated response to the user terminal.
[1112] 6. Terminal: The terminal receives the response from the server and displays "This shirt is made of 100% cotton" on the user screen.
[1113] Prompt Sentence Examples
[1114] For example, if a user asks, "Does this shampoo contain any allergens?", the generative AI model might generate a prompt like this:
[1115] Question analysis:
[1116] Product: Shampoo
[1117] Question: Are there any ingredients that can cause allergies?
[1118] Produces response: This shampoo is allergen-free. However, if you would like to know more about the ingredients, please see the back of the package.
[1119] As described above, this system allows users to receive quick and accurate responses to questions about products and services in physical stores.
[1120] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1121] Step 1:
[1122] The user enters a question into the input form on the terminal and presses the send button, which causes the terminal to capture the user's question.
[1123] Input: User question (e.g., "What material is this shirt made of?")
[1124] Output: Captured question information
[1125] Specific operation: The user enters a question into the input field on the device's UI and presses the "Send" button.
[1126] Step 2:
[1127] The device converts the captured questions into JSON format and sends it as an HTTP POST request to the server's API endpoint.
[1128] Input: Captured question information
[1129] Output: HTTP POST request to the server
[1130] Specific operation: The terminal encodes the question into JSON format and sends an HTTP POST request to the server using the requests library.
[1131] Step 3:
[1132] The server receives a question from a user, which is then passed to a natural language processing engine for analysis.
[1133] Input: Question sent as an HTTP POST request
[1134] Output: Parsing request passed to the natural language processing engine
[1135] Specific operation: The server retrieves the question data received and uses the requests library to send an analysis request to the natural language processing engine.
[1136] Step 4:
[1137] The server receives the results analyzed by the natural language processing engine, which include the user's intent and keywords.
[1138] Input: Parsing request to the natural language processing engine
[1139] Output: Analysis results (e.g., keywords "shirt" and "material")
[1140] Specific behavior: Receives analysis results returned from a natural language processing engine.
[1141] Step 5:
[1142] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[1143] Input: Analysis results
[1144] Output: The generated response (e.g., "This shirt is made of 100% cotton")
[1145] Specific operation: The analysis results are input into the generative AI model as a prompt sentence to obtain an appropriate response.
[1146] Step 6:
[1147] The server returns the generated response to the user terminal.
[1148] Input: Generated response
[1149] Output: HTTP response to the user's device
[1150] Specific operation: The generated response is encoded in JSON format and sent to the user's terminal as an HTTP response.
[1151] Step 7:
[1152] The terminal receives the response sent back from the server and displays it on the user screen.
[1153] Input: HTTP response from the server
[1154] Output: The displayed response (e.g., "This shirt is made of 100% cotton")
[1155] Specific behavior: Displaying the received response in the UI (e.g., displaying it in a text box or alert).
[1156] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1157] This invention is a system that receives questions about products or services from users, analyzes the emotions contained in the questions using a natural language processing engine and an emotion engine, and generates a response using a generative AI model based on the analysis results. Specific program processing and examples are described below.
[1158] System Configuration
[1159] server
[1160] The server receives questions from users, analyzes them using a natural language processing engine and an emotion engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are as follows:
[1161] Receiving questions:
[1162] The server receives a question sent from the user's device, which is sent as an HTTP POST request.
[1163] Natural language analysis:
[1164] The server analyzes the received question using a natural language processing engine to understand the user's intent and emotions, and the analysis results are used to generate a response.
[1165] Emotion analysis:
[1166] The emotion engine recognizes the user's emotions in response to questions analyzed by the natural language processing engine. The results of the emotion analysis are used to generate more detailed responses.
[1167] Response generation:
[1168] The server uses a generative AI model based on the results of natural language analysis and sentiment analysis to generate an appropriate response, such as "We recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1169] Sending a response:
[1170] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response.
[1171] Terminal
[1172] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[1173] Capturing input:
[1174] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[1175] Sending an API request:
[1176] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1177] Receiving and displaying the response:
[1178] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[1179] Specific examples
[1180] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[1181] 1. User:
[1182] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1183] 2. Terminal:
[1184] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1185] 3. Server:
[1186] The server receives the question and analyzes it using a natural language processing engine and an emotion engine. The natural language processing engine identifies the problem of "low battery," and the emotion engine analyzes the user's emotion.
[1187] 4. Server:
[1188] The generative AI model is then used to generate a response such as "We recommend you review the apps that are using a lot of battery power and delete any unnecessary apps." This response is adjusted based on the user's emotions.
[1189] 5. Server:
[1190] The generated response is sent back to the terminal in JSON format.
[1191] 6. Terminal:
[1192] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1193] In this way, users can quickly and accurately obtain information to solve their problems. This system improves the customer experience by eliminating waiting times and the hassle of visiting a physical store. It also reduces the burden on counter staff at physical stores, enabling more efficient support. The addition of an emotion engine enables more personalized responses, further improving user satisfaction.
[1194] The processing flow will be explained below.
[1195] Step 1:
[1196] User: The user uses a device such as a smartphone or PC to enter a question into an input form. For example, they might enter, "My battery is draining quickly. What should I do?"
[1197] Step 2:
[1198] Terminal: When the user presses the "Submit" button, the terminal captures this input, converts the captured question into JSON format, and sends it as an HTTP POST request to the server's API endpoint. The content of the request is JSON data containing the question text.
[1199] Step 3:
[1200] Server: The server receives the API request sent from the device, parses the body of the received request, and extracts the user's question.
[1201] Step 4:
[1202] Server: The server sends the extracted question to a natural language processing engine, which analyzes the intent and content of the question. As a result of the analysis, the subject of the question and related keywords are identified.
[1203] Step 5:
[1204] Server: Based on the parsed question, the emotion engine is used to recognize the user's emotions, for example, emotions such as "anxiety" or "dissatisfaction."
[1205] Step 6:
[1206] Server: Based on the analysis results of the emotion engine, the analysis results (question content and emotion) are input into the generative AI model and a response is generated. The generated response takes into account the user's emotions, and might be something like, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1207] Step 7:
[1208] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[1209] Step 8:
[1210] Device: The device receives the response sent back from the server. It parses the received JSON data and displays it on the screen in a user-friendly format. The displayed message is, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1211] Step 9:
[1212] User: The user checks the response displayed on the device and takes the suggested action. For example, the user opens the smartphone settings, checks which apps are using a lot of battery power, and resolves the battery issue by deleting unnecessary apps.
[1213] In this way, a series of steps is completed: the user's question is captured on the device, sent to the server, analyzed by the natural language processing engine and emotion engine, an appropriate response is generated by the generative AI model, and the response is sent back to the device. Through specific actions at each step, the user can quickly and effectively obtain information to solve their problem.
[1214] Example 2
[1215] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1216] Conventional systems have had difficulty generating appropriate responses to user questions. This is due to the lack of accurate analysis of the question content or the provision of personalized responses that reflect the user's emotions. This has made it difficult to improve user satisfaction and provide efficient support.
[1217] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1218] In this invention, the server includes means for capturing input from a user, converting it into JSON format, and transmitting it, means for using a natural language processing engine to analyze the question received from the user, means for using an emotion engine to analyze emotions based on the analysis results obtained by the natural language processing, means for generating a response using a generative AI model based on the results of the natural language analysis and the emotion analysis, and means for converting the generated response into JSON format and returning it to the user. This makes it possible to accurately analyze the content of the user's question and quickly provide a personalized response that reflects the user's emotions.
[1219] "Means for capturing user input, converting it into JSON format, and sending it" refers to a method for obtaining information entered by a user into an input form, converting it into the JSON format commonly used for data exchange, and sending it to a server.
[1220] The "means using a natural language processing engine" is a software module that analyzes questions and text data received from users to understand their meaning and intent.
[1221] The "means for using an emotion engine" is a software module for identifying and analyzing a user's emotion from text data analyzed by a natural language processing engine.
[1222] "Means for generating responses using a generative AI model" refers to a method for automatically generating appropriate responses using AI technology based on the results of natural language analysis and sentiment analysis.
[1223] "Means for converting the generated response into JSON format and returning it to the user" refers to a method for converting the response generated by the generative AI model into JSON format and returning it to the user.
[1224] "Means for inputting a prompt sentence and generating a response in an interactive format" refers to a method for inputting a prompt sentence containing a question or instruction to an AI model and generating a natural, interactive response based on the prompt.
[1225] This invention is a system that receives questions about products and services from users, analyzes the content and sentiment contained in the questions, and generates responses using a generative AI model based on the analysis results. The specific hardware and software configurations and processing of this system are described below.
[1226] System Configuration
[1227] server
[1228] The server receives queries from users, analyzes their content, and generates and sends back appropriate responses. It has the following main functions:
[1229] 1. Receiving Questions:
[1230] The server receives questions sent from the user's device as HTTP POST requests using common web server technologies (e.g., Node.js and the Express framework).
[1231] 2. Natural language analysis:
[1232] The received question is analyzed using a natural language processing engine, using the Google Cloud Natural Language API to extract the user's intent and keywords.
[1233] 3. Emotion analysis:
[1234] The system uses an emotion engine to recognize the user's emotions in response to questions analyzed by a natural language processing engine. An example of an emotion engine used is the IBM Watson Tone Analyzer.
[1235] 4. Response Generation:
[1236] Based on the results of natural language analysis and sentiment analysis, an appropriate response is generated using a generative AI model such as OpenAI GPT-4.
[1237] 5. Sending the response:
[1238] The generated response is converted to JSON format and sent back to the terminal as an HTTP response.
[1239] Terminal
[1240] The terminal is responsible for capturing input from the user, sending it to the server, and displaying the responses sent back from the server. It has the following main functions:
[1241] 1. Capturing input:
[1242] When a user enters a question into the input form and presses the submit button, the terminal captures this input. The form, which runs on a web browser, is implemented using JavaScript.
[1243] 2. Sending API requests:
[1244] The captured input is converted to JSON format and sent as an HTTP POST request to the server's API endpoint. A common method for sending requests is to use the JavaScript "axios" library.
[1245] 3. Receiving and displaying responses:
[1246] It receives the response sent back from the server and displays it on the screen in a user-friendly format. You can use "Vue.js" or "React" for front-end processing.
[1247] Specific examples
[1248] For example, if a user types the question "My battery is draining quickly. What should I do?", here's what happens:
[1249] 1. User:
[1250] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1251] 2. Terminal:
[1252] The device captures the question, converts it into JSON format, and sends it to the server's API endpoint.
[1253] 3. Server:
[1254] The server receives the question, uses a natural language processing engine to identify the problem of "low battery," and uses an emotion engine to analyze the user's emotions, such as "dissatisfaction" or "confusion."
[1255] 4. Server:
[1256] The generative AI model generates a response such as "Check which apps are using a lot of battery power and recommend deleting any unnecessary apps." This response is adjusted based on the user's emotions.
[1257] 5. Server:
[1258] The generated response is sent back to the terminal in JSON format.
[1259] 6. Terminal:
[1260] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1261] Prompt Sentence Examples
[1262] User Question: "My battery is draining quickly. What should I do?"
[1263] Prompt for generative AI model: "A user complains that their battery is draining too quickly. What would be a good solution?"
[1264] This system allows users to quickly and accurately obtain information to solve their problems, and by adding sentiment analysis, it enables more personalized responses, improving user satisfaction.
[1265] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1266] Step 1: User enters question
[1267] Input: The user inputs a question into the input form on the terminal.
[1268] Specific actions: The user opens a browser on their smartphone or computer, accesses the support page, types in "My battery is draining quickly. What should I do?", and presses the send button.
[1269] Output: The user sends the question entered in the input form to the terminal.
[1270] Step 2: The device captures the question and sends it to the server
[1271] Input: The question the user entered into the input form.
[1272] Specific operation: The device uses JavaScript to retrieve questions from the input form, convert them to JSON format, and then sends the JSON data as an HTTP POST request to the server's API endpoint, using the "axios" library, for example.
[1273] Output: The device sends the question data in JSON format to the server.
[1274] Step 3: The server receives the query
[1275] Input: Question data in JSON format sent from the terminal.
[1276] How it works: The server sets up an endpoint that listens for HTTP POST requests, and when a request arrives, it extracts the question data in JSON format from the request body. This process is done using Node.js and Express, for example.
[1277] Output: The server receives the question data in JSON format and stores it as parseable text data.
[1278] Step 4: The server performs natural language analysis
[1279] Input: Text data of the received question.
[1280] How it works: The server calls the Google Cloud Natural Language API and sends the question data. The natural language processing engine analyzes the text and extracts the user's intent and keywords. For example, the keyword "low battery" is identified.
[1281] Output: The server receives the analysis results from the natural language processing engine and obtains data including the user's intent and important keywords.
[1282] Step 5: The server performs sentiment analysis
[1283] Input: Analysis results obtained from the natural language processing engine.
[1284] How it works: The server calls the IBM Watson Tone Analyzer API and sends the analysis results. The emotion engine identifies emotions in the text and detects emotions such as "frustrated" or "confused."
[1285] Output: The server receives the emotion analysis results and obtains data that indicates the user's emotional state.
[1286] Step 6: Server Generates Response
[1287] Input: Results of natural language analysis and sentiment analysis.
[1288] Specific operation: The server calls the OpenAI GPT-4 API and sends a prompt message. The prompt message is in the form of "The user complains that the battery is draining quickly. Please tell us what to do." Based on this prompt, the generative AI model generates a response such as "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1289] Output: The server receives the response text from the generative AI model and stores it.
[1290] Step 7: The server converts the response into JSON format and sends it back to the device.
[1291] Input: The generated response text.
[1292] Specific operation: The server formats the response text into JSON format and sends the formatted JSON data to the terminal as an HTTP response.
[1293] Output: The server sends the generated response back to the device in JSON format.
[1294] Step 8: The device receives the response from the server and displays it to the user
[1295] Input: JSON formatted response data returned from the server.
[1296] Specific operation: The device receives the HTTP response and parses the JSON data. The parsed response is inserted into an HTML element and displayed on the user's screen. For example, the response might say, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1297] Output: The terminal displays the parsed response to the user.
[1298] (Application example 2)
[1299] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1300] In recent years, in online shopping and virtual stores, users expect their questions and problems to be resolved quickly and accurately. However, conventional systems have difficulty providing appropriate responses to user questions, particularly in terms of personalized responses that reflect the user's emotions. Furthermore, voice interfaces are inadequate, making it difficult to achieve natural interactions, especially through wearable devices such as smart glasses. This has resulted in the challenge of not fully improving the customer experience.
[1301] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for analyzing the emotion of the question using an emotion analysis engine, means for generating a response using a generative AI model based on the analysis result, means for returning the generated response to the user, means for capturing the user's question as voice and converting it into text using voice recognition technology, and means for notifying the user of the generated response visually or audibly. This enables real-time analysis and response to the user's question, and further enables generation of a personalized response according to the emotion, thereby increasing user satisfaction.
[1302] The "means for receiving questions about products and services from users" is a function that allows users to input questions about products and services into the system and transmit them to the server.
[1303] "Means for analysis using a natural language processing engine" refers to a function that analyzes text data received from a user using machine learning and statistical methods to understand the intent and content of the question.
[1304] The "means for analyzing emotions using an emotion analysis engine" is a function for analyzing the emotions contained in the user's question text and determining whether the emotion corresponds to positive, negative, neutral, or the like.
[1305] "Means for generating responses using a generative AI model" refers to a function for using an artificial intelligence model that automatically generates appropriate responses based on the results of natural language processing and sentiment analysis.
[1306] The "means for returning to the user" is a function for sending the generated response to the user and displaying or reproducing it on the screen or through audio.
[1307] "Means for converting to text using speech recognition technology" is a function for using a speech recognition algorithm to convert a question input by a user into text data.
[1308] "Visual or audio notification means" refers to a function for visually displaying or audibly playing the generated response to the user.
[1309] In order to put the present invention into practice, it is necessary to build a system in which the elements of the user, terminal, and server function in cooperation with each other.
[1310] Program Overview
[1311] The server has the following features:
[1312] 1. Receiving a question: This is a function to receive a question sent from the user's terminal. This question is received as voice or text.
[1313] 2. Natural Language Analysis: This function analyzes the received question using a natural language processing engine to understand the user's intent. Natural language processing technologies such as Google Cloud Natural Language API are used for the analysis.
[1314] 3. Sentiment analysis: This function analyzes the emotions contained in questions using a sentiment analysis engine. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[1315] 4. Response Generation: This function generates an appropriate response using a generative AI model (such as OpenAI GPT-3) based on the analysis results.
[1316] 5. Sending the response: This function converts the generated response into JSON format and sends it to the user's device.
[1317] The terminal has the following features:
[1318] 1. Input capture: When a user types a question by voice, the smart glasses' microphone is used to capture the voice and convert it into text using voice recognition technology (such as Google Cloud Speech-to-Text).
[1319] 2. Send API request: The converted text is converted into JSON format and sent as an HTTP POST request to the server's API endpoint.
[1320] 3. Receiving and displaying response: The response sent back from the server is received and displayed on the smart glasses display for the user to visually confirm, or output as an audio output.
[1321] Specific examples
[1322] For example, if a user uses smart glasses and asks, "I'm not sleeping well these days. How can I improve it?", the following steps will occur:
[1323] 1. Input capture: Audio is captured using the microphone on the smart glasses and converted to text using Google Cloud Speech-to-Text.
[1324] 2. Natural Language Analysis: We use the Google Cloud Natural Language API to analyze the question "Poor quality of sleep."
[1325] 3. Sentiment Analysis: Analyze emotions using IBM Watson Tone Analyzer to identify user frustrations and concerns.
[1326] 4. Response Generation: Using Open AI GPT-3, we generate a response like, "First, I recommend reducing your screen time before bed and trying yoga or meditation as a way to relax."
[1327] 5. Send and display response: The generated response is sent back to the terminal in JSON format and displayed on the smart glasses or read aloud.
[1328] Prompt Sentence Examples
[1329] If the user's question is "I'm not sleeping well lately. What can I do to improve it?", the prompt might look like this:
[1330] User's query: I've been having trouble sleeping lately. How can I improve it?
[1331] Sentiment: -0.3
[1332] Emotions: {"anger": 0.1, "sadness": 0.5, "joy": 0.2}
[1333] Response:
[1334] As described above, the system works in cooperation with the server and the terminal, generating appropriate responses to user questions in real time. This system improves the user's customer experience and provides efficient support.
[1335] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1336] Step 1:
[1337] The user uses the smart glasses to input a question by voice, and the user's voice is captured by the microphone of the smart glasses. The input voice data is collected.
[1338] Step 2:
[1339] The device converts the captured audio into text using the Google Cloud Speech-to-Text API. During this conversion process, the audio data is converted into text data, and the user's question is obtained in text format.
[1340] Step 3:
[1341] The terminal converts the converted text data into JSON format and sends it to the server's API endpoint as an HTTP POST request. The input information is text data, and a JSON format request is generated as the output.
[1342] Step 4:
[1343] The server analyzes the question received from the device using a natural language processing engine (for example, Google Cloud Natural Language API). During this analysis, data calculations are performed to understand the intent and content of the question. The analysis results are output as data indicating the intent.
[1344] Step 5:
[1345] The server analyzes emotions using a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) based on the analysis results of the natural language processing engine. The input is data containing intent, and the output is a sentiment score. During this process, data calculations are performed to identify sentiment categories (positive, negative, neutral, etc.).
[1346] Step 6:
[1347] The server combines the results of natural language processing and sentiment analysis to generate a response using a generative AI model (e.g., OpenAI GPT-3). The input is data containing intent and sentiment scores, and the output is an appropriate response text. The generated response is then processed based on the prompt.
[1348] Step 7:
[1349] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response. The input is the generated response text and the output is the JSON formatted response.
[1350] Step 8:
[1351] The device parses the response received from the server and displays or reads it out loud to the user in a user-friendly format. The input is the response data in JSON format, and the output is visual or audio feedback. During this process, the data is reformatted to make it easier for the user to understand.
[1352] These steps result in a system that generates real-time responses to user questions and provides feedback in an appropriate format.
[1353] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1354] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1355] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1356] [Fourth embodiment]
[1357] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1358] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1359] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1360] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1361] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1362] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1363] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1364] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1365] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1366] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1367] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1368] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1369] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1370] This invention is a system that receives questions about products and services from users, analyzes them using a natural language processing engine, and generates appropriate responses using a generative AI model. Specific program processing and examples are described below.
[1371] System Configuration
[1372] server
[1373] The server receives a question from the user, analyzes it using a natural language processing engine, and then generates an appropriate response using a generative AI model based on the analysis results and sends it back to the user. The server's main functions are:
[1374] Receiving questions:
[1375] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[1376] Natural language analysis:
[1377] The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[1378] Response generation:
[1379] The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for delivery back to the user.
[1380] Sending a response:
[1381] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[1382] Terminal
[1383] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[1384] Capturing input:
[1385] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[1386] Sending an API request:
[1387] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1388] Receiving and displaying the response:
[1389] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[1390] Specific examples
[1391] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[1392] 1. User:
[1393] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1394] 2. Terminal:
[1395] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1396] 3. Server:
[1397] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[1398] 4. Server:
[1399] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1400] 5. Server:
[1401] The generated response is sent back to the terminal in JSON format.
[1402] 6. Terminal:
[1403] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1404] In this way, users can quickly and accurately obtain information to resolve their issues. This system improves the customer experience by eliminating waiting times and trips to physical stores. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[1405] The processing flow will be explained below.
[1406] Step 1:
[1407] User: The user uses a device such as a smartphone or PC to enter a question into an input form. Example: "My battery is draining quickly. What should I do?"
[1408] Step 2:
[1409] Terminal: When the user presses the "Submit" button, the terminal captures this input and creates an API request, formats the question in JSON format, and sends it as an HTTP POST request to the server's API endpoint.
[1410] Step 3:
[1411] Server: The server receives the API request sent from the device, parses the JSON data included in the request body, and extracts the user's question.
[1412] Step 4:
[1413] Server: The server sends the extracted questions to a natural language processing engine, which analyzes the questions to understand their main content and sentiment, and obtains the analysis results.
[1414] Step 5:
[1415] Server: Based on the results of natural language analysis, the server inputs the analysis results into a generative AI model and generates a response. Example: "We recommend that you check which apps are using a lot of battery and delete any unnecessary apps."
[1416] Step 6:
[1417] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[1418] Step 7:
[1419] Device: The device receives the response sent back from the server. It analyzes the response content and displays it on the screen in a user-friendly format. Example: "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1420] Step 8:
[1421] User: The user checks the response displayed on the device and follows the instructions to take appropriate measures to resolve the problem.
[1422] Through the above steps, the system responds promptly and appropriately to the user's questions and provides information for resolving the problem.
[1423] Example 1
[1424] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1425] Modern consumers expect quick responses to general product and service-related questions and problem-solving, necessitating an efficient and effective support system for end users. However, traditional inquiry systems have limitations in accurately analyzing user intent and emotions and providing appropriate responses, which can lead to a decline in customer satisfaction. In addition, individual responses require significant resources, creating cost challenges.
[1426] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1427] In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for generating a response using a generative AI model based on the analysis result, means for a terminal to capture the received question, convert it into JSON format, and send it to the server, means for the server to receive the converted data and generate a response based on the analysis result, and means for returning the generated response to the user and for the terminal to display the response to the user, thereby enabling the user to quickly and accurately obtain information for solving their problem.
[1428] The "means for receiving questions" is a function for receiving questions about products or services sent by users from the terminal.
[1429] A "natural language processing engine" is software that analyzes questions received from users and understands their intentions and emotions.
[1430] A "generative AI model" is an artificial intelligence model that generates appropriate responses based on the results analyzed by a natural language processing engine.
[1431] The "means for generating a response" is a function that uses a generative AI model to create a response based on the analysis results.
[1432] The "means for capturing questions" is a function for recognizing questions entered by users into the terminal and capturing them as data.
[1433] The "means for converting into JSON format" is a function for converting the captured user question into JSON format in order to send the data to the server.
[1434] An "API endpoint" is a specific URL or URI on a server that a device accesses to send a question or inquiry to the server.
[1435] The "means for receiving data" is a function that allows the server to receive data sent from the terminal.
[1436] A "terminal" is a device that a user uses to enter and display questions and responses.
[1437] A "user" is someone who enters a question about a product or service.
[1438] This invention is a system that receives questions about products and services from users, analyzes them with a natural language processing engine, generates appropriate responses using a generative AI model, and returns them to the users. Specific program processing and examples are described below.
[1439] System Configuration
[1440] server
[1441] The server receives a question from the user, analyzes the question using a natural language processing engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are:
[1442] Receiving questions:
[1443] The server receives the question sent from the user's terminal, which is sent as an HTTP POST request.
[1444] Natural language analysis:
[1445] The server analyzes the received question using a natural language processing engine (e.g., Google NLP or SpaCy) to understand the user's intent and emotions. The results of this analysis are used to generate a response.
[1446] Response generation:
[1447] The server uses a generative AI model (e.g., OpenAI GPT-4) based on the parsed question to generate an appropriate response.
[1448] Sending a response:
[1449] The server then sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[1450] Terminal
[1451] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[1452] Capturing input:
[1453] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[1454] Sending an API request:
[1455] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1456] Receiving and displaying the response:
[1457] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[1458] Specific examples
[1459] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[1460] 1. User:
[1461] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1462] 2. Terminal:
[1463] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1464] 3. Server:
[1465] The server receives the question and analyzes it using a natural language processing engine. As a result, the problem of "low battery" is identified.
[1466] 4. Server:
[1467] Next, the generative AI model is used to generate a response such as, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1468] 5. Server:
[1469] The generated response is sent back to the terminal in JSON format.
[1470] 6. Terminal:
[1471] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1472] This system allows users to quickly and accurately obtain information to resolve their issues, eliminating waiting times and the hassle of visiting a physical store, improving the customer experience. It also reduces the burden on counter staff at physical stores, enabling more efficient support.
[1473] Examples of prompt statements
[1474] Here are some examples of prompts:
[1475] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the submit button.
[1476] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1477] Step 1:
[1478] User:
[1479] The user enters a question into the input form on the device and presses the send button. For example, the user might enter, "My battery is draining quickly. What should I do?"
[1480] Input: The question the user entered into the input form.
[1481] Output: None.
[1482] Specific operation: The user enters their question or problem directly into the input form on the device.
[1483] Step 2:
[1484] Device:
[1485] The device captures the questions entered by the user, converts them into JSON format, and sends them as an HTTP POST request to the server's API endpoint.
[1486] Input: Questions entered by the user in the input form.
[1487] Output: The converted questions in JSON format.
[1488] Specific operation: The terminal captures the entered question, converts it into the appropriate format (JSON), and sends it to the server.
[1489] Step 3:
[1490] server:
[1491] The server receives the HTTP POST request sent from the terminal.
[1492] Input: Question sent from the terminal in JSON format.
[1493] Output: The received question data.
[1494] Specific operation: The server receives the question sent from the terminal and prepares for analysis.
[1495] Step 4:
[1496] server:
[1497] The server analyzes the received question using a natural language processing engine, specifically by performing grammatical analysis, keyword extraction, and intent recognition.
[1498] Input: Received question data.
[1499] Output: Analysis results (including user intent and sentiment).
[1500] Specific operation: The natural language processing engine analyzes the question and extracts important keywords and the user's intent.
[1501] Step 5:
[1502] server:
[1503] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[1504] Input: Analysis results of the natural language processing engine.
[1505] Output: The response generated by the generative AI model.
[1506] Specific operation: Based on the analysis results, the generative AI model generates an appropriate response sentence.
[1507] Step 6:
[1508] server:
[1509] The server converts the generated response into JSON format and sends it to the terminal as an HTTP response.
[1510] Input: The generated response.
[1511] Output: The response formatted as JSON.
[1512] Specific operation: The server converts the generated response into an appropriate format and sends it to the terminal.
[1513] Step 7:
[1514] Device:
[1515] The terminal receives the response sent by the server and displays it to the user.
[1516] Input: The response sent by the server in JSON format.
[1517] Output: The response that is displayed to the user.
[1518] Specific operation: The terminal analyzes the received response and displays it in a format that is easy for the user to understand.
[1519] This series of steps allows users to quickly and accurately find a solution to their question.
[1520] (Application example 1)
[1521] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1522] In modern brick-and-mortar stores, customers need to be able to quickly and accurately resolve their product and service-related questions. However, traditional brick-and-mortar stores have a limited number of customer service representatives, making it difficult to respond to all customers in real time. This has led to problems such as a decline in customer satisfaction and an increased burden on store staff. It is also difficult to guarantee the accuracy and consistency of the information provided to customers. To solve these problems, there is a need to provide effective technology.
[1523] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1524] In this invention, the server includes a means for receiving questions about products or services from users, a means for analyzing the questions using a natural language processing engine, and a means for generating responses using a generative AI model based on the analysis results, thereby enabling the server to provide quick and accurate responses to customer questions.
[1525] "Products and services" means goods or activities offered for commercial purposes that are purchased or used by consumers.
[1526] "User" means any person or entity that uses the System or Services.
[1527] "Means for receiving" refers to the device or process used to obtain information or data from a user.
[1528] A "natural language processing engine" refers to a computer program or algorithm that analyzes human language and understands its meaning and structure.
[1529] "Means for analysis" refers to the methods and techniques used to break down received data and extract meaning and significance.
[1530] A "generative AI model" refers to an artificial intelligence algorithm or program that generates appropriate responses or suggestions based on training data.
[1531] "Means for generating a response" refers to the methods and techniques for generating an answer to a question based on the analysis results.
[1532] "Returning means" refers to the method or process for sending the generated response to the user.
[1533] "Means for generating API requests" refers to the code and protocols used to send requests to a server through an application interface.
[1534] "User terminal" refers to an electronic device used by a user, such as a smartphone or tablet.
[1535] "Means for displaying" refers to a display device and its control program for visually presenting information to a user.
[1536] A "server" refers to a central processing unit that receives requests from clients via a network and returns responses.
[1537] This invention provides a system for quickly and accurately responding to questions about products and services when a user asks a question in a physical store. This system includes a server and a user terminal (e.g., a smartphone) and operates in the following steps.
[1538] System Configuration
[1539] server
[1540] The server receives a question from the user and analyzes it using a natural language processing engine. It then generates an appropriate response using a generative AI model based on the analysis results. The generated response is then sent back to the user's device. The server's main functions are as follows:
[1541] Receiving a question: The server receives a question sent from the user's terminal. The question is sent as an HTTP POST request.
[1542] Natural language analysis: The server analyzes the received question using a natural language processing engine to understand the user's intentions and emotions. The results of this analysis are used to generate a response.
[1543] Response Generation: The server uses a generative AI model based on the parsed question to generate an appropriate response, which is then prepared for sending back to the user.
[1544] Sending a response: The server sends the generated response back to the user's terminal, allowing the user to quickly obtain a solution to their problem.
[1545] Terminal
[1546] The user terminal captures input from the user, sends it to the server, and displays the responses sent back from the server. The terminal's main functions are:
[1547] Input capture: When a user enters a question into an input form and presses the submit button, the device captures this input.
[1548] Send API request: Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1549] Receive and display the response: Receive the response sent back from the server and display it on the screen in a user-friendly format.
[1550] Specific examples
[1551] For example, if a user types the question "What material is this shirt made of?", the question is processed as follows:
[1552] 1. User: The user enters "What material is this shirt made of?" into the input form on the device and presses the send button.
[1553] 2. Terminal: The terminal captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1554] 3. Server: The server receives the query and analyzes it using a natural language processing engine. The keywords "shirt" and "material" are identified as the analysis results.
[1555] 4. Server: Then, using a generative AI model, it generates a response such as "This shirt is made of 100% cotton."
[1556] 5. Server: Returns the generated response to the user terminal.
[1557] 6. Terminal: The terminal receives the response from the server and displays "This shirt is made of 100% cotton" on the user screen.
[1558] Prompt Sentence Examples
[1559] For example, if a user asks, "Does this shampoo contain any allergens?", the generative AI model might generate a prompt like this:
[1560] Question analysis:
[1561] Product: Shampoo
[1562] Question: Are there any ingredients that can cause allergies?
[1563] Produces response: This shampoo is allergen-free. However, if you would like to know more about the ingredients, please see the back of the package.
[1564] As described above, this system allows users to receive quick and accurate responses to questions about products and services in physical stores.
[1565] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1566] Step 1:
[1567] The user enters a question into the input form on the terminal and presses the send button, which causes the terminal to capture the user's question.
[1568] Input: User question (e.g., "What material is this shirt made of?")
[1569] Output: Captured question information
[1570] Specific operation: The user enters a question into the input field on the device's UI and presses the "Send" button.
[1571] Step 2:
[1572] The device converts the captured questions into JSON format and sends it as an HTTP POST request to the server's API endpoint.
[1573] Input: Captured question information
[1574] Output: HTTP POST request to the server
[1575] Specific operation: The terminal encodes the question into JSON format and sends an HTTP POST request to the server using the requests library.
[1576] Step 3:
[1577] The server receives a question from a user, which is then passed to a natural language processing engine for analysis.
[1578] Input: Question sent as an HTTP POST request
[1579] Output: Parsing request passed to the natural language processing engine
[1580] Specific operation: The server retrieves the question data received and uses the requests library to send an analysis request to the natural language processing engine.
[1581] Step 4:
[1582] The server receives the results analyzed by the natural language processing engine, which include the user's intent and keywords.
[1583] Input: Parsing request to the natural language processing engine
[1584] Output: Analysis results (e.g., keywords "shirt" and "material")
[1585] Specific behavior: Receives analysis results returned from a natural language processing engine.
[1586] Step 5:
[1587] The server uses a generative AI model based on the analysis results to generate an appropriate response.
[1588] Input: Analysis results
[1589] Output: The generated response (e.g., "This shirt is made of 100% cotton")
[1590] Specific operation: The analysis results are input into the generative AI model as a prompt sentence to obtain an appropriate response.
[1591] Step 6:
[1592] The server returns the generated response to the user terminal.
[1593] Input: Generated response
[1594] Output: HTTP response to the user's device
[1595] Specific operation: The generated response is encoded in JSON format and sent to the user's terminal as an HTTP response.
[1596] Step 7:
[1597] The terminal receives the response sent back from the server and displays it on the user screen.
[1598] Input: HTTP response from the server
[1599] Output: The displayed response (e.g., "This shirt is made of 100% cotton")
[1600] Specific behavior: Displaying the received response in the UI (e.g., displaying it in a text box or alert).
[1601] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1602] This invention is a system that receives questions about products or services from users, analyzes the emotions contained in the questions using a natural language processing engine and an emotion engine, and generates a response using a generative AI model based on the analysis results. Specific program processing and examples are described below.
[1603] System Configuration
[1604] server
[1605] The server receives questions from users, analyzes them using a natural language processing engine and an emotion engine, generates an appropriate response using a generative AI model, and sends it back to the user. The server's main functions are as follows:
[1606] Receiving questions:
[1607] The server receives a question sent from the user's device, which is sent as an HTTP POST request.
[1608] Natural language analysis:
[1609] The server analyzes the received question using a natural language processing engine to understand the user's intent and emotions, and the analysis results are used to generate a response.
[1610] Emotion analysis:
[1611] The emotion engine recognizes the user's emotions in response to questions analyzed by the natural language processing engine. The results of the emotion analysis are used to generate more detailed responses.
[1612] Response generation:
[1613] The server uses a generative AI model based on the results of natural language analysis and sentiment analysis to generate an appropriate response, such as "We recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1614] Sending a response:
[1615] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response.
[1616] Terminal
[1617] The terminal captures input from the user, sends it to the server, and displays the responses sent back by the server. The terminal's main functions are:
[1618] Capturing input:
[1619] When a user enters a question into the input form and presses the submit button, the terminal captures this input.
[1620] Sending an API request:
[1621] Convert the captured input into JSON format and send it as an HTTP POST request to the server's API endpoint.
[1622] Receiving and displaying the response:
[1623] The response sent back from the server is received and displayed on the screen in a user-friendly format.
[1624] Specific examples
[1625] For example, if a user types the question "My battery is draining quickly. What should I do?", this question is processed as follows:
[1626] 1. User:
[1627] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1628] 2. Terminal:
[1629] The device captures this question, converts it into JSON format, and sends it to the server's API endpoint.
[1630] 3. Server:
[1631] The server receives the question and analyzes it using a natural language processing engine and an emotion engine. The natural language processing engine identifies the problem of "low battery," and the emotion engine analyzes the user's emotion.
[1632] 4. Server:
[1633] The generative AI model is then used to generate a response such as "We recommend you review the apps that are using a lot of battery power and delete any unnecessary apps." This response is adjusted based on the user's emotions.
[1634] 5. Server:
[1635] The generated response is sent back to the terminal in JSON format.
[1636] 6. Terminal:
[1637] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1638] In this way, users can quickly and accurately obtain information to solve their problems. This system improves the customer experience by eliminating waiting times and the hassle of visiting a physical store. It also reduces the burden on counter staff at physical stores, enabling more efficient support. The addition of an emotion engine enables more personalized responses, further improving user satisfaction.
[1639] The processing flow will be explained below.
[1640] Step 1:
[1641] User: The user uses a device such as a smartphone or PC to enter a question into an input form. For example, they might enter, "My battery is draining quickly. What should I do?"
[1642] Step 2:
[1643] Terminal: When the user presses the "Submit" button, the terminal captures this input, converts the captured question into JSON format, and sends it as an HTTP POST request to the server's API endpoint. The content of the request is JSON data containing the question text.
[1644] Step 3:
[1645] Server: The server receives the API request sent from the device, parses the body of the received request, and extracts the user's question.
[1646] Step 4:
[1647] Server: The server sends the extracted question to a natural language processing engine, which analyzes the intent and content of the question. As a result of the analysis, the subject of the question and related keywords are identified.
[1648] Step 5:
[1649] Server: Based on the parsed question, the emotion engine is used to recognize the user's emotions, for example, emotions such as "anxiety" or "dissatisfaction."
[1650] Step 6:
[1651] Server: Based on the analysis results of the emotion engine, the analysis results (question content and emotion) are input into the generative AI model and a response is generated. The generated response takes into account the user's emotions, and might be something like, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1652] Step 7:
[1653] Server: Converts the generated response into JSON format and returns it to the terminal as an HTTP response.
[1654] Step 8:
[1655] Device: The device receives the response sent back from the server. It parses the received JSON data and displays it on the screen in a user-friendly format. The displayed message is, "Don't worry, we recommend checking which apps are using a lot of battery and deleting any unnecessary apps."
[1656] Step 9:
[1657] User: The user checks the response displayed on the device and takes the suggested action. For example, the user opens the smartphone settings, checks which apps are using a lot of battery power, and resolves the battery issue by deleting unnecessary apps.
[1658] In this way, a series of steps is completed: the user's question is captured on the device, sent to the server, analyzed by the natural language processing engine and emotion engine, an appropriate response is generated by the generative AI model, and the response is sent back to the device. Through specific actions at each step, the user can quickly and effectively obtain information to solve their problem.
[1659] Example 2
[1660] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1661] Conventional systems have had difficulty generating appropriate responses to user questions. This is due to the lack of accurate analysis of the question content or the provision of personalized responses that reflect the user's emotions. This has made it difficult to improve user satisfaction and provide efficient support.
[1662] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1663] In this invention, the server includes means for capturing input from a user, converting it into JSON format, and transmitting it, means for using a natural language processing engine to analyze the question received from the user, means for using an emotion engine to analyze emotions based on the analysis results obtained by the natural language processing, means for generating a response using a generative AI model based on the results of the natural language analysis and the emotion analysis, and means for converting the generated response into JSON format and returning it to the user. This makes it possible to accurately analyze the content of the user's question and quickly provide a personalized response that reflects the user's emotions.
[1664] "Means for capturing user input, converting it into JSON format, and sending it" refers to a method for obtaining information entered by a user into an input form, converting it into the JSON format commonly used for data exchange, and sending it to a server.
[1665] The "means using a natural language processing engine" is a software module that analyzes questions and text data received from users to understand their meaning and intent.
[1666] The "means for using an emotion engine" is a software module for identifying and analyzing a user's emotion from text data analyzed by a natural language processing engine.
[1667] "Means for generating responses using a generative AI model" refers to a method for automatically generating appropriate responses using AI technology based on the results of natural language analysis and sentiment analysis.
[1668] "Means for converting the generated response into JSON format and returning it to the user" refers to a method for converting the response generated by the generative AI model into JSON format and returning it to the user.
[1669] "Means for inputting a prompt sentence and generating a response in an interactive format" refers to a method for inputting a prompt sentence containing a question or instruction to an AI model and generating a natural, interactive response based on the prompt.
[1670] This invention is a system that receives questions about products and services from users, analyzes the content and sentiment contained in the questions, and generates responses using a generative AI model based on the analysis results. The specific hardware and software configurations and processing of this system are described below.
[1671] System Configuration
[1672] server
[1673] The server receives queries from users, analyzes their content, and generates and sends back appropriate responses. It has the following main functions:
[1674] 1. Receiving Questions:
[1675] The server receives questions sent from the user's device as HTTP POST requests using common web server technologies (e.g., Node.js and the Express framework).
[1676] 2. Natural language analysis:
[1677] The received question is analyzed using a natural language processing engine, using the Google Cloud Natural Language API to extract the user's intent and keywords.
[1678] 3. Emotion analysis:
[1679] The system uses an emotion engine to recognize the user's emotions in response to questions analyzed by a natural language processing engine. An example of an emotion engine used is the IBM Watson Tone Analyzer.
[1680] 4. Response Generation:
[1681] Based on the results of natural language analysis and sentiment analysis, an appropriate response is generated using a generative AI model such as OpenAI GPT-4.
[1682] 5. Sending the response:
[1683] The generated response is converted to JSON format and sent back to the terminal as an HTTP response.
[1684] Terminal
[1685] The terminal is responsible for capturing input from the user, sending it to the server, and displaying the responses sent back from the server. It has the following main functions:
[1686] 1. Capturing input:
[1687] When a user enters a question into the input form and presses the submit button, the terminal captures this input. The form, which runs on a web browser, is implemented using JavaScript.
[1688] 2. Sending API requests:
[1689] The captured input is converted to JSON format and sent as an HTTP POST request to the server's API endpoint. A common method for sending requests is to use the JavaScript "axios" library.
[1690] 3. Receiving and displaying responses:
[1691] It receives the response sent back from the server and displays it on the screen in a user-friendly format. You can use "Vue.js" or "React" for front-end processing.
[1692] Specific examples
[1693] For example, if a user types the question "My battery is draining quickly. What should I do?", here's what happens:
[1694] 1. User:
[1695] The user enters "My battery is draining quickly. What should I do?" into the input form on the device and presses the send button.
[1696] 2. Terminal:
[1697] The device captures the question, converts it into JSON format, and sends it to the server's API endpoint.
[1698] 3. Server:
[1699] The server receives the question, uses a natural language processing engine to identify the problem of "low battery," and uses an emotion engine to analyze the user's emotions, such as "dissatisfaction" or "confusion."
[1700] 4. Server:
[1701] The generative AI model generates a response such as "Check which apps are using a lot of battery power and recommend deleting any unnecessary apps." This response is adjusted based on the user's emotions.
[1702] 5. Server:
[1703] The generated response is sent back to the terminal in JSON format.
[1704] 6. Terminal:
[1705] The device receives the response from the server and displays the message "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps" on the user's screen.
[1706] Prompt Sentence Examples
[1707] User Question: "My battery is draining quickly. What should I do?"
[1708] Prompt for generative AI model: "A user complains that their battery is draining too quickly. What would be a good solution?"
[1709] This system allows users to quickly and accurately obtain information to solve their problems, and by adding sentiment analysis, it enables more personalized responses, improving user satisfaction.
[1710] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1711] Step 1: User enters question
[1712] Input: The user inputs a question into the input form on the terminal.
[1713] Specific actions: The user opens a browser on their smartphone or computer, accesses the support page, types in "My battery is draining quickly. What should I do?", and presses the send button.
[1714] Output: The user sends the question entered in the input form to the terminal.
[1715] Step 2: The device captures the question and sends it to the server
[1716] Input: The question the user entered into the input form.
[1717] Specific operation: The device uses JavaScript to retrieve questions from the input form, convert them to JSON format, and then sends the JSON data as an HTTP POST request to the server's API endpoint, using the "axios" library, for example.
[1718] Output: The device sends the question data in JSON format to the server.
[1719] Step 3: The server receives the query
[1720] Input: Question data in JSON format sent from the terminal.
[1721] How it works: The server sets up an endpoint that listens for HTTP POST requests, and when a request arrives, it extracts the question data in JSON format from the request body. This process is done using Node.js and Express, for example.
[1722] Output: The server receives the question data in JSON format and stores it as parseable text data.
[1723] Step 4: The server performs natural language analysis
[1724] Input: Text data of the received question.
[1725] How it works: The server calls the Google Cloud Natural Language API and sends the question data. The natural language processing engine analyzes the text and extracts the user's intent and keywords. For example, the keyword "low battery" is identified.
[1726] Output: The server receives the analysis results from the natural language processing engine and obtains data including the user's intent and important keywords.
[1727] Step 5: The server performs sentiment analysis
[1728] Input: Analysis results obtained from the natural language processing engine.
[1729] How it works: The server calls the IBM Watson Tone Analyzer API and sends the analysis results. The emotion engine identifies emotions in the text and detects emotions such as "frustrated" or "confused."
[1730] Output: The server receives the emotion analysis results and obtains data that indicates the user's emotional state.
[1731] Step 6: Server Generates Response
[1732] Input: Results of natural language analysis and sentiment analysis.
[1733] Specific operation: The server calls the OpenAI GPT-4 API and sends a prompt message. The prompt message is in the form of "The user complains that the battery is draining quickly. Please tell us what to do." Based on this prompt, the generative AI model generates a response such as "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1734] Output: The server receives the response text from the generative AI model and stores it.
[1735] Step 7: The server converts the response into JSON format and sends it back to the device.
[1736] Input: The generated response text.
[1737] Specific operation: The server formats the response text into JSON format and sends the formatted JSON data to the terminal as an HTTP response.
[1738] Output: The server sends the generated response back to the device in JSON format.
[1739] Step 8: The device receives the response from the server and displays it to the user
[1740] Input: JSON formatted response data returned from the server.
[1741] Specific operation: The device receives the HTTP response and parses the JSON data. The parsed response is inserted into an HTML element and displayed on the user's screen. For example, the response might say, "We recommend that you check which apps are using a lot of battery power and delete any unnecessary apps."
[1742] Output: The terminal displays the parsed response to the user.
[1743] (Application example 2)
[1744] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1745] In recent years, in online shopping and virtual stores, users expect their questions and problems to be resolved quickly and accurately. However, conventional systems have difficulty providing appropriate responses to user questions, particularly in terms of personalized responses that reflect the user's emotions. Furthermore, voice interfaces are inadequate, making it difficult to achieve natural interactions, especially through wearable devices such as smart glasses. This has resulted in the challenge of not fully improving the customer experience.
[1746] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a question about a product or service from a user, means for analyzing the question using a natural language processing engine, means for analyzing the emotion of the question using an emotion analysis engine, means for generating a response using a generative AI model based on the analysis result, means for returning the generated response to the user, means for capturing the user's question as voice and converting it into text using voice recognition technology, and means for notifying the user of the generated response visually or audibly. This enables real-time analysis and response to the user's question, and further enables generation of a personalized response according to the emotion, thereby increasing user satisfaction.
[1747] The "means for receiving questions about products and services from users" is a function that allows users to input questions about products and services into the system and transmit them to the server.
[1748] "Means for analysis using a natural language processing engine" refers to a function that analyzes text data received from a user using machine learning and statistical methods to understand the intent and content of the question.
[1749] The "means for analyzing emotions using an emotion analysis engine" is a function for analyzing the emotions contained in the user's question text and determining whether the emotion corresponds to positive, negative, neutral, or the like.
[1750] "Means for generating responses using a generative AI model" refers to a function for using an artificial intelligence model that automatically generates appropriate responses based on the results of natural language processing and sentiment analysis.
[1751] The "means for returning to the user" is a function for sending the generated response to the user and displaying or reproducing it on the screen or through audio.
[1752] "Means for converting to text using speech recognition technology" is a function for using a speech recognition algorithm to convert a question input by a user into text data.
[1753] "Visual or audio notification means" refers to a function for visually displaying or audibly playing the generated response to the user.
[1754] In order to put the present invention into practice, it is necessary to build a system in which the elements of the user, terminal, and server function in cooperation with each other.
[1755] Program Overview
[1756] The server has the following features:
[1757] 1. Receiving a question: This is a function to receive a question sent from the user's terminal. This question is received as voice or text.
[1758] 2. Natural Language Analysis: This function analyzes the received question using a natural language processing engine to understand the user's intent. Natural language processing technologies such as Google Cloud Natural Language API are used for the analysis.
[1759] 3. Sentiment analysis: This function analyzes the emotions contained in questions using a sentiment analysis engine. Sentiment analysis is performed using tools such as IBM Watson Tone Analyzer.
[1760] 4. Response Generation: This function generates an appropriate response using a generative AI model (such as OpenAI GPT-3) based on the analysis results.
[1761] 5. Sending the response: This function converts the generated response into JSON format and sends it to the user's device.
[1762] The terminal has the following features:
[1763] 1. Input capture: When a user types a question by voice, the smart glasses' microphone is used to capture the voice and convert it into text using voice recognition technology (such as Google Cloud Speech-to-Text).
[1764] 2. Send API request: The converted text is converted into JSON format and sent as an HTTP POST request to the server's API endpoint.
[1765] 3. Receiving and displaying response: The response sent back from the server is received and displayed on the smart glasses display for the user to visually confirm, or output as an audio output.
[1766] Specific examples
[1767] For example, if a user uses smart glasses and asks, "I'm not sleeping well these days. How can I improve it?", the following steps will occur:
[1768] 1. Input capture: Audio is captured using the microphone on the smart glasses and converted to text using Google Cloud Speech-to-Text.
[1769] 2. Natural Language Analysis: We use the Google Cloud Natural Language API to analyze the question "Poor quality of sleep."
[1770] 3. Sentiment Analysis: Analyze emotions using IBM Watson Tone Analyzer to identify user frustrations and concerns.
[1771] 4. Response Generation: Using Open AI GPT-3, we generate a response like, "First, I recommend reducing your screen time before bed and trying yoga or meditation as a way to relax."
[1772] 5. Send and display response: The generated response is sent back to the terminal in JSON format and displayed on the smart glasses or read aloud.
[1773] Prompt Sentence Examples
[1774] If the user's question is "I'm not sleeping well lately. What can I do to improve it?", the prompt might look like this:
[1775] User's query: I've been having trouble sleeping lately. How can I improve it?
[1776] Sentiment: -0.3
[1777] Emotions: {"anger": 0.1, "sadness": 0.5, "joy": 0.2}
[1778] Response:
[1779] As described above, the system works in cooperation with the server and the terminal, generating appropriate responses to user questions in real time. This system improves the user's customer experience and provides efficient support.
[1780] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1781] Step 1:
[1782] The user uses the smart glasses to input a question by voice, and the user's voice is captured by the microphone of the smart glasses. The input voice data is collected.
[1783] Step 2:
[1784] The device converts the captured audio into text using the Google Cloud Speech-to-Text API. During this conversion process, the audio data is converted into text data, and the user's question is obtained in text format.
[1785] Step 3:
[1786] The terminal converts the converted text data into JSON format and sends it to the server's API endpoint as an HTTP POST request. The input information is text data, and a JSON format request is generated as the output.
[1787] Step 4:
[1788] The server analyzes the question received from the device using a natural language processing engine (for example, Google Cloud Natural Language API). During this analysis, data calculations are performed to understand the intent and content of the question. The analysis results are output as data indicating the intent.
[1789] Step 5:
[1790] The server analyzes emotions using a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) based on the analysis results of the natural language processing engine. The input is data containing intent, and the output is a sentiment score. During this process, data calculations are performed to identify sentiment categories (positive, negative, neutral, etc.).
[1791] Step 6:
[1792] The server combines the results of natural language processing and sentiment analysis to generate a response using a generative AI model (e.g., OpenAI GPT-3). The input is data containing intent and sentiment scores, and the output is an appropriate response text. The generated response is then processed based on the prompt.
[1793] Step 7:
[1794] The server converts the generated response into JSON format and sends it back to the terminal as an HTTP response. The input is the generated response text and the output is the JSON formatted response.
[1795] Step 8:
[1796] The device parses the response received from the server and displays or reads it out loud to the user in a user-friendly format. The input is the response data in JSON format, and the output is visual or audio feedback. During this process, the data is reformatted to make it easier for the user to understand.
[1797] These steps result in a system that generates real-time responses to user questions and provides feedback in an appropriate format.
[1798] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1799] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1800] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1801] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1802] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1803] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1804] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1805] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1806] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1807] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1808] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1809] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1810] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1811] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1812] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1813] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1814] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1815] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1816] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1817] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1818] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1819] The following is further disclosed regarding the above embodiment.
[1820] (Claim 1)
[1821] means for receiving product or service related questions from users;
[1822] means for analyzing the question using a natural language processing engine;
[1823] means for generating a response using a generative AI model based on the analysis results;
[1824] means for returning the generated response to a user;
[1825] A system including:
[1826] (Claim 2)
[1827] 10. The system of claim 1, wherein the natural language processing engine further comprises means for analyzing sentiment from a user's question.
[1828] (Claim 3)
[1829] 10. The system of claim 1, wherein the generative AI model includes means for interactively generating responses.
[1830] "Example 1"
[1831] (Claim 1)
[1832] means for receiving product or service related questions from users;
[1833] means for analyzing the question using a natural language processing engine;
[1834] means for generating a response using a generative AI model based on the analysis results;
[1835] means for returning the generated response to a user;
[1836] A means for the terminal to capture the received question, convert it into a JSON format, and transmit it to a server;
[1837] a server receiving the converted data and generating a response based on the analysis result;
[1838] means for returning the generated response to the terminal and for the terminal to display the response to the user;
[1839] A system including:
[1840] (Claim 2)
[1841] 10. The system of claim 1, wherein the natural language processing engine further comprises means for analyzing sentiment from a user's question.
[1842] (Claim 3)
[1843] 10. The system of claim 1, wherein the generative AI model includes means for interactively generating responses.
[1844] "Application Example 1"
[1845] (Claim 1)
[1846] means for receiving product or service related questions from users;
[1847] means for analyzing the question using a natural language processing engine;
[1848] means for generating a response using a generative AI model based on the analysis results;
[1849] means for returning the generated response to a user;
[1850] A means for generating an API request for sending a question from a user terminal to a server;
[1851] means for displaying the generated response on a user terminal;
[1852] A system including:
[1853] (Claim 2)
[1854] 10. The system of claim 1, wherein the natural language processing engine further comprises means for analyzing sentiment from a user's question.
[1855] (Claim 3)
[1856] 10. The system of claim 1, wherein the generative AI model includes means for interactively generating responses.
[1857] "Example 2: Combining Emotion Engines"
[1858] (Claim 1)
[1859] A means of capturing user input, converting it to JSON format, and sending it.
[1860] means for using a natural language processing engine to analyze questions received from a user;
[1861] a means for using an emotion engine to analyze emotions based on the analysis results obtained by natural language processing;
[1862] A means for generating responses using a generative AI model based on the results of natural language analysis and sentiment analysis;
[1863] A means of converting the generated response into JSON format and sending it back to the user;
[1864] A system including:
[1865] (Claim 2)
[1866] 10. The system of claim 1, further comprising: means for analyzing emotions from a user's question using an emotion engine; and means for adjusting the generated response according to the user's emotion.
[1867] (Claim 3)
[1868] 10. The system of claim 1, further comprising means for inputting a prompt sentence using a generative AI model to interactively generate a response.
[1869] "Application example 2 when combining emotion engines"
[1870] (Claim 1)
[1871] means for receiving product or service related questions from users;
[1872] means for analyzing the question using a natural language processing engine;
[1873] means for analyzing emotions from the question using a sentiment analysis engine;
[1874] means for generating a response using a generative AI model based on the analysis results;
[1875] means for returning the generated response to a user;
[1876] a means for capturing a user's voice query and converting it into text using speech recognition technology;
[1877] means for visually or audibly notifying a user of the generated response;
[1878] A system including:
[1879] (Claim 2)
[1880] 10. The system of claim 1, wherein the natural language processing engine further comprises means for analyzing sentiment from a user's question.
[1881] (Claim 3)
[1882] 10. The system of claim 1, wherein the generative AI model includes means for interactively generating responses. [Explanation of symbols]
[1883] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving product or service related questions from users; means for analyzing the question using a natural language processing engine; means for generating a response using a generative AI model based on the analysis results; means for returning the generated response to a user; A system including:
2. The system of claim 1 , wherein the natural language processing engine further comprises means for analyzing sentiment from a user's question.
3. The system of claim 1 , wherein the generative AI model includes means for interactively generating responses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A