system

A system utilizing natural language processing and generative AI provides quick and accurate answers to user questions, addressing the inefficiencies of traditional information access methods by automating question analysis and response generation.

JP2026064825APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Users face challenges in accessing quick and accurate information for procedures due to the need to browse multiple web pages and the inadequacy of FAQ pages and customer support in providing timely and appropriate answers.

Method used

A system that allows users to input questions using a terminal, which are analyzed by a server using natural language processing, and generate answers through a generative AI based on trained data, enabling quick and accurate responses.

Benefits of technology

Enables users to easily understand and perform procedures quickly by providing sophisticated answers to a wide variety of questions, reducing the time and effort required for information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064825000001_ABST
    Figure 2026064825000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means by which the user inputs a question using a terminal, A means of sending a question to the server, A server receives a question and uses natural language processing to analyze it, A method by which the server requests a generative AI to generate an answer based on the analysis results, A method by which a generative AI generates an answer and sends it back to the server, A means by which the server sends the generated response to the terminal, A system that includes means for displaying the responses sent by the terminal to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, many procedures are carried out online, but in many cases, users need to browse multiple web pages to understand the specific procedure method. Therefore, it not only takes time and effort for users, but there is also difficulty in accessing accurate information. Also, even if there are FAQ pages or customer support, they do not necessarily provide answers quickly and appropriately. Against this background, there is a need for a system that allows users to easily understand and quickly perform procedures.

Means for Solving the Problems

[0005] The present invention provides a system that includes means for a user to input a question using a terminal and send the question to a server; means for the server to receive the question and analyze it using natural language processing; means for the server to request a generative AI to generate an answer based on the analysis results; means for the generative AI to generate an answer and send it back to the server; means for the server to send the generated answer to the terminal; and means for the terminal to display the transmitted answer to the user. This system allows users to obtain quick and accurate answers simply by inputting a question. In particular, by including means for natural language processing to analyze keywords in the question content and means for the generative AI to generate an answer using trained data, the system can provide sophisticated answers that can handle a wide variety of questions from users.

[0006] "Device" refers to electronic devices such as computers, smartphones, and tablets that are operated by the user.

[0007] A "server" refers to a computer system that receives, processes, and transmits information over a network.

[0008] A "user" refers to a person who uses the system to input questions and receive answers.

[0009] "Entering a question" refers to the user entering text through the input interface of their device.

[0010] "Sending a question" refers to the process where a device transmits the entered question to a server via the network.

[0011] "Receiving a question and analyzing it using natural language processing" refers to the server using a specific algorithm to analyze the received question as text.

[0012] "Requesting a generative AI to generate an answer based on the analysis results" means that the server uses the analyzed information to request an answer from the generative AI.

[0013] "Generative AI" refers to artificial intelligence that generates text based on large-scale datasets that have been collected.

[0014] "Generating an answer" refers to a generative AI creating an appropriate response in natural language format based on analyzed information.

[0015] "Sending back to the server" refers to sending the response generated by the generative AI back to the server.

[0016] "Sending the response to the terminal" refers to the server transmitting the generated response to the user's terminal via the network.

[0017] "Displaying responses" refers to the process of displaying the responses received by the device in a format that the user can view on the interface.

[0018] "Natural language processing" refers to a field of machine learning techniques used to analyze, understand, and generate text data.

[0019] "Analyzing keywords" refers to extracting important words and phrases from text using natural language processing techniques.

[0020] "Training data" refers to the large dataset of text that a generative AI uses for training.

[0021] "Means of generating answers" refers to methods by which generative AI creates natural language answers to user questions based on the data it has been trained on. [Brief explanation of the drawing]

[0022] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0023] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0024] First, the language used in the following description will be explained.

[0025] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0026] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0027] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0028] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0030] [First Embodiment]

[0031] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0032] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0035] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0038] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0042] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0043] The present invention is a system in which a user inputs a question using a terminal, the input question is sent to a server, the server receives the question and analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0044] First, the user enters a question using their device. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is then sent from the device to the server.

[0045] Next, the server receives the question and analyzes it using natural language processing. Keywords from the question are extracted through natural language processing. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0046] The server then requests a generative AI to generate an answer based on the analysis results. The generative AI generates an appropriate answer from the large dataset collected based on the analyzed keywords. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0047] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed on the screen and becomes viewable by the user.

[0048] This system is designed to allow users to easily understand and quickly perform procedures. By utilizing natural language processing and generative AI, it enables advanced question analysis and answer generation, allowing it to respond quickly and accurately to a wide range of user questions.

[0049] The following describes the processing flow.

[0050] Step 1:

[0051] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0052] Step 2:

[0053] The user submits the question to the server. When the user clicks the "Submit" button, the question is sent to the server as an HTTP POST request.

[0054] Step 3:

[0055] The server receives the question. The server receives the request at the API endpoint and extracts the question content from the request body.

[0056] Step 4:

[0057] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0058] Step 5:

[0059] The server requests the generative AI to generate an answer based on the analysis results. The server uses an API to send a request to the generative AI that includes the analyzed keyword information.

[0060] Step 6:

[0061] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates the response: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0062] Step 7:

[0063] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0064] Step 8:

[0065] The server sends the generated response to the terminal. The server then sends the generated response back to the terminal as an HTTP response.

[0066] Step 9:

[0067] The device displays the received response to the user. The device (e.g., a browser) parses the response received from the server and displays it on the screen. The user's screen displays the text: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0068] (Example 1)

[0069] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0070] The problem that this invention aims to solve is to provide a system that allows users to obtain quick and accurate answers when they have questions. Specifically, with conventional manual search methods, it is often time-consuming for users to find the information they need, and it is often difficult to access the appropriate information. To solve this problem, the invention aims to develop a system that automatically analyzes questions by combining natural language processing and generative AI models, and generates and provides answers immediately.

[0071] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0072] In this invention, the server includes means for receiving a question and analyzing it using natural language processing, means for requesting an AI model to generate an answer based on the analysis results, and means for transmitting the generated answer to a terminal. This makes it possible to provide highly accurate answers immediately to questions entered by users using an information terminal.

[0073] An "information terminal" is an electronic device used by users to input questions and send them to a server.

[0074] A "question" is the content of an inquiry that a user enters using an information terminal.

[0075] A "data server" is a central control unit that receives questions sent from information terminals and requests analysis and response generation.

[0076] "Natural language processing" is a technology that analyzes questions received by a data server and extracts key information from the content of those questions.

[0077] "Key information" refers to important words and phrases extracted from the question content through natural language processing.

[0078] A "generative AI model" is an artificial intelligence model used to generate appropriate responses based on the analysis results of a data server.

[0079] "Answer generation" is the process by which a generative AI model generates appropriate answers to questions based on analysis results.

[0080] "Transmission" refers to the act of a data server transferring a generated response to an information terminal.

[0081] "Display" refers to the process by which an information terminal visually presents the received response to the user.

[0082] The present invention is a system in which a user inputs a question using an information terminal, sends the question to a data server, the data server analyzes the received question using natural language processing, requests an AI model to generate an answer based on the analysis results, sends the answer generated by the AI ​​model back to the information terminal via the data server, and displays the result to the user.

[0083] First, the user enters a question using an information terminal. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is sent from the information terminal to the data server using an HTTP POST request.

[0084] Next, the data server receives the question and analyzes it using natural language processing (NLP). Specifically, it uses Google's NLP library to extract key information from the question. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0085] Subsequently, the data server requests a generative AI model (for example, OpenAI®'s GPT-3®) to generate an answer based on the analysis results. The generative AI model generates an appropriate answer from a large collected dataset (for example, publicly available data on the internet) based on the analyzed key information. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0086] The generated response is sent back to the data server, which then sends it to the user's information terminal. Finally, the information terminal displays the received response to the user. Specifically, a response such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed in an appropriate format, making it easy for the user to view.

[0087] This system enables users to understand procedures and take swift action by providing quick and accurate information when they have questions. By utilizing natural language processing and generative AI models, it enables advanced question analysis and answer generation, allowing for rapid responses to a wide range of user inquiries. Furthermore, the system's design allows users to instantly obtain necessary information without manually performing internet searches, significantly improving user convenience.

[0088] Example of a prompt:

[0089] "The user has entered the following question: 'How do I get a replacement passport?' Please generate the appropriate information to answer this question."

[0090] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0091] Step 1:

[0092] The user enters and submits a question.

[0093] The user enters a question into a text box on the information terminal interface and presses the "Send" button. For example, they might enter "Please tell me how to obtain a resident registration certificate" and click the send button.

[0094] Input: Question text entered by the user on the information terminal.

[0095] Output: The question text is sent as an HTTP POST request.

[0096] Step 2:

[0097] The device sends a question, and the server receives it.

[0098] The terminal sends the user-entered question to the server via an HTTP POST request. The server receives this question, logs it, and prepares a response.

[0099] Input: HTTP POST request generated in Step 1

[0100] Output: Question text stored on the server

[0101] Step 3:

[0102] The server analyzes the received question using natural language processing.

[0103] The server analyzes the received question using Google's NLP library. This analysis extracts key information from the question. For example, keywords such as "resident registration" and "how to obtain" may be identified.

[0104] Input: Question text stored on the server

[0105] Output: Key information extracted through analysis

[0106] Step 4:

[0107] The server sends prompt messages to the AI ​​model based on the analysis results.

[0108] The server uses the key information obtained from the analysis results to create and send prompt messages to the generated AI model. For example, it might create a prompt message such as, "Please answer the question based on the following keywords: resident registration certificate, how to obtain it."

[0109] Input: Key information extracted through analysis

[0110] Output: Prompt message sent to the generated AI model

[0111] Step 5:

[0112] The generative AI model generates the answer and sends it back to the server.

[0113] The generation AI model generates an appropriate response based on the prompt and sends that response back to the server. For example, it might generate the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0114] Input: Prompt sent to the generated AI model

[0115] Output: Generated response text sent back to the server

[0116] Step 6:

[0117] The server sends the generated response to the terminal.

[0118] The server sends the generated response to the terminal via an HTTP POST response.

[0119] Input: Response text sent from the generating AI model to the server.

[0120] Output: Response text sent to the terminal as an HTTP POST response

[0121] Step 7:

[0122] The device displays the received response to the user.

[0123] The device displays the response received from the server on the screen. The user can view the displayed response. For example, the response might say, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city / ward / town / village office. A fee may be charged."

[0124] Input: Response text received from the server

[0125] Output: Response text displayed to the user

[0126] (Application Example 1)

[0127] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0128] Traditional in-store customer support systems often made it difficult for customers to quickly and accurately obtain the information they needed, as information retrieval was cumbersome. Furthermore, reliance on interaction with staff increased their workload and led to inconsistent service quality. Additionally, there was a lack of mechanisms to provide detailed, real-time information about specific products. To address these challenges, technology that allows customers to access information smoothly is necessary.

[0129] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0130] In this invention, the server includes means for the user to input a question using a user interface, means for sending the question to the server, means for the server to receive the question and analyze it using natural language processing, means for the server to request a generative AI to generate an answer based on the analysis results, means for the generative AI to generate an answer and send it back to the server, means for the server to send the generated answer to a user device, means for the user device to display the transmitted answer to the user, and means for the user interface to be configured using a smart wearable device. This makes it possible for customers to input questions in real time through a smart wearable device and to quickly and accurately obtain relevant product information and guidance.

[0131] "User interface" is a general term for devices or software that allow users to directly operate or input information.

[0132] A "question" is the content that a user enters into the system to request information or guidance.

[0133] A "server" is a computer system that provides specific services over a network and processes data in response to user requests.

[0134] "Natural language processing" is a general term for technologies that allow computers to analyze and understand human language.

[0135] "Generative AI" refers to artificial intelligence that uses large datasets to produce appropriate responses or products in response to specified inputs.

[0136] "User device" is a general term for terminals and devices used by users. Examples include smartphones and smart glasses.

[0137] A "smart wearable device" is a general term for portable electronic devices that provide various functions when worn by the user. Specific examples include smart glasses and smartwatches.

[0138] "Keywords" are important words or phrases extracted from the question content that are used to generate answers or perform searches.

[0139] A "large-scale dataset" refers to the vast amount of data used to train generative AI, including text, images, and audio.

[0140] "Real-time" refers to a temporal attribute that means responding immediately to user requests, resulting in extremely low latency.

[0141] This invention is a system in which a user uses a smart wearable device to input questions in real time while in a store and receive appropriate answers.

[0142] First, the user puts on a smart wearable device such as smart glasses and inputs a question using the device's built-in user interface. For example, the user might input a question using the voice input function of the smart glasses, such as "Please tell me how to use this product." This question is then sent from the device to the server.

[0143] The server receives a question and analyzes its content using natural language processing techniques (such as libraries like spaCy or NLTK). Through this analysis, important keywords are extracted from the question. For example, the keyword "how to use the product" might be extracted.

[0144] Based on the analysis results, the server requests a generative AI (for example, OpenAI's GPT-3) to generate an answer. The generative AI uses a large, trained dataset to generate an appropriate answer. For example, it might generate an answer such as, "First, turn on the power to this product and follow the instructions to set it up. Then, download the dedicated app to use it."

[0145] The generated response is sent back to the server, which then transmits it to the user's smart wearable device. The device receives the transmitted response and displays it in the user's field of view. This allows the user to obtain the necessary information in real time.

[0146] For example, if a customer asks in a store, "What's the difference between this TV and this speaker?", the server analyzes the question and generates an answer such as, "This TV has 4K resolution and is HDR compatible. On the other hand, this speaker has Bluetooth functionality and provides clear sound throughout the room," which is then displayed on the smart glasses.

[0147] Example of a prompt:

[0148] "Please generate an answer to the following question using the AI ​​generator. The question is: 'What is the difference between this TV and this speaker?'"

[0149] "User question: 'Is this product waterproof?' Please have the AI ​​generate an answer."

[0150] This allows customers to easily obtain product information in-store, providing an efficient shopping experience.

[0151] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0152] Step 1:

[0153] The user enters a question via voice or touch input through a smart wearable device (e.g., smart glasses). The entered question is first stored in the internal memory of the smart wearable device through its interface, and then sent to a server over the network.

[0154] Input: User's question (e.g., "How do I use this product?")

[0155] Output: Question data is sent to the server via the network.

[0156] Step 2:

[0157] The server receives the question and analyzes its content using a natural language processing (NLP) engine. An NLP engine (e.g., spaCy or NLTK) is then used to extract important keywords from the question.

[0158] Input: Submitted question data

[0159] Output: Analyzed keywords (e.g., "How to use the product")

[0160] Step 3:

[0161] The server requests a generative AI (e.g., OpenAI's GPT-3) to generate an answer based on the extracted keywords. The generative AI model generates an appropriate answer based on a large dataset in which it has been trained. For example, it can generate a detailed answer about how to use a product based on a prompt.

[0162] Input: Analyzed keywords, prompt text (e.g., "Please tell me how to use this product.")

[0163] Output: Generated response (Example: "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0164] Step 4:

[0165] The generated responses are sent back to the server, which then transmits them to the user's smart wearable device. During this process, the server converts the response data into an appropriate format and prepares it for transmission.

[0166] Input: Generated answer

[0167] Output: Response data sent to smart wearable devices

[0168] Step 5:

[0169] The user's smart wearable device acquires the received response data and displays the response on its screen. The user can view the generated response in real time through the smart wearable device.

[0170] Input: Response data sent from the server

[0171] Output: A response displayed on the smart wearable device's screen (e.g., "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0172] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0173] The present invention is a system that recognizes not only the questions entered by the user but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system includes a process in which the user enters a question using a terminal, sends the question to a server, the server receives the question, analyzes the user's emotions using an emotion engine, then analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results and emotion information, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0174] First, the user enters a question using a device. At this stage, text information and voice data entered by the user are also collected. For example, the user might enter the question, "How do I obtain a resident registration certificate?"

[0175] Next, the voice and text data, along with the questions entered by the user, are sent to the server. The server receives this data and first analyzes the user's emotional state using an emotion engine. This emotion engine analyzes emotions from the context of the text entered by the user and from the voice data. For example, it may detect emotional information such as the user being anxious, angry, or troubled.

[0176] The server then passes the sentiment information and the question to a natural language processing module, which analyzes the question. Important keywords (e.g., resident registration, acquisition method) are extracted at this stage.

[0177] The server requests a generative AI to generate a response based on the analysis results and sentiment information. The generative AI then generates an appropriate response from a large dataset based on the analyzed information and sentiment information. For example, if the user is anxious, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office."

[0178] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office," might be displayed and become viewable by the user.

[0179] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0180] The following describes the processing flow.

[0181] Step 1:

[0182] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0183] Step 2:

[0184] The user sends sentiment data along with the question to the server. When the user clicks the "Submit" button, the question content and sentiment data (e.g., facial recognition and voice tone) are sent to the server as an HTTP POST request.

[0185] Step 3:

[0186] The server receives the question and sentiment data. The server receives the request at the API endpoint and extracts the question content and sentiment data from the request body.

[0187] Step 4:

[0188] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the text and voice data entered by the user to identify the user's emotions (e.g., anxious, angry, troubled).

[0189] Step 5:

[0190] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0191] Step 6:

[0192] The server requests the generative AI to generate an answer based on the analysis results and sentiment information. The server uses an API to send a request to the generative AI that includes the analyzed keyword information and the user's sentiment information.

[0193] Step 7:

[0194] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates an appropriate response (e.g., "Don't worry, to obtain a resident registration certificate...") based on analyzed keywords and sentiment information.

[0195] Step 8:

[0196] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0197] Step 9:

[0198] The server sends the generated response to the terminal. The server returns the generated response to the terminal as an HTTP response, providing an appropriate response that matches the user's emotions.

[0199] Step 10:

[0200] The terminal displays the received response to the user. The terminal analyzes the response received from the server and displays it on the screen. The user's screen displays a response such as "Don't worry, to obtain your resident registration certificate..." in an appropriate format.

[0201] (Example 2)

[0202] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0203] Traditional question-answering systems focused on analyzing user input and generating appropriate answers, but they were unable to recognize the user's emotional state and personalize answers accordingly. As a result, they provided uniform answers that ignored user emotions, leading to a poor user experience. Furthermore, they lacked sentiment analysis using voice data, making advanced analysis that reflected the user's input modality difficult. A new system is needed to address these issues.

[0204] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0205] In this invention, the server includes means for analyzing received questions and voice data with an emotion analysis engine, means for passing emotion information and question content to a natural language processing module for analysis, and means for requesting a generative AI to generate an answer based on the analysis results and emotion information. This makes it possible to provide personalized answers that take into account the user's emotional state.

[0206] A "terminal" is a device used by a user to input questions or voice data, and includes smartphones, PCs, tablets, and other similar devices.

[0207] A "server" is a central processing unit that analyzes received data and utilizes various engines and modules to generate appropriate responses.

[0208] A "sentiment analysis engine" is a software program that analyzes text and voice data sent by the user to detect the user's emotional state.

[0209] A "natural language processing module" is a software program that analyzes the content of a question, extracts important keywords, and understands the context.

[0210] "Generative AI" refers to artificial intelligence models that generate appropriate responses based on analyzed information and emotional information.

[0211] "Question and audio data" refers to the collective text and audio information entered by the user using their device.

[0212] "Analysis results" refers to the collective term for the question content analyzed by the natural language processing module and the sentiment information analyzed by the sentiment analysis engine.

[0213] "Answer" refers to response information generated by a generative AI based on its analysis results, and is presented to the user.

[0214] Modes for carrying out the invention

[0215] This invention is a system that recognizes not only the questions entered by the user, but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system utilizes an emotion analysis engine, a natural language processing module, and a generative AI.

[0216] The user first enters their question using a device, such as a smartphone, PC, or tablet. The text information and voice data entered by the user are also collected by the device. For example, a user might enter, "Please tell me how to obtain a resident registration certificate."

[0217] The entered questions and voice data are sent from the terminal to the server. The server receives this data and then uses an emotion analysis engine to analyze the user's emotional state. This emotion analysis engine detects the user's emotions from the context of the text, the tone of the voice data, the speed, etc. For example, it extracts emotional information such as whether the user is anxious, angry, or troubled.

[0218] Next, the server passes the sentiment information and question content to a natural language processing module for analysis. The natural language processing module extracts important keywords (e.g., resident registration, acquisition method) from the input question and analyzes the meaning of the sentences. This analysis result, along with the sentiment information, is then passed back to the server.

[0219] The server requests a generative AI to generate an answer based on the analysis results and sentiment information. This generative AI uses a generative AI model such as OpenAI GPT-4 (registered trademark). Based on the analyzed information and sentiment information, the generative AI generates an appropriate answer from a large collected dataset. For example, if the user is anxious, it might generate an answer such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0220] The generated response is sent back to the server, which then sends it to the user's device. The device then displays the received response to the user. For example, the user's device screen might display a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0221] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0222] Specific example

[0223] For example, if a user enters the prompt "Please tell me how to obtain a resident registration certificate. I'm in a hurry," the sentiment analysis engine will analyze the emotion of "hurry," and the generative AI will generate a response that takes this into consideration. As a result, the user will be provided with the response, "Don't worry, to obtain a resident registration certificate, you need to submit an application form at your nearest city or town hall and present identification documents."

[0224] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0225] Step 1:

[0226] The user enters questions and voice data using a device.

[0227] Users input questions or voice data using devices such as smartphones, PCs, and tablets. For example, they might input, "Please tell me how to obtain a resident registration certificate." During this process, the entered text and voice data are converted into corresponding input formats.

[0228] Input: User's question text and audio data

[0229] Output: Text and audio data stored on the device

[0230] Step 2:

[0231] The terminal sends user input data to the server.

[0232] The terminal sends the collected text and audio data to the server. This transmission is done via an HTTP request or other appropriate communication protocol.

[0233] Input: Text and audio data stored on the device

[0234] Output: Text and audio data sent to the server

[0235] Step 3:

[0236] The server analyzes the received data using an emotion analysis engine.

[0237] The server receives text and audio data sent by the user and passes it to the sentiment analysis engine. The sentiment analysis engine analyzes the context of the text and the tone and speed of the audio data to detect the user's emotional state. For example, if the tone of the audio data is high and fast, it will detect an emotion of anxiety.

[0238] Input: Text data and audio data

[0239] Output: Analyzed emotion information (e.g., "anxious")

[0240] Step 4:

[0241] The server passes sentiment information and questions to a natural language processing module for analysis.

[0242] The sentiment analysis results and the user's question are sent to a natural language processing (NLP) module. The NLP module analyzes the context of the question and extracts important keywords (e.g., "resident registration," "how to obtain"). This allows the meaning of the question to be represented in a structured data format.

[0243] Input: Analyzed sentiment information and text data

[0244] Output: Analysis results with key keywords extracted.

[0245] Step 5:

[0246] The server requests a generative AI to generate an answer based on the analysis results and emotional information.

[0247] The server requests the generative AI to generate an answer based on the analysis results and sentiment information obtained from the NLP module. Specifically, it sends the generative AI a prompt message saying, "The user is anxious. Please answer the question, 'How do I obtain a resident registration certificate?'"

[0248] Input: Analysis results and sentiment information

[0249] Output: Prompt message for generative AI

[0250] Step 6:

[0251] Generative AI generates appropriate answers.

[0252] The generative AI generates a response based on the received prompt text. Here, the generated response is made with consideration for emotions. For example, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city or town hall."

[0253] Input: Prompt message for generative AI

[0254] Output: Appropriate answer

[0255] Step 7:

[0256] The generative AI sends the answer back to the server.

[0257] The generated response is sent back to the server. The server receives this response and prepares to send it to the terminal.

[0258] Input: Appropriate answer

[0259] Output: Response sent back to the server

[0260] Step 8:

[0261] The server sends the response to the terminal, and the terminal displays it to the user.

[0262] The server sends the received response to the user's device. The device receives this response and displays it to the user in an appropriate format. For example, the device screen might display: "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0263] Input: Response sent back to the server

[0264] Output: The answer displayed on the user's device.

[0265] (Application Example 2)

[0266] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0267] In recent years, improving customer satisfaction in physical stores has become increasingly important. However, traditional customer service systems have struggled to recognize customer emotions and provide real-time responses and product suggestions accordingly. Furthermore, they lacked the means to provide appropriate and personalized answers to customer questions. Because of these challenges, there is a need for systems that can achieve more sophisticated customer service.

[0268] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0269] In this invention, the server includes means for the user to input a question using an information processing device, means for transmitting the question to a communication device, means for the communication device to receive the question and analyze it using natural language processing, means for analyzing the user's emotions using an emotion analysis engine, means for requesting a generative AI to generate an answer based on the analysis results and emotion information, means for the generative AI to generate an answer considering the emotion information and send it back to the communication device, and means for transmitting the generated answer to the information processing device and displaying it to the user. This enables the recognition of customer emotions in real time and allows for appropriate responses, product suggestions, and personalized answers to questions accordingly.

[0270] An "information processing device" is an electronic device used to input questions from users and display the submitted answers.

[0271] A "communication device" is a device that receives questions from users and works in conjunction with a server to perform natural language processing and sentiment analysis.

[0272] "Natural language processing" is a technology that analyzes language data entered by a user to understand the meaning of the question.

[0273] A "sentiment analysis engine" is a system that implements technologies and algorithms for analyzing emotional information from text and voice data entered by the user.

[0274] "Generative AI" is an artificial intelligence that generates answers considering the content of a user's question and sentiment information based on trained data.

[0275] "Question" refers to the content input by the user using an information processing device and transmitted to the server via a communication device.

[0276] "Analysis result" is the output result of the question content and sentiment information analyzed by the natural language processing and sentiment analysis engines.

[0277] "User's sentiment" means the sentiment state detected by the sentiment analysis engine from the user's input data.

[0278] "Answer" is the answer or information generated by the generative AI based on the question content and the user's sentiment.

[0279] "Display" is the act of the information processing device visually providing an answer to the user.

[0280] "System" refers to the overall configuration and functions that combine a series of means to implement the present invention.

[0281] The present invention is a sentiment recognition customer service assistant system for enhancing customer service in physical stores. Specific embodiments for implementing this system will be described below.

[0282] Configuration of the System

[0283] This system includes smart glasses and other information processing devices, communication devices, servers, and generative AI. The user (here referring to a salesperson) wears smart glasses and faces the customer.

[0284] Hardware and Software Used

[0285] Information processing device: Smart glasses or personal computers [[ID=​ Communication device: A network module for communicating with a server

[0287] Server: A central device that analyzes question and sentiment information and requests a generative AI

[0288] Generative AI: A natural language generation model such as OpenAI GPT-3

[0289] Sentiment analysis engine: Face recognition and sentiment analysis software such as DeepFace

[0290] Natural language processing: An NLP module for analyzing the content of questions

[0291] Data processing and data calculation

[0292] The server performs the following data processing.

[0293] 1. Sentiment analysis: The server receives the customer's face image captured by the camera of the smart glasses and uses a sentiment analysis engine (e.g., DeepFace) to analyze the customer's sentiment in real time. This detects the customer's sentiment such as being confused, angry, satisfied, etc.

[0294] 2. Natural language processing: The server receives the text data of the question issued by the customer, analyzes the content of the question using a natural language processing module (e.g., spaCy or NLTK), and extracts keywords.

[0295] 3. Answer generation: Based on the results of sentiment analysis and natural language processing, a generative AI (e.g., OpenAI GPT-3) generates an appropriate answer. The generative AI provides a personalized answer considering the user's question and sentiment information.

[0296] 4. Result display: The generated answer is sent from the server to the information processing device and displayed on the display of the smart glasses.

[0297] Specific example

[0298] For example, if a customer has a confused expression and enters a question through smart glasses saying, "Please tell me the features of this TV," the server first analyzes that confused emotion. Next, it analyzes the content of the question and inputs the following prompt into the generative AI.

[0299] Example of a prompt:

[0300] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[0301] Based on this, the generative AI generates responses such as, "This TV has 4K resolution and excellent color reproduction. It also features the latest HDR technology, making it ideal for watching movies and sports. You'll immediately notice the difference when you actually use it." This response is displayed on the smart glasses, and the salesperson conveys it to the customer.

[0302] This system allows for real-time recognition of customer emotions and the provision of appropriate responses and product suggestions, which is expected to improve customer satisfaction.

[0303] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0304] Step 1:

[0305] The user puts on smart glasses and begins face-to-face interaction with the customer. The smart glasses' camera captures the customer's face and sends it to the information processing device. The input is the customer's facial image, and the output is real-time image data.

[0306] Step 2:

[0307] The information processing device transmits the captured face image of the customer to the server via the communication device. The server analyzes the customer's emotion using an emotion analysis engine (e.g., DeepFace). The input is the face image data, and customer emotion information (e.g., confusion, satisfaction, anger, etc.) is generated as the output. As a specific operation, the emotion analysis algorithm by DeepFace is applied.

[0308] Step 3:

[0309] The user receives the customer's question through the smart glasses and inputs the text information. The information processing device transmits this text data to the server via the communication device. The input is the question text, and the transmission of the question data to the server is performed as the output.

[0310] Step 4:

[0311] The server analyzes the received question text using a natural language processing module (e.g., spaCy or NLTK) and extracts the keywords of the question content. The input is the question text, and keyword data is generated as the output. As specific operations, the application of the natural language processing algorithm and the extraction of keywords are performed.

[0312] Step 5:

[0313] Based on the analyzed question content and emotion information, the server requests the generative AI (e.g., OpenAI GPT-3) to generate an answer. Here, a prompt text is generated and sent to the generative AI. The input is the question content and emotion information, and the creation of the prompt text is performed. An appropriate answer is generated as the output. As specific operations, there are the generation of the prompt text and the query transmission to the generative AI.

[0314] Example of prompt text:

[0315] The customer has an emotion of "confusion". The question content is: "Please tell me the features of this TV." Please propose an appropriate response.

[0316] Step 6:

[0317] The generative AI sends the generated response back to the server. The server then transmits the received response to an information processing device via a communication device. The input is the generated response text, and the output is the transmission of the response data.

[0318] Step 7:

[0319] The information processing device receives the response and displays it on the smart glasses' display. This allows the user to provide appropriate answers and product suggestions to customers in real time. The input is the response text sent from the server, and the output is the display of the response to the user. Specifically, this operation includes displaying text on the display.

[0320] The above processing steps will enhance customer service in physical stores, enabling real-time emotion recognition and appropriate responses.

[0321] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0322] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0323] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0324] [Second Embodiment]

[0325] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0326] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0327] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0328] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0329] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0330] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0331] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0332] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0333] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0334] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0335] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0336] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0337] The present invention is a system in which a user inputs a question using a terminal, the input question is sent to a server, the server receives the question and analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0338] First, the user enters a question using their device. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is then sent from the device to the server.

[0339] Next, the server receives the question and analyzes it using natural language processing. Keywords from the question are extracted through natural language processing. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0340] The server then requests a generative AI to generate an answer based on the analysis results. The generative AI generates an appropriate answer from the large dataset collected based on the analyzed keywords. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0341] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed on the screen and becomes viewable by the user.

[0342] This system is designed to allow users to easily understand and quickly perform procedures. By utilizing natural language processing and generative AI, it enables advanced question analysis and answer generation, allowing it to respond quickly and accurately to a wide range of user questions.

[0343] The following describes the processing flow.

[0344] Step 1:

[0345] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0346] Step 2:

[0347] The user submits the question to the server. When the user clicks the "Submit" button, the question is sent to the server as an HTTP POST request.

[0348] Step 3:

[0349] The server receives the question. The server receives the request at the API endpoint and extracts the question content from the request body.

[0350] Step 4:

[0351] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0352] Step 5:

[0353] The server requests the generative AI to generate an answer based on the analysis results. The server uses an API to send a request to the generative AI that includes the analyzed keyword information.

[0354] Step 6:

[0355] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates the response: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0356] Step 7:

[0357] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0358] Step 8:

[0359] The server sends the generated response to the terminal. The server then sends the generated response back to the terminal as an HTTP response.

[0360] Step 9:

[0361] The device displays the received response to the user. The device (e.g., a browser) parses the response received from the server and displays it on the screen. The user's screen displays the text: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0362] (Example 1)

[0363] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0364] The problem that this invention aims to solve is to provide a system that allows users to obtain quick and accurate answers when they have questions. Specifically, with conventional manual search methods, it is often time-consuming for users to find the information they need, and it is often difficult to access the appropriate information. To solve this problem, the invention aims to develop a system that automatically analyzes questions by combining natural language processing and generative AI models, and generates and provides answers immediately.

[0365] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0366] In this invention, the server includes means for receiving a question and analyzing it using natural language processing, means for requesting an AI model to generate an answer based on the analysis results, and means for transmitting the generated answer to a terminal. This makes it possible to provide highly accurate answers immediately to questions entered by users using an information terminal.

[0367] An "information terminal" is an electronic device used by users to input questions and send them to a server.

[0368] A "question" is the content of an inquiry that a user enters using an information terminal.

[0369] A "data server" is a central control unit that receives questions sent from information terminals and requests analysis and response generation.

[0370] "Natural language processing" is a technology that analyzes questions received by a data server and extracts key information from the content of those questions.

[0371] "Key information" refers to important words and phrases extracted from the question content through natural language processing.

[0372] A "generative AI model" is an artificial intelligence model used to generate appropriate responses based on the analysis results of a data server.

[0373] "Answer generation" is the process by which a generative AI model generates appropriate answers to questions based on analysis results.

[0374] "Transmission" refers to the act of a data server transferring a generated response to an information terminal.

[0375] "Display" refers to the process by which an information terminal visually presents the received response to the user.

[0376] The present invention is a system in which a user inputs a question using an information terminal, sends the question to a data server, the data server analyzes the received question using natural language processing, requests an AI model to generate an answer based on the analysis results, sends the answer generated by the AI ​​model back to the information terminal via the data server, and displays the result to the user.

[0377] First, the user enters a question using an information terminal. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is sent from the information terminal to the data server using an HTTP POST request.

[0378] Next, the data server receives the question and analyzes it using natural language processing (NLP). Specifically, it uses Google's NLP library to extract key information from the question. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0379] Subsequently, the data server requests a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the analysis results. The generative AI model generates an appropriate answer from the collected large dataset (e.g., publicly available data on the internet) based on the analyzed key information. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0380] The generated response is sent back to the data server, which then sends it to the user's information terminal. Finally, the information terminal displays the received response to the user. Specifically, a response such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed in an appropriate format, making it easy for the user to view.

[0381] This system enables users to understand procedures and take swift action by providing quick and accurate information when they have questions. By utilizing natural language processing and generative AI models, it enables advanced question analysis and answer generation, allowing for rapid responses to a wide range of user inquiries. Furthermore, the system's design allows users to instantly obtain necessary information without manually performing internet searches, significantly improving user convenience.

[0382] Example of a prompt:

[0383] "The user has entered the following question: 'How do I get a replacement passport?' Please generate the appropriate information to answer this question."

[0384] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0385] Step 1:

[0386] The user enters and submits a question.

[0387] The user enters a question into a text box on the information terminal interface and presses the "Send" button. For example, they might enter "Please tell me how to obtain a resident registration certificate" and click the send button.

[0388] Input: Question text entered by the user on the information terminal.

[0389] Output: The question text is sent as an HTTP POST request.

[0390] Step 2:

[0391] The device sends a question, and the server receives it.

[0392] The terminal sends the user-entered question to the server via an HTTP POST request. The server receives this question, logs it, and prepares a response.

[0393] Input: HTTP POST request generated in Step 1

[0394] Output: Question text stored on the server

[0395] Step 3:

[0396] The server analyzes the received question using natural language processing.

[0397] The server analyzes the received question using Google's NLP library. This analysis extracts key information from the question. For example, keywords such as "resident registration" and "how to obtain" may be identified.

[0398] Input: Question text stored on the server

[0399] Output: Key information extracted through analysis

[0400] Step 4:

[0401] The server sends prompt messages to the AI ​​model based on the analysis results.

[0402] The server uses the key information obtained from the analysis results to create and send prompt messages to the generated AI model. For example, it might create a prompt message such as, "Please answer the question based on the following keywords: resident registration certificate, how to obtain it."

[0403] Input: Key information extracted through analysis

[0404] Output: Prompt message sent to the generated AI model

[0405] Step 5:

[0406] The generative AI model generates the answer and sends it back to the server.

[0407] The generation AI model generates an appropriate response based on the prompt and sends that response back to the server. For example, it might generate the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0408] Input: Prompt sent to the generated AI model

[0409] Output: Generated response text sent back to the server

[0410] Step 6:

[0411] The server sends the generated response to the terminal.

[0412] The server sends the generated response to the terminal via an HTTP POST response.

[0413] Input: Response text sent from the generating AI model to the server.

[0414] Output: Response text sent to the terminal as an HTTP POST response

[0415] Step 7:

[0416] The device displays the received response to the user.

[0417] The device displays the response received from the server on the screen. The user can view the displayed response. For example, the response might say, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city / ward / town / village office. A fee may be charged."

[0418] Input: Response text received from the server

[0419] Output: Response text displayed to the user

[0420] (Application Example 1)

[0421] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0422] Traditional in-store customer support systems often made it difficult for customers to quickly and accurately obtain the information they needed, as information retrieval was cumbersome. Furthermore, reliance on interaction with staff increased their workload and led to inconsistent service quality. Additionally, there was a lack of mechanisms to provide detailed, real-time information about specific products. To address these challenges, technology that allows customers to access information smoothly is necessary.

[0423] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0424] In this invention, the server includes means for the user to input a question using a user interface, means for sending the question to the server, means for the server to receive the question and analyze it using natural language processing, means for the server to request a generative AI to generate an answer based on the analysis results, means for the generative AI to generate an answer and send it back to the server, means for the server to send the generated answer to a user device, means for the user device to display the transmitted answer to the user, and means for the user interface to be configured using a smart wearable device. This makes it possible for customers to input questions in real time through a smart wearable device and to quickly and accurately obtain relevant product information and guidance.

[0425] "User interface" is a general term for devices or software that allow users to directly operate or input information.

[0426] A "question" is the content that a user enters into the system to request information or guidance.

[0427] A "server" is a computer system that provides specific services over a network and processes data in response to user requests.

[0428] "Natural language processing" is a general term for technologies that allow computers to analyze and understand human language.

[0429] "Generative AI" refers to artificial intelligence that uses large datasets to produce appropriate responses or products in response to specified inputs.

[0430] "User device" is a general term for terminals and devices used by users. Examples include smartphones and smart glasses.

[0431] A "smart wearable device" is a general term for portable electronic devices that provide various functions when worn by the user. Specific examples include smart glasses and smartwatches.

[0432] "Keywords" are important words or phrases extracted from the question content that are used to generate answers or perform searches.

[0433] A "large-scale dataset" refers to the vast amount of data used to train generative AI, including text, images, and audio.

[0434] "Real-time" refers to a temporal attribute that means responding immediately to user requests, resulting in extremely low latency.

[0435] This invention is a system in which a user uses a smart wearable device to input questions in real time while in a store and receive appropriate answers.

[0436] First, the user puts on a smart wearable device such as smart glasses and inputs a question using the device's built-in user interface. For example, the user might input a question using the voice input function of the smart glasses, such as "Please tell me how to use this product." This question is then sent from the device to the server.

[0437] The server receives a question and analyzes its content using natural language processing techniques (such as libraries like spaCy or NLTK). Through this analysis, important keywords are extracted from the question. For example, the keyword "how to use the product" might be extracted.

[0438] Based on the analysis results, the server requests a generative AI (for example, OpenAI's GPT-3) to generate an answer. The generative AI uses a large, trained dataset to generate an appropriate answer. For example, it might generate an answer such as, "First, turn on the power to this product and follow the instructions to set it up. Then, download the dedicated app to use it."

[0439] The generated response is sent back to the server, which then transmits it to the user's smart wearable device. The device receives the transmitted response and displays it in the user's field of view. This allows the user to obtain the necessary information in real time.

[0440] For example, if a customer asks in a store, "What's the difference between this TV and this speaker?", the server analyzes the question and generates an answer such as, "This TV has 4K resolution and is HDR compatible. On the other hand, this speaker has Bluetooth functionality and provides clear sound throughout the room," which is then displayed on the smart glasses.

[0441] Example of a prompt:

[0442] "Please generate an answer to the following question using the AI ​​generator. The question is: 'What is the difference between this TV and this speaker?'"

[0443] "User question: 'Is this product waterproof?' Please have the AI ​​generate an answer."

[0444] This allows customers to easily obtain product information in-store, providing an efficient shopping experience.

[0445] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0446] Step 1:

[0447] The user enters a question via voice or touch input through a smart wearable device (e.g., smart glasses). The entered question is first stored in the internal memory of the smart wearable device through its interface, and then sent to a server over the network.

[0448] Input: User's question (e.g., "How do I use this product?")

[0449] Output: Question data is sent to the server via the network.

[0450] Step 2:

[0451] The server receives the question and analyzes its content using a natural language processing (NLP) engine. An NLP engine (e.g., spaCy or NLTK) is then used to extract important keywords from the question.

[0452] Input: Submitted question data

[0453] Output: Analyzed keywords (e.g., "How to use the product")

[0454] Step 3:

[0455] The server requests a generative AI (e.g., OpenAI's GPT-3) to generate an answer based on the extracted keywords. The generative AI model generates an appropriate answer based on a large dataset in which it has been trained. For example, it can generate a detailed answer about how to use a product based on a prompt.

[0456] Input: Analyzed keywords, prompt text (e.g., "Please tell me how to use this product.")

[0457] Output: Generated response (Example: "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0458] Step 4:

[0459] The generated responses are sent back to the server, which then transmits them to the user's smart wearable device. During this process, the server converts the response data into an appropriate format and prepares it for transmission.

[0460] Input: Generated answer

[0461] Output: Response data sent to smart wearable devices

[0462] Step 5:

[0463] The user's smart wearable device acquires the received response data and displays the response on its screen. The user can view the generated response in real time through the smart wearable device.

[0464] Input: Response data sent from the server

[0465] Output: A response displayed on the smart wearable device's screen (e.g., "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0466] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0467] The present invention is a system that recognizes not only the questions entered by the user but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system includes a process in which the user enters a question using a terminal, sends the question to a server, the server receives the question, analyzes the user's emotions using an emotion engine, then analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results and emotion information, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0468] First, the user enters a question using a device. At this stage, text information and voice data entered by the user are also collected. For example, the user might enter the question, "How do I obtain a resident registration certificate?"

[0469] Next, the voice and text data, along with the questions entered by the user, are sent to the server. The server receives this data and first analyzes the user's emotional state using an emotion engine. This emotion engine analyzes emotions from the context of the text entered by the user and from the voice data. For example, it may detect emotional information such as the user being anxious, angry, or troubled.

[0470] The server then passes the sentiment information and the question to a natural language processing module, which analyzes the question. Important keywords (e.g., resident registration, acquisition method) are extracted at this stage.

[0471] The server requests a generative AI to generate a response based on the analysis results and sentiment information. The generative AI then generates an appropriate response from a large dataset based on the analyzed information and sentiment information. For example, if the user is anxious, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office."

[0472] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office," might be displayed and become viewable by the user.

[0473] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0474] The following describes the processing flow.

[0475] Step 1:

[0476] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0477] Step 2:

[0478] The user sends sentiment data along with the question to the server. When the user clicks the "Submit" button, the question content and sentiment data (e.g., facial recognition and voice tone) are sent to the server as an HTTP POST request.

[0479] Step 3:

[0480] The server receives the question and sentiment data. The server receives the request at the API endpoint and extracts the question content and sentiment data from the request body.

[0481] Step 4:

[0482] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the text and voice data entered by the user to identify the user's emotions (e.g., anxious, angry, troubled).

[0483] Step 5:

[0484] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0485] Step 6:

[0486] The server requests the generative AI to generate an answer based on the analysis results and sentiment information. The server uses an API to send a request to the generative AI that includes the analyzed keyword information and the user's sentiment information.

[0487] Step 7:

[0488] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates an appropriate response (e.g., "Don't worry, to obtain a resident registration certificate...") based on analyzed keywords and sentiment information.

[0489] Step 8:

[0490] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0491] Step 9:

[0492] The server sends the generated response to the terminal. The server returns the generated response to the terminal as an HTTP response, providing an appropriate response that matches the user's emotions.

[0493] Step 10:

[0494] The terminal displays the received response to the user. The terminal analyzes the response received from the server and displays it on the screen. The user's screen displays a response such as "Don't worry, to obtain your resident registration certificate..." in an appropriate format.

[0495] (Example 2)

[0496] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0497] Traditional question-answering systems focused on analyzing user input and generating appropriate answers, but they were unable to recognize the user's emotional state and personalize answers accordingly. As a result, they provided uniform answers that ignored user emotions, leading to a poor user experience. Furthermore, they lacked sentiment analysis using voice data, making advanced analysis that reflected the user's input modality difficult. A new system is needed to address these issues.

[0498] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0499] In this invention, the server includes means for analyzing received questions and voice data with an emotion analysis engine, means for passing emotion information and question content to a natural language processing module for analysis, and means for requesting a generative AI to generate an answer based on the analysis results and emotion information. This makes it possible to provide personalized answers that take into account the user's emotional state.

[0500] A "terminal" is a device used by a user to input questions or voice data, and includes smartphones, PCs, tablets, and other similar devices.

[0501] A "server" is a central processing unit that analyzes received data and utilizes various engines and modules to generate appropriate responses.

[0502] A "sentiment analysis engine" is a software program that analyzes text and voice data sent by the user to detect the user's emotional state.

[0503] A "natural language processing module" is a software program that analyzes the content of a question, extracts important keywords, and understands the context.

[0504] "Generative AI" refers to artificial intelligence models that generate appropriate responses based on analyzed information and emotional information.

[0505] "Question and audio data" refers to the collective text and audio information entered by the user using their device.

[0506] "Analysis results" refers to the collective term for the question content analyzed by the natural language processing module and the sentiment information analyzed by the sentiment analysis engine.

[0507] "Answer" refers to response information generated by a generative AI based on its analysis results, and is presented to the user.

[0508] Modes for carrying out the invention

[0509] This invention is a system that recognizes not only the questions entered by the user, but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system utilizes an emotion analysis engine, a natural language processing module, and a generative AI.

[0510] The user first enters their question using a device, such as a smartphone, PC, or tablet. The text information and voice data entered by the user are also collected by the device. For example, a user might enter, "Please tell me how to obtain a resident registration certificate."

[0511] The entered questions and voice data are sent from the terminal to the server. The server receives this data and then uses an emotion analysis engine to analyze the user's emotional state. This emotion analysis engine detects the user's emotions from the context of the text, the tone of the voice data, the speed, etc. For example, it extracts emotional information such as whether the user is anxious, angry, or troubled.

[0512] Next, the server passes the sentiment information and question content to a natural language processing module for analysis. The natural language processing module extracts important keywords (e.g., resident registration, acquisition method) from the input question and analyzes the meaning of the sentences. This analysis result, along with the sentiment information, is then passed back to the server.

[0513] The server requests a generative AI to generate an answer based on the analysis results and sentiment information. This generative AI uses a generative AI model such as OpenAI GPT-4. Based on the analyzed information and sentiment information, the generative AI generates an appropriate answer from a large collected dataset. For example, if the user is anxious, it might generate an answer such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0514] The generated response is sent back to the server, which then sends it to the user's device. The device then displays the received response to the user. For example, the user's device screen might display a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0515] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0516] Specific example

[0517] For example, if a user enters the prompt "Please tell me how to obtain a resident registration certificate. I'm in a hurry," the sentiment analysis engine will analyze the emotion of "hurry," and the generative AI will generate a response that takes this into consideration. As a result, the user will be provided with the response, "Don't worry, to obtain a resident registration certificate, you need to submit an application form at your nearest city or town hall and present identification documents."

[0518] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0519] Step 1:

[0520] The user enters questions and voice data using a device.

[0521] Users input questions or voice data using devices such as smartphones, PCs, and tablets. For example, they might input, "Please tell me how to obtain a resident registration certificate." During this process, the entered text and voice data are converted into corresponding input formats.

[0522] Input: User's question text and audio data

[0523] Output: Text and audio data stored on the device

[0524] Step 2:

[0525] The terminal sends user input data to the server.

[0526] The terminal sends the collected text and audio data to the server. This transmission is done via an HTTP request or other appropriate communication protocol.

[0527] Input: Text and audio data stored on the device

[0528] Output: Text and audio data sent to the server

[0529] Step 3:

[0530] The server analyzes the received data using an emotion analysis engine.

[0531] The server receives text and audio data sent by the user and passes it to the sentiment analysis engine. The sentiment analysis engine analyzes the context of the text and the tone and speed of the audio data to detect the user's emotional state. For example, if the tone of the audio data is high and fast, it will detect an emotion of anxiety.

[0532] Input: Text data and audio data

[0533] Output: Analyzed emotion information (e.g., "anxious")

[0534] Step 4:

[0535] The server passes sentiment information and questions to a natural language processing module for analysis.

[0536] The sentiment analysis results and the user's question are sent to a natural language processing (NLP) module. The NLP module analyzes the context of the question and extracts important keywords (e.g., "resident registration," "how to obtain"). This allows the meaning of the question to be represented in a structured data format.

[0537] Input: Analyzed sentiment information and text data

[0538] Output: Analysis results with key keywords extracted.

[0539] Step 5:

[0540] The server requests a generative AI to generate an answer based on the analysis results and emotional information.

[0541] The server requests the generative AI to generate an answer based on the analysis results and sentiment information obtained from the NLP module. Specifically, it sends the generative AI a prompt message saying, "The user is anxious. Please answer the question, 'How do I obtain a resident registration certificate?'"

[0542] Input: Analysis results and sentiment information

[0543] Output: Prompt message for generative AI

[0544] Step 6:

[0545] Generative AI generates appropriate answers.

[0546] The generative AI generates a response based on the received prompt text. Here, the generated response is made with consideration for emotions. For example, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city or town hall."

[0547] Input: Prompt message for generative AI

[0548] Output: Appropriate answer

[0549] Step 7:

[0550] The generative AI sends the answer back to the server.

[0551] The generated response is sent back to the server. The server receives this response and prepares to send it to the terminal.

[0552] Input: Appropriate answer

[0553] Output: Response sent back to the server

[0554] Step 8:

[0555] The server sends the response to the terminal, and the terminal displays it to the user.

[0556] The server sends the received response to the user's device. The device receives this response and displays it to the user in an appropriate format. For example, the device screen might display: "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0557] Input: Response sent back to the server

[0558] Output: The answer displayed on the user's device.

[0559] (Application Example 2)

[0560] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0561] In recent years, improving customer satisfaction in physical stores has become increasingly important. However, traditional customer service systems have struggled to recognize customer emotions and provide real-time responses and product suggestions accordingly. Furthermore, they lacked the means to provide appropriate and personalized answers to customer questions. Because of these challenges, there is a need for systems that can achieve more sophisticated customer service.

[0562] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0563] In this invention, the server includes means for the user to input a question using an information processing device, means for transmitting the question to a communication device, means for the communication device to receive the question and analyze it using natural language processing, means for analyzing the user's emotions using an emotion analysis engine, means for requesting a generative AI to generate an answer based on the analysis results and emotion information, means for the generative AI to generate an answer considering the emotion information and send it back to the communication device, and means for transmitting the generated answer to the information processing device and displaying it to the user. This enables the recognition of customer emotions in real time and allows for appropriate responses, product suggestions, and personalized answers to questions accordingly.

[0564] An "information processing device" is an electronic device used to input questions from users and display the submitted answers.

[0565] A "communication device" is a device that receives questions from users and works in conjunction with a server to perform natural language processing and sentiment analysis.

[0566] "Natural language processing" is a technology that analyzes language data entered by a user to understand the meaning of the question.

[0567] A "sentiment analysis engine" is a system that implements technologies and algorithms for analyzing emotional information from text and voice data entered by the user.

[0568] "Generative AI" is artificial intelligence that generates answers based on trained data, taking into account the user's question content and emotional information.

[0569] A "question" refers to the content that a user inputs using an information processing device and sends to a server via a communication device.

[0570] "Analysis results" refer to the output of question content and sentiment information analyzed by the natural language processing and sentiment analysis engines.

[0571] "User emotion" refers to the emotional state detected by the emotion analysis engine from the user's input data.

[0572] An "answer" refers to the response or information generated by a generative AI based on the question and the user's emotions.

[0573] "Display" refers to the act of an information processing device visually providing an answer to the user.

[0574] "System" refers to the overall configuration and function of a series of means combined to realize the present invention.

[0575] This invention is an emotion-recognition customer service assistant system for enhancing customer service in physical stores. The specific implementation of this system is described below.

[0576] System Configuration

[0577] This system includes smart glasses and other information processing devices, communication devices, servers, and generative AI. The user (referring to a salesperson in this case) wears smart glasses and interacts with customers.

[0578] Hardware and software to be used

[0579] Information processing devices: Smart glasses and personal computers

[0580] Communication device: A network module for communicating with the server.

[0581] Server: The central device that analyzes questions and sentiment information and requests it to the generative AI.

[0582] Generative AI: Natural language generation models such as OpenAI GPT-3

[0583] Emotion analysis engine: Face recognition and emotion analysis software such as DeepFace

[0584] Natural Language Processing: NLP module for analyzing question content

[0585] Data processing and data calculation

[0586] The server performs the following data processing.

[0587] 1. Emotion Analysis: The customer's facial image, captured by the smart glasses' camera, is sent to a server, where an emotion analysis engine (e.g., DeepFace) is used to analyze the customer's emotions in real time. This allows for the detection of emotions such as confusion, anger, or satisfaction.

[0588] 2. Natural Language Processing: The server receives text data of questions asked by customers, analyzes the content of the questions using a natural language processing module (e.g., spaCy or NLTK), and extracts keywords.

[0589] 3. Answer Generation: Based on the results of sentiment analysis and natural language processing, a generative AI (e.g., OpenAI GPT-3) generates an appropriate answer. The generative AI takes into account the user's question and sentiment information to provide a personalized answer.

[0590] 4. Display of results: The generated answers are sent from the server to the information processing device and displayed on the smart glasses' screen.

[0591] Specific example

[0592] For example, if a customer has a confused expression and enters a question through smart glasses saying, "Please tell me the features of this TV," the server first analyzes that confused emotion. Next, it analyzes the content of the question and inputs the following prompt into the generative AI.

[0593] Example of a prompt:

[0594] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[0595] Based on this, the generative AI generates responses such as, "This TV has 4K resolution and excellent color reproduction. It also features the latest HDR technology, making it ideal for watching movies and sports. You'll immediately notice the difference when you actually use it." This response is displayed on the smart glasses, and the salesperson conveys it to the customer.

[0596] This system allows for real-time recognition of customer emotions and the provision of appropriate responses and product suggestions, which is expected to improve customer satisfaction.

[0597] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0598] Step 1:

[0599] The user puts on smart glasses and begins face-to-face interaction with the customer. The smart glasses' camera captures the customer's face and sends it to the information processing device. The input is the customer's facial image, and the output is real-time image data.

[0600] Step 2:

[0601] The information processing device transmits the captured customer's facial image to the server via a communication device. The server analyzes the customer's emotions using an emotion analysis engine (e.g., DeepFace). The input is facial image data, and the output is customer emotion information (e.g., confused, satisfied, angry). Specifically, the DeepFace emotion analysis algorithm is applied.

[0602] Step 3:

[0603] The user receives customer questions through smart glasses and inputs the text information. The information processing device transmits this text data to a server via a communication device. The input is the question text, and the output is the transmission of the question data to the server.

[0604] Step 4:

[0605] The server analyzes the received question text using a natural language processing module (e.g., spaCy or NLTK) and extracts keywords from the question. The input is the question text, and keyword data is generated as output. Specifically, the process involves applying a natural language processing algorithm and extracting keywords.

[0606] Step 5:

[0607] The server requests a generative AI (e.g., OpenAI GPT-3) to generate an answer based on the analyzed question content and sentiment information. A prompt is generated and sent to the generative AI. The input consists of the question content and sentiment information, and the prompt is created. An appropriate answer is generated as output. Specifically, the process involves generating a prompt and sending a query to the generative AI.

[0608] Example of a prompt:

[0609] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[0610] Step 6:

[0611] The generative AI sends the generated response back to the server. The server then transmits the received response to an information processing device via a communication device. The input is the generated response text, and the output is the transmission of the response data.

[0612] Step 7:

[0613] The information processing device receives the response and displays it on the smart glasses' display. This allows the user to provide appropriate answers and product suggestions to customers in real time. The input is the response text sent from the server, and the output is the display of the response to the user. Specifically, this operation includes displaying text on the display.

[0614] The above processing steps will enhance customer service in physical stores, enabling real-time emotion recognition and appropriate responses.

[0615] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0616] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0617] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0618] [Third Embodiment]

[0619] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0620] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0621] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0622] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0623] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0624] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0625] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0626] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0627] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0628] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0629] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0630] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0631] The present invention is a system in which a user inputs a question using a terminal, the input question is sent to a server, the server receives the question and analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0632] First, the user enters a question using their device. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is then sent from the device to the server.

[0633] Next, the server receives the question and analyzes it using natural language processing. Keywords from the question are extracted through natural language processing. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0634] The server then requests a generative AI to generate an answer based on the analysis results. The generative AI generates an appropriate answer from the large dataset collected based on the analyzed keywords. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0635] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed on the screen and becomes viewable by the user.

[0636] This system is designed to allow users to easily understand and quickly perform procedures. By utilizing natural language processing and generative AI, it enables advanced question analysis and answer generation, allowing it to respond quickly and accurately to a wide range of user questions.

[0637] The following describes the processing flow.

[0638] Step 1:

[0639] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0640] Step 2:

[0641] The user submits the question to the server. When the user clicks the "Submit" button, the question is sent to the server as an HTTP POST request.

[0642] Step 3:

[0643] The server receives the question. The server receives the request at the API endpoint and extracts the question content from the request body.

[0644] Step 4:

[0645] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0646] Step 5:

[0647] The server requests the generative AI to generate an answer based on the analysis results. The server uses an API to send a request to the generative AI that includes the analyzed keyword information.

[0648] Step 6:

[0649] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates the response: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0650] Step 7:

[0651] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0652] Step 8:

[0653] The server sends the generated response to the terminal. The server then sends the generated response back to the terminal as an HTTP response.

[0654] Step 9:

[0655] The device displays the received response to the user. The device (e.g., a browser) parses the response received from the server and displays it on the screen. The user's screen displays the text: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0656] (Example 1)

[0657] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0658] The problem that this invention aims to solve is to provide a system that allows users to obtain quick and accurate answers when they have questions. Specifically, with conventional manual search methods, it is often time-consuming for users to find the information they need, and it is often difficult to access the appropriate information. To solve this problem, the invention aims to develop a system that automatically analyzes questions by combining natural language processing and generative AI models, and generates and provides answers immediately.

[0659] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0660] In this invention, the server includes means for receiving a question and analyzing it using natural language processing, means for requesting an AI model to generate an answer based on the analysis results, and means for transmitting the generated answer to a terminal. This makes it possible to provide highly accurate answers immediately to questions entered by users using an information terminal.

[0661] An "information terminal" is an electronic device used by users to input questions and send them to a server.

[0662] A "question" is the content of an inquiry that a user enters using an information terminal.

[0663] A "data server" is a central control unit that receives questions sent from information terminals and requests analysis and response generation.

[0664] "Natural language processing" is a technology that analyzes questions received by a data server and extracts key information from the content of those questions.

[0665] "Key information" refers to important words and phrases extracted from the question content through natural language processing.

[0666] A "generative AI model" is an artificial intelligence model used to generate appropriate responses based on the analysis results of a data server.

[0667] "Answer generation" is the process by which a generative AI model generates appropriate answers to questions based on analysis results.

[0668] "Transmission" refers to the act of a data server transferring a generated response to an information terminal.

[0669] "Display" refers to the process by which an information terminal visually presents the received response to the user.

[0670] The present invention is a system in which a user inputs a question using an information terminal, sends the question to a data server, the data server analyzes the received question using natural language processing, requests an AI model to generate an answer based on the analysis results, sends the answer generated by the AI ​​model back to the information terminal via the data server, and displays the result to the user.

[0671] First, the user enters a question using an information terminal. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is sent from the information terminal to the data server using an HTTP POST request.

[0672] Next, the data server receives the question and analyzes it using natural language processing (NLP). Specifically, it uses Google's NLP library to extract key information from the question. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0673] Subsequently, the data server requests a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the analysis results. The generative AI model generates an appropriate answer from the collected large dataset (e.g., publicly available data on the internet) based on the analyzed key information. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0674] The generated response is sent back to the data server, which then sends it to the user's information terminal. Finally, the information terminal displays the received response to the user. Specifically, a response such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed in an appropriate format, making it easy for the user to view.

[0675] This system enables users to understand procedures and take swift action by providing quick and accurate information when they have questions. By utilizing natural language processing and generative AI models, it enables advanced question analysis and answer generation, allowing for rapid responses to a wide range of user inquiries. Furthermore, the system's design allows users to instantly obtain necessary information without manually performing internet searches, significantly improving user convenience.

[0676] Example of a prompt:

[0677] "The user has entered the following question: 'How do I get a replacement passport?' Please generate the appropriate information to answer this question."

[0678] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0679] Step 1:

[0680] The user enters and submits a question.

[0681] The user enters a question into a text box on the information terminal interface and presses the "Send" button. For example, they might enter "Please tell me how to obtain a resident registration certificate" and click the send button.

[0682] Input: Question text entered by the user on the information terminal.

[0683] Output: The question text is sent as an HTTP POST request.

[0684] Step 2:

[0685] The device sends a question, and the server receives it.

[0686] The terminal sends the user-entered question to the server via an HTTP POST request. The server receives this question, logs it, and prepares a response.

[0687] Input: HTTP POST request generated in Step 1

[0688] Output: Question text stored on the server

[0689] Step 3:

[0690] The server analyzes the received question using natural language processing.

[0691] The server analyzes the received question using Google's NLP library. This analysis extracts key information from the question. For example, keywords such as "resident registration" and "how to obtain" may be identified.

[0692] Input: Question text stored on the server

[0693] Output: Key information extracted through analysis

[0694] Step 4:

[0695] The server sends prompt messages to the AI ​​model based on the analysis results.

[0696] The server uses the key information obtained from the analysis results to create and send prompt messages to the generated AI model. For example, it might create a prompt message such as, "Please answer the question based on the following keywords: resident registration certificate, how to obtain it."

[0697] Input: Key information extracted through analysis

[0698] Output: Prompt message sent to the generated AI model

[0699] Step 5:

[0700] The generative AI model generates the answer and sends it back to the server.

[0701] The generation AI model generates an appropriate response based on the prompt and sends that response back to the server. For example, it might generate the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0702] Input: Prompt sent to the generated AI model

[0703] Output: Generated response text sent back to the server

[0704] Step 6:

[0705] The server sends the generated response to the terminal.

[0706] The server sends the generated response to the terminal via an HTTP POST response.

[0707] Input: Response text sent from the generating AI model to the server.

[0708] Output: Response text sent to the terminal as an HTTP POST response

[0709] Step 7:

[0710] The device displays the received response to the user.

[0711] The device displays the response received from the server on the screen. The user can view the displayed response. For example, the response might say, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city / ward / town / village office. A fee may be charged."

[0712] Input: Response text received from the server

[0713] Output: Response text displayed to the user

[0714] (Application Example 1)

[0715] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0716] Traditional in-store customer support systems often made it difficult for customers to quickly and accurately obtain the information they needed, as information retrieval was cumbersome. Furthermore, reliance on interaction with staff increased their workload and led to inconsistent service quality. Additionally, there was a lack of mechanisms to provide detailed, real-time information about specific products. To address these challenges, technology that allows customers to access information smoothly is necessary.

[0717] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0718] In this invention, the server includes means for the user to input a question using a user interface, means for sending the question to the server, means for the server to receive the question and analyze it using natural language processing, means for the server to request a generative AI to generate an answer based on the analysis results, means for the generative AI to generate an answer and send it back to the server, means for the server to send the generated answer to a user device, means for the user device to display the transmitted answer to the user, and means for the user interface to be configured using a smart wearable device. This makes it possible for customers to input questions in real time through a smart wearable device and to quickly and accurately obtain relevant product information and guidance.

[0719] "User interface" is a general term for devices or software that allow users to directly operate or input information.

[0720] A "question" is the content that a user enters into the system to request information or guidance.

[0721] A "server" is a computer system that provides specific services over a network and processes data in response to user requests.

[0722] "Natural language processing" is a general term for technologies that allow computers to analyze and understand human language.

[0723] "Generative AI" refers to artificial intelligence that uses large datasets to produce appropriate responses or products in response to specified inputs.

[0724] "User device" is a general term for terminals and devices used by users. Examples include smartphones and smart glasses.

[0725] A "smart wearable device" is a general term for portable electronic devices that provide various functions when worn by the user. Specific examples include smart glasses and smartwatches.

[0726] "Keywords" are important words or phrases extracted from the question content that are used to generate answers or perform searches.

[0727] A "large-scale dataset" refers to the vast amount of data used to train generative AI, including text, images, and audio.

[0728] "Real-time" refers to a temporal attribute that means responding immediately to user requests, resulting in extremely low latency.

[0729] This invention is a system in which a user uses a smart wearable device to input questions in real time while in a store and receive appropriate answers.

[0730] First, the user puts on a smart wearable device such as smart glasses and inputs a question using the device's built-in user interface. For example, the user might input a question using the voice input function of the smart glasses, such as "Please tell me how to use this product." This question is then sent from the device to the server.

[0731] The server receives a question and analyzes its content using natural language processing techniques (such as libraries like spaCy or NLTK). Through this analysis, important keywords are extracted from the question. For example, the keyword "how to use the product" might be extracted.

[0732] Based on the analysis results, the server requests a generative AI (for example, OpenAI's GPT-3) to generate an answer. The generative AI uses a large, trained dataset to generate an appropriate answer. For example, it might generate an answer such as, "First, turn on the power to this product and follow the instructions to set it up. Then, download the dedicated app to use it."

[0733] The generated response is sent back to the server, which then transmits it to the user's smart wearable device. The device receives the transmitted response and displays it in the user's field of view. This allows the user to obtain the necessary information in real time.

[0734] For example, if a customer asks in a store, "What's the difference between this TV and this speaker?", the server analyzes the question and generates an answer such as, "This TV has 4K resolution and is HDR compatible. On the other hand, this speaker has Bluetooth functionality and provides clear sound throughout the room," which is then displayed on the smart glasses.

[0735] Example of a prompt:

[0736] "Please generate an answer to the following question using the AI ​​generator. The question is: 'What is the difference between this TV and this speaker?'"

[0737] "User question: 'Is this product waterproof?' Please have the AI ​​generate an answer."

[0738] This allows customers to easily obtain product information in-store, providing an efficient shopping experience.

[0739] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0740] Step 1:

[0741] The user enters a question via voice or touch input through a smart wearable device (e.g., smart glasses). The entered question is first stored in the internal memory of the smart wearable device through its interface, and then sent to a server over the network.

[0742] Input: User's question (e.g., "How do I use this product?")

[0743] Output: Question data is sent to the server via the network.

[0744] Step 2:

[0745] The server receives the question and analyzes its content using a natural language processing (NLP) engine. An NLP engine (e.g., spaCy or NLTK) is then used to extract important keywords from the question.

[0746] Input: Submitted question data

[0747] Output: Analyzed keywords (e.g., "How to use the product")

[0748] Step 3:

[0749] The server requests a generative AI (e.g., OpenAI's GPT-3) to generate an answer based on the extracted keywords. The generative AI model generates an appropriate answer based on a large dataset in which it has been trained. For example, it can generate a detailed answer about how to use a product based on a prompt.

[0750] Input: Analyzed keywords, prompt text (e.g., "Please tell me how to use this product.")

[0751] Output: Generated response (Example: "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0752] Step 4:

[0753] The generated responses are sent back to the server, which then transmits them to the user's smart wearable device. During this process, the server converts the response data into an appropriate format and prepares it for transmission.

[0754] Input: Generated answer

[0755] Output: Response data sent to smart wearable devices

[0756] Step 5:

[0757] The user's smart wearable device acquires the received response data and displays the response on its screen. The user can view the generated response in real time through the smart wearable device.

[0758] Input: Response data sent from the server

[0759] Output: A response displayed on the smart wearable device's screen (e.g., "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[0760] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0761] The present invention is a system that recognizes not only the questions entered by the user but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system includes a process in which the user enters a question using a terminal, sends the question to a server, the server receives the question, analyzes the user's emotions using an emotion engine, then analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results and emotion information, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0762] First, the user enters a question using a device. At this stage, text information and voice data entered by the user are also collected. For example, the user might enter the question, "How do I obtain a resident registration certificate?"

[0763] Next, the voice and text data, along with the questions entered by the user, are sent to the server. The server receives this data and first analyzes the user's emotional state using an emotion engine. This emotion engine analyzes emotions from the context of the text entered by the user and from the voice data. For example, it may detect emotional information such as the user being anxious, angry, or troubled.

[0764] The server then passes the sentiment information and the question to a natural language processing module, which analyzes the question. Important keywords (e.g., resident registration, acquisition method) are extracted at this stage.

[0765] The server requests a generative AI to generate a response based on the analysis results and sentiment information. The generative AI then generates an appropriate response from a large dataset based on the analyzed information and sentiment information. For example, if the user is anxious, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office."

[0766] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office," might be displayed and become viewable by the user.

[0767] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0768] The following describes the processing flow.

[0769] Step 1:

[0770] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0771] Step 2:

[0772] The user sends sentiment data along with the question to the server. When the user clicks the "Submit" button, the question content and sentiment data (e.g., facial recognition and voice tone) are sent to the server as an HTTP POST request.

[0773] Step 3:

[0774] The server receives the question and sentiment data. The server receives the request at the API endpoint and extracts the question content and sentiment data from the request body.

[0775] Step 4:

[0776] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the text and voice data entered by the user to identify the user's emotions (e.g., anxious, angry, troubled).

[0777] Step 5:

[0778] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0779] Step 6:

[0780] The server requests the generative AI to generate an answer based on the analysis results and sentiment information. The server uses an API to send a request to the generative AI that includes the analyzed keyword information and the user's sentiment information.

[0781] Step 7:

[0782] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates an appropriate response (e.g., "Don't worry, to obtain a resident registration certificate...") based on analyzed keywords and sentiment information.

[0783] Step 8:

[0784] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0785] Step 9:

[0786] The server sends the generated response to the terminal. The server returns the generated response to the terminal as an HTTP response, providing an appropriate response that matches the user's emotions.

[0787] Step 10:

[0788] The terminal displays the received response to the user. The terminal analyzes the response received from the server and displays it on the screen. The user's screen displays a response such as "Don't worry, to obtain your resident registration certificate..." in an appropriate format.

[0789] (Example 2)

[0790] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0791] Traditional question-answering systems focused on analyzing user input and generating appropriate answers, but they were unable to recognize the user's emotional state and personalize answers accordingly. As a result, they provided uniform answers that ignored user emotions, leading to a poor user experience. Furthermore, they lacked sentiment analysis using voice data, making advanced analysis that reflected the user's input modality difficult. A new system is needed to address these issues.

[0792] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0793] In this invention, the server includes means for analyzing received questions and voice data with an emotion analysis engine, means for passing emotion information and question content to a natural language processing module for analysis, and means for requesting a generative AI to generate an answer based on the analysis results and emotion information. This makes it possible to provide personalized answers that take into account the user's emotional state.

[0794] A "terminal" is a device used by a user to input questions or voice data, and includes smartphones, PCs, tablets, and other similar devices.

[0795] A "server" is a central processing unit that analyzes received data and utilizes various engines and modules to generate appropriate responses.

[0796] A "sentiment analysis engine" is a software program that analyzes text and voice data sent by the user to detect the user's emotional state.

[0797] A "natural language processing module" is a software program that analyzes the content of a question, extracts important keywords, and understands the context.

[0798] "Generative AI" refers to artificial intelligence models that generate appropriate responses based on analyzed information and emotional information.

[0799] "Question and audio data" refers to the collective text and audio information entered by the user using their device.

[0800] "Analysis results" refers to the collective term for the question content analyzed by the natural language processing module and the sentiment information analyzed by the sentiment analysis engine.

[0801] "Answer" refers to response information generated by a generative AI based on its analysis results, and is presented to the user.

[0802] Modes for carrying out the invention

[0803] This invention is a system that recognizes not only the questions entered by the user, but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system utilizes an emotion analysis engine, a natural language processing module, and a generative AI.

[0804] The user first enters their question using a device, such as a smartphone, PC, or tablet. The text information and voice data entered by the user are also collected by the device. For example, a user might enter, "Please tell me how to obtain a resident registration certificate."

[0805] The entered questions and voice data are sent from the terminal to the server. The server receives this data and then uses an emotion analysis engine to analyze the user's emotional state. This emotion analysis engine detects the user's emotions from the context of the text, the tone of the voice data, the speed, etc. For example, it extracts emotional information such as whether the user is anxious, angry, or troubled.

[0806] Next, the server passes the sentiment information and question content to a natural language processing module for analysis. The natural language processing module extracts important keywords (e.g., resident registration, acquisition method) from the input question and analyzes the meaning of the sentences. This analysis result, along with the sentiment information, is then passed back to the server.

[0807] The server requests a generative AI to generate an answer based on the analysis results and sentiment information. This generative AI uses a generative AI model such as OpenAI GPT-4. Based on the analyzed information and sentiment information, the generative AI generates an appropriate answer from a large collected dataset. For example, if the user is anxious, it might generate an answer such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0808] The generated response is sent back to the server, which then sends it to the user's device. The device then displays the received response to the user. For example, the user's device screen might display a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0809] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[0810] Specific example

[0811] For example, if a user enters the prompt "Please tell me how to obtain a resident registration certificate. I'm in a hurry," the sentiment analysis engine will analyze the emotion of "hurry," and the generative AI will generate a response that takes this into consideration. As a result, the user will be provided with the response, "Don't worry, to obtain a resident registration certificate, you need to submit an application form at your nearest city or town hall and present identification documents."

[0812] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0813] Step 1:

[0814] The user enters questions and voice data using a device.

[0815] Users input questions or voice data using devices such as smartphones, PCs, and tablets. For example, they might input, "Please tell me how to obtain a resident registration certificate." During this process, the entered text and voice data are converted into corresponding input formats.

[0816] Input: User's question text and audio data

[0817] Output: Text and audio data stored on the device

[0818] Step 2:

[0819] The terminal sends user input data to the server.

[0820] The terminal sends the collected text and audio data to the server. This transmission is done via an HTTP request or other appropriate communication protocol.

[0821] Input: Text and audio data stored on the device

[0822] Output: Text and audio data sent to the server

[0823] Step 3:

[0824] The server analyzes the received data using an emotion analysis engine.

[0825] The server receives text and audio data sent by the user and passes it to the sentiment analysis engine. The sentiment analysis engine analyzes the context of the text and the tone and speed of the audio data to detect the user's emotional state. For example, if the tone of the audio data is high and fast, it will detect an emotion of anxiety.

[0826] Input: Text data and audio data

[0827] Output: Analyzed emotion information (e.g., "anxious")

[0828] Step 4:

[0829] The server passes sentiment information and questions to a natural language processing module for analysis.

[0830] The sentiment analysis results and the user's question are sent to a natural language processing (NLP) module. The NLP module analyzes the context of the question and extracts important keywords (e.g., "resident registration," "how to obtain"). This allows the meaning of the question to be represented in a structured data format.

[0831] Input: Analyzed sentiment information and text data

[0832] Output: Analysis results with key keywords extracted.

[0833] Step 5:

[0834] The server requests a generative AI to generate an answer based on the analysis results and emotional information.

[0835] The server requests the generative AI to generate an answer based on the analysis results and sentiment information obtained from the NLP module. Specifically, it sends the generative AI a prompt message saying, "The user is anxious. Please answer the question, 'How do I obtain a resident registration certificate?'"

[0836] Input: Analysis results and sentiment information

[0837] Output: Prompt message for generative AI

[0838] Step 6:

[0839] Generative AI generates appropriate answers.

[0840] The generative AI generates a response based on the received prompt text. Here, the generated response is made with consideration for emotions. For example, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city or town hall."

[0841] Input: Prompt message for generative AI

[0842] Output: Appropriate answer

[0843] Step 7:

[0844] The generative AI sends the answer back to the server.

[0845] The generated response is sent back to the server. The server receives this response and prepares to send it to the terminal.

[0846] Input: Appropriate answer

[0847] Output: Response sent back to the server

[0848] Step 8:

[0849] The server sends the response to the terminal, and the terminal displays it to the user.

[0850] The server sends the received response to the user's device. The device receives this response and displays it to the user in an appropriate format. For example, the device screen might display: "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[0851] Input: Response sent back to the server

[0852] Output: The answer displayed on the user's device.

[0853] (Application Example 2)

[0854] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0855] In recent years, improving customer satisfaction in physical stores has become increasingly important. However, traditional customer service systems have struggled to recognize customer emotions and provide real-time responses and product suggestions accordingly. Furthermore, they lacked the means to provide appropriate and personalized answers to customer questions. Because of these challenges, there is a need for systems that can achieve more sophisticated customer service.

[0856] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0857] In this invention, the server includes means for the user to input a question using an information processing device, means for transmitting the question to a communication device, means for the communication device to receive the question and analyze it using natural language processing, means for analyzing the user's emotions using an emotion analysis engine, means for requesting a generative AI to generate an answer based on the analysis results and emotion information, means for the generative AI to generate an answer considering the emotion information and send it back to the communication device, and means for transmitting the generated answer to the information processing device and displaying it to the user. This enables the recognition of customer emotions in real time and allows for appropriate responses, product suggestions, and personalized answers to questions accordingly.

[0858] An "information processing device" is an electronic device used to input questions from users and display the submitted answers.

[0859] A "communication device" is a device that receives questions from users and works in conjunction with a server to perform natural language processing and sentiment analysis.

[0860] "Natural language processing" is a technology that analyzes language data entered by a user to understand the meaning of the question.

[0861] A "sentiment analysis engine" is a system that implements technologies and algorithms for analyzing emotional information from text and voice data entered by the user.

[0862] "Generative AI" is artificial intelligence that generates answers based on trained data, taking into account the user's question content and emotional information.

[0863] A "question" refers to the content that a user inputs using an information processing device and sends to a server via a communication device.

[0864] "Analysis results" refer to the output of question content and sentiment information analyzed by the natural language processing and sentiment analysis engines.

[0865] "User emotion" refers to the emotional state detected by the emotion analysis engine from the user's input data.

[0866] An "answer" refers to the response or information generated by a generative AI based on the question and the user's emotions.

[0867] "Display" refers to the act of an information processing device visually providing an answer to the user.

[0868] "System" refers to the overall configuration and function of a series of means combined to realize the present invention.

[0869] This invention is an emotion-recognition customer service assistant system for enhancing customer service in physical stores. The specific implementation of this system is described below.

[0870] System Configuration

[0871] This system includes smart glasses and other information processing devices, communication devices, servers, and generative AI. The user (referring to a salesperson in this case) wears smart glasses and interacts with customers.

[0872] Hardware and software to be used

[0873] Information processing devices: Smart glasses and personal computers

[0874] Communication device: A network module for communicating with the server.

[0875] Server: The central device that analyzes questions and sentiment information and requests it to the generative AI.

[0876] Generative AI: Natural language generation models such as OpenAI GPT-3

[0877] Emotion analysis engine: Face recognition and emotion analysis software such as DeepFace

[0878] Natural Language Processing: NLP module for analyzing question content

[0879] Data processing and data calculation

[0880] The server performs the following data processing.

[0881] 1. Emotion Analysis: The customer's facial image, captured by the smart glasses' camera, is sent to a server, where an emotion analysis engine (e.g., DeepFace) is used to analyze the customer's emotions in real time. This allows for the detection of emotions such as confusion, anger, or satisfaction.

[0882] 2. Natural Language Processing: The server receives text data of questions asked by customers, analyzes the content of the questions using a natural language processing module (e.g., spaCy or NLTK), and extracts keywords.

[0883] 3. Answer Generation: Based on the results of sentiment analysis and natural language processing, a generative AI (e.g., OpenAI GPT-3) generates an appropriate answer. The generative AI takes into account the user's question and sentiment information to provide a personalized answer.

[0884] 4. Display of results: The generated answers are sent from the server to the information processing device and displayed on the smart glasses' screen.

[0885] Specific example

[0886] For example, if a customer has a confused expression and enters a question through smart glasses saying, "Please tell me the features of this TV," the server first analyzes that confused emotion. Next, it analyzes the content of the question and inputs the following prompt into the generative AI.

[0887] Example of a prompt:

[0888] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[0889] Based on this, the generative AI generates responses such as, "This TV has 4K resolution and excellent color reproduction. It also features the latest HDR technology, making it ideal for watching movies and sports. You'll immediately notice the difference when you actually use it." This response is displayed on the smart glasses, and the salesperson conveys it to the customer.

[0890] This system allows for real-time recognition of customer emotions and the provision of appropriate responses and product suggestions, which is expected to improve customer satisfaction.

[0891] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0892] Step 1:

[0893] The user puts on smart glasses and begins face-to-face interaction with the customer. The smart glasses' camera captures the customer's face and sends it to the information processing device. The input is the customer's facial image, and the output is real-time image data.

[0894] Step 2:

[0895] The information processing device transmits the captured customer's facial image to the server via a communication device. The server analyzes the customer's emotions using an emotion analysis engine (e.g., DeepFace). The input is facial image data, and the output is customer emotion information (e.g., confused, satisfied, angry). Specifically, the DeepFace emotion analysis algorithm is applied.

[0896] Step 3:

[0897] The user receives customer questions through smart glasses and inputs the text information. The information processing device transmits this text data to a server via a communication device. The input is the question text, and the output is the transmission of the question data to the server.

[0898] Step 4:

[0899] The server analyzes the received question text using a natural language processing module (e.g., spaCy or NLTK) and extracts keywords from the question. The input is the question text, and keyword data is generated as output. Specifically, the process involves applying a natural language processing algorithm and extracting keywords.

[0900] Step 5:

[0901] The server requests a generative AI (e.g., OpenAI GPT-3) to generate an answer based on the analyzed question content and sentiment information. A prompt is generated and sent to the generative AI. The input consists of the question content and sentiment information, and the prompt is created. An appropriate answer is generated as output. Specifically, the process involves generating a prompt and sending a query to the generative AI.

[0902] Example of a prompt:

[0903] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[0904] Step 6:

[0905] The generative AI sends the generated response back to the server. The server then transmits the received response to an information processing device via a communication device. The input is the generated response text, and the output is the transmission of the response data.

[0906] Step 7:

[0907] The information processing device receives the response and displays it on the smart glasses' display. This allows the user to provide appropriate answers and product suggestions to customers in real time. The input is the response text sent from the server, and the output is the display of the response to the user. Specifically, this operation includes displaying text on the display.

[0908] The above processing steps will enhance customer service in physical stores, enabling real-time emotion recognition and appropriate responses.

[0909] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0910] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0911] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0912] [Fourth Embodiment]

[0913] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0914] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0915] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0916] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0917] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0918] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0919] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0920] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0921] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0922] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0923] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0924] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0925] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0926] The present invention is a system in which a user inputs a question using a terminal, the input question is sent to a server, the server receives the question and analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[0927] First, the user enters a question using their device. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is then sent from the device to the server.

[0928] Next, the server receives the question and analyzes it using natural language processing. Keywords from the question are extracted through natural language processing. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0929] The server then requests a generative AI to generate an answer based on the analysis results. The generative AI generates an appropriate answer from the large dataset collected based on the analyzed keywords. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0930] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed on the screen and becomes viewable by the user.

[0931] This system is designed to allow users to easily understand and quickly perform procedures. By utilizing natural language processing and generative AI, it enables advanced question analysis and answer generation, allowing it to respond quickly and accurately to a wide range of user questions.

[0932] The following describes the processing flow.

[0933] Step 1:

[0934] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[0935] Step 2:

[0936] The user submits the question to the server. When the user clicks the "Submit" button, the question is sent to the server as an HTTP POST request.

[0937] Step 3:

[0938] The server receives the question. The server receives the request at the API endpoint and extracts the question content from the request body.

[0939] Step 4:

[0940] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[0941] Step 5:

[0942] The server requests the generative AI to generate an answer based on the analysis results. The server uses an API to send a request to the generative AI that includes the analyzed keyword information.

[0943] Step 6:

[0944] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates the response: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0945] Step 7:

[0946] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[0947] Step 8:

[0948] The server sends the generated response to the terminal. The server then sends the generated response back to the terminal as an HTTP response.

[0949] Step 9:

[0950] The device displays the received response to the user. The device (e.g., a browser) parses the response received from the server and displays it on the screen. The user's screen displays the text: "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0951] (Example 1)

[0952] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0953] The problem that this invention aims to solve is to provide a system that allows users to obtain quick and accurate answers when they have questions. Specifically, with conventional manual search methods, it is often time-consuming for users to find the information they need, and it is often difficult to access the appropriate information. To solve this problem, the invention aims to develop a system that automatically analyzes questions by combining natural language processing and generative AI models, and generates and provides answers immediately.

[0954] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0955] In this invention, the server includes means for receiving a question and analyzing it using natural language processing, means for requesting an AI model to generate an answer based on the analysis results, and means for transmitting the generated answer to a terminal. This makes it possible to provide highly accurate answers immediately to questions entered by users using an information terminal.

[0956] An "information terminal" is an electronic device used by users to input questions and send them to a server.

[0957] A "question" is the content of an inquiry that a user enters using an information terminal.

[0958] A "data server" is a central control unit that receives questions sent from information terminals and requests analysis and response generation.

[0959] "Natural language processing" is a technology that analyzes questions received by a data server and extracts key information from the content of those questions.

[0960] "Key information" refers to important words and phrases extracted from the question content through natural language processing.

[0961] A "generative AI model" is an artificial intelligence model used to generate appropriate responses based on the analysis results of a data server.

[0962] "Answer generation" is the process by which a generative AI model generates appropriate answers to questions based on analysis results.

[0963] "Transmission" refers to the act of a data server transferring a generated response to an information terminal.

[0964] "Display" refers to the process by which an information terminal visually presents the received response to the user.

[0965] The present invention is a system in which a user inputs a question using an information terminal, sends the question to a data server, the data server analyzes the received question using natural language processing, requests an AI model to generate an answer based on the analysis results, sends the answer generated by the AI ​​model back to the information terminal via the data server, and displays the result to the user.

[0966] First, the user enters a question using an information terminal. For example, the user might enter the question, "How do I obtain a resident registration certificate?" The question entered by the user is sent from the information terminal to the data server using an HTTP POST request.

[0967] Next, the data server receives the question and analyzes it using natural language processing (NLP). Specifically, it uses Google's NLP library to extract key information from the question. For example, keywords related to "resident registration" and "how to obtain" are extracted.

[0968] Subsequently, the data server requests a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the analysis results. The generative AI model generates an appropriate answer from the collected large dataset (e.g., publicly available data on the internet) based on the analyzed key information. For example, it might generate an answer such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged."

[0969] The generated response is sent back to the data server, which then sends it to the user's information terminal. Finally, the information terminal displays the received response to the user. Specifically, a response such as, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. A fee may be charged," is displayed in an appropriate format, making it easy for the user to view.

[0970] This system enables users to understand procedures and take swift action by providing quick and accurate information when they have questions. By utilizing natural language processing and generative AI models, it enables advanced question analysis and answer generation, allowing for rapid responses to a wide range of user inquiries. Furthermore, the system's design allows users to instantly obtain necessary information without manually performing internet searches, significantly improving user convenience.

[0971] Example of a prompt:

[0972] "The user has entered the following question: 'How do I get a replacement passport?' Please generate the appropriate information to answer this question."

[0973] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0974] Step 1:

[0975] The user enters and submits a question.

[0976] The user enters a question into a text box on the information terminal interface and presses the "Send" button. For example, they might enter "Please tell me how to obtain a resident registration certificate" and click the send button.

[0977] Input: Question text entered by the user on the information terminal.

[0978] Output: The question text is sent as an HTTP POST request.

[0979] Step 2:

[0980] The device sends a question, and the server receives it.

[0981] The terminal sends the user-entered question to the server via an HTTP POST request. The server receives this question, logs it, and prepares a response.

[0982] Input: HTTP POST request generated in Step 1

[0983] Output: Question text stored on the server

[0984] Step 3:

[0985] The server analyzes the received question using natural language processing.

[0986] The server analyzes the received question using Google's NLP library. This analysis extracts key information from the question. For example, keywords such as "resident registration" and "how to obtain" may be identified.

[0987] Input: Question text stored on the server

[0988] Output: Key information extracted through analysis

[0989] Step 4:

[0990] The server sends prompt messages to the AI ​​model based on the analysis results.

[0991] The server uses the key information obtained from the analysis results to create and send prompt messages to the generated AI model. For example, it might create a prompt message such as, "Please answer the question based on the following keywords: resident registration certificate, how to obtain it."

[0992] Input: Key information extracted through analysis

[0993] Output: Prompt message sent to the generated AI model

[0994] Step 5:

[0995] The generative AI model generates the answer and sends it back to the server.

[0996] The generation AI model generates an appropriate response based on the prompt and sends that response back to the server. For example, it might generate the response, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved."

[0997] Input: Prompt sent to the generated AI model

[0998] Output: Generated response text sent back to the server

[0999] Step 6:

[1000] The server sends the generated response to the terminal.

[1001] The server sends the generated response to the terminal via an HTTP POST response.

[1002] Input: Response text sent from the generating AI model to the server.

[1003] Output: Response text sent to the terminal as an HTTP POST response

[1004] Step 7:

[1005] The device displays the received response to the user.

[1006] The device displays the response received from the server on the screen. The user can view the displayed response. For example, the response might say, "To obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city / ward / town / village office. A fee may be charged."

[1007] Input: Response text received from the server

[1008] Output: Response text displayed to the user

[1009] (Application Example 1)

[1010] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1011] Traditional in-store customer support systems often made it difficult for customers to quickly and accurately obtain the information they needed, as information retrieval was cumbersome. Furthermore, reliance on interaction with staff increased their workload and led to inconsistent service quality. Additionally, there was a lack of mechanisms to provide detailed, real-time information about specific products. To address these challenges, technology that allows customers to access information smoothly is necessary.

[1012] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1013] In this invention, the server includes means for the user to input a question using a user interface, means for sending the question to the server, means for the server to receive the question and analyze it using natural language processing, means for the server to request a generative AI to generate an answer based on the analysis results, means for the generative AI to generate an answer and send it back to the server, means for the server to send the generated answer to a user device, means for the user device to display the transmitted answer to the user, and means for the user interface to be configured using a smart wearable device. This makes it possible for customers to input questions in real time through a smart wearable device and to quickly and accurately obtain relevant product information and guidance.

[1014] "User interface" is a general term for devices or software that allow users to directly operate or input information.

[1015] A "question" is the content that a user enters into the system to request information or guidance.

[1016] A "server" is a computer system that provides specific services over a network and processes data in response to user requests.

[1017] "Natural language processing" is a general term for technologies that allow computers to analyze and understand human language.

[1018] "Generative AI" refers to artificial intelligence that uses large datasets to produce appropriate responses or products in response to specified inputs.

[1019] "User device" is a general term for terminals and devices used by users. Examples include smartphones and smart glasses.

[1020] A "smart wearable device" is a general term for portable electronic devices that provide various functions when worn by the user. Specific examples include smart glasses and smartwatches.

[1021] "Keywords" are important words or phrases extracted from the question content that are used to generate answers or perform searches.

[1022] A "large-scale dataset" refers to the vast amount of data used to train generative AI, including text, images, and audio.

[1023] "Real-time" refers to a temporal attribute that means responding immediately to user requests, resulting in extremely low latency.

[1024] This invention is a system in which a user uses a smart wearable device to input questions in real time while in a store and receive appropriate answers.

[1025] First, the user puts on a smart wearable device such as smart glasses and inputs a question using the device's built-in user interface. For example, the user might input a question using the voice input function of the smart glasses, such as "Please tell me how to use this product." This question is then sent from the device to the server.

[1026] The server receives a question and analyzes its content using natural language processing techniques (such as libraries like spaCy or NLTK). Through this analysis, important keywords are extracted from the question. For example, the keyword "how to use the product" might be extracted.

[1027] Based on the analysis results, the server requests a generative AI (for example, OpenAI's GPT-3) to generate an answer. The generative AI uses a large, trained dataset to generate an appropriate answer. For example, it might generate an answer such as, "First, turn on the power to this product and follow the instructions to set it up. Then, download the dedicated app to use it."

[1028] The generated response is sent back to the server, which then transmits it to the user's smart wearable device. The device receives the transmitted response and displays it in the user's field of view. This allows the user to obtain the necessary information in real time.

[1029] For example, if a customer asks in a store, "What's the difference between this TV and this speaker?", the server analyzes the question and generates an answer such as, "This TV has 4K resolution and is HDR compatible. On the other hand, this speaker has Bluetooth functionality and provides clear sound throughout the room," which is then displayed on the smart glasses.

[1030] Example of a prompt:

[1031] "Please generate an answer to the following question using the AI ​​generator. The question is: 'What is the difference between this TV and this speaker?'"

[1032] "User question: 'Is this product waterproof?' Please have the AI ​​generate an answer."

[1033] This allows customers to easily obtain product information in-store, providing an efficient shopping experience.

[1034] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1035] Step 1:

[1036] The user enters a question via voice or touch input through a smart wearable device (e.g., smart glasses). The entered question is first stored in the internal memory of the smart wearable device through its interface, and then sent to a server over the network.

[1037] Input: User's question (e.g., "How do I use this product?")

[1038] Output: Question data is sent to the server via the network.

[1039] Step 2:

[1040] The server receives the question and analyzes its content using a natural language processing (NLP) engine. An NLP engine (e.g., spaCy or NLTK) is then used to extract important keywords from the question.

[1041] Input: Submitted question data

[1042] Output: Analyzed keywords (e.g., "How to use the product")

[1043] Step 3:

[1044] The server requests a generative AI (e.g., OpenAI's GPT-3) to generate an answer based on the extracted keywords. The generative AI model generates an appropriate answer based on a large dataset in which it has been trained. For example, it can generate a detailed answer about how to use a product based on a prompt.

[1045] Input: Analyzed keywords, prompt text (e.g., "Please tell me how to use this product.")

[1046] Output: Generated response (Example: "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[1047] Step 4:

[1048] The generated responses are sent back to the server, which then transmits them to the user's smart wearable device. During this process, the server converts the response data into an appropriate format and prepares it for transmission.

[1049] Input: Generated answer

[1050] Output: Response data sent to smart wearable devices

[1051] Step 5:

[1052] The user's smart wearable device acquires the received response data and displays the response on its screen. The user can view the generated response in real time through the smart wearable device.

[1053] Input: Response data sent from the server

[1054] Output: A response displayed on the smart wearable device's screen (e.g., "First, power on this product and follow the instructions to set it up. Then, download the dedicated app to use it.")

[1055] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1056] The present invention is a system that recognizes not only the questions entered by the user but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system includes a process in which the user enters a question using a terminal, sends the question to a server, the server receives the question, analyzes the user's emotions using an emotion engine, then analyzes it using natural language processing, requests a generative AI to generate an answer based on the analysis results and emotion information, the generative AI generates an answer and sends it back to the server, the server sends the generated answer to the user's terminal, and the terminal displays the sent answer to the user.

[1057] First, the user enters a question using a device. At this stage, text information and voice data entered by the user are also collected. For example, the user might enter the question, "How do I obtain a resident registration certificate?"

[1058] Next, the voice and text data, along with the questions entered by the user, are sent to the server. The server receives this data and first analyzes the user's emotional state using an emotion engine. This emotion engine analyzes emotions from the context of the text entered by the user and from the voice data. For example, it may detect emotional information such as the user being anxious, angry, or troubled.

[1059] The server then passes the sentiment information and the question to a natural language processing module, which analyzes the question. Important keywords (e.g., resident registration, acquisition method) are extracted at this stage.

[1060] The server requests a generative AI to generate a response based on the analysis results and sentiment information. The generative AI then generates an appropriate response from a large dataset based on the analyzed information and sentiment information. For example, if the user is anxious, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office."

[1061] The generated response is sent back to the server, which then sends it to the user's device. Finally, the device displays the received response to the user. The generated response is displayed on the user's screen in the appropriate format. For example, a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office. There may be a fee involved. We can also provide you with information on the location of the office," might be displayed and become viewable by the user.

[1062] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[1063] The following describes the processing flow.

[1064] Step 1:

[1065] The user enters a question using a terminal. The user types the question "How do I obtain a resident registration certificate?" into the input form on the terminal.

[1066] Step 2:

[1067] The user sends sentiment data along with the question to the server. When the user clicks the "Submit" button, the question content and sentiment data (e.g., facial recognition and voice tone) are sent to the server as an HTTP POST request.

[1068] Step 3:

[1069] The server receives the question and sentiment data. The server receives the request at the API endpoint and extracts the question content and sentiment data from the request body.

[1070] Step 4:

[1071] The server uses an emotion engine to analyze the user's emotional state. The emotion engine analyzes the text and voice data entered by the user to identify the user's emotions (e.g., anxious, angry, troubled).

[1072] Step 5:

[1073] The server passes the question to a natural language processing module for analysis. The natural language processing module analyzes the text "Please tell me how to obtain a resident registration certificate" and extracts important keywords (e.g., resident registration certificate, how to obtain).

[1074] Step 6:

[1075] The server requests the generative AI to generate an answer based on the analysis results and sentiment information. The server uses an API to send a request to the generative AI that includes the analyzed keyword information and the user's sentiment information.

[1076] Step 7:

[1077] The generative AI generates a response based on the received request. Based on a large dataset collected, the generative AI creates an appropriate response (e.g., "Don't worry, to obtain a resident registration certificate...") based on analyzed keywords and sentiment information.

[1078] Step 8:

[1079] The generative AI sends the generated answer back to the server. The response containing the generated answer is then sent back to the server.

[1080] Step 9:

[1081] The server sends the generated response to the terminal. The server returns the generated response to the terminal as an HTTP response, providing an appropriate response that matches the user's emotions.

[1082] Step 10:

[1083] The terminal displays the received response to the user. The terminal analyzes the response received from the server and displays it on the screen. The user's screen displays a response such as "Don't worry, to obtain your resident registration certificate..." in an appropriate format.

[1084] (Example 2)

[1085] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1086] Traditional question-answering systems focused on analyzing user input and generating appropriate answers, but they were unable to recognize the user's emotional state and personalize answers accordingly. As a result, they provided uniform answers that ignored user emotions, leading to a poor user experience. Furthermore, they lacked sentiment analysis using voice data, making advanced analysis that reflected the user's input modality difficult. A new system is needed to address these issues.

[1087] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1088] In this invention, the server includes means for analyzing received questions and voice data with an emotion analysis engine, means for passing emotion information and question content to a natural language processing module for analysis, and means for requesting a generative AI to generate an answer based on the analysis results and emotion information. This makes it possible to provide personalized answers that take into account the user's emotional state.

[1089] A "terminal" is a device used by a user to input questions or voice data, and includes smartphones, PCs, tablets, and other similar devices.

[1090] A "server" is a central processing unit that analyzes received data and utilizes various engines and modules to generate appropriate responses.

[1091] A "sentiment analysis engine" is a software program that analyzes text and voice data sent by the user to detect the user's emotional state.

[1092] A "natural language processing module" is a software program that analyzes the content of a question, extracts important keywords, and understands the context.

[1093] "Generative AI" refers to artificial intelligence models that generate appropriate responses based on analyzed information and emotional information.

[1094] "Question and audio data" refers to the collective text and audio information entered by the user using their device.

[1095] "Analysis results" refers to the collective term for the question content analyzed by the natural language processing module and the sentiment information analyzed by the sentiment analysis engine.

[1096] "Answer" refers to response information generated by a generative AI based on its analysis results, and is presented to the user.

[1097] Modes for carrying out the invention

[1098] This invention is a system that recognizes not only the questions entered by the user, but also the user's emotions through those questions, and provides more appropriate and personalized answers. This system utilizes an emotion analysis engine, a natural language processing module, and a generative AI.

[1099] The user first enters their question using a device, such as a smartphone, PC, or tablet. The text information and voice data entered by the user are also collected by the device. For example, a user might enter, "Please tell me how to obtain a resident registration certificate."

[1100] The entered questions and voice data are sent from the terminal to the server. The server receives this data and then uses an emotion analysis engine to analyze the user's emotional state. This emotion analysis engine detects the user's emotions from the context of the text, the tone of the voice data, the speed, etc. For example, it extracts emotional information such as whether the user is anxious, angry, or troubled.

[1101] Next, the server passes the sentiment information and question content to a natural language processing module for analysis. The natural language processing module extracts important keywords (e.g., resident registration, acquisition method) from the input question and analyzes the meaning of the sentences. This analysis result, along with the sentiment information, is then passed back to the server.

[1102] The server requests a generative AI to generate an answer based on the analysis results and sentiment information. This generative AI uses a generative AI model such as OpenAI GPT-4. Based on the analyzed information and sentiment information, the generative AI generates an appropriate answer from a large collected dataset. For example, if the user is anxious, it might generate an answer such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[1103] The generated response is sent back to the server, which then sends it to the user's device. The device then displays the received response to the user. For example, the user's device screen might display a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[1104] This system allows users not only to obtain answers to their questions, but also to receive appropriate and personalized responses tailored to their emotional state. This enables users to understand and execute procedures more confidently and quickly.

[1105] Specific example

[1106] For example, if a user enters the prompt "Please tell me how to obtain a resident registration certificate. I'm in a hurry," the sentiment analysis engine will analyze the emotion of "hurry," and the generative AI will generate a response that takes this into consideration. As a result, the user will be provided with the response, "Don't worry, to obtain a resident registration certificate, you need to submit an application form at your nearest city or town hall and present identification documents."

[1107] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1108] Step 1:

[1109] The user enters questions and voice data using a device.

[1110] Users input questions or voice data using devices such as smartphones, PCs, and tablets. For example, they might input, "Please tell me how to obtain a resident registration certificate." During this process, the entered text and voice data are converted into corresponding input formats.

[1111] Input: User's question text and audio data

[1112] Output: Text and audio data stored on the device

[1113] Step 2:

[1114] The terminal sends user input data to the server.

[1115] The terminal sends the collected text and audio data to the server. This transmission is done via an HTTP request or other appropriate communication protocol.

[1116] Input: Text and audio data stored on the device

[1117] Output: Text and audio data sent to the server

[1118] Step 3:

[1119] The server analyzes the received data using an emotion analysis engine.

[1120] The server receives text and audio data sent by the user and passes it to the sentiment analysis engine. The sentiment analysis engine analyzes the context of the text and the tone and speed of the audio data to detect the user's emotional state. For example, if the tone of the audio data is high and fast, it will detect an emotion of anxiety.

[1121] Input: Text data and audio data

[1122] Output: Analyzed emotion information (e.g., "anxious")

[1123] Step 4:

[1124] The server passes sentiment information and questions to a natural language processing module for analysis.

[1125] The sentiment analysis results and the user's question are sent to a natural language processing (NLP) module. The NLP module analyzes the context of the question and extracts important keywords (e.g., "resident registration," "how to obtain"). This allows the meaning of the question to be represented in a structured data format.

[1126] Input: Analyzed sentiment information and text data

[1127] Output: Analysis results with key keywords extracted.

[1128] Step 5:

[1129] The server requests a generative AI to generate an answer based on the analysis results and emotional information.

[1130] The server requests the generative AI to generate an answer based on the analysis results and sentiment information obtained from the NLP module. Specifically, it sends the generative AI a prompt message saying, "The user is anxious. Please answer the question, 'How do I obtain a resident registration certificate?'"

[1131] Input: Analysis results and sentiment information

[1132] Output: Prompt message for generative AI

[1133] Step 6:

[1134] Generative AI generates appropriate answers.

[1135] The generative AI generates a response based on the received prompt text. Here, the generated response is made with consideration for emotions. For example, it might generate a response such as, "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest city or town hall."

[1136] Input: Prompt message for generative AI

[1137] Output: Appropriate answer

[1138] Step 7:

[1139] The generative AI sends the answer back to the server.

[1140] The generated response is sent back to the server. The server receives this response and prepares to send it to the terminal.

[1141] Input: Appropriate answer

[1142] Output: Response sent back to the server

[1143] Step 8:

[1144] The server sends the response to the terminal, and the terminal displays it to the user.

[1145] The server sends the received response to the user's device. The device receives this response and displays it to the user in an appropriate format. For example, the device screen might display: "Don't worry, to obtain a resident registration certificate, you need to submit an application form and present identification documents at your nearest municipal office."

[1146] Input: Response sent back to the server

[1147] Output: The answer displayed on the user's device.

[1148] (Application Example 2)

[1149] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1150] In recent years, improving customer satisfaction in physical stores has become increasingly important. However, traditional customer service systems have struggled to recognize customer emotions and provide real-time responses and product suggestions accordingly. Furthermore, they lacked the means to provide appropriate and personalized answers to customer questions. Because of these challenges, there is a need for systems that can achieve more sophisticated customer service.

[1151] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1152] In this invention, the server includes means for the user to input a question using an information processing device, means for transmitting the question to a communication device, means for the communication device to receive the question and analyze it using natural language processing, means for analyzing the user's emotions using an emotion analysis engine, means for requesting a generative AI to generate an answer based on the analysis results and emotion information, means for the generative AI to generate an answer considering the emotion information and send it back to the communication device, and means for transmitting the generated answer to the information processing device and displaying it to the user. This enables the recognition of customer emotions in real time and allows for appropriate responses, product suggestions, and personalized answers to questions accordingly.

[1153] An "information processing device" is an electronic device used to input questions from users and display the submitted answers.

[1154] A "communication device" is a device that receives questions from users and works in conjunction with a server to perform natural language processing and sentiment analysis.

[1155] "Natural language processing" is a technology that analyzes language data entered by a user to understand the meaning of the question.

[1156] A "sentiment analysis engine" is a system that implements technologies and algorithms for analyzing emotional information from text and voice data entered by the user.

[1157] "Generative AI" is artificial intelligence that generates answers based on trained data, taking into account the user's question content and emotional information.

[1158] A "question" refers to the content that a user inputs using an information processing device and sends to a server via a communication device.

[1159] "Analysis results" refer to the output of question content and sentiment information analyzed by the natural language processing and sentiment analysis engines.

[1160] "User emotion" refers to the emotional state detected by the emotion analysis engine from the user's input data.

[1161] An "answer" refers to the response or information generated by a generative AI based on the question and the user's emotions.

[1162] "Display" refers to the act of an information processing device visually providing an answer to the user.

[1163] "System" refers to the overall configuration and function of a series of means combined to realize the present invention.

[1164] This invention is an emotion-recognition customer service assistant system for enhancing customer service in physical stores. The specific implementation of this system is described below.

[1165] System Configuration

[1166] This system includes smart glasses and other information processing devices, communication devices, servers, and generative AI. The user (referring to a salesperson in this case) wears smart glasses and interacts with customers.

[1167] Hardware and software to be used

[1168] Information processing devices: Smart glasses and personal computers

[1169] Communication device: A network module for communicating with the server.

[1170] Server: The central device that analyzes questions and sentiment information and requests it to the generative AI.

[1171] Generative AI: Natural language generation models such as OpenAI GPT-3

[1172] Emotion analysis engine: Face recognition and emotion analysis software such as DeepFace

[1173] Natural Language Processing: NLP module for analyzing question content

[1174] Data processing and data calculation

[1175] The server performs the following data processing.

[1176] 1. Emotion Analysis: The customer's facial image, captured by the smart glasses' camera, is sent to a server, where an emotion analysis engine (e.g., DeepFace) is used to analyze the customer's emotions in real time. This allows for the detection of emotions such as confusion, anger, or satisfaction.

[1177] 2. Natural Language Processing: The server receives text data of questions asked by customers, analyzes the content of the questions using a natural language processing module (e.g., spaCy or NLTK), and extracts keywords.

[1178] 3. Answer Generation: Based on the results of sentiment analysis and natural language processing, a generative AI (e.g., OpenAI GPT-3) generates an appropriate answer. The generative AI takes into account the user's question and sentiment information to provide a personalized answer.

[1179] 4. Display of results: The generated answers are sent from the server to the information processing device and displayed on the smart glasses' screen.

[1180] Specific example

[1181] For example, if a customer has a confused expression and enters a question through smart glasses saying, "Please tell me the features of this TV," the server first analyzes that confused emotion. Next, it analyzes the content of the question and inputs the following prompt into the generative AI.

[1182] Example of a prompt:

[1183] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[1184] Based on this, the generative AI generates responses such as, "This TV has 4K resolution and excellent color reproduction. It also features the latest HDR technology, making it ideal for watching movies and sports. You'll immediately notice the difference when you actually use it." This response is displayed on the smart glasses, and the salesperson conveys it to the customer.

[1185] This system allows for real-time recognition of customer emotions and the provision of appropriate responses and product suggestions, which is expected to improve customer satisfaction.

[1186] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1187] Step 1:

[1188] The user puts on smart glasses and begins face-to-face interaction with the customer. The smart glasses' camera captures the customer's face and sends it to the information processing device. The input is the customer's facial image, and the output is real-time image data.

[1189] Step 2:

[1190] The information processing device transmits the captured customer's facial image to the server via a communication device. The server analyzes the customer's emotions using an emotion analysis engine (e.g., DeepFace). The input is facial image data, and the output is customer emotion information (e.g., confused, satisfied, angry). Specifically, the DeepFace emotion analysis algorithm is applied.

[1191] Step 3:

[1192] The user receives customer questions through smart glasses and inputs the text information. The information processing device transmits this text data to a server via a communication device. The input is the question text, and the output is the transmission of the question data to the server.

[1193] Step 4:

[1194] The server analyzes the received question text using a natural language processing module (e.g., spaCy or NLTK) and extracts keywords from the question. The input is the question text, and keyword data is generated as output. Specifically, the process involves applying a natural language processing algorithm and extracting keywords.

[1195] Step 5:

[1196] The server requests a generative AI (e.g., OpenAI GPT-3) to generate an answer based on the analyzed question content and sentiment information. A prompt is generated and sent to the generative AI. The input consists of the question content and sentiment information, and the prompt is created. An appropriate answer is generated as output. Specifically, the process involves generating a prompt and sending a query to the generative AI.

[1197] Example of a prompt:

[1198] The customer is feeling confused. Their question is: "What are the features of this television?" Please suggest an appropriate response.

[1199] Step 6:

[1200] The generative AI sends the generated response back to the server. The server then transmits the received response to an information processing device via a communication device. The input is the generated response text, and the output is the transmission of the response data.

[1201] Step 7:

[1202] The information processing device receives the response and displays it on the smart glasses' display. This allows the user to provide appropriate answers and product suggestions to customers in real time. The input is the response text sent from the server, and the output is the display of the response to the user. Specifically, this operation includes displaying text on the display.

[1203] The above processing steps will enhance customer service in physical stores, enabling real-time emotion recognition and appropriate responses.

[1204] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1205] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1206] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1207] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1208] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1209] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1210] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1211] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1212] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1213] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1214] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1215] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1216] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1217] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1218] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1219] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1220] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1221] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1222] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1223] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1224] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1225] The following is further disclosed regarding the embodiments described above.

[1226] (Claim 1)

[1227] [Means for the user to input questions using a terminal,

[1228] [Means of sending questions to the server,

[1229] [Methods for receiving a question on a server and analyzing it using natural language processing,

[1230] [A method by which the server requests a generative AI to generate an answer based on the analysis results,

[1231] [A method in which a generative AI generates an answer and sends it back to the server,

[1232] [Means for the server to send the generated response to the terminal,

[1233] A system that includes means for displaying the submitted response to the user.

[1234] (Claim 2)

[1235] [The system according to claim 1, which includes means for natural language processing to analyze keywords in the content of a question.

[1236] (Claim 3)

[1237] The system according to claim 1, comprising means for a generative AI to generate an answer using trained data.

[1238] "Example 1"

[1239] (Claim 1)

[1240] [Means by which the user inputs questions using an information terminal,

[1241] [Means by which an information terminal sends a question to a data server,

[1242] [A means by which a data server receives a query and analyzes it using natural language processing,

[1243] [A method by which the data server requests the AI ​​model to generate an answer based on the analysis results,

[1244] [A means by which a generative AI model generates an answer and sends it back to the data server,

[1245] [Means for the data server to send the generated response to the information terminal,

[1246] A system that includes means for displaying the responses sent by an information terminal to the user.

[1247] (Claim 2)

[1248] [The system according to claim 1, wherein natural language processing analyzes key information of the question content.

[1249] (Claim 3)

[1250] [The system according to claim 1, in which a generative AI model generates an answer using the data on which it was trained.

[1251] "Application Example 1"

[1252] (Claim 1)

[1253] [Means by which the user inputs a question using a user interface,

[1254] [Means of sending questions to the server,

[1255] [Methods for receiving a question on a server and analyzing it using natural language processing,

[1256] [A method by which the server requests a generative AI to generate an answer based on the analysis results,

[1257] [A method in which a generative AI generates an answer and sends it back to the server,

[1258] [Means for the server to send the generated response to the user device,

[1259] [Means for displaying the user's response to the user,

[1260] [Means by which the user interface is configured using a smart wearable device,

[1261] A system that includes this.

[1262] (Claim 2)

[1263] [Natural language processing is a means of analyzing keywords in the question content,

[1264] [The system according to claim 1, in which a generative AI generates an answer based on analyzed keywords.

[1265] (Claim 3)

[1266] [The system according to claim 1, in which a generative AI generates an answer using a dataset.

[1267] "Example 2 of combining an emotion engine"

[1268] (Claim 1)

[1269] [Means for the user to input questions and voice data using a terminal,

[1270] [Means for sending questions and audio data to the server,

[1271] [Means for a server to receive questions and voice data and analyze the emotional state using an emotion analysis engine,

[1272] [Methods by which the server passes sentiment information and question content to a natural language processing module for analysis,

[1273] [A means by which the server requests a generative AI to generate an answer based on the analysis results and emotional information,

[1274] [A method in which a generative AI generates an answer and sends it back to the server,

[1275] A system that includes means by which a server sends a generated response to a terminal, and the terminal displays the response to the user.

[1276] (Claim 2)

[1277] [The system according to claim 1, which includes means for an emotion analysis engine to analyze emotions from the context of text and audio data.

[1278] (Claim 3)

[1279] [The system according to claim 1, which includes means for generating an appropriate response using analyzed information and emotional information by a generative AI.

[1280] "Application example 2 when combining with an emotional engine"

[1281] (Claim 1)

[1282] [Means by which the user inputs a question using an information processing device,

[1283] [Means for sending a question to a communication device,

[1284] [A means by which a communication device receives a question and analyzes it using natural language processing,

[1285] [A means by which a communication device analyzes the user's emotions using an emotion analysis engine,

[1286] [A means by which a communication device requests a generative AI to generate an answer based on analysis results and emotional information,

[1287] [A means by which a generative AI generates an answer considering emotional information and sends it back to a communication device,

[1288] [Means for transmitting the generated response from the communication device to the information processing device,

[1289] A system that includes means for displaying a transmitted response to the user using an information processing device.

[1290] (Claim 2)

[1291] [The system according to claim 1, wherein natural language processing analyzes keywords in the question content.

[1292] (Claim 3)

[1293] [The system according to claim 1, in which a generative AI generates an answer while taking emotional information into consideration. [Explanation of Symbols]

[1294] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for the user to input questions using a device, A means of sending a question to the server, A server receives a question and uses natural language processing to analyze it, A method by which the server requests a generative AI to generate an answer based on the analysis results, A method by which a generative AI generates an answer and sends it back to the server, A means by which the server sends the generated response to the terminal, A system that includes means for displaying the responses sent by the terminal to the user.

2. The system according to claim 1, which includes means for natural language processing to analyze keywords in the content of a question.

3. The system according to claim 1, comprising means for generating an answer using data on which a generative AI has been trained.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A