system
The system addresses the inconvenience of face-to-face consultations by using generative AI and natural language processing to provide timely, personalized, and emotionally sensitive online responses.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Face-to-face consultations are inconvenient due to time and location restrictions, and the need for avoiding physical contact has increased with the spread of infectious diseases, necessitating an online consultation system that provides timely and personalized information and support.
A system that receives user input, analyzes it to generate appropriate responses, and sends them securely, utilizing generative AI and natural language processing to understand user intent and emotions, and references a database for tailored answers.
Enables users to obtain necessary information quickly and personally without location constraints, providing emotionally sensitive and accurate responses in real-time.
Smart Images

Figure 2026071592000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Face-to-face consultations at physical windows have the problem of low convenience because they are restricted by time and location. In addition, with the spread of the novel coronavirus infection, the need to avoid face-to-face contact is increasing. Therefore, there is a need for a system that can provide a consultation service equivalent to a window in an online environment and enable users to quickly obtain necessary information and support without physically moving.
Means for Solving the Problems
[0005] This invention solves the problem by providing a system that receives input data from a user, analyzes that data to generate an answer suitable for the consultation content, and sends it to the user. Specifically, the system includes a receiving means for receiving user input, an answer generation means for analyzing the input data and generating an answer, a transmitting means for sending the generated answer, and a data storage means for storing the input data and the generated answer, thereby providing a service equivalent to face-to-face consultation online. As a result, users can obtain the information they need in a timely manner without being restricted by time or location. The system also includes a function to convert voice data into text and generate an answer by referring to database information, making it a system that can respond to diverse user needs.
[0006] "Input data" refers to data that includes information and questions that the user provides to the system.
[0007] "Receiving means" refers to a function that acquires input data sent by the user and converts it into a format that can be processed within the system.
[0008] The "response generation means" is a function that analyzes the content of the consultation based on the received input data and generates an appropriate response.
[0009] "Transmission means" refers to a function that sends the generated response to the user's terminal and presents it in a format usable by the user.
[0010] "Data storage means" refers to a function that stores input data and generated responses and saves them in a reusable format as needed.
[0011] "Voice data" refers to data that includes information and questions that a user provides to the system via voice.
[0012] "Text data" refers to data constructed from characters after audio data has been converted into a string of characters.
[0013] "Database information" refers to a collection of information that aggregates past consultation history and related knowledge, and is referenced when generating responses. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to a system that utilizes generative AI to provide online consultation services. This system enables users to receive services equivalent to in-person consultations without having to visit a physical location.
[0036] Specifically, users input their inquiries from an internet-connected device. User input is in text format, and voice input is also possible if necessary. This input data is temporarily stored on the device and sent to the server via a secure communication protocol. The server has a generation AI installed that analyzes the received input data. The AI uses natural language processing technology to understand the user's intent and retrieves relevant answers from a database or generates new ones.
[0037] The server adjusts its response based on the context and sends the final answer back to the user's terminal. The user can view this answer on their terminal screen and can also choose to have it output as audio. For example, if the user asks, "What should I do next?", the server's AI will explain the necessary steps based on relevant information in an easy-to-understand manner.
[0038] Thus, because this system can provide consultations in real time, users can smoothly obtain the necessary information without being restricted by time or location. Furthermore, since the system utilizes an extensive database that includes historical data, it provides answers optimized for each user's situation. This technology allows users to receive prompt and appropriate support tailored to their individual needs.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The user enters their inquiry into the device as text or voice. If voice input is used, the device uses its built-in speech recognition module to convert the voice into text data.
[0042] Step 2:
[0043] The terminal sends the entered text data to the server using an encryption protocol (e.g., SSL / TLS).
[0044] Step 3:
[0045] The server analyzes the received text data, and the natural language processing engine performs analysis to understand the intent of the user's question, taking context and relevance into consideration.
[0046] Step 4:
[0047] The server uses AI generation to create the optimal answer based on the analysis results. This answer utilizes a database of past consultations and related information.
[0048] Step 5:
[0049] The server formats the response and converts it into a format that is easy for the user to understand.
[0050] Step 6:
[0051] The server generates a response and sends it back to the device, which then receives the response. The data is encrypted, ensuring secure delivery.
[0052] Step 7:
[0053] The device displays the received response on the screen. If necessary, it uses an audio output module to play the response aloud for confirmation.
[0054] Step 8:
[0055] If the user wants to ask further questions, this process can be repeated, allowing for further interaction.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] There is a need for a system that allows users to receive consultation services efficiently and safely online without having to visit a physical office, and that can accurately understand the user's intentions and provide appropriate responses. Furthermore, secure transmission of input information and improvement of response accuracy are key challenges.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes an acquisition means for receiving input information from the user, a response generation means for analyzing the input information and generating a response according to the content of the consultation, and a transmission means for sending the generated response to the user. This makes it possible for users to conduct consultations securely over the internet while obtaining highly accurate responses using natural language processing technology.
[0061] "Acquisition means" refers to a function for accurately receiving and temporarily storing input information provided by the user.
[0062] The "response generation means" is a function that analyzes the received input information and generates an appropriate response based on the content of the consultation.
[0063] A "means of communication" refers to a function that ensures the generated response is reliably sent to the user and allows them to confirm the result.
[0064] "Recording means" refers to a function for securely storing input information and generated responses, and making them available for reference as needed.
[0065] "Secure communication means" refers to a function that uses encryption technology to ensure secure communication so that input information cannot be accessed illegally during transmission.
[0066] "Analysis means" refers to a function that uses natural language processing technology to analyze input information and understand the user's intent.
[0067] A "language conversion means" is a function that converts audio information into text information and makes it usable for response generation.
[0068] This system allows users to receive online consultation services using internet-connected devices. At the heart of the system is a generative AI model installed on a server, which provides automated responses to user inquiries.
[0069] On the terminal, users can input their consultation details via text or voice. Text input uses the keyboard, and voice input uses the microphone. The terminal temporarily stores this input information and sends it to the server via a secure communication protocol (e.g., HTTPS) using encryption technology.
[0070] On the server side, software equipped with natural language processing technology runs as a generative AI model. A general-purpose natural language processing model can be used as a specific example. The server analyzes the received input information and quickly and accurately understands the user's intent. The AI decides whether to retrieve the answer from an existing database or generate new information, and sends the generated response to the user. The server ensures the context and accuracy of the response and adjusts it to provide useful information for the user.
[0071] As a concrete example, consider a scenario where a user enters the question, "How do I renew my passport?" into the terminal. The server's AI analyzes this question, retrieves the necessary procedural information from the database, and provides it to the user. The generative AI model understands the user's intent and generates an appropriate and rapid response.
[0072] Examples of prompts include direct questions such as, "Could you please tell me more about the passport renewal process?"
[0073] This system allows users to ask questions in real time, without being restricted by time or location, and to obtain the necessary information with peace of mind.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] Users input their inquiries as text or voice using an internet-connected device. The device temporarily stores the entered data. The input data is in text format; in the case of voice input, the device converts the voice to text. This ensures a consistent format for the input data.
[0077] Step 2:
[0078] The device sends stored input data to the server using a secure communication protocol (e.g., HTTPS). Before sending the data, the device encrypts the input content to ensure secure transmission. This process reduces the risk of unauthorized access to the data during transit.
[0079] Step 3:
[0080] The server analyzes the received data. Using a generative AI model, the server decodes the content of the input data using natural language processing techniques to understand the user's intent. This analysis identifies keywords in the input information and recognizes what kind of information is needed. The server uses these results to prepare a response.
[0081] Step 4:
[0082] The server's AI generates answers based on the analysis results. This process retrieves relevant information from a database and constructs responses appropriate to the user's questions. The server reviews the generated responses and adjusts the context and content as needed. This ensures that valuable information is provided to the user.
[0083] Step 5:
[0084] The server sends the final response to the user's device. The device displays this response on the screen and also provides an audio response if voice output is selected. Based on the information presented on the device, the user can decide on their next course of action. This allows users to quickly obtain necessary information whether they are in the office or on the go.
[0085] (Application Example 1)
[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] Traditional customer support on e-commerce sites has struggled to provide immediate responses, which has been a factor in lowering user satisfaction. Furthermore, there is a need to provide appropriate and timely answers to consumers' specific questions about products. Therefore, a system is needed that can provide accurate, real-time responses and improve the user experience.
[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0089] In this invention, the server includes receiving means for receiving input information from the user, response generating means for analyzing the input information and generating a response corresponding to the proposed content, and information presentation means for providing real-time support via a smart device. This enables the user to immediately receive accurate answers to questions about the product.
[0090] A "user" is an individual or legal entity that uses the system and provides input information.
[0091] "Input information" refers to voice or text data provided by the user as an inquiry or question.
[0092] "Receiving means" refers to a function or device for receiving input information from a user.
[0093] "Response generation means" refers to a function or algorithm for analyzing received input information and creating a response based on it.
[0094] "Transmission means" refers to a function or device for sending the generated response back to the user.
[0095] "Data storage means" refers to a data storage system for storing input information and generated responses.
[0096] "Voice conversion means" refers to a function or algorithm for converting voice input into text data.
[0097] A "smart device" is an electronic device connected to the internet, such as a smartphone or smart glasses, used by a user.
[0098] "Information presentation means" refers to a function that displays the generated response through a user interface and also outputs it as audio.
[0099] This invention provides a system that offers accurate, real-time responses when a user makes an inquiry about a product through a smart device. The system is configured as follows:
[0100] First, the user provides input information using a smart device (such as a smartphone running iOS or Android®, or smart glasses). This input information is provided as text or voice and is received by an application on the smart device. In the case of voice input, it is converted into text data by a speech-to-text conversion tool (such as Google® Speech-to-Text API).
[0101] Next, the received text data is sent to the server using a secure communication protocol (HTTPS). The server is equipped with a generative AI model (for example, OpenAI®'s GPT series) and functions as a response generation tool. The server analyzes the input information and generates an appropriate response by referring to past database information and related information. This response may include product attributes, usage instructions, and purchase procedures.
[0102] The generated response is sent from the server to the smart device and displayed to the user by an information display device. The user can view the displayed text information and, if necessary, also obtain the information through audio output.
[0103] As a concrete example, consider a scenario where a user types "How can I ship this product?" on their smartphone. In this case, the server's generated AI model uses a prompt such as "User's question: 'How can I ship this product?' Please provide a summary and detailed instructions." to provide the user with detailed shipping procedure information as a response. In this way, the user can obtain useful information in real time.
[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0105] Step 1:
[0106] The user launches an application on a smart device and provides input information. This input can be either voice or text. In the case of voice input, the terminal uses a voice conversion mechanism to convert the voice data into text data. As a result, text data is obtained as input information.
[0107] Step 2:
[0108] The terminal securely sends the generated text data to the server using the HTTPS protocol. During this process, the received text data is converted into a secure communication packet before being transmitted, thus performing data processing. The server receives this text data via the network.
[0109] Step 3:
[0110] The server analyzes the received text data and activates a generative AI model based on its content. It inputs a specific prompt sentence (e.g., "User question: 'How do I ship this product?' Please provide a summary and detailed instructions.") into the generative AI model and generates a relevant response. At this time, the server performs data calculations by referring to past database information and related information to obtain an appropriate response that is in line with the context.
[0111] Step 4:
[0112] The generated response is sent from the server to the terminal. The terminal receives this response and displays it to the user using an information display device. The user can choose whether to display it in text format or as audio output. In this process, the terminal converts the response data obtained from the server into a user interface format and performs data processing to present it appropriately.
[0113] Step 5:
[0114] The user can review the presented information and ask further questions if necessary. A new session begins when the user provides additional information at this stage. This iterative process ensures continued interaction between the user and the system.
[0115] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0116] This invention relates to a system that combines generative AI and emotion recognition technology to provide online consultation services in a more personalized manner. This system allows users to receive emotionally sensitive responses in addition to online consultations, without having to visit a physical service center.
[0117] Users input their inquiries through their device, and can also use voice input if necessary. The device receives the input text or voice and sends it to the server. When voice is input, the device's speech recognition function converts the voice to text. The server uses its internal natural language processing engine to understand the user's inquiries and intentions in order to analyze the received text.
[0118] During the analysis process, the server's emotion engine recognizes the user's emotions, and the generating AI adjusts the response based on that emotion information. Specifically, it analyzes the tone and context of the user's text and identifies the emotional elements within it. The emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's word choice and expressions, and takes this into consideration in the response generation process.
[0119] The server uses the results of emotion recognition, referencing the database, to construct a response best suited to the user's current feelings. If the user is nervous or confused, it provides a response that includes more polite and reassuring language. For example, if the user expresses anxiety such as "the procedure is complicated and I don't understand it," the server gently explains the specific steps, adds necessary links and support information, and assists the user in a way that calms them down.
[0120] This system allows users to receive not just information, but a more deeply understood service tailored to their individual emotional state. As a result, it delivers a more satisfying user experience. This form of invention provides users with a more human-centered interaction and improves the quality of online consultation services.
[0121] The following describes the processing flow.
[0122] Step 1:
[0123] The user enters their inquiry into the device as text or voice. In the case of voice input, the device uses a speech recognition module to convert the voice into text.
[0124] Step 2:
[0125] When a terminal sends text data received from a user to a server, it uses an encryption protocol for security purposes.
[0126] Step 3:
[0127] The server receives text data and analyzes the content of the consultation using a natural language processing engine. This analysis includes understanding the context and intent.
[0128] Step 4:
[0129] The server sends the analyzed data to the emotion engine to recognize the user's emotions. The emotion engine extracts emotional information by analyzing the tone and wording of the text.
[0130] Step 5:
[0131] The server generates an appropriate response based on emotional information obtained from the emotion engine, referencing database information. In this process, it selects words that take the user's emotions into consideration.
[0132] Step 6:
[0133] The server formats the generated response, converts it into a user-friendly format, and then sends it to the terminal.
[0134] Step 7:
[0135] The device displays the received response on the screen and presents it to the user. If necessary, it plays the response aloud using speech synthesis.
[0136] Step 8:
[0137] If the user asks further questions, the entire process restarts, and the interaction with the user continues.
[0138] (Example 2)
[0139] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0140] In online counseling services, a challenge is that individual emotional states are often not adequately considered, resulting in a lack of personalized responses. Specifically, it is difficult to understand users' emotions and provide appropriate answers. This tends to lead to lower satisfaction rates in online counseling.
[0141] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0142] In this invention, the server includes receiving means for receiving input data, conversion means for converting audio data into text data, and recognition means for recognizing the user's emotions. This makes it possible to generate personalized responses that are tailored to the user's emotions.
[0143] A "receiving means" is a function that acquires input data from the user and is responsible for the initial steps necessary to process that data within the system.
[0144] A "conversion means" is a function that handles the process of converting audio data into text data, making the audio information easier to recognize as text.
[0145] "Recognition means" refers to a function that analyzes emotional elements contained in the user's input data and identifies those emotions.
[0146] A "response generation means" is a function that forms an appropriate response to the user based on input data and recognized emotional information.
[0147] "Transmission method" refers to a function that sends the generated response to the user in an appropriate format.
[0148] "Data storage means" refers to a function that stores input data and generated responses, making them available for later reference and analysis.
[0149] This invention relates to a system that provides personalized responses to users during online individual consultations, taking into account their emotions at the time. The system mainly consists of a user terminal and a server.
[0150] The user uses the device to input the content they wish to discuss. Both text input and voice input are available, and the device is equipped with a keyboard and microphone. If voice input is selected, the device acquires the voice data and converts it into text data using specific speech recognition technology. Common software (e.g., a speech recognition API) is used for speech recognition.
[0151] The converted text data is sent from the user's terminal to the server. The server first analyzes the received text data using a natural language processing engine to understand the user's inquiry and intentions. Then, it uses an emotion recognition engine to extract emotional elements contained in the text. Various emotion analysis technologies (e.g., emotion analysis APIs) are used for emotion recognition.
[0152] When the server recognizes the user's emotional information, it uses a generative AI model to generate a response that is tailored to the user's inquiry and emotions. A general generative model is used for this AI, and an example of a prompt is, "Explain the procedure that the user finds complicated, and add support information to reassure them."
[0153] Finally, the generated response is sent to the device, where the user can view it on the screen. For example, if a user expresses concern such as "I don't understand the new procedure," the server will gently and clearly explain the procedure and provide relevant information to alleviate the user's anxiety.
[0154] Through this entire process, the system provides users not merely with information, but with a deep understanding and satisfaction tailored to their individual emotions and circumstances. This improves the quality of online consultation services and enables more human-centered interactions.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] The user uses the terminal to input the information they wish to discuss. Input can be in text or voice format. In the case of voice input, the terminal uses its built-in microphone to acquire voice data. Then, speech recognition software is used to convert the voice data into text data. In this process, the input is voice data, and the output is text data.
[0158] Step 2:
[0159] The terminal sends the converted text data to the server. This data transmission is performed using a secure communication method. Specifically, text data is sent. The input is the converted text data, and the output is the text data received by the server.
[0160] Step 3:
[0161] The server analyzes the received text data using a natural language processing engine. This allows it to understand the grammatical structure and meaning of the text, and extract the user's intent and requests. The input is the transmitted text data, and the output is the analyzed intent information and keywords.
[0162] Step 4:
[0163] The server uses an emotion recognition engine to analyze text data and identify various emotional elements. In this process, it recognizes emotions (e.g., joy, anxiety, anger, etc.) from the user's expressions. The input for the analysis is text data, and the output is identified emotion information.
[0164] Step 5:
[0165] The server uses a generative AI model to generate responses that are appropriate to the user's emotions based on the analysis results. In this process, emotional information is taken into account, and the tone and content of the responses are adjusted accordingly. The input consists of analyzed intent and emotional information, and the output is the generated personalized response text.
[0166] Step 6:
[0167] The server sends the generated response to the terminal. A fast and secure protocol is used for transmission to the user. The input is the generated response text, and the output is the text displayed on the user's terminal.
[0168] Step 7:
[0169] Users can view the received responses on their device. Here, information tailored to the user's situation and emotions is displayed, leading to a deeper and more satisfying consultation. The output is the appropriate response displayed on the screen.
[0170] (Application Example 2)
[0171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0172] In online consultations, there is a need to move away from traditional, monotonous, and uniform services and improve the user experience by providing personalized answers that are tailored to the individual emotional state of the user. However, existing systems have difficulty generating answers that take into account the user's emotional state, and this is a particular challenge when it comes to security-related consultations, as they cannot effectively alleviate the user's anxiety and fear.
[0173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0174] In this invention, the server includes an acquisition means for receiving input information from the user, a response processing means for analyzing the input information and generating an answer based on the consultation content and the user's emotional state, and an output means for providing the generated answer to the user. This makes it possible to provide an optimal answer that takes the user's emotional state into consideration.
[0175] A "receiving mechanism" is a mechanism necessary to acquire information entered by the user and process it within the system.
[0176] A "response processing means" is a mechanism that analyzes received input information and processes it to generate the optimal response based on the content of the consultation and the user's emotions.
[0177] An "output mechanism" is a device that presents the generated response to the user and completes the system interaction.
[0178] An "emotion identification means" is a mechanism that detects emotions from user input information and reflects that information in response generation.
[0179] An "information storage means" is a mechanism for storing input information and generated responses, in preparation for future use or reference.
[0180] A "speech-to-text conversion means" is a mechanism for converting speech information into text information and preparing it in a format that can be processed by natural language processing.
[0181] The "information data repository" is a database that stores past consultation details and related information, and is used to reference this data when generating responses.
[0182] To realize this invention, the user's smartphone or tablet functions as a receiving terminal. When the user inputs their inquiry as text or voice, the terminal uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text. The converted text is then sent to the server.
[0183] On the server side, text data is analyzed by a natural language processing engine (e.g., SpaCy). This allows the user's inquiry to be understood, and an emotion recognition API (e.g., IBM Watson® Tone Analyzer) is used as a means of emotion identification to evaluate the user's emotional state.
[0184] Based on this information, the server's response processing mechanism uses a generative AI model (e.g., GPT-3(registered trademark).5) to create a response appropriate to the user's emotional state. The generated response is returned to the terminal as text or audio as an output.
[0185] Furthermore, accumulated consultation history information is stored in an information database and referenced during future consultations. This process allows responses to be tailored to the user's history data and emotional state, enabling a more personalized service.
[0186] For example, if a user consults the system saying, "I've recently seen a suspicious person near my home and I'm worried," the system will identify the user's anxiety and suggest specific measures that align with their feelings, such as, "First, walk in well-lit areas and install security cameras and lights. It would also be good to spread awareness through local community apps."
[0187] An example of a prompt for a generative AI model would be: "The user has expressed anxiety. Please create advice to calm them down and provide specific security measures."
[0188] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0189] Step 1:
[0190] Users input their inquiries as text or voice using a device such as a smartphone or tablet. If voice input is used, the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and sends it to the server as text data. Input is either text or voice information, and output is text information.
[0191] Step 2:
[0192] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy). It analyzes the user's inquiry content and extracts the purpose and target from the text. The input here is text data from the user, and the output is the analyzed inquiry content data.
[0193] Step 3:
[0194] The server's emotion recognition mechanism analyzes emotions from user text using an emotion recognition API (e.g., IBM Watson Tone Analyzer). It identifies emotions such as joy, anxiety, and fear from the user's words and uses the analysis results as indicators for generating responses using AI. The input is parsed text data, and the output is the emotion analysis result.
[0195] Step 4:
[0196] The server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to generate a response tailored to the user's inquiry and emotional state. The generated response aims to alleviate the user's emotions and provide appropriate solutions. The input consists of the inquiry data and the emotion analysis results, while the output is the generated response.
[0197] Step 5:
[0198] The server sends the generated response to the terminal and presents it to the user. The user receives the response via text or voice through the terminal. The input is the generated response, and the output is the information presented to the user.
[0199] Step 6:
[0200] The server stores input information and response content using information storage means. The accumulated data is used as reference for future consultations, enabling the generation of more comprehensive responses. The input is all the user's consultation and response data, and the output is an updated information data repository.
[0201] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0202] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0203] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0204] [Second Embodiment]
[0205] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0206] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0207] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0208] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0209] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0210] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0211] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0212] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0213] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0214] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0215] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0216] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0217] This invention relates to a system that utilizes generative AI to provide online consultation services. This system enables users to receive services equivalent to in-person consultations without having to visit a physical location.
[0218] Specifically, users input their inquiries from an internet-connected device. User input is in text format, and voice input is also possible if necessary. This input data is temporarily stored on the device and sent to the server via a secure communication protocol. The server has a generation AI installed that analyzes the received input data. The AI uses natural language processing technology to understand the user's intent and retrieves relevant answers from a database or generates new ones.
[0219] The server adjusts its response based on the context and sends the final answer back to the user's terminal. The user can view this answer on their terminal screen and can also choose to have it output as audio. For example, if the user asks, "What should I do next?", the server's AI will explain the necessary steps based on relevant information in an easy-to-understand manner.
[0220] Thus, because this system can provide consultations in real time, users can smoothly obtain the necessary information without being restricted by time or location. Furthermore, since the system utilizes an extensive database that includes historical data, it provides answers optimized for each user's situation. This technology allows users to receive prompt and appropriate support tailored to their individual needs.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The user enters their inquiry into the device as text or voice. If voice input is used, the device uses its built-in speech recognition module to convert the voice into text data.
[0224] Step 2:
[0225] The terminal sends the entered text data to the server using an encryption protocol (e.g., SSL / TLS).
[0226] Step 3:
[0227] The server analyzes the received text data, and the natural language processing engine performs analysis to understand the intent of the user's question, taking context and relevance into consideration.
[0228] Step 4:
[0229] The server uses AI generation to create the optimal answer based on the analysis results. This answer utilizes a database of past consultations and related information.
[0230] Step 5:
[0231] The server formats the response and converts it into a format that is easy for the user to understand.
[0232] Step 6:
[0233] The server generates a response and sends it back to the device, which then receives the response. The data is encrypted, ensuring secure delivery.
[0234] Step 7:
[0235] The device displays the received response on the screen. If necessary, it uses an audio output module to play the response aloud for confirmation.
[0236] Step 8:
[0237] If the user wants to ask further questions, this process can be repeated, allowing for further interaction.
[0238] (Example 1)
[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0240] There is a need for a system that allows users to receive consultation services efficiently and safely online without having to visit a physical office, and that can accurately understand the user's intentions and provide appropriate responses. Furthermore, secure transmission of input information and improvement of response accuracy are key challenges.
[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0242] In this invention, the server includes an acquisition means for receiving input information from the user, a response generation means for analyzing the input information and generating a response according to the content of the consultation, and a transmission means for sending the generated response to the user. This makes it possible for users to conduct consultations securely over the internet while obtaining highly accurate responses using natural language processing technology.
[0243] "Acquisition means" refers to a function for accurately receiving and temporarily storing input information provided by the user.
[0244] The "response generation means" is a function that analyzes the received input information and generates an appropriate response based on the content of the consultation.
[0245] A "means of communication" refers to a function that ensures the generated response is reliably sent to the user and allows them to confirm the result.
[0246] "Recording means" refers to a function for securely storing input information and generated responses, and making them available for reference as needed.
[0247] "Secure communication means" refers to a function that uses encryption technology to ensure secure communication so that input information cannot be accessed illegally during transmission.
[0248] "Analysis means" refers to a function that uses natural language processing technology to analyze input information and understand the user's intent.
[0249] A "language conversion means" is a function that converts audio information into text information and makes it available for use in generating responses.
[0250] This system allows users to receive online consultation services using internet-connected devices. At the heart of the system is a generative AI model installed on a server, which provides automated responses to user inquiries.
[0251] On the terminal, users can input their consultation details via text or voice. Text input uses the keyboard, and voice input uses the microphone. The terminal temporarily stores this input information and sends it to the server via a secure communication protocol (e.g., HTTPS) using encryption technology.
[0252] On the server side, software equipped with natural language processing technology runs as a generative AI model. A general-purpose natural language processing model can be used as a specific example. The server analyzes the received input information and quickly and accurately understands the user's intent. The AI decides whether to retrieve the answer from an existing database or generate new information, and sends the generated response to the user. The server ensures the context and accuracy of the response and adjusts it to provide useful information for the user.
[0253] As a concrete example, consider a scenario where a user enters the question, "How do I renew my passport?" into the terminal. The server's AI analyzes this question, retrieves the necessary procedural information from the database, and provides it to the user. The generative AI model understands the user's intent and generates an appropriate and rapid response.
[0254] Examples of prompts include direct questions such as, "Could you please tell me more about the passport renewal process?"
[0255] This system allows users to ask questions in real time, without being restricted by time or location, and to obtain the necessary information with peace of mind.
[0256] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0257] Step 1:
[0258] Users input their inquiries as text or voice using an internet-connected device. The device temporarily stores the entered data. The input data is in text format; in the case of voice input, the device converts the voice to text. This ensures a consistent format for the input data.
[0259] Step 2:
[0260] The terminal sends stored input data to the server using a secure communication protocol (e.g., HTTPS). Before sending the data, the terminal encrypts the input content and transmits it securely. This process reduces the risk of unauthorized access to the data during transit.
[0261] Step 3:
[0262] The server analyzes the received data. Using a generative AI model, the server decodes the content of the input data using natural language processing techniques to understand the user's intent. This analysis identifies keywords in the input information and recognizes what kind of information is needed. The server uses these results to prepare a response.
[0263] Step 4:
[0264] The server's AI generates answers based on the analysis results. This process retrieves relevant information from a database and constructs responses appropriate to the user's questions. The server reviews the generated responses and adjusts the context and content as needed. This ensures that valuable information is provided to the user.
[0265] Step 5:
[0266] The server sends the final response to the user's device. The device displays this response on the screen and also provides an audio response if voice output is selected. Based on the information presented on the device, the user can decide on their next course of action. This allows users to quickly obtain necessary information whether they are in the office or on the go.
[0267] (Application Example 1)
[0268] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0269] Traditional customer support on e-commerce sites has struggled to provide immediate responses, which has been a factor in lowering user satisfaction. Furthermore, there is a need to provide appropriate and timely answers to consumers' specific questions about products. Therefore, a system is needed that can provide accurate, real-time responses and improve the user experience.
[0270] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0271] In this invention, the server includes receiving means for receiving input information from the user, response generating means for analyzing the input information and generating a response corresponding to the proposed content, and information presentation means for providing real-time support via a smart device. This enables the user to immediately receive accurate answers to questions about the product.
[0272] A "user" is an individual or legal entity that uses the system and provides input information.
[0273] "Input information" refers to voice or text data provided by the user as an inquiry or question.
[0274] "Receiving means" refers to a function or device for receiving input information from a user.
[0275] "Response generation means" refers to a function or algorithm for analyzing received input information and creating a response based on it.
[0276] "Transmission means" refers to a function or device for sending the generated response back to the user.
[0277] "Data storage means" refers to a data storage system for storing input information and generated responses.
[0278] "Voice conversion means" refers to a function or algorithm for converting voice input into text data.
[0279] A "smart device" is an electronic device connected to the internet, such as a smartphone or smart glasses, used by a user.
[0280] "Information presentation means" refers to a function that displays the generated response through a user interface and also outputs it as audio.
[0281] This invention provides a system that offers accurate, real-time responses when a user makes an inquiry about a product through a smart device. The system is configured as follows:
[0282] First, the user provides input information using a smart device (such as a smartphone running iOS or Android, or smart glasses). This input information is provided as text or voice and is received by an application on the smart device. In the case of voice input, it is converted into text data by a speech-to-text conversion tool (such as the Google Speech-to-Text API).
[0283] Next, the received text data is sent to the server using a secure communication protocol (HTTPS). The server is equipped with a generative AI model (e.g., OpenAI's GPT series) and functions as a response generation means. The server analyzes the input information and generates an appropriate response by referring to past database information and related information. This response may include product attributes, usage methods, purchase procedures, etc.
[0284] The generated response is sent from the server to the smart device and displayed to the user by the information presentation means. The user can view the information presented in text and can also obtain the information by voice output if necessary.
[0285] As a specific example, consider the case where a user inputs "Please tell me the delivery method of this product" on a smartphone. In this case, the generative AI model of the server uses a prompt sentence such as "Question from user: 'Please tell me the delivery method of this product'. Please explain the summary and procedure in detail." and provides detailed delivery procedure information for the user as a response. In this way, the user can obtain useful information in real time.
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] The user launches an application on the smart device and provides input information. At this time, the input is made by voice or text. In the case of voice input, the terminal uses voice conversion means to convert the voice data into text data. As a result, text data is obtained as the input information.
[0289] Step 2:
[0290] [[ID=I27]] The terminal securely sends the generated text data to the server using the HTTPS protocol. During this process, the received text data is converted into a secure communication packet before being transmitted, thus performing data processing. The server receives this text data via the network.
[0291] Step 3:
[0292] The server analyzes the received text data and activates a generative AI model based on its content. It inputs a specific prompt sentence (e.g., "User question: 'How do I ship this product?' Please provide a summary and detailed instructions.") into the generative AI model and generates a relevant response. At this time, the server performs data calculations by referring to past database information and related information to obtain an appropriate response that is in line with the context.
[0293] Step 4:
[0294] The generated response is sent from the server to the terminal. The terminal receives this response and displays it to the user using an information display device. The user can choose whether to display it in text format or as audio output. In this process, the terminal converts the response data obtained from the server into a user interface format and performs data processing to present it appropriately.
[0295] Step 5:
[0296] The user can review the presented information and ask further questions if necessary. A new session begins when the user provides additional information at this stage. This iterative process ensures continued interaction between the user and the system.
[0297] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0298] This invention relates to a system that combines generative AI and emotion recognition technology to provide online consultation services in a more personalized manner. This system allows users to receive emotionally sensitive responses in addition to online consultations, without having to visit a physical service center.
[0299] Users input their inquiries through their device, and can also use voice input if necessary. The device receives the input text or voice and sends it to the server. When voice is input, the device's speech recognition function converts the voice to text. The server uses its internal natural language processing engine to understand the user's inquiries and intentions in order to analyze the received text.
[0300] During the analysis process, the server's emotion engine recognizes the user's emotions, and the generating AI adjusts the response based on that emotion information. Specifically, it analyzes the tone and context of the user's text and identifies the emotional elements within it. The emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's word choice and expressions, and takes this into consideration in the response generation process.
[0301] The server uses the results of emotion recognition, referencing the database, to construct a response best suited to the user's current feelings. If the user is nervous or confused, it provides a response that includes more polite and reassuring language. For example, if the user expresses anxiety such as "the procedure is complicated and I don't understand it," the server gently explains the specific steps, adds necessary links and support information, and assists the user in a way that calms them down.
[0302] This system allows users to receive not just information, but a more deeply understood service tailored to their individual emotional state. As a result, it delivers a more satisfying user experience. This form of invention provides users with a more human-centered interaction and improves the quality of online consultation services.
[0303] The process flow will be described below.
[0304] Step 1:
[0305] The user inputs the consultation content into the terminal in text or voice. In the case of voice input, the terminal uses a voice recognition module to convert the voice into text.
[0306] Step 2:
[0307] When the terminal sends the text data received from the user to the server, an encryption protocol is used for security.
[0308] Step 3:
[0309] The server receives the text data and analyzes the consultation content using a natural language processing engine. This analysis includes grasping the context and understanding the intention.
[0310] Step 4:
[0311] The server sends the analyzed data to an emotion engine to recognize the user's emotion. The emotion engine analyzes the tone and word usage of the text to extract emotion information.
[0312] Step 5:
[0313] Based on the emotion information obtained by the server from the emotion engine, the server refers to the database information to generate an appropriate answer. At this time, words that take into account the user's emotion are selected.
[0314] Step 6:
[0315] The server formats the generated answer, converts it into an easy-to-understand form for the user, and then sends it to the terminal.
[0316] Step 7:
[0317] The device displays the received response on the screen and presents it to the user. If necessary, it plays the response aloud using speech synthesis.
[0318] Step 8:
[0319] If the user asks further questions, the entire process restarts, and the interaction with the user continues.
[0320] (Example 2)
[0321] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0322] In online counseling services, a challenge is that individual emotional states are often not adequately considered, resulting in a lack of personalized responses. Specifically, it is difficult to understand users' emotions and provide appropriate answers. This tends to lead to lower satisfaction rates in online counseling.
[0323] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0324] In this invention, the server includes receiving means for receiving input data, conversion means for converting audio data into text data, and recognition means for recognizing the user's emotions. This makes it possible to generate personalized responses that are tailored to the user's emotions.
[0325] A "receiving means" is a function that acquires input data from the user and is responsible for the initial steps necessary to process that data within the system.
[0326] A "conversion means" is a function that handles the process of converting audio data into text data, making the audio information easier to recognize as text.
[0327] "Recognition means" refers to a function that analyzes emotional elements contained in the user's input data and identifies those emotions.
[0328] A "response generation means" is a function that forms an appropriate response to the user based on input data and recognized emotional information.
[0329] "Transmission method" refers to a function that sends the generated response to the user in an appropriate format.
[0330] "Data storage means" refers to a function that stores input data and generated responses, making them available for later reference and analysis.
[0331] This invention relates to a system that provides personalized responses to users during online individual consultations, taking into account their emotions at the time. The system mainly consists of a user terminal and a server.
[0332] The user uses the device to input the content they wish to discuss. Both text input and voice input are available, and the device is equipped with a keyboard and microphone. If voice input is selected, the device acquires the voice data and converts it into text data using specific speech recognition technology. Common software (e.g., a speech recognition API) is used for speech recognition.
[0333] The converted text data is sent from the user's terminal to the server. The server first analyzes the received text data using a natural language processing engine to understand the user's inquiry and intentions. Then, it uses an emotion recognition engine to extract emotional elements contained in the text. Various emotion analysis technologies (e.g., emotion analysis APIs) are used for emotion recognition.
[0334] When the server recognizes the user's emotional information, it uses a generative AI model to generate a response that is tailored to the user's inquiry and emotions. A general generative model is used for this AI, and an example of a prompt is, "Explain the procedure that the user finds complicated, and add support information to reassure them."
[0335] Finally, the generated response is sent to the device, where the user can view it on the screen. For example, if a user expresses concern such as "I don't understand the new procedure," the server will gently and clearly explain the procedure and provide relevant information to alleviate the user's anxiety.
[0336] Through this entire process, the system provides users not merely with information, but with a deep understanding and satisfaction tailored to their individual emotions and circumstances. This improves the quality of online consultation services and enables more human-centered interactions.
[0337] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0338] Step 1:
[0339] The user uses the terminal to input the information they wish to discuss. Input can be in text or voice format. In the case of voice input, the terminal uses its built-in microphone to acquire voice data. Then, speech recognition software is used to convert the voice data into text data. In this process, the input is voice data, and the output is text data.
[0340] Step 2:
[0341] The terminal sends the converted text data to the server. This data transmission is performed using a secure communication method. Specifically, text data is sent. The input is the converted text data, and the output is the text data received by the server.
[0342] Step 3:
[0343] The server analyzes the received text data using a natural language processing engine. This allows it to understand the grammatical structure and meaning of the text, and extract the user's intent and requests. The input is the transmitted text data, and the output is the analyzed intent information and keywords.
[0344] Step 4:
[0345] The server uses an emotion recognition engine to analyze text data and identify various emotional elements. In this process, it recognizes emotions (e.g., joy, anxiety, anger, etc.) from the user's expressions. The input for the analysis is text data, and the output is identified emotion information.
[0346] Step 5:
[0347] The server uses a generative AI model to generate responses that are appropriate to the user's emotions based on the analysis results. In this process, emotional information is taken into account, and the tone and content of the responses are adjusted accordingly. The input consists of analyzed intent and emotional information, and the output is the generated personalized response text.
[0348] Step 6:
[0349] The server sends the generated response to the terminal. A fast and secure protocol is used for transmission to the user. The input is the generated response text, and the output is the text displayed on the user's terminal.
[0350] Step 7:
[0351] Users can view the received responses on their device. Here, information tailored to the user's situation and emotions is displayed, leading to a deeper and more satisfying consultation. The output is the appropriate response displayed on the screen.
[0352] (Application Example 2)
[0353] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0354] In online consultations, there is a need to move away from traditional, monotonous, and uniform services and improve the user experience by providing personalized answers that are tailored to the individual emotional state of the user. However, existing systems have difficulty generating answers that take into account the user's emotional state, and this is a particular challenge when it comes to security-related consultations, as they cannot effectively alleviate the user's anxiety and fear.
[0355] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0356] In this invention, the server includes an acquisition means for receiving input information from the user, a response processing means for analyzing the input information and generating an answer based on the consultation content and the user's emotional state, and an output means for providing the generated answer to the user. This makes it possible to provide an optimal answer that takes the user's emotional state into consideration.
[0357] A "receiving mechanism" is a mechanism necessary to acquire information entered by the user and process it within the system.
[0358] A "response processing means" is a mechanism that analyzes received input information and processes it to generate the optimal response based on the content of the consultation and the user's emotions.
[0359] An "output mechanism" is a device that presents the generated response to the user and completes the system interaction.
[0360] An "emotion identification means" is a mechanism that detects emotions from user input information and reflects that information in response generation.
[0361] An "information storage means" is a mechanism for storing input information and generated responses, in preparation for future use or reference.
[0362] A "speech-to-text conversion means" is a mechanism for converting speech information into text information and preparing it in a format that can be processed by natural language processing.
[0363] The "information data repository" is a database that stores past consultation details and related information, and is used to reference this data when generating responses.
[0364] To realize this invention, the user's smartphone or tablet functions as a receiving terminal. When the user inputs their inquiry as text or voice, the terminal uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text. The converted text is then sent to the server.
[0365] On the server side, text data is analyzed by a natural language processing engine (e.g., SpaCy). This allows the user's inquiry to be understood, and an emotion recognition API (e.g., IBM Watson Tone Analyzer) is used as a means of emotion identification to evaluate the user's emotional state.
[0366] Based on this information, the server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to create a response appropriate to the user's emotional state. The generated response is returned to the terminal as text or audio.
[0367] Furthermore, accumulated consultation history information is stored in an information database and referenced during future consultations. This process allows responses to be tailored to the user's history data and emotional state, enabling a more personalized service.
[0368] For example, if a user consults the system saying, "I've recently seen a suspicious person near my home and I'm worried," the system will identify the user's anxiety and suggest specific measures that align with their feelings, such as, "First, walk in well-lit areas and install security cameras and lights. It would also be good to spread awareness through local community apps."
[0369] An example of a prompt for a generative AI model would be: "The user has expressed anxiety. Please create advice to calm them down and provide specific security measures."
[0370] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0371] Step 1:
[0372] Users input their inquiries as text or voice using a device such as a smartphone or tablet. If voice input is used, the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and sends it to the server as text data. Input is either text or voice information, and output is text information.
[0373] Step 2:
[0374] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy). It analyzes the user's inquiry content and extracts the purpose and target from the text. The input here is text data from the user, and the output is the analyzed inquiry content data.
[0375] Step 3:
[0376] The server's emotion recognition mechanism analyzes emotions from user text using an emotion recognition API (e.g., IBM Watson Tone Analyzer). It identifies emotions such as joy, anxiety, and fear from the user's words and uses the analysis results as indicators for generating responses using AI. The input is parsed text data, and the output is the emotion analysis result.
[0377] Step 4:
[0378] The server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to generate a response tailored to the user's inquiry and emotional state. The generated response aims to alleviate the user's emotions and provide appropriate solutions. The input consists of the inquiry data and the emotion analysis results, while the output is the generated response.
[0379] Step 5:
[0380] The server sends the generated response to the terminal and presents it to the user. The user receives the response via text or voice through the terminal. The input is the generated response, and the output is the information presented to the user.
[0381] Step 6:
[0382] The server stores input information and response content using information storage means. The accumulated data is used as reference for future consultations, enabling the generation of more comprehensive responses. The input is all the user's consultation and response data, and the output is an updated information data repository.
[0383] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0384] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0385] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0386] [Third Embodiment]
[0387] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0388] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0389] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0390] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0391] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0392] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0393] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0394] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0395] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0396] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0397] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0398] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0399] This invention relates to a system that utilizes generative AI to provide online consultation services. This system enables users to receive services equivalent to in-person consultations without having to visit a physical location.
[0400] Specifically, users input their inquiries from an internet-connected device. User input is in text format, and voice input is also possible if necessary. This input data is temporarily stored on the device and sent to the server via a secure communication protocol. The server has a generation AI installed that analyzes the received input data. The AI uses natural language processing technology to understand the user's intent and retrieves relevant answers from a database or generates new ones.
[0401] The server adjusts its response based on the context and sends the final answer back to the user's terminal. The user can view this answer on their terminal screen and can also choose to have it output as audio. For example, if the user asks, "What should I do next?", the server's AI will explain the necessary steps based on relevant information in an easy-to-understand manner.
[0402] Thus, because this system can provide consultations in real time, users can smoothly obtain the necessary information without being restricted by time or location. Furthermore, since the system utilizes an extensive database that includes historical data, it provides answers optimized for each user's situation. This technology allows users to receive prompt and appropriate support tailored to their individual needs.
[0403] The following describes the processing flow.
[0404] Step 1:
[0405] The user enters their inquiry into the device as text or voice. If voice input is used, the device uses its built-in speech recognition module to convert the voice into text data.
[0406] Step 2:
[0407] The terminal sends the entered text data to the server using an encryption protocol (e.g., SSL / TLS).
[0408] Step 3:
[0409] The server analyzes the received text data, and the natural language processing engine performs analysis to understand the intent of the user's question, taking context and relevance into consideration.
[0410] Step 4:
[0411] The server uses AI generation to create the optimal answer based on the analysis results. This answer utilizes a database of past consultations and related information.
[0412] Step 5:
[0413] The server formats the response and converts it into a format that is easy for the user to understand.
[0414] Step 6:
[0415] The server generates a response and sends it back to the device, which then receives the response. The data is encrypted, ensuring secure delivery.
[0416] Step 7:
[0417] The device displays the received response on the screen. If necessary, it uses an audio output module to play the response aloud for confirmation.
[0418] Step 8:
[0419] If the user wants to ask further questions, this process can be repeated, allowing for further interaction.
[0420] (Example 1)
[0421] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0422] There is a need for a system that allows users to receive consultation services efficiently and safely online without having to visit a physical office, and that can accurately understand the user's intentions and provide appropriate responses. Furthermore, secure transmission of input information and improvement of response accuracy are key challenges.
[0423] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0424] In this invention, the server includes an acquisition means for receiving input information from the user, a response generation means for analyzing the input information and generating a response according to the content of the consultation, and a transmission means for sending the generated response to the user. This makes it possible for users to conduct consultations securely over the internet while obtaining highly accurate responses using natural language processing technology.
[0425] "Acquisition means" refers to a function for accurately receiving and temporarily storing input information provided by the user.
[0426] The "response generation means" is a function that analyzes the received input information and generates an appropriate response based on the content of the consultation.
[0427] A "means of communication" refers to a function that ensures the generated response is reliably sent to the user and allows them to confirm the result.
[0428] "Recording means" refers to a function for securely storing input information and generated responses, and making them available for reference as needed.
[0429] "Secure communication means" refers to a function that uses encryption technology to ensure secure communication so that input information cannot be accessed illegally during transmission.
[0430] "Analysis means" refers to a function that uses natural language processing technology to analyze input information and understand the user's intent.
[0431] A "language conversion means" is a function that converts audio information into text information and makes it available for use in generating responses.
[0432] This system allows users to receive online consultation services using internet-connected devices. At the heart of the system is a generative AI model installed on a server, which provides automated responses to user inquiries.
[0433] On the terminal, users can input their consultation details via text or voice. Text input uses the keyboard, and voice input uses the microphone. The terminal temporarily stores this input information and sends it to the server via a secure communication protocol (e.g., HTTPS) using encryption technology.
[0434] On the server side, software equipped with natural language processing technology runs as a generative AI model. A general-purpose natural language processing model can be used as a specific example. The server analyzes the received input information and quickly and accurately understands the user's intent. The AI decides whether to retrieve the answer from an existing database or generate new information, and sends the generated response to the user. The server ensures the context and accuracy of the response and adjusts it to provide useful information for the user.
[0435] As a concrete example, consider a scenario where a user enters the question, "How do I renew my passport?" into the terminal. The server's AI analyzes this question, retrieves the necessary procedural information from the database, and provides it to the user. The generative AI model understands the user's intent and generates an appropriate and rapid response.
[0436] Examples of prompts include direct questions such as, "Could you please tell me more about the passport renewal process?"
[0437] This system allows users to ask questions in real time, without being restricted by time or location, and to obtain the necessary information with peace of mind.
[0438] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0439] Step 1:
[0440] Users input their inquiries as text or voice using an internet-connected device. The device temporarily stores the entered data. The input data is in text format; in the case of voice input, the device converts the voice to text. This ensures a consistent format for the input data.
[0441] Step 2:
[0442] The terminal sends stored input data to the server using a secure communication protocol (e.g., HTTPS). Before sending the data, the terminal encrypts the input content and transmits it securely. This process reduces the risk of unauthorized access to the data during transit.
[0443] Step 3:
[0444] The server analyzes the received data. Using a generative AI model, the server decodes the content of the input data using natural language processing techniques to understand the user's intent. This analysis identifies keywords in the input information and recognizes what kind of information is needed. The server uses these results to prepare a response.
[0445] Step 4:
[0446] The server's AI generates answers based on the analysis results. This process retrieves relevant information from a database and constructs responses appropriate to the user's questions. The server reviews the generated responses and adjusts the context and content as needed. This ensures that valuable information is provided to the user.
[0447] Step 5:
[0448] The server sends the final response to the user's device. The device displays this response on the screen and also provides an audio response if voice output is selected. Based on the information presented on the device, the user can decide on their next course of action. This allows users to quickly obtain necessary information whether they are in the office or on the go.
[0449] (Application Example 1)
[0450] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0451] Traditional customer support on e-commerce sites has struggled to provide immediate responses, which has been a factor in lowering user satisfaction. Furthermore, there is a need to provide appropriate and timely answers to consumers' specific questions about products. Therefore, a system is needed that can provide accurate, real-time responses and improve the user experience.
[0452] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0453] In this invention, the server includes receiving means for receiving input information from the user, response generating means for analyzing the input information and generating a response corresponding to the proposed content, and information presentation means for providing real-time support via a smart device. This enables the user to immediately receive accurate answers to questions about the product.
[0454] A "user" is an individual or legal entity that uses the system and provides input information.
[0455] "Input information" refers to voice or text data provided by the user as an inquiry or question.
[0456] "Receiving means" refers to a function or device for receiving input information from a user.
[0457] "Response generation means" refers to a function or algorithm for analyzing received input information and creating a response based on it.
[0458] "Transmission means" refers to a function or device for sending the generated response back to the user.
[0459] "Data storage means" refers to a data storage system for storing input information and generated responses.
[0460] "Voice conversion means" refers to a function or algorithm for converting voice input into text data.
[0461] A "smart device" is an electronic device connected to the internet, such as a smartphone or smart glasses, used by a user.
[0462] "Information presentation means" refers to a function that displays the generated response through a user interface and also outputs it as audio.
[0463] This invention provides a system that offers accurate, real-time responses when a user makes an inquiry about a product through a smart device. The system is configured as follows:
[0464] First, the user provides input information using a smart device (such as a smartphone running iOS or Android, or smart glasses). This input information is provided as text or voice and is received by an application on the smart device. In the case of voice input, it is converted into text data by a speech-to-text conversion tool (such as the Google Speech-to-Text API).
[0465] Next, the received text data is sent to the server using a secure communication protocol (HTTPS). The server is equipped with a generative AI model (for example, OpenAI's GPT series) and functions as a response generation tool. The server analyzes the input information and generates an appropriate response by referring to past database information and related information. This response may include product attributes, usage instructions, and purchase procedures.
[0466] The generated response is sent from the server to the smart device and displayed to the user by an information display device. The user can view the displayed text information and, if necessary, also obtain the information through audio output.
[0467] As a concrete example, consider a scenario where a user types "How can I ship this product?" on their smartphone. In this case, the server's generated AI model uses a prompt such as "User's question: 'How can I ship this product?' Please provide a summary and detailed instructions." to provide the user with detailed shipping procedure information as a response. In this way, the user can obtain useful information in real time.
[0468] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0469] Step 1:
[0470] The user launches an application on a smart device and provides input information. This input can be either voice or text. In the case of voice input, the terminal uses a voice conversion device to convert the voice data into text data. As a result, text data is obtained as input information.
[0471] Step 2:
[0472] The terminal securely sends the generated text data to the server using the HTTPS protocol. During this process, the received text data is converted into a secure communication packet before being transmitted, thus performing data processing. The server receives this text data via the network.
[0473] Step 3:
[0474] The server analyzes the received text data and activates a generative AI model based on its content. It inputs a specific prompt sentence (e.g., "User question: 'How do I ship this product?' Please provide a summary and detailed instructions.") into the generative AI model and generates a relevant response. At this time, the server performs data calculations by referring to past database information and related information to obtain an appropriate response that is in line with the context.
[0475] Step 4:
[0476] The generated response is sent from the server to the terminal. The terminal receives this response and displays it to the user using an information display device. The user can choose whether to display it in text format or as audio output. In this process, the terminal converts the response data obtained from the server into a user interface format and performs data processing to present it appropriately.
[0477] Step 5:
[0478] The user can review the presented information and ask further questions if necessary. A new session begins when the user provides additional information at this stage. This iterative process ensures continued interaction between the user and the system.
[0479] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0480] This invention relates to a system that combines generative AI and emotion recognition technology to provide online consultation services in a more personalized manner. This system allows users to receive emotionally sensitive responses in addition to online consultations, without having to visit a physical service center.
[0481] Users input their inquiries through their device, and can also use voice input if necessary. The device receives the input text or voice and sends it to the server. When voice is input, the device's speech recognition function converts the voice to text. The server uses its internal natural language processing engine to understand the user's inquiries and intentions in order to analyze the received text.
[0482] During the analysis process, the server's emotion engine recognizes the user's emotions, and the generating AI adjusts the response based on that emotion information. Specifically, it analyzes the tone and context of the user's text and identifies the emotional elements within it. The emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's word choice and expressions, and takes this into consideration in the response generation process.
[0483] The server uses the results of emotion recognition, referencing the database, to construct a response best suited to the user's current feelings. If the user is nervous or confused, it provides a response that includes more polite and reassuring language. For example, if the user expresses anxiety such as "the procedure is complicated and I don't understand it," the server gently explains the specific steps, adds necessary links and support information, and assists the user in a way that calms them down.
[0484] This system allows users to receive not just information, but a more deeply understood service tailored to their individual emotional state. As a result, it delivers a more satisfying user experience. This form of invention provides users with a more human-centered interaction and improves the quality of online consultation services.
[0485] The following describes the processing flow.
[0486] Step 1:
[0487] The user enters their inquiry into the device as text or voice. In the case of voice input, the device uses a speech recognition module to convert the voice into text.
[0488] Step 2:
[0489] When a terminal sends text data received from a user to a server, it uses an encryption protocol for security purposes.
[0490] Step 3:
[0491] The server receives text data and analyzes the content of the consultation using a natural language processing engine. This analysis includes understanding the context and intent.
[0492] Step 4:
[0493] The server sends the analyzed data to the emotion engine to recognize the user's emotions. The emotion engine extracts emotional information by analyzing the tone and wording of the text.
[0494] Step 5:
[0495] The server generates an appropriate response based on emotional information obtained from the emotion engine, referencing database information. In this process, it selects words that take the user's emotions into consideration.
[0496] Step 6:
[0497] The server formats the generated response, converts it into a user-friendly format, and then sends it to the terminal.
[0498] Step 7:
[0499] The device displays the received response on the screen and presents it to the user. If necessary, it plays the response aloud using speech synthesis.
[0500] Step 8:
[0501] If the user asks further questions, the entire process restarts, and the interaction with the user continues.
[0502] (Example 2)
[0503] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0504] In online counseling services, a challenge is that individual emotional states are often not adequately considered, resulting in a lack of personalized responses. Specifically, it is difficult to understand users' emotions and provide appropriate responses. This tends to lead to lower satisfaction levels in online counseling.
[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0506] In this invention, the server includes receiving means for receiving input data, conversion means for converting audio data into text data, and recognition means for recognizing the user's emotions. This makes it possible to generate personalized responses that are tailored to the user's emotions.
[0507] A "receiving means" is a function that acquires input data from the user and is responsible for the initial steps necessary to process that data within the system.
[0508] A "conversion means" is a function that handles the process of converting audio data into text data, making the audio information easier to recognize as text.
[0509] "Recognition means" refers to a function that analyzes emotional elements contained in the user's input data and identifies those emotions.
[0510] A "response generation means" is a function that forms an appropriate response to the user based on input data and recognized emotional information.
[0511] "Transmission method" refers to a function that sends the generated response to the user in an appropriate format.
[0512] "Data storage means" refers to a function that stores input data and generated responses, making them available for later reference and analysis.
[0513] This invention relates to a system that provides personalized responses to users during online individual consultations, taking into account their emotions at the time. The system mainly consists of a user terminal and a server.
[0514] The user uses the device to input the content they wish to discuss. Both text input and voice input are available, and the device is equipped with a keyboard and microphone. If voice input is selected, the device acquires the voice data and converts it into text data using specific speech recognition technology. Common software (e.g., a speech recognition API) is used for speech recognition.
[0515] The converted text data is sent from the user's terminal to the server. The server first analyzes the received text data using a natural language processing engine to understand the user's inquiry and intentions. Then, it uses an emotion recognition engine to extract emotional elements contained in the text. Various emotion analysis technologies (e.g., emotion analysis APIs) are used for emotion recognition.
[0516] When the server recognizes the user's emotional information, it uses a generative AI model to generate a response that is tailored to the user's inquiry and emotions. A general generative model is used for this AI, and an example of a prompt is, "Explain the procedure that the user finds complicated, and add support information to reassure them."
[0517] Finally, the generated response is sent to the device, where the user can view it on the screen. For example, if a user expresses concern such as "I don't understand the new procedure," the server will gently and clearly explain the procedure and provide relevant information to alleviate the user's anxiety.
[0518] Through this entire process, the system provides users not merely with information, but with a deep understanding and satisfaction tailored to their individual emotions and circumstances. This improves the quality of online consultation services and enables more human-centered interactions.
[0519] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0520] Step 1:
[0521] The user uses the terminal to input the information they wish to discuss. Input can be in text or voice format. In the case of voice input, the terminal uses its built-in microphone to acquire voice data. Then, speech recognition software is used to convert the voice data into text data. In this process, the input is voice data, and the output is text data.
[0522] Step 2:
[0523] The terminal sends the converted text data to the server. This data transmission is performed using a secure communication method. Specifically, text data is sent. The input is the converted text data, and the output is the text data received by the server.
[0524] Step 3:
[0525] The server analyzes the received text data using a natural language processing engine. This allows it to understand the grammatical structure and meaning of the text, and extract the user's intent and requests. The input is the transmitted text data, and the output is the analyzed intent information and keywords.
[0526] Step 4:
[0527] The server uses an emotion recognition engine to analyze text data and identify various emotional elements. In this process, it recognizes emotions (e.g., joy, anxiety, anger, etc.) from the user's expressions. The input for the analysis is text data, and the output is identified emotion information.
[0528] Step 5:
[0529] The server uses a generative AI model to generate responses that are appropriate to the user's emotions based on the analysis results. In this process, emotional information is taken into account, and the tone and content of the responses are adjusted accordingly. The input consists of analyzed intent and emotional information, and the output is the generated personalized response text.
[0530] Step 6:
[0531] The server sends the generated response to the terminal. A fast and secure protocol is used for transmission to the user. The input is the generated response text, and the output is the text displayed on the user's terminal.
[0532] Step 7:
[0533] Users can view the received responses on their device. Here, information tailored to the user's situation and emotions is displayed, leading to a deeper and more satisfying consultation. The output is the appropriate response displayed on the screen.
[0534] (Application Example 2)
[0535] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] In online consultations, there is a need to move away from traditional, monotonous, and uniform services and improve the user experience by providing personalized answers that are tailored to the individual emotional state of the user. However, existing systems have difficulty generating answers that take into account the user's emotional state, and this is a particular challenge when it comes to security-related consultations, as they cannot effectively alleviate the user's anxiety and fear.
[0537] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0538] In this invention, the server includes an acquisition means for receiving input information from the user, a response processing means for analyzing the input information and generating an answer based on the consultation content and the user's emotional state, and an output means for providing the generated answer to the user. This makes it possible to provide an optimal answer that takes the user's emotional state into consideration.
[0539] A "receiving mechanism" is a mechanism necessary to acquire information entered by the user and process it within the system.
[0540] A "response processing means" is a mechanism that analyzes received input information and processes it to generate the optimal response based on the content of the consultation and the user's emotions.
[0541] An "output mechanism" is a device that presents the generated response to the user and completes the system interaction.
[0542] An "emotion identification means" is a mechanism that detects emotions from user input information and reflects that information in response generation.
[0543] An "information storage means" is a mechanism for storing input information and generated responses, in preparation for future use or reference.
[0544] A "speech-to-text conversion means" is a mechanism for converting speech information into text information and preparing it in a format that can be processed by natural language processing.
[0545] The "information data repository" is a database that stores past consultation details and related information, and is used to reference this data when generating responses.
[0546] To realize this invention, the user's smartphone or tablet functions as a receiving terminal. When the user inputs their inquiry as text or voice, the terminal uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text. The converted text is then sent to the server.
[0547] On the server side, text data is analyzed by a natural language processing engine (e.g., SpaCy). This allows the user's inquiry to be understood, and an emotion recognition API (e.g., IBM Watson Tone Analyzer) is used as a means of emotion identification to evaluate the user's emotional state.
[0548] Based on this information, the server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to create a response appropriate to the user's emotional state. The generated response is returned to the terminal as text or audio.
[0549] Furthermore, accumulated consultation history information is stored in an information database and referenced during future consultations. This process allows responses to be tailored to the user's history data and emotional state, enabling a more personalized service.
[0550] For example, if a user consults the system saying, "I've recently seen a suspicious person near my home and I'm worried," the system will identify the user's anxiety and suggest specific measures that align with their feelings, such as, "First, walk in well-lit areas and install security cameras and lights. It would also be good to spread awareness through local community apps."
[0551] An example of a prompt for a generative AI model would be: "The user has expressed anxiety. Please create advice to calm them down and provide specific security measures."
[0552] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0553] Step 1:
[0554] Users input their inquiries as text or voice using a device such as a smartphone or tablet. If voice input is used, the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and sends it to the server as text data. Input is either text or voice information, and output is text information.
[0555] Step 2:
[0556] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy). It analyzes the user's inquiry content and extracts the purpose and target from the text. The input here is text data from the user, and the output is the analyzed inquiry content data.
[0557] Step 3:
[0558] The server's emotion recognition mechanism analyzes emotions from user text using an emotion recognition API (e.g., IBM Watson Tone Analyzer). It identifies emotions such as joy, anxiety, and fear from the user's words and uses the analysis results as indicators for generating responses using AI. The input is parsed text data, and the output is the emotion analysis result.
[0559] Step 4:
[0560] The server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to generate a response tailored to the user's inquiry and emotional state. The generated response aims to alleviate the user's emotions and provide appropriate solutions. The input consists of the inquiry data and the emotion analysis results, while the output is the generated response.
[0561] Step 5:
[0562] The server sends the generated response to the terminal and presents it to the user. The user receives the response via text or voice through the terminal. The input is the generated response, and the output is the information presented to the user.
[0563] Step 6:
[0564] The server stores input information and response content using information storage means. The accumulated data is used as reference for future consultations, enabling the generation of more comprehensive responses. The input is all the user's consultation and response data, and the output is an updated information data repository.
[0565] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0566] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0567] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0568] [Fourth Embodiment]
[0569] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0570] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0571] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0572] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0573] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0574] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0575] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0576] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0577] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0578] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0579] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0580] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0581] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0582] This invention relates to a system that utilizes generative AI to provide online consultation services. This system enables users to receive services equivalent to in-person consultations without having to visit a physical location.
[0583] Specifically, users input their inquiries from an internet-connected device. User input is in text format, and voice input is also possible if necessary. This input data is temporarily stored on the device and sent to the server via a secure communication protocol. The server has a generation AI installed that analyzes the received input data. The AI uses natural language processing technology to understand the user's intent and retrieves relevant answers from a database or generates new ones.
[0584] The server adjusts its response based on the context and sends the final answer back to the user's terminal. The user can view this answer on their terminal screen and can also choose to have it output as audio. For example, if the user asks, "What should I do next?", the server's AI will explain the necessary steps based on relevant information in an easy-to-understand manner.
[0585] Thus, because this system can provide consultations in real time, users can smoothly obtain the necessary information without being restricted by time or location. Furthermore, since the system utilizes an extensive database that includes historical data, it provides answers optimized for each user's situation. This technology allows users to receive prompt and appropriate support tailored to their individual needs.
[0586] The following describes the processing flow.
[0587] Step 1:
[0588] The user enters their inquiry into the device as text or voice. If voice input is used, the device uses its built-in speech recognition module to convert the voice into text data.
[0589] Step 2:
[0590] The terminal sends the entered text data to the server using an encryption protocol (e.g., SSL / TLS).
[0591] Step 3:
[0592] The server analyzes the received text data, and the natural language processing engine performs analysis to understand the intent of the user's question, taking context and relevance into consideration.
[0593] Step 4:
[0594] The server uses AI generation to create the optimal answer based on the analysis results. This answer utilizes a database of past consultations and related information.
[0595] Step 5:
[0596] The server formats the response and converts it into a format that is easy for the user to understand.
[0597] Step 6:
[0598] The server generates a response and sends it back to the device, which then receives the response. The data is encrypted, ensuring secure delivery.
[0599] Step 7:
[0600] The device displays the received response on the screen. If necessary, it uses an audio output module to play the response aloud for confirmation.
[0601] Step 8:
[0602] If the user wants to ask further questions, this process can be repeated, allowing for further interaction.
[0603] (Example 1)
[0604] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0605] There is a need for a system that allows users to receive consultation services efficiently and safely online without having to visit a physical office, and that can accurately understand the user's intentions and provide appropriate responses. Furthermore, secure transmission of input information and improvement of response accuracy are key challenges.
[0606] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0607] In this invention, the server includes an acquisition means for receiving input information from the user, a response generation means for analyzing the input information and generating a response according to the content of the consultation, and a transmission means for sending the generated response to the user. This makes it possible for users to conduct consultations securely over the internet while obtaining highly accurate responses using natural language processing technology.
[0608] "Acquisition means" refers to a function for accurately receiving and temporarily storing input information provided by the user.
[0609] The "response generation means" is a function that analyzes the received input information and generates an appropriate response based on the content of the consultation.
[0610] A "means of communication" refers to a function that ensures the generated response is reliably sent to the user and allows them to confirm the result.
[0611] "Recording means" refers to a function for securely storing input information and generated responses, and making them available for reference as needed.
[0612] "Secure communication means" refers to a function that uses encryption technology to ensure secure communication so that input information cannot be accessed illegally during transmission.
[0613] "Analysis means" refers to a function that uses natural language processing technology to analyze input information and understand the user's intent.
[0614] A "language conversion means" is a function that converts audio information into text information and makes it available for use in generating responses.
[0615] This system allows users to receive online consultation services using internet-connected devices. At the heart of the system is a generative AI model installed on a server, which provides automated responses to user inquiries.
[0616] On the terminal, users can input their consultation details via text or voice. Text input uses the keyboard, and voice input uses the microphone. The terminal temporarily stores this input information and sends it to the server via a secure communication protocol (e.g., HTTPS) using encryption technology.
[0617] On the server side, software equipped with natural language processing technology runs as a generative AI model. A general-purpose natural language processing model can be used as a specific example. The server analyzes the received input information and quickly and accurately understands the user's intent. The AI decides whether to retrieve the answer from an existing database or generate new information, and sends the generated response to the user. The server ensures the context and accuracy of the response and adjusts it to provide useful information for the user.
[0618] As a concrete example, consider a scenario where a user enters the question, "How do I renew my passport?" into the terminal. The server's AI analyzes this question, retrieves the necessary procedural information from the database, and provides it to the user. The generative AI model understands the user's intent and generates an appropriate and rapid response.
[0619] Examples of prompts include direct questions such as, "Could you please tell me more about the passport renewal process?"
[0620] This system allows users to ask questions in real time, without being restricted by time or location, and to obtain the necessary information with peace of mind.
[0621] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0622] Step 1:
[0623] Users input their inquiries as text or voice using an internet-connected device. The device temporarily stores the entered data. The input data is in text format; in the case of voice input, the device converts the voice to text. This ensures a consistent format for the input data.
[0624] Step 2:
[0625] The terminal sends stored input data to the server using a secure communication protocol (e.g., HTTPS). Before sending the data, the terminal encrypts the input content and transmits it securely. This process reduces the risk of unauthorized access to the data during transit.
[0626] Step 3:
[0627] The server analyzes the received data. Using a generative AI model, the server decodes the content of the input data using natural language processing techniques to understand the user's intent. This analysis identifies keywords in the input information and recognizes what kind of information is needed. The server uses these results to prepare a response.
[0628] Step 4:
[0629] The server's AI generates answers based on the analysis results. This process retrieves relevant information from a database and constructs responses appropriate to the user's questions. The server reviews the generated responses and adjusts the context and content as needed. This ensures that valuable information is provided to the user.
[0630] Step 5:
[0631] The server sends the final response to the user's device. The device displays this response on the screen and also provides an audio response if voice output is selected. Based on the information presented on the device, the user can decide on their next course of action. This allows users to quickly obtain necessary information whether they are in the office or on the go.
[0632] (Application Example 1)
[0633] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0634] Traditional customer support on e-commerce sites has struggled to provide immediate responses, which has been a factor in lowering user satisfaction. Furthermore, there is a need to provide appropriate and timely answers to consumers' specific questions about products. Therefore, a system is needed that can provide accurate, real-time responses and improve the user experience.
[0635] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0636] In this invention, the server includes receiving means for receiving input information from the user, response generating means for analyzing the input information and generating a response corresponding to the proposed content, and information presentation means for providing real-time support via a smart device. This enables the user to immediately receive accurate answers to questions about the product.
[0637] A "user" is an individual or legal entity that uses the system and provides input information.
[0638] "Input information" refers to voice or text data provided by the user as an inquiry or question.
[0639] "Receiving means" refers to a function or device for receiving input information from a user.
[0640] "Response generation means" refers to a function or algorithm for analyzing received input information and creating a response based on it.
[0641] "Transmission means" refers to a function or device for sending the generated response back to the user.
[0642] "Data storage means" refers to a data storage system for storing input information and generated responses.
[0643] "Voice conversion means" refers to a function or algorithm for converting voice input into text data.
[0644] A "smart device" is an electronic device connected to the internet, such as a smartphone or smart glasses, used by a user.
[0645] "Information presentation means" refers to a function that displays the generated response through a user interface and also outputs it as audio.
[0646] This invention provides a system that offers accurate, real-time responses when a user makes an inquiry about a product through a smart device. The system is configured as follows:
[0647] First, the user provides input information using a smart device (such as a smartphone running iOS or Android, or smart glasses). This input information is provided as text or voice and is received by an application on the smart device. In the case of voice input, it is converted into text data by a speech-to-text conversion tool (such as the Google Speech-to-Text API).
[0648] Next, the received text data is sent to the server using a secure communication protocol (HTTPS). The server is equipped with a generative AI model (for example, OpenAI's GPT series) and functions as a response generation tool. The server analyzes the input information and generates an appropriate response by referring to past database information and related information. This response may include product attributes, usage instructions, and purchase procedures.
[0649] The generated response is sent from the server to the smart device and displayed to the user by an information display device. The user can view the displayed text information and, if necessary, also obtain the information through audio output.
[0650] As a concrete example, consider a scenario where a user types "How can I ship this product?" on their smartphone. In this case, the server's generated AI model uses a prompt such as "User's question: 'How can I ship this product?' Please provide a summary and detailed instructions." to provide the user with detailed shipping procedure information as a response. In this way, the user can obtain useful information in real time.
[0651] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0652] Step 1:
[0653] The user launches an application on a smart device and provides input information. This input can be either voice or text. In the case of voice input, the terminal uses a voice conversion device to convert the voice data into text data. As a result, text data is obtained as input information.
[0654] Step 2:
[0655] The terminal securely sends the generated text data to the server using the HTTPS protocol. During this process, the received text data is converted into a secure communication packet before being transmitted, thus performing data processing. The server receives this text data via the network.
[0656] Step 3:
[0657] The server analyzes the received text data and activates a generative AI model based on its content. It inputs a specific prompt sentence (e.g., "User question: 'How do I ship this product?' Please provide a summary and detailed instructions.") into the generative AI model and generates a relevant response. At this time, the server performs data calculations by referring to past database information and related information to obtain an appropriate response that is in line with the context.
[0658] Step 4:
[0659] The generated response is sent from the server to the terminal. The terminal receives this response and displays it to the user using an information display device. The user can choose whether to display it in text format or as audio output. In this process, the terminal converts the response data obtained from the server into a user interface format and performs data processing to present it appropriately.
[0660] Step 5:
[0661] The user can review the presented information and ask further questions if necessary. A new session begins when the user provides additional information at this stage. This iterative process ensures continued interaction between the user and the system.
[0662] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0663] This invention relates to a system that combines generative AI and emotion recognition technology to provide online consultation services in a more personalized manner. This system allows users to receive emotionally sensitive responses in addition to online consultations, without having to visit a physical service center.
[0664] Users input their inquiries through their device, and can also use voice input if necessary. The device receives the input text or voice and sends it to the server. When voice is input, the device's speech recognition function converts the voice to text. The server uses its internal natural language processing engine to understand the user's inquiries and intentions in order to analyze the received text.
[0665] During the analysis process, the server's emotion engine recognizes the user's emotions, and the generating AI adjusts the response based on that emotion information. Specifically, it analyzes the tone and context of the user's text and identifies the emotional elements within it. The emotion engine recognizes emotions such as joy, sadness, anger, and surprise from the user's word choice and expressions, and takes this into consideration in the response generation process.
[0666] The server uses the results of emotion recognition, referencing the database, to construct a response best suited to the user's current feelings. If the user is nervous or confused, it provides a response that includes more polite and reassuring language. For example, if the user expresses anxiety such as "the procedure is complicated and I don't understand it," the server gently explains the specific steps, adds necessary links and support information, and assists the user in a way that calms them down.
[0667] This system allows users to receive not just information, but a more deeply understood service tailored to their individual emotional state. As a result, it delivers a more satisfying user experience. This form of invention provides users with a more human-centered interaction and improves the quality of online consultation services.
[0668] The following describes the processing flow.
[0669] Step 1:
[0670] The user enters their inquiry into the device as text or voice. In the case of voice input, the device uses a speech recognition module to convert the voice into text.
[0671] Step 2:
[0672] When a terminal sends text data received from a user to a server, it uses an encryption protocol for security purposes.
[0673] Step 3:
[0674] The server receives text data and analyzes the content of the consultation using a natural language processing engine. This analysis includes understanding the context and intent.
[0675] Step 4:
[0676] The server sends the analyzed data to the emotion engine to recognize the user's emotions. The emotion engine extracts emotional information by analyzing the tone and wording of the text.
[0677] Step 5:
[0678] The server generates an appropriate response based on emotional information obtained from the emotion engine, referencing database information. In this process, it selects words that take the user's emotions into consideration.
[0679] Step 6:
[0680] The server formats the generated response, converts it into a user-friendly format, and then sends it to the terminal.
[0681] Step 7:
[0682] The device displays the received response on the screen and presents it to the user. If necessary, it plays the response aloud using speech synthesis.
[0683] Step 8:
[0684] If the user asks further questions, the entire process restarts, and the interaction with the user continues.
[0685] (Example 2)
[0686] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0687] In online counseling services, a challenge is that individual emotional states are often not adequately considered, resulting in a lack of personalized responses. Specifically, it is difficult to understand users' emotions and provide appropriate responses. This tends to lead to lower satisfaction levels in online counseling.
[0688] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0689] In this invention, the server includes receiving means for receiving input data, conversion means for converting audio data into text data, and recognition means for recognizing the user's emotions. This makes it possible to generate personalized responses that are tailored to the user's emotions.
[0690] A "receiving means" is a function that acquires input data from the user and is responsible for the initial steps necessary to process that data within the system.
[0691] A "conversion means" is a function that handles the process of converting audio data into text data, making the audio information easier to recognize as text.
[0692] "Recognition means" refers to a function that analyzes emotional elements contained in the user's input data and identifies those emotions.
[0693] A "response generation means" is a function that forms an appropriate response to the user based on input data and recognized emotional information.
[0694] "Transmission method" refers to a function that sends the generated response to the user in an appropriate format.
[0695] "Data storage means" refers to a function that stores input data and generated responses, making them available for later reference and analysis.
[0696] This invention relates to a system that provides personalized responses to users during online individual consultations, taking into account their emotions at the time. The system mainly consists of a user terminal and a server.
[0697] The user uses the device to input the content they wish to discuss. Both text input and voice input are available, and the device is equipped with a keyboard and microphone. If voice input is selected, the device acquires the voice data and converts it into text data using specific speech recognition technology. Common software (e.g., a speech recognition API) is used for speech recognition.
[0698] The converted text data is sent from the user's terminal to the server. The server first analyzes the received text data using a natural language processing engine to understand the user's inquiry and intentions. Then, it uses an emotion recognition engine to extract emotional elements contained in the text. Various emotion analysis technologies (e.g., emotion analysis APIs) are used for emotion recognition.
[0699] When the server recognizes the user's emotional information, it uses a generative AI model to generate a response that is tailored to the user's inquiry and emotions. A general generative model is used for this AI, and an example of a prompt is, "Explain the procedure that the user finds complicated, and add support information to reassure them."
[0700] Finally, the generated response is sent to the device, where the user can view it on the screen. For example, if a user expresses concern such as "I don't understand the new procedure," the server will gently and clearly explain the procedure and provide relevant information to alleviate the user's anxiety.
[0701] Through this entire process, the system provides users not merely with information, but with a deep understanding and satisfaction tailored to their individual emotions and circumstances. This improves the quality of online consultation services and enables more human-centered interactions.
[0702] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0703] Step 1:
[0704] The user uses the terminal to input the information they wish to discuss. Input can be in text or voice format. In the case of voice input, the terminal uses its built-in microphone to acquire voice data. Then, speech recognition software is used to convert the voice data into text data. In this process, the input is voice data, and the output is text data.
[0705] Step 2:
[0706] The terminal sends the converted text data to the server. This data transmission is performed using a secure communication method. Specifically, text data is sent. The input is the converted text data, and the output is the text data received by the server.
[0707] Step 3:
[0708] The server analyzes the received text data using a natural language processing engine. This allows it to understand the grammatical structure and meaning of the text, and extract the user's intent and requests. The input is the transmitted text data, and the output is the analyzed intent information and keywords.
[0709] Step 4:
[0710] The server uses an emotion recognition engine to analyze text data and identify various emotional elements. In this process, it recognizes emotions (e.g., joy, anxiety, anger, etc.) from the user's expressions. The input for the analysis is text data, and the output is identified emotion information.
[0711] Step 5:
[0712] The server uses a generative AI model to generate responses that are appropriate to the user's emotions based on the analysis results. In this process, emotional information is taken into account, and the tone and content of the responses are adjusted accordingly. The input consists of analyzed intent and emotional information, and the output is the generated personalized response text.
[0713] Step 6:
[0714] The server sends the generated response to the terminal. A fast and secure protocol is used for transmission to the user. The input is the generated response text, and the output is the text displayed on the user's terminal.
[0715] Step 7:
[0716] Users can view the received responses on their device. Here, information tailored to the user's situation and emotions is displayed, leading to a deeper and more satisfying consultation. The output is the appropriate response displayed on the screen.
[0717] (Application Example 2)
[0718] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0719] In online consultations, there is a need to move away from traditional, monotonous, and uniform services and improve the user experience by providing personalized answers that are tailored to the individual emotional state of the user. However, existing systems have difficulty generating answers that take into account the user's emotional state, and this is a particular challenge when it comes to security-related consultations, as they cannot effectively alleviate the user's anxiety and fear.
[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0721] In this invention, the server includes an acquisition means for receiving input information from the user, a response processing means for analyzing the input information and generating an answer based on the consultation content and the user's emotional state, and an output means for providing the generated answer to the user. This makes it possible to provide an optimal answer that takes the user's emotional state into consideration.
[0722] A "receiving mechanism" is a mechanism necessary to acquire information entered by the user and process it within the system.
[0723] A "response processing means" is a mechanism that analyzes received input information and processes it to generate the optimal response based on the content of the consultation and the user's emotions.
[0724] An "output mechanism" is a device that presents the generated response to the user and completes the system interaction.
[0725] An "emotion identification means" is a mechanism that detects emotions from user input information and reflects that information in response generation.
[0726] An "information storage means" is a mechanism for storing input information and generated responses, in preparation for future use or reference.
[0727] A "speech-to-text conversion means" is a mechanism for converting speech information into text information and preparing it in a format that can be processed by natural language processing.
[0728] The "information data repository" is a database that stores past consultation details and related information, and is used to reference this data when generating responses.
[0729] To realize this invention, the user's smartphone or tablet functions as a receiving terminal. When the user inputs their inquiry as text or voice, the terminal uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text. The converted text is then sent to the server.
[0730] On the server side, text data is analyzed by a natural language processing engine (e.g., SpaCy). This allows the user's inquiry to be understood, and an emotion recognition API (e.g., IBM Watson Tone Analyzer) is used as a means of emotion identification to evaluate the user's emotional state.
[0731] Based on this information, the server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to create a response appropriate to the user's emotional state. The generated response is returned to the terminal as text or audio.
[0732] Furthermore, accumulated consultation history information is stored in an information database and referenced during future consultations. This process allows responses to be tailored to the user's history data and emotional state, enabling a more personalized service.
[0733] For example, if a user consults the system saying, "I've recently seen a suspicious person near my home and I'm worried," the system will identify the user's anxiety and suggest specific measures that align with their feelings, such as, "First, walk in well-lit areas and install security cameras and lights. It would also be good to spread awareness through local community apps."
[0734] An example of a prompt for a generative AI model would be: "The user has expressed anxiety. Please create advice to calm them down and provide specific security measures."
[0735] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0736] Step 1:
[0737] Users input their inquiries as text or voice using a device such as a smartphone or tablet. If voice input is used, the device uses a speech recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and sends it to the server as text data. Input is either text or voice information, and output is text information.
[0738] Step 2:
[0739] The server analyzes the received text data using a natural language processing engine (e.g., SpaCy). It analyzes the user's inquiry content and extracts the purpose and target from the text. The input here is text data from the user, and the output is the analyzed inquiry content data.
[0740] Step 3:
[0741] The server's emotion recognition mechanism analyzes emotions from user text using an emotion recognition API (e.g., IBM Watson Tone Analyzer). It identifies emotions such as joy, anxiety, and fear from the user's words and uses the analysis results as indicators for generating responses using AI. The input is parsed text data, and the output is the emotion analysis result.
[0742] Step 4:
[0743] The server's response processing mechanism uses a generative AI model (e.g., GPT-3.5) to generate a response tailored to the user's inquiry and emotional state. The generated response aims to alleviate the user's emotions and provide appropriate solutions. The input consists of the inquiry data and the emotion analysis results, while the output is the generated response.
[0744] Step 5:
[0745] The server sends the generated response to the terminal and presents it to the user. The user receives the response via text or voice through the terminal. The input is the generated response, and the output is the information presented to the user.
[0746] Step 6:
[0747] The server stores input information and response content using information storage means. The accumulated data is used as reference for future consultations, enabling the generation of more comprehensive responses. The input is all the user's consultation and response data, and the output is an updated information data repository.
[0748] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0749] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0750] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0751] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0752] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0753] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0754] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0755] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0756] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0757] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0758] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0759] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0760] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0761] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0762] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0763] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0764] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0765] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0766] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0767] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0768] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0769] The following is further disclosed regarding the embodiments described above.
[0770] (Claim 1)
[0771] A receiving means for receiving input data from the user,
[0772] The aforementioned input data is analyzed and a response generation means is generated to produce a response corresponding to the content of the consultation.
[0773] A transmission means for sending the generated response to the user,
[0774] A data storage means for storing the input data and the generated response,
[0775] A system that includes this.
[0776] (Claim 2)
[0777] The system according to claim 1, further comprising a text conversion means for converting the audio data into text data when the input data is audio data.
[0778] (Claim 3)
[0779] The system according to claim 1, wherein the response generation means generates the response by referring to past database information and related information.
[0780] "Example 1"
[0781] (Claim 1)
[0782] A means for receiving input information from the user,
[0783] A response generation means that analyzes the input information and generates a response corresponding to the content of the consultation,
[0784] A means for transmitting the generated response to the user,
[0785] Recording means for storing the input information and the generated response,
[0786] A secure communication means that transmits the input information to the server using secure communication technology,
[0787] The response generation means includes an analysis means that uses natural language processing technology,
[0788] A system that includes this.
[0789] (Claim 2)
[0790] The system according to claim 1, further comprising a language conversion means for converting the input information into text information when the input information is audio information.
[0791] (Claim 3)
[0792] The system according to claim 1, wherein the response generation means refers to past data and related information, understands the user's intent, and generates a response.
[0793] "Application Example 1"
[0794] (Claim 1)
[0795] A receiving means for receiving input information from the user,
[0796] A response generation means that analyzes the input information and generates a response corresponding to the proposed content,
[0797] A transmission means for sending the generated response to the user,
[0798] A data storage means for storing the input information and the generated response,
[0799] A voice conversion method that converts voice input to text,
[0800] Information presentation means that provides real-time support via smart devices,
[0801] A system that includes this.
[0802] (Claim 2)
[0803] The system according to claim 1, wherein the response generation means generates the response by referring to past database information and related information.
[0804] (Claim 3)
[0805] The system according to claim 1, wherein the voice conversion means converts the voice input into text data using voice recognition technology.
[0806] "Example 2 of combining an emotion engine"
[0807] (Claim 1)
[0808] A receiving means for receiving input data from the user,
[0809] A conversion means for converting audio data into text data when the input data is audio data,
[0810] A recognition method that analyzes input data and recognizes the user's emotions,
[0811] A response generation means that generates a response according to the content of the consultation using information on emotion recognition,
[0812] A means for sending the generated response to the user,
[0813] A data storage means for storing input data and generated responses,
[0814] A system that includes this.
[0815] (Claim 2)
[0816] The system according to claim 1, wherein the response generation means refers to past information data and related information and adjusts the response based on sentiment information.
[0817] (Claim 3)
[0818] The system according to claim 1, wherein the recognition means analyzes the context and tone of the text and identifies a plurality of emotional elements.
[0819] "Application example 2 when combining with an emotional engine"
[0820] (Claim 1)
[0821] A means for receiving input information from the user,
[0822] A response processing means that analyzes the input information and generates an answer based on the consultation content and the user's emotional state,
[0823] An output means for providing the generated response to the user,
[0824] The response processing means includes an emotion identification means for detecting the user's emotions,
[0825] Information storage means for storing the input information and the generated response,
[0826] A system that includes this.
[0827] (Claim 2)
[0828] The system according to claim 1, further comprising a speech-to-text conversion means for converting the input information into text information when the input information is audio information.
[0829] (Claim 3)
[0830] The system according to claim 1, wherein the response processing means generates the response by referring to a past information data repository and related information, and provides a response optimized for the user's emotional state. [Explanation of symbols]
[0831] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A receiving means for receiving input data from the user, The aforementioned input data is analyzed and a response generation means is generated to produce a response corresponding to the content of the consultation. A transmission means for sending the generated response to the user, A data storage means for storing the input data and the generated response, A system that includes this.
2. The system according to claim 1, further comprising a text conversion means for converting the audio data into text data when the input data is audio data.
3. The system according to claim 1, wherein the response generation means generates the response by referring to past database information and related information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A