System
A system with a chat-style interface and generative AI model addresses the lack of easy access to smartphone instruction manuals by providing real-time answers, enhancing user satisfaction.
Patent Information
- Application Number
- JP2024128420
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
Smartphones lack easy access to instruction manuals, leading to user frustration and increased support inquiries, resulting in delayed responses.
A system that allows users to input inquiries through a chat-style interface, analyzed by a natural language processing engine, with a generative AI model providing real-time answers based on the inquiry content.
Enables users to quickly resolve questions about smartphone usage, improving customer satisfaction by providing reliable information.
Smart Images

Figure 2026025611000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current smartphones do not come with instruction manuals, leaving users with few easy ways to obtain information about how to use their smartphones. This often leaves users frustrated with the lack of official means to resolve questions or problems they may have about using their smartphones. Furthermore, this can lead to an increase in inquiries to support centers, resulting in delayed responses. Therefore, there is a need for a system that allows users to quickly and easily learn how to use their smartphones. [Means for solving the problem]
[0005] This invention relates to a system in which a server receives inquiries about how to use a smartphone when the user inputs them through a chat-style interface and analyzes the inquiries using a natural language processing engine. Based on the analysis results, a generative AI model generates appropriate answers and sends them to the user's device, allowing the user to obtain the information they need in real time. This system allows users to quickly resolve questions about how to use their smartphone, reducing stress and improving customer satisfaction. Furthermore, by providing answers based on the contents of the instruction manual, the system maintains its credibility as an official information source.
[0006] A "user" is an individual or organization that uses a smartphone.
[0007] "Terminal" refers to a smartphone or similar device used by a user.
[0008] A "chat-style interface" is a form of user interface for an application that allows users to enter text queries and obtain information interactively.
[0009] An "inquiry" is a question entered in text format by a user to provide information or problems they would like to know about using their smartphone.
[0010] "Receiving" refers to the action of the server taking in the text data entered by the user.
[0011] A "server" is a computer system that receives queries from users, analyzes them, and generates responses.
[0012] A "natural language processing engine" is a software system that analyzes the content of a user's inquiry and understands its meaning and intent.
[0013] "Analysis" refers to the process of processing input text data and understanding its content and intent.
[0014] A "generative AI model" is an artificial intelligence model that generates appropriate answers based on analysis results.
[0015] An "answer" is the information or explanation that a generative AI model provides in response to a user's inquiry.
[0016] "Sending" refers to the action of sending the generated answer from the server to the user's terminal.
[0017] "System" refers to the entire technical framework through which users can make inquiries and receive answers in accordance with the present invention. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that allows users to ask questions about how to use their smartphones in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. The system's main components are the user, the device, and a server.
[0040] System configuration
[0041] User operations
[0042] 1. The user launches the official app and goes to the help section.
[0043] 2. The user opens a chat box and types in a text message with a usage question.
[0044] Example: "How do I set up Wi-Fi?"
[0045] Device behavior
[0046] 1. The device receives the text entered by the user and sends it to the server.
[0047] The data is sent via an API.
[0048] Server Operation
[0049] 1. The server receives a query from a user.
[0050] 2. The server sends the received text to a natural language processing engine for analysis.
[0051] 3. The server receives the analysis results from the natural language processing engine and sends them to the generative AI model.
[0052] Natural Language Processing and AI Model Behavior
[0053] 1. The natural language processing engine analyzes the user's inquiry and understands its intent.
[0054] Example: Identifying the query intent as "How do I set up Wi-Fi?"
[0055] 2. The generative AI model generates an appropriate answer based on the analysis results.
[0056] For example, generate an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0057] Server response processing
[0058] 1. The server receives the generated answer and sends it back to the user's device.
[0059] Terminal display
[0060] 1. The device displays the received response in the chat box for the user to view.
[0061] 2. The user checks the answers and follows the instructions to set up the smartphone.
[0062] Specific examples
[0063] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0064] This system allows users to instantly obtain the information they need in a chat format, quickly resolving any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0065] The processing flow will be explained below.
[0066] Step 1:
[0067] A user launches the official app and navigates to the help section. The user opens the chat box and enters their inquiry in text format. For example, they might type, "How do I set up Wi-Fi?"
[0068] Step 2:
[0069] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0070] Step 3:
[0071] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0072] Step 4:
[0073] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0074] Step 5:
[0075] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0076] Step 6:
[0077] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, also via an API.
[0078] Step 7:
[0079] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0080] Step 8:
[0081] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0082] Step 9:
[0083] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0084] Example 1
[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0086] Users of modern smartphones and other electronic devices often have questions about how to operate their devices. However, traditional methods offer limited options for getting fast, accurate answers when users encounter problems. Manually searching for information in instruction manuals or online help takes time and effort. Therefore, a new system is needed to help users solve problems efficiently.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes means for a user to input an inquiry via a chat-style interface of the electronic device, means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, and means for transmitting the generated answer to the user's electronic device. This allows the user to quickly receive an accurate answer in chat format and efficiently resolve questions about how to operate the device.
[0089] "User" refers to a user who operates an electronic device.
[0090] "Electronic devices" refers to computing devices such as smartphones, tablets, and personal computers.
[0091] "Chat-style interface" refers to a user interface that allows users to enter text messages and communicate information interactively.
[0092] An "inquiry" refers to a user inputting a question or doubt about how to operate an electronic device.
[0093] A "server" refers to a computer system that receives queries from users, processes them, and generates and transmits responses.
[0094] A "natural language processing engine" refers to software that analyzes text input from a user and understands their intent.
[0095] A "generative AI model" refers to an artificial intelligence model that automatically generates appropriate answers based on the results analyzed by a natural language processing engine.
[0096] "Answer" refers to information generated by a generative AI model that includes answers or instructions to a user's inquiry.
[0097] "Instructions" refers to documents or sources of information that contain detailed instructions on the operation and functionality of an electronic device.
[0098] This invention is a system that allows users to ask questions about how to use electronic devices in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. This system mainly consists of a user, a terminal, and a server.
[0099] System configuration
[0100] User operations
[0101] A user launches the official app and navigates to the help section. The user opens the chat box and enters a usage question in text format. For example, if the user types "How do I set up Wi-Fi?", the following process occurs:
[0102] Device behavior
[0103] The device takes the text entered by the user and sends it to the server, where it is packaged in JSON format and protected by SSL / TLS encryption via an API.
[0104] Server Operation
[0105] The server receives a query from a user. The server passes the received request to a processing program via a web server such as Apache or Nginx. For example, a Flask application written in Python handles this request. The server then sends the received text to a natural language processing engine. This analysis is performed using the Google Cloud Natural Language API or an NLP library (e.g., spaCy, NLTK). After the analysis is complete, the server sends the results to a generative AI model (e.g., OpenAI's GPT-3 or GPT-4).
[0106] Natural Language Processing and AI Model Behavior
[0107] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, in response to the inquiry "How do I set up Wi-Fi?", the keyword "How to set up Wi-Fi" is extracted. Next, the generative AI model generates an appropriate answer based on the analysis results. For example, it generates an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password."
[0108] Server response processing
[0109] The server receives the generated answer and sends it back to the user's device, where the data is again sent via an HTTP response.
[0110] Terminal display
[0111] The device displays the received answer in a chat box for the user to view. The user can then check the answer and follow the instructions to set up their smartphone.
[0112] Specific examples
[0113] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0114] Example prompt sentence:
[0115] The user types "How do I set up Wi-Fi?" into the chat box. Generate an appropriate response.
[0116] This system allows users to instantly obtain the information they need in a chat format, enabling them to quickly resolve any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0118] Step 1:
[0119] A user launches the official app and navigates to the help section. The user unlocks the smartphone and taps the app icon. The app launches and displays the home screen. The user clicks the help or support tab within the app.
[0120] Input: Smartphone operation
[0121] Output: Help section on screen
[0122] Step 2:
[0123] The user opens the chat box, types a usage question in text format, for example, "How do I set up Wi-Fi?", and taps the send button.
[0124] Input: User question text
[0125] Output: The terminal displays the user's input text.
[0126] Step 3:
[0127] The device retrieves the user's input text and sends it to the server. The device generates an HTTP POST request, packages the user's input text in JSON format, and sends it to the SSL / TLS encrypted API endpoint.
[0128] Input: User-entered text
[0129] Output: Encrypted HTTP POST request
[0130] Step 4:
[0131] The server receives a query from a user. The server receives the request through a web server (e.g., Nginx, Apache) and passes the request to its internal processing program (e.g., a Python Flask application).
[0132] Input: Encrypted HTTP POST request
[0133] Output: The request data passed to the processing program
[0134] Step 5:
[0135] The server sends the received text to a natural language processing engine, which extracts the text from the request data, sends it to the natural language processing engine (e.g., Google Cloud Natural Language API), and makes an API call for analysis.
[0136] Input: The text portion of the request data
[0137] Output: API request to the natural language processing engine
[0138] Step 6:
[0139] A natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it analyzes the text "How do I set up Wi-Fi?" and identifies the intent as "How do I set up Wi-Fi?"
[0140] Input: The text to be parsed
[0141] Output: Analysis results (intent, keywords, etc.)
[0142] Step 7:
[0143] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model. The server then generates an API request to pass the analysis results to the generative AI model (e.g., OpenAI GPT-3). This request includes a prompt.
[0144] Input: Analysis results of the natural language processing engine
[0145] Output: API request to the generative AI model
[0146] Step 8:
[0147] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0148] Input: prompt sentence for generative AI model
[0149] Output: Generated answer text
[0150] Step 9:
[0151] The server receives the generated answer and returns it to the user's device. The server packages the generated answer in JSON format and sends it to the user's device as an HTTP response.
[0152] Input: Generated answer text
[0153] Output: Answer sent to the terminal as an HTTP response
[0154] Step 10:
[0155] The device displays the received answer in the chat box for the user to view. The device deserializes the received JSON data and displays the answer text in the chat box. The user checks the answer and follows the instructions to configure their smartphone.
[0156] Input: JSON data of the HTTP response
[0157] Output: Answer text displayed in the chat box
[0158] (Application example 1)
[0159] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0160] In the past, it was difficult for workers to receive prompt and appropriate support when operating or troubleshooting robots used in factories. This often resulted in reduced work efficiency and lost productivity. In such situations, a system was needed that would enable workers to quickly understand how to use robots and immediately provide reliable information to solve problems.
[0161] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0162] In this invention, the server includes means for a user to input an inquiry via a chat-style interface on a terminal, server means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, means for sending the generated answer to the user's terminal, and means for analyzing the inquiry specifically regarding the operation of a robot in a factory and generating appropriate instructions, thereby enabling the user to obtain detailed information on how to operate and troubleshoot the robot in a factory in real time.
[0163] "User" refers to any individual or entity that uses the System to make an inquiry.
[0164] "Terminal" refers to a computer or device for entering queries and displaying responses.
[0165] A "chat-style interface" refers to a user interface in which a user inputs a text-based inquiry and receives a response from the system.
[0166] "Server" refers to a central processing unit for receiving queries from users, analyzing them, and generating responses.
[0167] A "natural language processing engine" refers to a machine learning or rule-based system that analyzes user queries and understands their intent.
[0168] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on the analysis results of a natural language processing engine.
[0169] "Inquiries regarding the operation of robots in factories" refers to questions regarding how to operate and troubleshoot robots used in factories.
[0170] An "answer" refers to the information or instructions generated by a generative AI model in response to a user's inquiry.
[0171] The system for realizing this application example operates in the following procedure.
[0172] First, the user accesses a chat-style interface using a device, such as a PC, tablet, or smartphone, and enters text-based inquiries about the robot's operation.
[0173] The device takes the entered text data and sends it over the Internet to a server using a standard HTTP request.
[0174] The server processes the received query. First, the query is analyzed by a natural language processing engine. This engine uses natural language processing libraries such as spaCy or NLTK. The analysis identifies the intent of the query and important keywords.
[0175] The server then uses a generative AI model, such as GPT-4, to generate an answer based on the analysis results. The AI model generates an appropriate and detailed answer based on the analyzed intent and keywords.
[0176] The generated answers are then sent back to the terminal via the server and displayed to the user, allowing the user to obtain appropriate operating instructions and troubleshooting methods in real time.
[0177] For example, if a user types, "Please tell me how to calibrate my robot," this query is analyzed by a natural language processing engine, and the generative AI model generates the answer, "Enter calibration mode and adjust each sensor according to the manual."
[0178] In this way, the system can provide users with immediate and reliable information regarding their robot operation questions.
[0179] Example prompt sentence:
[0180] User: "How do I calibrate my robot?"
[0181] Answer: "Enter calibration mode and adjust each sensor according to the manual."
[0182] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0183] Step 1:
[0184] The user uses a terminal to access a chat-style interface and enters a text inquiry about the operation of the robot, for example, "Tell me how to calibrate the robot."
[0185] Input: User's inquiry (text format)
[0186] Output: Query text data
[0187] Step 2:
[0188] The device receives the text data entered by the user and sends it to the server as an API request, using the standard HTTP protocol.
[0189] Input: User query text data
[0190] Output: HTTP request to the server
[0191] Step 3:
[0192] The server sends the received query to a natural language processing engine for analysis, which analyzes the query content and identifies its intent and important keywords.
[0193] Input: Query text data as an HTTP request
[0194] Output: Analysis results (intent and keywords)
[0195] What it does: Analyzes text using natural language processing libraries such as spaCy and NLTK
[0196] Step 4:
[0197] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, which then generates an appropriate answer based on the analysis results.
[0198] Input: Analysis results (intent and keywords)
[0199] Output: Generated answer text
[0200] Specific operation: Generate text using generative AI models such as GPT-4
[0201] Step 5:
[0202] The server retrieves the generated answer and sends it to the user's device using an HTTP response.
[0203] Input: Generated answer text
[0204] Output: HTTP response to the user's device
[0205] Step 6:
[0206] The terminal displays the received response on the chat interface so that it can be seen by the user, who can then confirm it and perform the necessary operations.
[0207] Input: HTTP response containing the answer text from the server
[0208] Output: Answer text displayed in the chat interface
[0209] Specific operation: The received response text is displayed on the screen. For example, "Enter calibration mode and adjust each sensor according to the manual."
[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0211] This invention combines a system in which a user inputs an inquiry about how to use a smartphone in chat format, analyzes the inquiry content with a natural language processing engine, and generates an answer with a generative AI model, with an emotion engine that recognizes the user's emotions.This system includes the user, device, server, natural language processing engine, generative AI model, and emotion engine as its main components.
[0212] System configuration
[0213] User operations
[0214] 1. The user launches the official app and goes to the help section.
[0215] 2. The user opens a chat box and types in a text message with a usage question.
[0216] Example: "How do I set up Wi-Fi?"
[0217] Device behavior
[0218] 1. The device receives the text entered by the user and sends it to the server via the internet via an API.
[0219] Server Operation
[0220] 1. The server receives a query from the user. The server temporarily stores the received data and prepares it for the next process.
[0221] 2. The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0222] How Natural Language Processing and Sentiment Analysis Work
[0223] 1. The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the user's intent as "How do I set up Wi-Fi?"
[0224] 2. The emotion engine analyzes the user's query text to determine their emotions, for example, identifying that the user is feeling "troubled" or "angry."
[0225] Server Processing
[0226] 1. The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API.
[0227] Answer generation behavior
[0228] 1. The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and content of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0229] Server response processing
[0230] 1. The server receives the generated answer and sends it back to the user's device, via the internet through an API.
[0231] Terminal display
[0232] 1. The terminal displays the received answer in the chat box. The user can view the displayed answer and follow the instructions to perform the operation.
[0233] Specific examples
[0234] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app and the text includes confusion or irritation, the device sends the query to a server. The server receives the text, sends it to a natural language processing engine to analyze the intent, and uses an emotion engine to analyze the user's emotions. The analysis results identify that the query's intent is "How do I set up Wi-Fi?" and that the user is having trouble. Based on this analysis, a generative AI model generates a specific, emotion-sensitive answer and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help," and can follow the instructions to perform the operation.
[0235] This system not only allows users to instantly obtain the information they need in a chat format, but also improves their understanding of how to use the device and the smooth operation of their smartphone by receiving responses that take into consideration their emotions at the time.In addition, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0236] The processing flow will be explained below.
[0237] Step 1:
[0238] A user launches the official app and navigates to the help section. The user opens the chat box and types a usage question in text format. For example, they type "How do I set up Wi-Fi?"
[0239] Step 2:
[0240] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0241] Step 3:
[0242] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0243] Step 4:
[0244] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0245] Step 5:
[0246] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0247] Step 6:
[0248] The server receives the analysis results from the natural language processing engine, and at the same time, sends the user's query text to the emotion engine for emotion analysis.
[0249] Step 7:
[0250] The emotion engine analyzes the user's text to identify their emotional state, for example detecting emotions such as "confused" or "frustrated."
[0251] Step 8:
[0252] The server integrates the analysis results from the natural language processing engine and the emotion engine and sends them to the generative AI model, also via an API.
[0253] Step 9:
[0254] The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and details of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0255] Step 10:
[0256] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0257] Step 11:
[0258] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0259] Step 12:
[0260] The user then performs the smartphone configuration based on the answers, for example opening the Settings app, tapping the 'Wi-Fi' tab, selecting an available network and entering the password.
[0261] Through this series of processes, users not only receive information but also receive instructions that take into consideration their emotions and state, thereby reducing stress and improving the user experience.
[0262] Example 2
[0263] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0264] With conventional interfaces, even when a user asks a question about how to operate a smartphone, the answer does not take into consideration the user's emotions at the time. As a result, even when the user is confused or angry, only mechanical answers are provided, which does not improve the user experience. There was also a need for a system that could provide quick and appropriate answers.
[0265] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0266] In this invention, the server includes an information processing server means for receiving the inquiry content, a means for analyzing the received inquiry content using a natural language processing engine, a sentiment analysis means for analyzing sentiment from text entered by the user, a means for generating an answer based on the analysis result and the sentiment analysis result using a generative AI model, and a means for transmitting the generated answer to the user's information processing device, thereby making it possible to provide an appropriate and detailed answer that takes into consideration the user's sentiment.
[0267] The term "user" refers to a person who makes an inquiry using an information processing device.
[0268] An "information processing device" is a device that a user can operate and input an inquiry into, and includes a smartphone, a tablet, and the like.
[0269] "Chat-style interface" refers to a screen and software that allows users to enter messages and interact in text format.
[0270] An "information processing server" refers to a computer system that receives inquiries from users and performs processing to analyze and generate answers.
[0271] "Natural language processing engine" refers to software and algorithms that analyze a user's text input to understand their intent.
[0272] "Sentiment analyzer" refers to software and algorithms for analyzing emotions from a user's input text and identifying their emotional state.
[0273] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on analysis results and sentiment analysis results.
[0274] "Answer generation means" refers to the means for communicating the answer generated by the generative AI model to the user.
[0275] The "transmission means" refers to a communication means for transmitting the generated answer to the user's information processing device.
[0276] The present invention is a system that allows a user to input an inquiry using a chat-style interface of an information processing device, analyzes the inquiry, and generates an appropriate answer. This system includes an information processing server, a natural language processing engine, a sentiment analysis means, a generative AI model, and an answer transmission means as its main components.
[0277] User operations
[0278] First, a user starts up an information processing device (e.g., a smartphone) and navigates to the help section. Next, the user opens a chat box and enters a question about how to use the smartphone in text format. For example, the user might enter, "How do I set up Wi-Fi?"
[0279] Device behavior
[0280] The text data entered by the user's information processing device is acquired and sent to an information processing server via the Internet. This transmission is performed via an API.
[0281] Server Operation
[0282] The information processing server receives a user's query and temporarily stores the data. The server prepares the data for subsequent processing by sending it to a natural language processing engine (e.g., a general natural language processing library).
[0283] Natural Language Processing and Sentiment Analysis
[0284] A natural language processing engine analyzes the received text and understands the intent of the user's inquiry (e.g., "How to set up Wi-Fi"). Furthermore, a sentiment analysis means (e.g., a general sentiment analysis tool) analyzes emotions from the user's text and identifies emotions such as "troubled" or "angry."
[0285] Server Processing
[0286] The server receives the analysis results from the natural language processing engine and sentiment analysis tool and sends them to a generative AI model (e.g., a general generative AI tool). This transmission is also done via an API.
[0287] Answer generation behavior
[0288] The generative AI model generates an appropriate answer based on the analysis results it receives. In particular, it generates answers with a tone and content that takes into account the user's emotions based on the results of sentiment analysis. For example, if the user is in trouble, the generated answer will include more careful and detailed explanations.
[0289] Specific examples
[0290] For example, if a user types "How do I set up Wi-Fi?", analysis results indicate that the intent is to ask "How do I set up Wi-Fi?", and sentiment analysis identifies that the user is confused. The generative AI model generates an answer such as, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. If you need more help, we can help you."
[0291] Server response processing
[0292] The generated answer is returned to the information processing server, which then transmits it to the user's information processing device. Transmission is again performed via the internet through an API.
[0293] Terminal display
[0294] The user's information processing device displays the received response in a chat box. The user can view the displayed response and perform operations according to the instructions. In this way, the user can receive a quick and thoughtful response.
[0295] Prompt Sentence Examples
[0296] "A user asks, 'How do I set up Wi-Fi?' This user seems confused. Please generate a thoughtful and detailed answer based on this intent."
[0297] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0298] Step 1:
[0299] A user launches the official app for their information processing device and goes to the help section. They open the chat box and type a question about how to use their smartphone in text format. For example, they type "How do I set up Wi-Fi?" This input data is sent to the device.
[0300] Step 2:
[0301] The terminal receives the text entered by the user, converts the text data into JSON format, and sends the converted JSON data to the server via the API over the Internet (input: user text, output: JSON data sent to the server).
[0302] Step 3:
[0303] The server receives user queries, temporarily stores them in a database, and prepares the data to send to a natural language processing engine for analysis (input: JSON data, output: data ready for analysis).
[0304] Step 4:
[0305] The server sends the JSON data to a natural language processing engine. The natural language processing engine analyzes the query and identifies the user's intent. In this case, the intent is identified as "Tell me how to set up Wi-Fi." (Input: JSON data, Output: Analysis results with intent identified).
[0306] Step 5:
[0307] The server receives the analysis results and sends the data to the emotion analysis means, which analyzes emotions from the text data and identifies emotions such as "confusion" or "anger" (input: analysis results, output: emotion analysis results).
[0308] Step 6:
[0309] The server integrates the analysis results from the natural language processing engine and the sentiment analysis method. The server then sends the integrated results to the generative AI model. This transmission is also done via API (input: integrated analysis results, output: data sent to the generative AI model).
[0310] Step 7:
[0311] The generative AI model generates an appropriate answer based on the analysis results it receives. It takes into account the results of sentiment analysis to generate an answer with a tone and content that takes the user's emotions into consideration. For example, it generates an answer like, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help." (Input: Analysis results, Output: Generated answer).
[0312] Step 8:
[0313] The server receives the answer received from the generative AI model, converts the data into JSON format, and sends the converted data to the user's information processing device (input: generated answer, output: JSON data sent to the information processing device).
[0314] Step 9:
[0315] The user's information processing device parses the JSON-formatted answer data received from the server and displays the answer text in the chat box. The user can view the displayed answer and perform operations according to the instructions (input: JSON-formatted answer data, output: answer displayed in the chat box).
[0316] (Application example 2)
[0317] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0318] Conventional systems often fail to provide appropriate responses to users when they make inquiries because they do not take their emotions into consideration. Furthermore, because responses are not generated based on emotions, users' frustration and confusion are often left unresolved, leading to a decrease in satisfaction. Furthermore, this problem is particularly pronounced on online shopping sites, where responses that do not take users' emotions into consideration ultimately result in a decrease in sales and an increase in the burden on customer support.
[0319] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input an inquiry via a chat-style interface on the terminal, server means for receiving the inquiry content, means for analyzing the received inquiry content with a natural language processing engine, means for generating an answer based on the analysis results using a generative AI model, means for combining with an emotion engine that recognizes the user's emotions in order to adjust the generated answer, and means for transmitting the generated answer to the user's terminal. This makes it possible to quickly provide an appropriate answer that takes the user's emotions into consideration.
[0320] "Terminal" refers to an electronic device operated by a user, and specifically includes smartphones, tablets, personal computers, etc.
[0321] A "chat-style interface" is a format in which users communicate with each other by entering text.
[0322] A "server" is a computer system that processes and stores data on a network.
[0323] A "natural language processing engine" is a technology that analyzes input text data and understands its meaning and intent.
[0324] A "generative AI model" is an artificial intelligence model that generates appropriate answers based on analysis results.
[0325] An "emotion engine" is a technology that recognizes emotions from a user's text and adjusts responses based on those emotions.
[0326] "Means for generating an answer" refers to the process of using a generative AI model to create an appropriate answer to a user's inquiry.
[0327] The "means for transmitting the answer to the user's terminal" is a function for transmitting the generated answer to the user's terminal via a network.
[0328] This invention is a system that, when a user inputs an inquiry in chat format, analyzes the inquiry content using a natural language processing engine and an emotion engine, and generates an answer using a generative AI model. This system includes the user's terminal, a server, a natural language processing engine, a generative AI model, and an emotion engine as its main components.
[0329] System Program
[0330] In this system, users input inquiries using a chat-style interface on their device. The text entered by the user is sent to a server via the Internet. The server receives the inquiry, sends it to a natural language processing engine to analyze its meaning, and simultaneously recognizes the user's emotions using an emotion engine. Based on the results of these analyses, a generative AI model generates an appropriate response, which is then sent back to the user's device via the server.
[0331] Processing Description
[0332] The server receives the query text entered by the user and then analyzes it using a natural language processing engine, such as the open-source Natural Language Toolkit (NLTK) or Google's Cloud Natural Language API. The emotion engine then recognizes the emotion in the user's text. Examples of emotion engines that can be used include Microsoft's Azure Text Analytics and IBM's Watson Tone Analyzer.
[0333] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. Generative AI models such as OpenAI's GPT-3 and GPT-4 can be used. The generative AI model generates a response based on the analysis results, adjusting the tone and content to take the user's emotions into consideration. The server then sends the generated response to the user's device, which displays the response in a chat box.
[0334] Specific examples
[0335] For example, if a user types "My order hasn't arrived yet, what should I do?" into the chat box, this text is sent over the internet to a server. The server then sends this text to a natural language processing engine and an emotion engine to analyze intent and emotion. The analysis results determine that the intent of the inquiry is about "delayed delivery of my order" and that the user is "confused." The generative AI model uses this analysis to generate a specific, emotion-sensitive response, providing the following answer: "We're sorry for the concern. To track your order, please go to the 'Order History' page, select the order, and view the tracking information. Please let us know if we need further assistance."
[0336] Prompt Sentence Examples
[0337] User Question: "My order hasn't arrived yet, what should I do?"
[0338] User Emotion: Confused
[0339] Response: "We're sorry for your concern. To track your order, please go to your 'Order History' page and select the order to view tracking information. Please let us know if we need further assistance."
[0340] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0341] Step 1:
[0342] A user launches the official app and enters an inquiry into the chat-style interface. For example, the user might enter, "My order hasn't arrived yet. What should I do?" This input is temporarily stored on the device as text data.
[0343] Input: User query text
[0344] Output: Save to the terminal as text data
[0345] Step 2:
[0346] The device sends the text entered by the user to the server via the API, which is transmitted over the Internet.
[0347] Input: User query text
[0348] Output: Data sent to the server
[0349] Step 3:
[0350] The server receives a user's inquiry and temporarily stores the received data. It then sends the received inquiry to the natural language processing engine. Specifically, it sends the text data to the API of the natural language processing engine.
[0351] Input: Query text
[0352] Output: Data sent to the natural language processing engine
[0353] Step 4:
[0354] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the inquiry as being about a "delayed delivery of an order." The analysis results are then sent back to the server.
[0355] Input: Query text
[0356] Output: Intention information as a result of analysis
[0357] Step 5:
[0358] The server receives the intent analysis results from the natural language processing engine and sends them to the emotion engine. The emotion engine analyzes the emotion from the query text and identifies emotions such as "confusion." The analysis results are then returned to the server.
[0359] Input: Query text and intent analysis results
[0360] Output: Emotion analysis results
[0361] Step 6:
[0362] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API. The generative AI model generates an appropriate answer based on the analysis results.
[0363] Input: Results of natural language processing and sentiment analysis
[0364] Output: The generated answer
[0365] Step 7:
[0366] The generative AI model returns the generated content to the server, which then sends the generated answer to the user's device via the Internet.
[0367] Input: Generated answer
[0368] Output: Data sent to the user's device
[0369] Step 8:
[0370] The terminal displays the answer received from the server in the chat box, and the user can view the generated answer and perform operations according to the instructions.
[0371] Input: Generated answer from the server
[0372] Output: Displayed in the user's chat box
[0373] Through the above steps, the user can quickly receive an appropriate response that takes into consideration their feelings.
[0374] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0375] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0376] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0377] [Second embodiment]
[0378] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0379] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0380] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0381] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0382] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0383] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0384] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0385] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0386] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0387] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0388] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0389] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0390] This invention is a system that allows users to ask questions about how to use their smartphones in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. The system's main components are the user, the device, and a server.
[0391] System configuration
[0392] User operations
[0393] 1. The user launches the official app and goes to the help section.
[0394] 2. The user opens a chat box and types in a text message with a usage question.
[0395] Example: "How do I set up Wi-Fi?"
[0396] Device behavior
[0397] 1. The device receives the text entered by the user and sends it to the server.
[0398] The data is sent via an API.
[0399] Server Operation
[0400] 1. The server receives a query from a user.
[0401] 2. The server sends the received text to a natural language processing engine for analysis.
[0402] 3. The server receives the analysis results from the natural language processing engine and sends them to the generative AI model.
[0403] Natural Language Processing and AI Model Behavior
[0404] 1. The natural language processing engine analyzes the user's inquiry and understands its intent.
[0405] Example: Identifying the query intent as "How do I set up Wi-Fi?"
[0406] 2. The generative AI model generates an appropriate answer based on the analysis results.
[0407] For example, generate an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0408] Server response processing
[0409] 1. The server receives the generated answer and sends it back to the user's device.
[0410] Terminal display
[0411] 1. The device displays the received response in the chat box for the user to view.
[0412] 2. The user checks the answers and follows the instructions to set up the smartphone.
[0413] Specific examples
[0414] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0415] This system allows users to instantly obtain the information they need in a chat format, quickly resolving any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0416] The processing flow will be explained below.
[0417] Step 1:
[0418] A user launches the official app and navigates to the help section. The user opens the chat box and enters their inquiry in text format. For example, they might type, "How do I set up Wi-Fi?"
[0419] Step 2:
[0420] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0421] Step 3:
[0422] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0423] Step 4:
[0424] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0425] Step 5:
[0426] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0427] Step 6:
[0428] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, also via an API.
[0429] Step 7:
[0430] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0431] Step 8:
[0432] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0433] Step 9:
[0434] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0435] Example 1
[0436] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0437] Users of modern smartphones and other electronic devices often have questions about how to operate their devices. However, traditional methods offer limited options for getting fast, accurate answers when users encounter problems. Manually searching for information in instruction manuals or online help takes time and effort. Therefore, a new system is needed to help users solve problems efficiently.
[0438] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0439] In this invention, the server includes means for a user to input an inquiry via a chat-style interface of the electronic device, means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, and means for transmitting the generated answer to the user's electronic device. This allows the user to quickly receive an accurate answer in chat format and efficiently resolve questions about how to operate the device.
[0440] "User" refers to a user who operates an electronic device.
[0441] "Electronic devices" refers to computing devices such as smartphones, tablets, and personal computers.
[0442] "Chat-style interface" refers to a user interface that allows users to enter text messages and communicate information interactively.
[0443] An "inquiry" refers to a user inputting a question or doubt about how to operate an electronic device.
[0444] A "server" refers to a computer system that receives queries from users, processes them, and generates and transmits responses.
[0445] A "natural language processing engine" refers to software that analyzes text input from a user and understands their intent.
[0446] A "generative AI model" refers to an artificial intelligence model that automatically generates appropriate answers based on the results analyzed by a natural language processing engine.
[0447] "Answer" refers to information generated by a generative AI model that includes answers or instructions to a user's inquiry.
[0448] "Instructions" refers to documents or sources of information that contain detailed instructions on the operation and functionality of an electronic device.
[0449] This invention is a system that allows users to ask questions about how to use electronic devices in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. This system mainly consists of a user, a terminal, and a server.
[0450] System configuration
[0451] User operations
[0452] A user launches the official app and navigates to the help section. The user opens the chat box and enters a usage question in text format. For example, if the user types "How do I set up Wi-Fi?", the following process occurs:
[0453] Device behavior
[0454] The device takes the text entered by the user and sends it to the server, where it is packaged in JSON format and protected by SSL / TLS encryption via an API.
[0455] Server Operation
[0456] The server receives a query from a user. The server passes the received request to a processing program via a web server such as Apache or Nginx. For example, a Flask application written in Python handles this request. The server then sends the received text to a natural language processing engine. This analysis is performed using the Google Cloud Natural Language API or an NLP library (e.g., spaCy, NLTK). After the analysis is complete, the server sends the results to a generative AI model (e.g., OpenAI's GPT-3 or GPT-4).
[0457] Natural Language Processing and AI Model Behavior
[0458] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, in response to the inquiry "How do I set up Wi-Fi?", the keyword "How to set up Wi-Fi" is extracted. Next, the generative AI model generates an appropriate answer based on the analysis results. For example, it generates an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password."
[0459] Server response processing
[0460] The server receives the generated answer and sends it back to the user's device, where the data is again sent via an HTTP response.
[0461] Terminal display
[0462] The device displays the received answer in a chat box for the user to view. The user can then check the answer and follow the instructions to set up their smartphone.
[0463] Specific examples
[0464] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0465] Example prompt sentence:
[0466] The user types "How do I set up Wi-Fi?" into the chat box. Generate an appropriate response.
[0467] This system allows users to instantly obtain the information they need in a chat format, enabling them to quickly resolve any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0468] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0469] Step 1:
[0470] A user launches the official app and navigates to the help section. The user unlocks the smartphone and taps the app icon. The app launches and displays the home screen. The user clicks the help or support tab within the app.
[0471] Input: Smartphone operation
[0472] Output: Help section on screen
[0473] Step 2:
[0474] The user opens the chat box, types a usage question in text format, for example, "How do I set up Wi-Fi?", and taps the send button.
[0475] Input: User question text
[0476] Output: The terminal displays the user's input text.
[0477] Step 3:
[0478] The device retrieves the user's input text and sends it to the server. The device generates an HTTP POST request, packages the user's input text in JSON format, and sends it to the SSL / TLS encrypted API endpoint.
[0479] Input: User-entered text
[0480] Output: Encrypted HTTP POST request
[0481] Step 4:
[0482] The server receives a query from a user. The server receives the request through a web server (e.g., Nginx, Apache) and passes the request to its internal processing program (e.g., a Python Flask application).
[0483] Input: Encrypted HTTP POST request
[0484] Output: The request data passed to the processing program
[0485] Step 5:
[0486] The server sends the received text to a natural language processing engine, which extracts the text from the request data, sends it to the natural language processing engine (e.g., Google Cloud Natural Language API), and makes an API call for analysis.
[0487] Input: The text portion of the request data
[0488] Output: API request to the natural language processing engine
[0489] Step 6:
[0490] A natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it analyzes the text "How do I set up Wi-Fi?" and identifies the intent as "How do I set up Wi-Fi?"
[0491] Input: The text to be parsed
[0492] Output: Analysis results (intent, keywords, etc.)
[0493] Step 7:
[0494] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model. The server then generates an API request to pass the analysis results to the generative AI model (e.g., OpenAI GPT-3). This request includes a prompt.
[0495] Input: Analysis results of the natural language processing engine
[0496] Output: API request to the generative AI model
[0497] Step 8:
[0498] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0499] Input: prompt sentence for generative AI model
[0500] Output: Generated answer text
[0501] Step 9:
[0502] The server receives the generated answer and returns it to the user's device. The server packages the generated answer in JSON format and sends it to the user's device as an HTTP response.
[0503] Input: Generated answer text
[0504] Output: Answer sent to the terminal as an HTTP response
[0505] Step 10:
[0506] The device displays the received answer in the chat box for the user to view. The device deserializes the received JSON data and displays the answer text in the chat box. The user checks the answer and follows the instructions to configure their smartphone.
[0507] Input: JSON data of the HTTP response
[0508] Output: Answer text displayed in the chat box
[0509] (Application example 1)
[0510] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0511] In the past, it was difficult for workers to receive prompt and appropriate support when operating or troubleshooting robots used in factories. This often resulted in reduced work efficiency and lost productivity. In such situations, a system was needed that would enable workers to quickly understand how to use robots and immediately provide reliable information to solve problems.
[0512] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0513] In this invention, the server includes means for a user to input an inquiry via a chat-style interface on a terminal, server means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, means for sending the generated answer to the user's terminal, and means for analyzing the inquiry specifically regarding the operation of a robot in a factory and generating appropriate instructions, thereby enabling the user to obtain detailed information on how to operate and troubleshoot the robot in a factory in real time.
[0514] "User" refers to any individual or entity that uses the System to make an inquiry.
[0515] "Terminal" refers to a computer or device for entering queries and displaying responses.
[0516] A "chat-style interface" refers to a user interface in which a user inputs a text-based inquiry and receives a response from the system.
[0517] "Server" refers to a central processing unit for receiving queries from users, analyzing them, and generating responses.
[0518] A "natural language processing engine" refers to a machine learning or rule-based system that analyzes user queries and understands their intent.
[0519] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on the analysis results of a natural language processing engine.
[0520] "Inquiries regarding the operation of robots in factories" refers to questions regarding how to operate and troubleshoot robots used in factories.
[0521] An "answer" refers to the information or instructions generated by a generative AI model in response to a user's inquiry.
[0522] The system for realizing this application example operates in the following procedure.
[0523] First, the user accesses a chat-style interface using a device, such as a PC, tablet, or smartphone, and enters text-based inquiries about the robot's operation.
[0524] The device takes the entered text data and sends it over the Internet to a server using a standard HTTP request.
[0525] The server processes the received query. First, the query is analyzed by a natural language processing engine. This engine uses natural language processing libraries such as spaCy or NLTK. The analysis identifies the intent of the query and important keywords.
[0526] The server then uses a generative AI model, such as GPT-4, to generate an answer based on the analysis results. The AI model generates an appropriate and detailed answer based on the analyzed intent and keywords.
[0527] The generated answers are then sent back to the terminal via the server and displayed to the user, allowing the user to obtain appropriate operating instructions and troubleshooting methods in real time.
[0528] For example, if a user types, "Please tell me how to calibrate my robot," this query is analyzed by a natural language processing engine, and the generative AI model generates the answer, "Enter calibration mode and adjust each sensor according to the manual."
[0529] In this way, the system can provide users with immediate and reliable information regarding their robot operation questions.
[0530] Example prompt sentence:
[0531] User: "How do I calibrate my robot?"
[0532] Answer: "Enter calibration mode and adjust each sensor according to the manual."
[0533] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0534] Step 1:
[0535] The user uses a terminal to access a chat-style interface and enters a text inquiry about the operation of the robot, for example, "Tell me how to calibrate the robot."
[0536] Input: User's inquiry (text format)
[0537] Output: Query text data
[0538] Step 2:
[0539] The device receives the text data entered by the user and sends it to the server as an API request, using the standard HTTP protocol.
[0540] Input: User query text data
[0541] Output: HTTP request to the server
[0542] Step 3:
[0543] The server sends the received query to a natural language processing engine for analysis, which analyzes the query content and identifies its intent and important keywords.
[0544] Input: Query text data as an HTTP request
[0545] Output: Analysis results (intent and keywords)
[0546] What it does: Analyzes text using natural language processing libraries such as spaCy and NLTK
[0547] Step 4:
[0548] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, which then generates an appropriate answer based on the analysis results.
[0549] Input: Analysis results (intent and keywords)
[0550] Output: Generated answer text
[0551] Specific operation: Generate text using generative AI models such as GPT-4
[0552] Step 5:
[0553] The server retrieves the generated answer and sends it to the user's device using an HTTP response.
[0554] Input: Generated answer text
[0555] Output: HTTP response to the user's device
[0556] Step 6:
[0557] The terminal displays the received response on the chat interface so that it can be seen by the user, who can then confirm it and perform the necessary operations.
[0558] Input: HTTP response containing the answer text from the server
[0559] Output: Answer text displayed in the chat interface
[0560] Specific operation: The received response text is displayed on the screen. For example, "Enter calibration mode and adjust each sensor according to the manual."
[0561] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0562] This invention combines a system in which a user inputs an inquiry about how to use a smartphone in chat format, analyzes the inquiry content with a natural language processing engine, and generates an answer with a generative AI model, with an emotion engine that recognizes the user's emotions.This system includes the user, device, server, natural language processing engine, generative AI model, and emotion engine as its main components.
[0563] System configuration
[0564] User operations
[0565] 1. The user launches the official app and goes to the help section.
[0566] 2. The user opens a chat box and types in a text message with a usage question.
[0567] Example: "How do I set up Wi-Fi?"
[0568] Device behavior
[0569] 1. The device receives the text entered by the user and sends it to the server via the internet via an API.
[0570] Server Operation
[0571] 1. The server receives a query from the user. The server temporarily stores the received data and prepares it for the next process.
[0572] 2. The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0573] How Natural Language Processing and Sentiment Analysis Work
[0574] 1. The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the user's intent as "How do I set up Wi-Fi?"
[0575] 2. The emotion engine analyzes the user's query text to determine their emotions, for example, identifying that the user is feeling "troubled" or "angry."
[0576] Server Processing
[0577] 1. The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API.
[0578] Answer generation behavior
[0579] 1. The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and content of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0580] Server response processing
[0581] 1. The server receives the generated answer and sends it back to the user's device, via the internet through an API.
[0582] Terminal display
[0583] 1. The terminal displays the received answer in the chat box. The user can view the displayed answer and follow the instructions to perform the operation.
[0584] Specific examples
[0585] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app and the text includes confusion or irritation, the device sends the query to a server. The server receives the text, sends it to a natural language processing engine to analyze the intent, and uses an emotion engine to analyze the user's emotions. The analysis results identify that the query's intent is "How do I set up Wi-Fi?" and that the user is having trouble. Based on this analysis, a generative AI model generates a specific, emotion-sensitive answer and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help," and can follow the instructions to perform the operation.
[0586] This system not only allows users to instantly obtain the information they need in a chat format, but also improves their understanding of how to use the device and the smooth operation of their smartphone by receiving responses that take into consideration their emotions at the time.In addition, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0587] The processing flow will be explained below.
[0588] Step 1:
[0589] A user launches the official app and navigates to the help section. The user opens the chat box and types a usage question in text format. For example, they type "How do I set up Wi-Fi?"
[0590] Step 2:
[0591] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0592] Step 3:
[0593] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0594] Step 4:
[0595] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0596] Step 5:
[0597] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0598] Step 6:
[0599] The server receives the analysis results from the natural language processing engine, and at the same time, sends the user's query text to the emotion engine for emotion analysis.
[0600] Step 7:
[0601] The emotion engine analyzes the user's text to identify their emotional state, for example detecting emotions such as "confused" or "frustrated."
[0602] Step 8:
[0603] The server integrates the analysis results from the natural language processing engine and the emotion engine and sends them to the generative AI model, also via an API.
[0604] Step 9:
[0605] The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and details of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0606] Step 10:
[0607] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0608] Step 11:
[0609] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0610] Step 12:
[0611] The user then performs the smartphone configuration based on the answers, for example opening the Settings app, tapping the 'Wi-Fi' tab, selecting an available network and entering the password.
[0612] Through this series of processes, users not only receive information but also receive instructions that take into consideration their emotions and state, thereby reducing stress and improving the user experience.
[0613] Example 2
[0614] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0615] With conventional interfaces, even when a user asks a question about how to operate a smartphone, the answer does not take into consideration the user's emotions at the time. As a result, even when the user is confused or angry, only mechanical answers are provided, which does not improve the user experience. There was also a need for a system that could provide quick and appropriate answers.
[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0617] In this invention, the server includes an information processing server means for receiving the inquiry content, a means for analyzing the received inquiry content using a natural language processing engine, a sentiment analysis means for analyzing sentiment from text entered by the user, a means for generating an answer based on the analysis result and the sentiment analysis result using a generative AI model, and a means for transmitting the generated answer to the user's information processing device, thereby making it possible to provide an appropriate and detailed answer that takes into consideration the user's sentiment.
[0618] The term "user" refers to a person who makes an inquiry using an information processing device.
[0619] An "information processing device" is a device that a user can operate and input an inquiry into, and includes a smartphone, a tablet, and the like.
[0620] "Chat-style interface" refers to a screen and software that allows users to enter messages and interact in text format.
[0621] An "information processing server" refers to a computer system that receives inquiries from users and performs processing to analyze and generate answers.
[0622] "Natural language processing engine" refers to software and algorithms that analyze a user's text input to understand their intent.
[0623] "Sentiment analyzer" refers to software and algorithms for analyzing emotions from a user's input text and identifying their emotional state.
[0624] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on analysis results and sentiment analysis results.
[0625] "Answer generation means" refers to the means for communicating the answer generated by the generative AI model to the user.
[0626] The "transmission means" refers to a communication means for transmitting the generated answer to the user's information processing device.
[0627] The present invention is a system that allows a user to input an inquiry using a chat-style interface of an information processing device, analyzes the inquiry, and generates an appropriate answer. This system includes an information processing server, a natural language processing engine, a sentiment analysis means, a generative AI model, and an answer transmission means as its main components.
[0628] User operations
[0629] First, a user starts up an information processing device (e.g., a smartphone) and navigates to the help section. Next, the user opens a chat box and enters a question about how to use the smartphone in text format. For example, the user might enter, "How do I set up Wi-Fi?"
[0630] Device behavior
[0631] The text data entered by the user's information processing device is acquired and sent to an information processing server via the Internet. This transmission is performed via an API.
[0632] Server Operation
[0633] The information processing server receives a user's query and temporarily stores the data. The server prepares the data for subsequent processing by sending it to a natural language processing engine (e.g., a general natural language processing library).
[0634] Natural Language Processing and Sentiment Analysis
[0635] A natural language processing engine analyzes the received text and understands the intent of the user's inquiry (e.g., "How to set up Wi-Fi"). Furthermore, a sentiment analysis means (e.g., a general sentiment analysis tool) analyzes emotions from the user's text and identifies emotions such as "troubled" or "angry."
[0636] Server Processing
[0637] The server receives the analysis results from the natural language processing engine and sentiment analysis tool and sends them to a generative AI model (e.g., a general generative AI tool). This transmission is also done via an API.
[0638] Answer generation behavior
[0639] The generative AI model generates an appropriate answer based on the analysis results it receives. In particular, it generates answers with a tone and content that takes into account the user's emotions based on the results of sentiment analysis. For example, if the user is in trouble, the generated answer will include more careful and detailed explanations.
[0640] Specific examples
[0641] For example, if a user types "How do I set up Wi-Fi?", analysis results indicate that the intent is to ask "How do I set up Wi-Fi?", and sentiment analysis identifies that the user is confused. The generative AI model generates an answer such as, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. If you need more help, we can help you."
[0642] Server response processing
[0643] The generated answer is returned to the information processing server, which then transmits it to the user's information processing device. Transmission is again performed via the internet through an API.
[0644] Terminal display
[0645] The user's information processing device displays the received response in a chat box. The user can view the displayed response and perform operations according to the instructions. In this way, the user can receive a quick and thoughtful response.
[0646] Prompt Sentence Examples
[0647] "A user asks, 'How do I set up Wi-Fi?' This user seems confused. Please generate a thoughtful and detailed answer based on this intent."
[0648] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0649] Step 1:
[0650] A user launches the official app for their information processing device and goes to the help section. They open the chat box and type a question about how to use their smartphone in text format. For example, they type "How do I set up Wi-Fi?" This input data is sent to the device.
[0651] Step 2:
[0652] The terminal receives the text entered by the user, converts the text data into JSON format, and sends the converted JSON data to the server via the API over the Internet (input: user text, output: JSON data sent to the server).
[0653] Step 3:
[0654] The server receives user queries, temporarily stores them in a database, and prepares the data to send to a natural language processing engine for analysis (input: JSON data, output: data ready for analysis).
[0655] Step 4:
[0656] The server sends the JSON data to a natural language processing engine. The natural language processing engine analyzes the query and identifies the user's intent. In this case, the intent is identified as "Tell me how to set up Wi-Fi." (Input: JSON data, Output: Analysis results with intent identified).
[0657] Step 5:
[0658] The server receives the analysis results and sends the data to the emotion analysis means, which analyzes emotions from the text data and identifies emotions such as "confusion" or "anger" (input: analysis results, output: emotion analysis results).
[0659] Step 6:
[0660] The server integrates the analysis results from the natural language processing engine and the sentiment analysis method. The server then sends the integrated results to the generative AI model. This transmission is also done via API (input: integrated analysis results, output: data sent to the generative AI model).
[0661] Step 7:
[0662] The generative AI model generates an appropriate answer based on the analysis results it receives. It takes into account the results of sentiment analysis to generate an answer with a tone and content that takes the user's emotions into consideration. For example, it generates an answer like, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help." (Input: Analysis results, Output: Generated answer).
[0663] Step 8:
[0664] The server receives the answer received from the generative AI model, converts the data into JSON format, and sends the converted data to the user's information processing device (input: generated answer, output: JSON data sent to the information processing device).
[0665] Step 9:
[0666] The user's information processing device parses the JSON-formatted answer data received from the server and displays the answer text in the chat box. The user can view the displayed answer and perform operations according to the instructions (input: JSON-formatted answer data, output: answer displayed in the chat box).
[0667] (Application example 2)
[0668] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0669] Conventional systems often fail to provide appropriate responses to users when they make inquiries because they do not take their emotions into consideration. Furthermore, because responses are not generated based on emotions, users' frustration and confusion are often left unresolved, leading to a decrease in satisfaction. Furthermore, this problem is particularly pronounced on online shopping sites, where responses that do not take users' emotions into consideration ultimately result in a decrease in sales and an increase in the burden on customer support.
[0670] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input an inquiry via a chat-style interface on the terminal, server means for receiving the inquiry content, means for analyzing the received inquiry content with a natural language processing engine, means for generating an answer based on the analysis results using a generative AI model, means for combining with an emotion engine that recognizes the user's emotions in order to adjust the generated answer, and means for transmitting the generated answer to the user's terminal. This makes it possible to quickly provide an appropriate answer that takes the user's emotions into consideration.
[0671] "Terminal" refers to an electronic device operated by a user, and specifically includes smartphones, tablets, personal computers, etc.
[0672] A "chat-style interface" is a format in which users communicate with each other by entering text.
[0673] A "server" is a computer system that processes and stores data on a network.
[0674] A "natural language processing engine" is a technology that analyzes input text data and understands its meaning and intent.
[0675] A "generative AI model" is an artificial intelligence model that generates appropriate answers based on analysis results.
[0676] An "emotion engine" is a technology that recognizes emotions from a user's text and adjusts responses based on those emotions.
[0677] "Means for generating an answer" refers to the process of using a generative AI model to create an appropriate answer to a user's inquiry.
[0678] The "means for transmitting the answer to the user's terminal" is a function for transmitting the generated answer to the user's terminal via a network.
[0679] This invention is a system that, when a user inputs an inquiry in chat format, analyzes the inquiry content using a natural language processing engine and an emotion engine, and generates an answer using a generative AI model. This system includes the user's terminal, a server, a natural language processing engine, a generative AI model, and an emotion engine as its main components.
[0680] System Program
[0681] In this system, users input inquiries using a chat-style interface on their device. The text entered by the user is sent to a server via the Internet. The server receives the inquiry, sends it to a natural language processing engine to analyze its meaning, and simultaneously recognizes the user's emotions using an emotion engine. Based on the results of these analyses, a generative AI model generates an appropriate response, which is then sent back to the user's device via the server.
[0682] Processing Description
[0683] The server receives the query text entered by the user and then analyzes it using a natural language processing engine, such as the open-source Natural Language Toolkit (NLTK) or Google's Cloud Natural Language API. The emotion engine then recognizes the emotion in the user's text. Examples of emotion engines that can be used include Microsoft's Azure Text Analytics and IBM's Watson Tone Analyzer.
[0684] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. Generative AI models such as OpenAI's GPT-3 and GPT-4 can be used. The generative AI model generates a response based on the analysis results, adjusting the tone and content to take the user's emotions into consideration. The server then sends the generated response to the user's device, which displays the response in a chat box.
[0685] Specific examples
[0686] For example, if a user types "My order hasn't arrived yet, what should I do?" into the chat box, this text is sent over the internet to a server. The server then sends this text to a natural language processing engine and an emotion engine to analyze intent and emotion. The analysis results determine that the intent of the inquiry is about "delayed delivery of my order" and that the user is "confused." The generative AI model uses this analysis to generate a specific, emotion-sensitive response, providing the following answer: "We're sorry for the concern. To track your order, please go to the 'Order History' page, select the order, and view the tracking information. Please let us know if we need further assistance."
[0687] Prompt Sentence Examples
[0688] User Question: "My order hasn't arrived yet, what should I do?"
[0689] User Emotion: Confused
[0690] Response: "We're sorry for your concern. To track your order, please go to your 'Order History' page and select the order to view tracking information. Please let us know if we need further assistance."
[0691] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0692] Step 1:
[0693] A user launches the official app and enters an inquiry into the chat-style interface. For example, the user might enter, "My order hasn't arrived yet. What should I do?" This input is temporarily stored on the device as text data.
[0694] Input: User query text
[0695] Output: Save to the terminal as text data
[0696] Step 2:
[0697] The device sends the text entered by the user to the server via the API, which is transmitted over the Internet.
[0698] Input: User query text
[0699] Output: Data sent to the server
[0700] Step 3:
[0701] The server receives a user's inquiry and temporarily stores the received data. It then sends the received inquiry to the natural language processing engine. Specifically, it sends the text data to the API of the natural language processing engine.
[0702] Input: Query text
[0703] Output: Data sent to the natural language processing engine
[0704] Step 4:
[0705] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the inquiry as being about a "delayed delivery of an order." The analysis results are then sent back to the server.
[0706] Input: Query text
[0707] Output: Intention information as a result of analysis
[0708] Step 5:
[0709] The server receives the intent analysis results from the natural language processing engine and sends them to the emotion engine. The emotion engine analyzes the emotion from the query text and identifies emotions such as "confusion." The analysis results are then returned to the server.
[0710] Input: Query text and intent analysis results
[0711] Output: Emotion analysis results
[0712] Step 6:
[0713] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API. The generative AI model generates an appropriate answer based on the analysis results.
[0714] Input: Results of natural language processing and sentiment analysis
[0715] Output: The generated answer
[0716] Step 7:
[0717] The generative AI model returns the generated content to the server, which then sends the generated answer to the user's device via the Internet.
[0718] Input: Generated answer
[0719] Output: Data sent to the user's device
[0720] Step 8:
[0721] The terminal displays the answer received from the server in the chat box, and the user can view the generated answer and perform operations according to the instructions.
[0722] Input: Generated answer from the server
[0723] Output: Displayed in the user's chat box
[0724] Through the above steps, the user can quickly receive an appropriate response that takes into consideration their feelings.
[0725] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0726] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0727] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0728] [Third embodiment]
[0729] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0730] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0731] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0732] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0733] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0734] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0735] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0736] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0737] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0738] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0739] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0740] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0741] This invention is a system that allows users to ask questions about how to use their smartphones in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. The system's main components are the user, the device, and a server.
[0742] System configuration
[0743] User operations
[0744] 1. The user launches the official app and goes to the help section.
[0745] 2. The user opens a chat box and types in a text message with a usage question.
[0746] Example: "How do I set up Wi-Fi?"
[0747] Device behavior
[0748] 1. The device receives the text entered by the user and sends it to the server.
[0749] The data is sent via an API.
[0750] Server Operation
[0751] 1. The server receives a query from a user.
[0752] 2. The server sends the received text to a natural language processing engine for analysis.
[0753] 3. The server receives the analysis results from the natural language processing engine and sends them to the generative AI model.
[0754] Natural Language Processing and AI Model Behavior
[0755] 1. The natural language processing engine analyzes the user's inquiry and understands its intent.
[0756] Example: Identifying the query intent as "How do I set up Wi-Fi?"
[0757] 2. The generative AI model generates an appropriate answer based on the analysis results.
[0758] For example, generate an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0759] Server response processing
[0760] 1. The server receives the generated answer and sends it back to the user's device.
[0761] Terminal display
[0762] 1. The device displays the received response in the chat box for the user to view.
[0763] 2. The user checks the answers and follows the instructions to set up the smartphone.
[0764] Specific examples
[0765] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0766] This system allows users to instantly obtain the information they need in a chat format, quickly resolving any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0767] The processing flow will be explained below.
[0768] Step 1:
[0769] A user launches the official app and navigates to the help section. The user opens the chat box and enters their inquiry in text format. For example, they might type, "How do I set up Wi-Fi?"
[0770] Step 2:
[0771] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0772] Step 3:
[0773] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0774] Step 4:
[0775] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0776] Step 5:
[0777] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0778] Step 6:
[0779] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, also via an API.
[0780] Step 7:
[0781] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0782] Step 8:
[0783] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0784] Step 9:
[0785] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0786] Example 1
[0787] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0788] Users of modern smartphones and other electronic devices often have questions about how to operate their devices. However, traditional methods offer limited options for getting fast, accurate answers when users encounter problems. Manually searching for information in instruction manuals or online help takes time and effort. Therefore, a new system is needed to help users solve problems efficiently.
[0789] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0790] In this invention, the server includes means for a user to input an inquiry via a chat-style interface of the electronic device, means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, and means for transmitting the generated answer to the user's electronic device. This allows the user to quickly receive an accurate answer in chat format and efficiently resolve questions about how to operate the device.
[0791] "User" refers to a user who operates an electronic device.
[0792] "Electronic devices" refers to computing devices such as smartphones, tablets, and personal computers.
[0793] "Chat-style interface" refers to a user interface that allows users to enter text messages and communicate information interactively.
[0794] An "inquiry" refers to a user inputting a question or doubt about how to operate an electronic device.
[0795] A "server" refers to a computer system that receives queries from users, processes them, and generates and transmits responses.
[0796] A "natural language processing engine" refers to software that analyzes text input from a user and understands their intent.
[0797] A "generative AI model" refers to an artificial intelligence model that automatically generates appropriate answers based on the results analyzed by a natural language processing engine.
[0798] "Answer" refers to information generated by a generative AI model that includes answers or instructions to a user's inquiry.
[0799] "Instructions" refers to documents or sources of information that contain detailed instructions on the operation and functionality of an electronic device.
[0800] This invention is a system that allows users to ask questions about how to use electronic devices in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. This system mainly consists of a user, a terminal, and a server.
[0801] System configuration
[0802] User operations
[0803] A user launches the official app and navigates to the help section. The user opens the chat box and enters a usage question in text format. For example, if the user types "How do I set up Wi-Fi?", the following process occurs:
[0804] Device behavior
[0805] The device takes the text entered by the user and sends it to the server, where it is packaged in JSON format and protected by SSL / TLS encryption via an API.
[0806] Server Operation
[0807] The server receives a query from a user. The server passes the received request to a processing program via a web server such as Apache or Nginx. For example, a Flask application written in Python handles this request. The server then sends the received text to a natural language processing engine. This analysis is performed using the Google Cloud Natural Language API or an NLP library (e.g., spaCy, NLTK). After the analysis is complete, the server sends the results to a generative AI model (e.g., OpenAI's GPT-3 or GPT-4).
[0808] Natural Language Processing and AI Model Behavior
[0809] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, in response to the inquiry "How do I set up Wi-Fi?", the keyword "How to set up Wi-Fi" is extracted. Next, the generative AI model generates an appropriate answer based on the analysis results. For example, it generates an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password."
[0810] Server response processing
[0811] The server receives the generated answer and sends it back to the user's device, where the data is again sent via an HTTP response.
[0812] Terminal display
[0813] The device displays the received answer in a chat box for the user to view. The user can then check the answer and follow the instructions to set up their smartphone.
[0814] Specific examples
[0815] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[0816] Example prompt sentence:
[0817] The user types "How do I set up Wi-Fi?" into the chat box. Generate an appropriate response.
[0818] This system allows users to instantly obtain the information they need in a chat format, enabling them to quickly resolve any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0819] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0820] Step 1:
[0821] A user launches the official app and navigates to the help section. The user unlocks the smartphone and taps the app icon. The app launches and displays the home screen. The user clicks the help or support tab within the app.
[0822] Input: Smartphone operation
[0823] Output: Help section on screen
[0824] Step 2:
[0825] The user opens the chat box, types a usage question in text format, for example, "How do I set up Wi-Fi?", and taps the send button.
[0826] Input: User question text
[0827] Output: The terminal displays the user's input text.
[0828] Step 3:
[0829] The device retrieves the user's input text and sends it to the server. The device generates an HTTP POST request, packages the user's input text in JSON format, and sends it to the SSL / TLS encrypted API endpoint.
[0830] Input: User-entered text
[0831] Output: Encrypted HTTP POST request
[0832] Step 4:
[0833] The server receives a query from a user. The server receives the request through a web server (e.g., Nginx, Apache) and passes the request to its internal processing program (e.g., a Python Flask application).
[0834] Input: Encrypted HTTP POST request
[0835] Output: The request data passed to the processing program
[0836] Step 5:
[0837] The server sends the received text to a natural language processing engine, which extracts the text from the request data, sends it to the natural language processing engine (e.g., Google Cloud Natural Language API), and makes an API call for analysis.
[0838] Input: The text portion of the request data
[0839] Output: API request to the natural language processing engine
[0840] Step 6:
[0841] A natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it analyzes the text "How do I set up Wi-Fi?" and identifies the intent as "How do I set up Wi-Fi?"
[0842] Input: The text to be parsed
[0843] Output: Analysis results (intent, keywords, etc.)
[0844] Step 7:
[0845] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model. The server then generates an API request to pass the analysis results to the generative AI model (e.g., OpenAI GPT-3). This request includes a prompt.
[0846] Input: Analysis results of the natural language processing engine
[0847] Output: API request to the generative AI model
[0848] Step 8:
[0849] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[0850] Input: prompt sentence for generative AI model
[0851] Output: Generated answer text
[0852] Step 9:
[0853] The server receives the generated answer and returns it to the user's device. The server packages the generated answer in JSON format and sends it to the user's device as an HTTP response.
[0854] Input: Generated answer text
[0855] Output: Answer sent to the terminal as an HTTP response
[0856] Step 10:
[0857] The device displays the received answer in the chat box for the user to view. The device deserializes the received JSON data and displays the answer text in the chat box. The user checks the answer and follows the instructions to configure their smartphone.
[0858] Input: JSON data of the HTTP response
[0859] Output: Answer text displayed in the chat box
[0860] (Application example 1)
[0861] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0862] In the past, it was difficult for workers to receive prompt and appropriate support when operating or troubleshooting robots used in factories. This often resulted in reduced work efficiency and lost productivity. In such situations, a system was needed that would enable workers to quickly understand how to use robots and immediately provide reliable information to solve problems.
[0863] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0864] In this invention, the server includes means for a user to input an inquiry via a chat-style interface on a terminal, server means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, means for sending the generated answer to the user's terminal, and means for analyzing the inquiry specifically regarding the operation of a robot in a factory and generating appropriate instructions, thereby enabling the user to obtain detailed information on how to operate and troubleshoot the robot in a factory in real time.
[0865] "User" refers to any individual or entity that uses the System to make an inquiry.
[0866] "Terminal" refers to a computer or device for entering queries and displaying responses.
[0867] A "chat-style interface" refers to a user interface in which a user inputs a text-based inquiry and receives a response from the system.
[0868] "Server" refers to a central processing unit for receiving queries from users, analyzing them, and generating responses.
[0869] A "natural language processing engine" refers to a machine learning or rule-based system that analyzes user queries and understands their intent.
[0870] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on the analysis results of a natural language processing engine.
[0871] "Inquiries regarding the operation of robots in factories" refers to questions regarding how to operate and troubleshoot robots used in factories.
[0872] An "answer" refers to the information or instructions generated by a generative AI model in response to a user's inquiry.
[0873] The system for realizing this application example operates in the following procedure.
[0874] First, the user accesses a chat-style interface using a device, such as a PC, tablet, or smartphone, and enters text-based inquiries about the robot's operation.
[0875] The device takes the entered text data and sends it over the Internet to a server using a standard HTTP request.
[0876] The server processes the received query. First, the query is analyzed by a natural language processing engine. This engine uses natural language processing libraries such as spaCy or NLTK. The analysis identifies the intent of the query and important keywords.
[0877] The server then uses a generative AI model, such as GPT-4, to generate an answer based on the analysis results. The AI model generates an appropriate and detailed answer based on the analyzed intent and keywords.
[0878] The generated answers are then sent back to the terminal via the server and displayed to the user, allowing the user to obtain appropriate operating instructions and troubleshooting methods in real time.
[0879] For example, if a user types, "Please tell me how to calibrate my robot," this query is analyzed by a natural language processing engine, and the generative AI model generates the answer, "Enter calibration mode and adjust each sensor according to the manual."
[0880] In this way, the system can provide users with immediate and reliable information regarding their robot operation questions.
[0881] Example prompt sentence:
[0882] User: "How do I calibrate my robot?"
[0883] Answer: "Enter calibration mode and adjust each sensor according to the manual."
[0884] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0885] Step 1:
[0886] The user uses a terminal to access a chat-style interface and enters a text inquiry about the operation of the robot, for example, "Tell me how to calibrate the robot."
[0887] Input: User's inquiry (text format)
[0888] Output: Query text data
[0889] Step 2:
[0890] The device receives the text data entered by the user and sends it to the server as an API request, using the standard HTTP protocol.
[0891] Input: User query text data
[0892] Output: HTTP request to the server
[0893] Step 3:
[0894] The server sends the received query to a natural language processing engine for analysis, which analyzes the query content and identifies its intent and important keywords.
[0895] Input: Query text data as an HTTP request
[0896] Output: Analysis results (intent and keywords)
[0897] What it does: Analyzes text using natural language processing libraries such as spaCy and NLTK
[0898] Step 4:
[0899] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, which then generates an appropriate answer based on the analysis results.
[0900] Input: Analysis results (intent and keywords)
[0901] Output: Generated answer text
[0902] Specific operation: Generate text using generative AI models such as GPT-4
[0903] Step 5:
[0904] The server retrieves the generated answer and sends it to the user's device using an HTTP response.
[0905] Input: Generated answer text
[0906] Output: HTTP response to the user's device
[0907] Step 6:
[0908] The terminal displays the received response on the chat interface so that it can be seen by the user, who can then confirm it and perform the necessary operations.
[0909] Input: HTTP response containing the answer text from the server
[0910] Output: Answer text displayed in the chat interface
[0911] Specific operation: The received response text is displayed on the screen. For example, "Enter calibration mode and adjust each sensor according to the manual."
[0912] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0913] This invention combines a system in which a user inputs an inquiry about how to use a smartphone in chat format, analyzes the inquiry content with a natural language processing engine, and generates an answer with a generative AI model, with an emotion engine that recognizes the user's emotions.This system includes the user, device, server, natural language processing engine, generative AI model, and emotion engine as its main components.
[0914] System configuration
[0915] User operations
[0916] 1. The user launches the official app and goes to the help section.
[0917] 2. The user opens a chat box and types in a text message with a usage question.
[0918] Example: "How do I set up Wi-Fi?"
[0919] Device behavior
[0920] 1. The device receives the text entered by the user and sends it to the server via the internet via an API.
[0921] Server Operation
[0922] 1. The server receives a query from the user. The server temporarily stores the received data and prepares it for the next process.
[0923] 2. The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0924] How Natural Language Processing and Sentiment Analysis Work
[0925] 1. The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the user's intent as "How do I set up Wi-Fi?"
[0926] 2. The emotion engine analyzes the user's query text to determine their emotions, for example, identifying that the user is feeling "troubled" or "angry."
[0927] Server Processing
[0928] 1. The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API.
[0929] Answer generation behavior
[0930] 1. The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and content of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0931] Server response processing
[0932] 1. The server receives the generated answer and sends it back to the user's device, via the internet through an API.
[0933] Terminal display
[0934] 1. The terminal displays the received answer in the chat box. The user can view the displayed answer and follow the instructions to perform the operation.
[0935] Specific examples
[0936] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app and the text includes confusion or irritation, the device sends the query to a server. The server receives the text, sends it to a natural language processing engine to analyze the intent, and uses an emotion engine to analyze the user's emotions. The analysis results identify that the query's intent is "How do I set up Wi-Fi?" and that the user is having trouble. Based on this analysis, a generative AI model generates a specific, emotion-sensitive answer and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help," and can follow the instructions to perform the operation.
[0937] This system not only allows users to instantly obtain the information they need in a chat format, but also improves their understanding of how to use the device and the smooth operation of their smartphone by receiving responses that take into consideration their emotions at the time.In addition, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[0938] The processing flow will be explained below.
[0939] Step 1:
[0940] A user launches the official app and navigates to the help section. The user opens the chat box and types a usage question in text format. For example, they type "How do I set up Wi-Fi?"
[0941] Step 2:
[0942] The device takes the text entered by the user and sends it to the server via the internet through an API.
[0943] Step 3:
[0944] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[0945] Step 4:
[0946] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[0947] Step 5:
[0948] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[0949] Step 6:
[0950] The server receives the analysis results from the natural language processing engine, and at the same time, sends the user's query text to the emotion engine for emotion analysis.
[0951] Step 7:
[0952] The emotion engine analyzes the user's text to identify their emotional state, for example detecting emotions such as "confused" or "frustrated."
[0953] Step 8:
[0954] The server integrates the analysis results from the natural language processing engine and the emotion engine and sends them to the generative AI model, also via an API.
[0955] Step 9:
[0956] The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and details of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[0957] Step 10:
[0958] The server receives the generated answer and sends it to the user's device via the internet through an API.
[0959] Step 11:
[0960] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[0961] Step 12:
[0962] The user then performs the smartphone configuration based on the answers, for example opening the Settings app, tapping the 'Wi-Fi' tab, selecting an available network and entering the password.
[0963] Through this series of processes, users not only receive information but also receive instructions that take into consideration their emotions and state, thereby reducing stress and improving the user experience.
[0964] Example 2
[0965] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0966] With conventional interfaces, even when a user asks a question about how to operate a smartphone, the answer does not take into consideration the user's emotions at the time. As a result, even when the user is confused or angry, only mechanical answers are provided, which does not improve the user experience. There was also a need for a system that could provide quick and appropriate answers.
[0967] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0968] In this invention, the server includes an information processing server means for receiving the inquiry content, a means for analyzing the received inquiry content using a natural language processing engine, a sentiment analysis means for analyzing sentiment from text entered by the user, a means for generating an answer based on the analysis result and the sentiment analysis result using a generative AI model, and a means for transmitting the generated answer to the user's information processing device, thereby making it possible to provide an appropriate and detailed answer that takes into consideration the user's sentiment.
[0969] The term "user" refers to a person who makes an inquiry using an information processing device.
[0970] An "information processing device" is a device that a user can operate and input an inquiry into, and includes a smartphone, a tablet, and the like.
[0971] "Chat-style interface" refers to a screen and software that allows users to enter messages and interact in text format.
[0972] An "information processing server" refers to a computer system that receives inquiries from users and performs processing to analyze and generate answers.
[0973] "Natural language processing engine" refers to software and algorithms that analyze a user's text input to understand their intent.
[0974] "Sentiment analyzer" refers to software and algorithms for analyzing emotions from a user's input text and identifying their emotional state.
[0975] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on analysis results and sentiment analysis results.
[0976] "Answer generation means" refers to the means for communicating the answer generated by the generative AI model to the user.
[0977] The "transmission means" refers to a communication means for transmitting the generated answer to the user's information processing device.
[0978] The present invention is a system that allows a user to input an inquiry using a chat-style interface of an information processing device, analyzes the inquiry, and generates an appropriate answer. This system includes an information processing server, a natural language processing engine, a sentiment analysis means, a generative AI model, and an answer transmission means as its main components.
[0979] User operations
[0980] First, a user starts up an information processing device (e.g., a smartphone) and navigates to the help section. Next, the user opens a chat box and enters a question about how to use the smartphone in text format. For example, the user might enter, "How do I set up Wi-Fi?"
[0981] Device behavior
[0982] The text data entered by the user's information processing device is acquired and sent to an information processing server via the Internet. This transmission is performed via an API.
[0983] Server Operation
[0984] The information processing server receives a user's query and temporarily stores the data. The server prepares the data for subsequent processing by sending it to a natural language processing engine (e.g., a general natural language processing library).
[0985] Natural Language Processing and Sentiment Analysis
[0986] A natural language processing engine analyzes the received text and understands the intent of the user's inquiry (e.g., "How to set up Wi-Fi"). Furthermore, a sentiment analysis means (e.g., a general sentiment analysis tool) analyzes emotions from the user's text and identifies emotions such as "troubled" or "angry."
[0987] Server Processing
[0988] The server receives the analysis results from the natural language processing engine and sentiment analysis tool and sends them to a generative AI model (e.g., a general generative AI tool). This transmission is also done via an API.
[0989] Answer generation behavior
[0990] The generative AI model generates an appropriate answer based on the analysis results it receives. In particular, it generates answers with a tone and content that takes into account the user's emotions based on the results of sentiment analysis. For example, if the user is in trouble, the generated answer will include more careful and detailed explanations.
[0991] Specific examples
[0992] For example, if a user types "How do I set up Wi-Fi?", analysis results indicate that the intent is to ask "How do I set up Wi-Fi?", and sentiment analysis identifies that the user is confused. The generative AI model generates an answer such as, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. If you need more help, we can help you."
[0993] Server response processing
[0994] The generated answer is returned to the information processing server, which then transmits it to the user's information processing device. Transmission is again performed via the internet through an API.
[0995] Terminal display
[0996] The user's information processing device displays the received response in a chat box. The user can view the displayed response and perform operations according to the instructions. In this way, the user can receive a quick and thoughtful response.
[0997] Prompt Sentence Examples
[0998] "A user asks, 'How do I set up Wi-Fi?' This user seems confused. Please generate a thoughtful and detailed answer based on this intent."
[0999] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1000] Step 1:
[1001] A user launches the official app for their information processing device and goes to the help section. They open the chat box and type a question about how to use their smartphone in text format. For example, they type "How do I set up Wi-Fi?" This input data is sent to the device.
[1002] Step 2:
[1003] The terminal receives the text entered by the user, converts the text data into JSON format, and sends the converted JSON data to the server via the API over the Internet (input: user text, output: JSON data sent to the server).
[1004] Step 3:
[1005] The server receives user queries, temporarily stores them in a database, and prepares the data to send to a natural language processing engine for analysis (input: JSON data, output: data ready for analysis).
[1006] Step 4:
[1007] The server sends the JSON data to a natural language processing engine. The natural language processing engine analyzes the query and identifies the user's intent. In this case, the intent is identified as "Tell me how to set up Wi-Fi." (Input: JSON data, Output: Analysis results with intent identified).
[1008] Step 5:
[1009] The server receives the analysis results and sends the data to the emotion analysis means, which analyzes emotions from the text data and identifies emotions such as "confusion" or "anger" (input: analysis results, output: emotion analysis results).
[1010] Step 6:
[1011] The server integrates the analysis results from the natural language processing engine and the sentiment analysis method. The server then sends the integrated results to the generative AI model. This transmission is also done via API (input: integrated analysis results, output: data sent to the generative AI model).
[1012] Step 7:
[1013] The generative AI model generates an appropriate answer based on the analysis results it receives. It takes into account the results of sentiment analysis to generate an answer with a tone and content that takes the user's emotions into consideration. For example, it generates an answer like, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help." (Input: Analysis results, Output: Generated answer).
[1014] Step 8:
[1015] The server receives the answer received from the generative AI model, converts the data into JSON format, and sends the converted data to the user's information processing device (input: generated answer, output: JSON data sent to the information processing device).
[1016] Step 9:
[1017] The user's information processing device parses the JSON-formatted answer data received from the server and displays the answer text in the chat box. The user can view the displayed answer and perform operations according to the instructions (input: JSON-formatted answer data, output: answer displayed in the chat box).
[1018] (Application example 2)
[1019] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1020] Conventional systems often fail to provide appropriate responses to users when they make inquiries because they do not take their emotions into consideration. Furthermore, because responses are not generated based on emotions, users' frustration and confusion are often left unresolved, leading to a decrease in satisfaction. Furthermore, this problem is particularly pronounced on online shopping sites, where responses that do not take users' emotions into consideration ultimately result in a decrease in sales and an increase in the burden on customer support.
[1021] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input an inquiry via a chat-style interface on the terminal, server means for receiving the inquiry content, means for analyzing the received inquiry content with a natural language processing engine, means for generating an answer based on the analysis results using a generative AI model, means for combining with an emotion engine that recognizes the user's emotions in order to adjust the generated answer, and means for transmitting the generated answer to the user's terminal. This makes it possible to quickly provide an appropriate answer that takes the user's emotions into consideration.
[1022] "Terminal" refers to an electronic device operated by a user, and specifically includes smartphones, tablets, personal computers, etc.
[1023] A "chat-style interface" is a format in which users communicate with each other by entering text.
[1024] A "server" is a computer system that processes and stores data on a network.
[1025] A "natural language processing engine" is a technology that analyzes input text data and understands its meaning and intent.
[1026] A "generative AI model" is an artificial intelligence model that generates appropriate answers based on analysis results.
[1027] An "emotion engine" is a technology that recognizes emotions from a user's text and adjusts responses based on those emotions.
[1028] "Means for generating an answer" refers to the process of using a generative AI model to create an appropriate answer to a user's inquiry.
[1029] The "means for transmitting the answer to the user's terminal" is a function for transmitting the generated answer to the user's terminal via a network.
[1030] This invention is a system that, when a user inputs an inquiry in chat format, analyzes the inquiry content using a natural language processing engine and an emotion engine, and generates an answer using a generative AI model. This system includes the user's terminal, a server, a natural language processing engine, a generative AI model, and an emotion engine as its main components.
[1031] System Program
[1032] In this system, users input inquiries using a chat-style interface on their device. The text entered by the user is sent to a server via the Internet. The server receives the inquiry, sends it to a natural language processing engine to analyze its meaning, and simultaneously recognizes the user's emotions using an emotion engine. Based on the results of these analyses, a generative AI model generates an appropriate response, which is then sent back to the user's device via the server.
[1033] Processing Description
[1034] The server receives the query text entered by the user and then analyzes it using a natural language processing engine, such as the open-source Natural Language Toolkit (NLTK) or Google's Cloud Natural Language API. The emotion engine then recognizes the emotion in the user's text. Examples of emotion engines that can be used include Microsoft's Azure Text Analytics and IBM's Watson Tone Analyzer.
[1035] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. Generative AI models such as OpenAI's GPT-3 and GPT-4 can be used. The generative AI model generates a response based on the analysis results, adjusting the tone and content to take the user's emotions into consideration. The server then sends the generated response to the user's device, which displays the response in a chat box.
[1036] Specific examples
[1037] For example, if a user types "My order hasn't arrived yet, what should I do?" into the chat box, this text is sent over the internet to a server. The server then sends this text to a natural language processing engine and an emotion engine to analyze intent and emotion. The analysis results determine that the intent of the inquiry is about "delayed delivery of my order" and that the user is "confused." The generative AI model uses this analysis to generate a specific, emotion-sensitive response, providing the following answer: "We're sorry for the concern. To track your order, please go to the 'Order History' page, select the order, and view the tracking information. Please let us know if we need further assistance."
[1038] Prompt Sentence Examples
[1039] User Question: "My order hasn't arrived yet, what should I do?"
[1040] User Emotion: Confused
[1041] Response: "We're sorry for your concern. To track your order, please go to your 'Order History' page and select the order to view tracking information. Please let us know if we need further assistance."
[1042] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1043] Step 1:
[1044] A user launches the official app and enters an inquiry into the chat-style interface. For example, the user might enter, "My order hasn't arrived yet. What should I do?" This input is temporarily stored on the device as text data.
[1045] Input: User query text
[1046] Output: Save to the terminal as text data
[1047] Step 2:
[1048] The device sends the text entered by the user to the server via the API, which is transmitted over the Internet.
[1049] Input: User query text
[1050] Output: Data sent to the server
[1051] Step 3:
[1052] The server receives a user's inquiry and temporarily stores the received data. It then sends the received inquiry to the natural language processing engine. Specifically, it sends the text data to the API of the natural language processing engine.
[1053] Input: Query text
[1054] Output: Data sent to the natural language processing engine
[1055] Step 4:
[1056] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the inquiry as being about a "delayed delivery of an order." The analysis results are then sent back to the server.
[1057] Input: Query text
[1058] Output: Intention information as a result of analysis
[1059] Step 5:
[1060] The server receives the intent analysis results from the natural language processing engine and sends them to the emotion engine. The emotion engine analyzes the emotion from the query text and identifies emotions such as "confusion." The analysis results are then returned to the server.
[1061] Input: Query text and intent analysis results
[1062] Output: Emotion analysis results
[1063] Step 6:
[1064] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API. The generative AI model generates an appropriate answer based on the analysis results.
[1065] Input: Results of natural language processing and sentiment analysis
[1066] Output: The generated answer
[1067] Step 7:
[1068] The generative AI model returns the generated content to the server, which then sends the generated answer to the user's device via the Internet.
[1069] Input: Generated answer
[1070] Output: Data sent to the user's device
[1071] Step 8:
[1072] The terminal displays the answer received from the server in the chat box, and the user can view the generated answer and perform operations according to the instructions.
[1073] Input: Generated answer from the server
[1074] Output: Displayed in the user's chat box
[1075] Through the above steps, the user can quickly receive an appropriate response that takes into consideration their feelings.
[1076] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1077] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1078] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1079] [Fourth embodiment]
[1080] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1081] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1082] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1083] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1084] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1085] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1086] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1087] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1088] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1089] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1090] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1091] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1092] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1093] This invention is a system that allows users to ask questions about how to use their smartphones in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. The system's main components are the user, the device, and a server.
[1094] System configuration
[1095] User operations
[1096] 1. The user launches the official app and goes to the help section.
[1097] 2. The user opens a chat box and types in a text message with a usage question.
[1098] Example: "How do I set up Wi-Fi?"
[1099] Device behavior
[1100] 1. The device receives the text entered by the user and sends it to the server.
[1101] The data is sent via an API.
[1102] Server Operation
[1103] 1. The server receives a query from a user.
[1104] 2. The server sends the received text to a natural language processing engine for analysis.
[1105] 3. The server receives the analysis results from the natural language processing engine and sends them to the generative AI model.
[1106] Natural Language Processing and AI Model Behavior
[1107] 1. The natural language processing engine analyzes the user's inquiry and understands its intent.
[1108] Example: Identifying the query intent as "How do I set up Wi-Fi?"
[1109] 2. The generative AI model generates an appropriate answer based on the analysis results.
[1110] For example, generate an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[1111] Server response processing
[1112] 1. The server receives the generated answer and sends it back to the user's device.
[1113] Terminal display
[1114] 1. The device displays the received response in the chat box for the user to view.
[1115] 2. The user checks the answers and follows the instructions to set up the smartphone.
[1116] Specific examples
[1117] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[1118] This system allows users to instantly obtain the information they need in a chat format, quickly resolving any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[1119] The processing flow will be explained below.
[1120] Step 1:
[1121] A user launches the official app and navigates to the help section. The user opens the chat box and enters their inquiry in text format. For example, they might type, "How do I set up Wi-Fi?"
[1122] Step 2:
[1123] The device takes the text entered by the user and sends it to the server via the internet through an API.
[1124] Step 3:
[1125] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[1126] Step 4:
[1127] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[1128] Step 5:
[1129] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[1130] Step 6:
[1131] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, also via an API.
[1132] Step 7:
[1133] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[1134] Step 8:
[1135] The server receives the generated answer and sends it to the user's device via the internet through an API.
[1136] Step 9:
[1137] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[1138] Example 1
[1139] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1140] Users of modern smartphones and other electronic devices often have questions about how to operate their devices. However, traditional methods offer limited options for getting fast, accurate answers when users encounter problems. Manually searching for information in instruction manuals or online help takes time and effort. Therefore, a new system is needed to help users solve problems efficiently.
[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1142] In this invention, the server includes means for a user to input an inquiry via a chat-style interface of the electronic device, means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, and means for transmitting the generated answer to the user's electronic device. This allows the user to quickly receive an accurate answer in chat format and efficiently resolve questions about how to operate the device.
[1143] "User" refers to a user who operates an electronic device.
[1144] "Electronic devices" refers to computing devices such as smartphones, tablets, and personal computers.
[1145] "Chat-style interface" refers to a user interface that allows users to enter text messages and communicate information interactively.
[1146] An "inquiry" refers to a user inputting a question or doubt about how to operate an electronic device.
[1147] A "server" refers to a computer system that receives queries from users, processes them, and generates and transmits responses.
[1148] A "natural language processing engine" refers to software that analyzes text input from a user and understands their intent.
[1149] A "generative AI model" refers to an artificial intelligence model that automatically generates appropriate answers based on the results analyzed by a natural language processing engine.
[1150] "Answer" refers to information generated by a generative AI model that includes answers or instructions to a user's inquiry.
[1151] "Instructions" refers to documents or sources of information that contain detailed instructions on the operation and functionality of an electronic device.
[1152] This invention is a system that allows users to ask questions about how to use electronic devices in a chat format, and a generative AI model provides answers in real time based on the content of the inquiry. This system mainly consists of a user, a terminal, and a server.
[1153] System configuration
[1154] User operations
[1155] A user launches the official app and navigates to the help section. The user opens the chat box and enters a usage question in text format. For example, if the user types "How do I set up Wi-Fi?", the following process occurs:
[1156] Device behavior
[1157] The device takes the text entered by the user and sends it to the server, where it is packaged in JSON format and protected by SSL / TLS encryption via an API.
[1158] Server Operation
[1159] The server receives a query from a user. The server passes the received request to a processing program via a web server such as Apache or Nginx. For example, a Flask application written in Python handles this request. The server then sends the received text to a natural language processing engine. This analysis is performed using the Google Cloud Natural Language API or an NLP library (e.g., spaCy, NLTK). After the analysis is complete, the server sends the results to a generative AI model (e.g., OpenAI's GPT-3 or GPT-4).
[1160] Natural Language Processing and AI Model Behavior
[1161] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, in response to the inquiry "How do I set up Wi-Fi?", the keyword "How to set up Wi-Fi" is extracted. Next, the generative AI model generates an appropriate answer based on the analysis results. For example, it generates an answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password."
[1162] Server response processing
[1163] The server receives the generated answer and sends it back to the user's device, where the data is again sent via an HTTP response.
[1164] Terminal display
[1165] The device displays the received answer in a chat box for the user to view. The user can then check the answer and follow the instructions to set up their smartphone.
[1166] Specific examples
[1167] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app, the device sends the query to a server. The server receives the text and sends it to a natural language processing engine to analyze the intent. The analysis identifies the query as being about "setting up Wi-Fi." The generative AI model then generates a specific answer based on the analysis and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password," and can follow the instructions to perform the operation.
[1168] Example prompt sentence:
[1169] The user types "How do I set up Wi-Fi?" into the chat box. Generate an appropriate response.
[1170] This system allows users to instantly obtain the information they need in a chat format, enabling them to quickly resolve any questions they may have about using their smartphone. Furthermore, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[1171] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1172] Step 1:
[1173] A user launches the official app and navigates to the help section. The user unlocks the smartphone and taps the app icon. The app launches and displays the home screen. The user clicks the help or support tab within the app.
[1174] Input: Smartphone operation
[1175] Output: Help section on screen
[1176] Step 2:
[1177] The user opens the chat box, types a usage question in text format, for example, "How do I set up Wi-Fi?", and taps the send button.
[1178] Input: User question text
[1179] Output: The terminal displays the user's input text.
[1180] Step 3:
[1181] The device retrieves the user's input text and sends it to the server. The device generates an HTTP POST request, packages the user's input text in JSON format, and sends it to the SSL / TLS encrypted API endpoint.
[1182] Input: User-entered text
[1183] Output: Encrypted HTTP POST request
[1184] Step 4:
[1185] The server receives a query from a user. The server receives the request through a web server (e.g., Nginx, Apache) and passes the request to its internal processing program (e.g., a Python Flask application).
[1186] Input: Encrypted HTTP POST request
[1187] Output: The request data passed to the processing program
[1188] Step 5:
[1189] The server sends the received text to a natural language processing engine, which extracts the text from the request data, sends it to the natural language processing engine (e.g., Google Cloud Natural Language API), and makes an API call for analysis.
[1190] Input: The text portion of the request data
[1191] Output: API request to the natural language processing engine
[1192] Step 6:
[1193] A natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it analyzes the text "How do I set up Wi-Fi?" and identifies the intent as "How do I set up Wi-Fi?"
[1194] Input: The text to be parsed
[1195] Output: Analysis results (intent, keywords, etc.)
[1196] Step 7:
[1197] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model. The server then generates an API request to pass the analysis results to the generative AI model (e.g., OpenAI GPT-3). This request includes a prompt.
[1198] Input: Analysis results of the natural language processing engine
[1199] Output: API request to the generative AI model
[1200] Step 8:
[1201] The generative AI model generates an appropriate answer based on the analysis results, such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password."
[1202] Input: prompt sentence for generative AI model
[1203] Output: Generated answer text
[1204] Step 9:
[1205] The server receives the generated answer and returns it to the user's device. The server packages the generated answer in JSON format and sends it to the user's device as an HTTP response.
[1206] Input: Generated answer text
[1207] Output: Answer sent to the terminal as an HTTP response
[1208] Step 10:
[1209] The device displays the received answer in the chat box for the user to view. The device deserializes the received JSON data and displays the answer text in the chat box. The user checks the answer and follows the instructions to configure their smartphone.
[1210] Input: JSON data of the HTTP response
[1211] Output: Answer text displayed in the chat box
[1212] (Application example 1)
[1213] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1214] In the past, it was difficult for workers to receive prompt and appropriate support when operating or troubleshooting robots used in factories. This often resulted in reduced work efficiency and lost productivity. In such situations, a system was needed that would enable workers to quickly understand how to use robots and immediately provide reliable information to solve problems.
[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1216] In this invention, the server includes means for a user to input an inquiry via a chat-style interface on a terminal, server means for receiving the inquiry, means for analyzing the received inquiry using a natural language processing engine, means for generating an answer based on the analysis result using a generative AI model, means for sending the generated answer to the user's terminal, and means for analyzing the inquiry specifically regarding the operation of a robot in a factory and generating appropriate instructions, thereby enabling the user to obtain detailed information on how to operate and troubleshoot the robot in a factory in real time.
[1217] "User" refers to any individual or entity that uses the System to make an inquiry.
[1218] "Terminal" refers to a computer or device for entering queries and displaying responses.
[1219] A "chat-style interface" refers to a user interface in which a user inputs a text-based inquiry and receives a response from the system.
[1220] "Server" refers to a central processing unit for receiving queries from users, analyzing them, and generating responses.
[1221] A "natural language processing engine" refers to a machine learning or rule-based system that analyzes user queries and understands their intent.
[1222] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on the analysis results of a natural language processing engine.
[1223] "Inquiries regarding the operation of robots in factories" refers to questions regarding how to operate and troubleshoot robots used in factories.
[1224] An "answer" refers to the information or instructions generated by a generative AI model in response to a user's inquiry.
[1225] The system for realizing this application example operates in the following procedure.
[1226] First, the user accesses a chat-style interface using a device, such as a PC, tablet, or smartphone, and enters text-based inquiries about the robot's operation.
[1227] The device takes the entered text data and sends it over the Internet to a server using a standard HTTP request.
[1228] The server processes the received query. First, the query is analyzed by a natural language processing engine. This engine uses natural language processing libraries such as spaCy or NLTK. The analysis identifies the intent of the query and important keywords.
[1229] The server then uses a generative AI model, such as GPT-4, to generate an answer based on the analysis results. The AI model generates an appropriate and detailed answer based on the analyzed intent and keywords.
[1230] The generated answers are then sent back to the terminal via the server and displayed to the user, allowing the user to obtain appropriate operating instructions and troubleshooting methods in real time.
[1231] For example, if a user types, "Please tell me how to calibrate my robot," this query is analyzed by a natural language processing engine, and the generative AI model generates the answer, "Enter calibration mode and adjust each sensor according to the manual."
[1232] In this way, the system can provide users with immediate and reliable information regarding their robot operation questions.
[1233] Example prompt sentence:
[1234] User: "How do I calibrate my robot?"
[1235] Answer: "Enter calibration mode and adjust each sensor according to the manual."
[1236] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1237] Step 1:
[1238] The user uses a terminal to access a chat-style interface and enters a text inquiry about the operation of the robot, for example, "Tell me how to calibrate the robot."
[1239] Input: User's inquiry (text format)
[1240] Output: Query text data
[1241] Step 2:
[1242] The device receives the text data entered by the user and sends it to the server as an API request, using the standard HTTP protocol.
[1243] Input: User query text data
[1244] Output: HTTP request to the server
[1245] Step 3:
[1246] The server sends the received query to a natural language processing engine for analysis, which analyzes the query content and identifies its intent and important keywords.
[1247] Input: Query text data as an HTTP request
[1248] Output: Analysis results (intent and keywords)
[1249] What it does: Analyzes text using natural language processing libraries such as spaCy and NLTK
[1250] Step 4:
[1251] The server receives the analysis results from the natural language processing engine and sends them to the generative AI model, which then generates an appropriate answer based on the analysis results.
[1252] Input: Analysis results (intent and keywords)
[1253] Output: Generated answer text
[1254] Specific operation: Generate text using generative AI models such as GPT-4
[1255] Step 5:
[1256] The server retrieves the generated answer and sends it to the user's device using an HTTP response.
[1257] Input: Generated answer text
[1258] Output: HTTP response to the user's device
[1259] Step 6:
[1260] The terminal displays the received response on the chat interface so that it can be seen by the user, who can then confirm it and perform the necessary operations.
[1261] Input: HTTP response containing the answer text from the server
[1262] Output: Answer text displayed in the chat interface
[1263] Specific operation: The received response text is displayed on the screen. For example, "Enter calibration mode and adjust each sensor according to the manual."
[1264] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1265] This invention combines a system in which a user inputs an inquiry about how to use a smartphone in chat format, analyzes the inquiry content with a natural language processing engine, and generates an answer with a generative AI model, with an emotion engine that recognizes the user's emotions.This system includes the user, device, server, natural language processing engine, generative AI model, and emotion engine as its main components.
[1266] System configuration
[1267] User operations
[1268] 1. The user launches the official app and goes to the help section.
[1269] 2. The user opens a chat box and types in a text message with a usage question.
[1270] Example: "How do I set up Wi-Fi?"
[1271] Device behavior
[1272] 1. The device receives the text entered by the user and sends it to the server via the internet via an API.
[1273] Server Operation
[1274] 1. The server receives a query from the user. The server temporarily stores the received data and prepares it for the next process.
[1275] 2. The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[1276] How Natural Language Processing and Sentiment Analysis Work
[1277] 1. The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the user's intent as "How do I set up Wi-Fi?"
[1278] 2. The emotion engine analyzes the user's query text to determine their emotions, for example, identifying that the user is feeling "troubled" or "angry."
[1279] Server Processing
[1280] 1. The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API.
[1281] Answer generation behavior
[1282] 1. The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and content of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[1283] Server response processing
[1284] 1. The server receives the generated answer and sends it back to the user's device, via the internet through an API.
[1285] Terminal display
[1286] 1. The terminal displays the received answer in the chat box. The user can view the displayed answer and follow the instructions to perform the operation.
[1287] Specific examples
[1288] For example, if a user types "How do I set up Wi-Fi?" into a chat box within an app and the text includes confusion or irritation, the device sends the query to a server. The server receives the text, sends it to a natural language processing engine to analyze the intent, and uses an emotion engine to analyze the user's emotions. The analysis results identify that the query's intent is "How do I set up Wi-Fi?" and that the user is having trouble. Based on this analysis, a generative AI model generates a specific, emotion-sensitive answer and sends it to the server. The user receives the answer, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help," and can follow the instructions to perform the operation.
[1289] This system not only allows users to instantly obtain the information they need in a chat format, but also improves their understanding of how to use the device and the smooth operation of their smartphone by receiving responses that take into consideration their emotions at the time.In addition, by providing reliable information based on the contents of the instruction manual, it is possible to significantly improve user satisfaction.
[1290] The processing flow will be explained below.
[1291] Step 1:
[1292] A user launches the official app and navigates to the help section. The user opens the chat box and types a usage question in text format. For example, they type "How do I set up Wi-Fi?"
[1293] Step 2:
[1294] The device takes the text entered by the user and sends it to the server via the internet through an API.
[1295] Step 3:
[1296] The server receives a query from the user, temporarily stores the received data, and prepares it for the next process.
[1297] Step 4:
[1298] The server sends the received text to the natural language processing engine for analysis. Specifically, it sends the text data to the NLP engine's API.
[1299] Step 5:
[1300] The natural language processing engine analyzes the content of the user's inquiry and understands its intent. For example, it identifies the intent as "How do I set up Wi-Fi?"
[1301] Step 6:
[1302] The server receives the analysis results from the natural language processing engine, and at the same time, sends the user's query text to the emotion engine for emotion analysis.
[1303] Step 7:
[1304] The emotion engine analyzes the user's text to identify their emotional state, for example detecting emotions such as "confused" or "frustrated."
[1305] Step 8:
[1306] The server integrates the analysis results from the natural language processing engine and the emotion engine and sends them to the generative AI model, also via an API.
[1307] Step 9:
[1308] The generative AI model generates an appropriate answer based on the analysis results. Furthermore, the tone and details of the answer are adjusted based on the analysis results of the emotion engine. For example, if the user is having trouble, a more polite and detailed explanation will be included. An answer such as "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter the password. We'll help you further if you need help" will be generated.
[1309] Step 10:
[1310] The server receives the generated answer and sends it to the user's device via the internet through an API.
[1311] Step 11:
[1312] The terminal displays the received response in the chat box, and the user can view the displayed response and perform operations according to the instructions.
[1313] Step 12:
[1314] The user then performs the smartphone configuration based on the answers, for example opening the Settings app, tapping the 'Wi-Fi' tab, selecting an available network and entering the password.
[1315] Through this series of processes, users not only receive information but also receive instructions that take into consideration their emotions and state, thereby reducing stress and improving the user experience.
[1316] Example 2
[1317] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1318] With conventional interfaces, even when a user asks a question about how to operate a smartphone, the answer does not take into consideration the user's emotions at the time. As a result, even when the user is confused or angry, only mechanical answers are provided, which does not improve the user experience. There was also a need for a system that could provide quick and appropriate answers.
[1319] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1320] In this invention, the server includes an information processing server means for receiving the inquiry content, a means for analyzing the received inquiry content using a natural language processing engine, a sentiment analysis means for analyzing sentiment from text entered by the user, a means for generating an answer based on the analysis result and the sentiment analysis result using a generative AI model, and a means for transmitting the generated answer to the user's information processing device, thereby making it possible to provide an appropriate and detailed answer that takes into consideration the user's sentiment.
[1321] The term "user" refers to a person who makes an inquiry using an information processing device.
[1322] An "information processing device" is a device that a user can operate and input an inquiry into, and includes a smartphone, a tablet, and the like.
[1323] "Chat-style interface" refers to a screen and software that allows users to enter messages and interact in text format.
[1324] An "information processing server" refers to a computer system that receives inquiries from users and performs processing to analyze and generate answers.
[1325] "Natural language processing engine" refers to software and algorithms that analyze a user's text input to understand their intent.
[1326] "Sentiment analyzer" refers to software and algorithms for analyzing emotions from a user's input text and identifying their emotional state.
[1327] A "generative AI model" refers to an artificial intelligence model that generates appropriate answers based on analysis results and sentiment analysis results.
[1328] "Answer generation means" refers to the means for communicating the answer generated by the generative AI model to the user.
[1329] The "transmission means" refers to a communication means for transmitting the generated answer to the user's information processing device.
[1330] The present invention is a system that allows a user to input an inquiry using a chat-style interface of an information processing device, analyzes the inquiry, and generates an appropriate answer. This system includes an information processing server, a natural language processing engine, a sentiment analysis means, a generative AI model, and an answer transmission means as its main components.
[1331] User operations
[1332] First, a user starts up an information processing device (e.g., a smartphone) and navigates to the help section. Next, the user opens a chat box and enters a question about how to use the smartphone in text format. For example, the user might enter, "How do I set up Wi-Fi?"
[1333] Device behavior
[1334] The text data entered by the user's information processing device is acquired and sent to an information processing server via the Internet. This transmission is performed via an API.
[1335] Server Operation
[1336] The information processing server receives a user's query and temporarily stores the data. The server prepares the data for subsequent processing by sending it to a natural language processing engine (e.g., a general natural language processing library).
[1337] Natural Language Processing and Sentiment Analysis
[1338] A natural language processing engine analyzes the received text and understands the intent of the user's inquiry (e.g., "How to set up Wi-Fi"). Furthermore, a sentiment analysis means (e.g., a general sentiment analysis tool) analyzes emotions from the user's text and identifies emotions such as "troubled" or "angry."
[1339] Server Processing
[1340] The server receives the analysis results from the natural language processing engine and sentiment analysis tool and sends them to a generative AI model (e.g., a general generative AI tool). This transmission is also done via an API.
[1341] Answer generation behavior
[1342] The generative AI model generates an appropriate answer based on the analysis results it receives. In particular, it generates answers with a tone and content that takes into account the user's emotions based on the results of sentiment analysis. For example, if the user is in trouble, the generated answer will include more careful and detailed explanations.
[1343] Specific examples
[1344] For example, if a user types "How do I set up Wi-Fi?", analysis results indicate that the intent is to ask "How do I set up Wi-Fi?", and sentiment analysis identifies that the user is confused. The generative AI model generates an answer such as, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. If you need more help, we can help you."
[1345] Server response processing
[1346] The generated answer is returned to the information processing server, which then transmits it to the user's information processing device. Transmission is again performed via the internet through an API.
[1347] Terminal display
[1348] The user's information processing device displays the received response in a chat box. The user can view the displayed response and perform operations according to the instructions. In this way, the user can receive a quick and thoughtful response.
[1349] Prompt Sentence Examples
[1350] "A user asks, 'How do I set up Wi-Fi?' This user seems confused. Please generate a thoughtful and detailed answer based on this intent."
[1351] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1352] Step 1:
[1353] A user launches the official app for their information processing device and goes to the help section. They open the chat box and type a question about how to use their smartphone in text format. For example, they type "How do I set up Wi-Fi?" This input data is sent to the device.
[1354] Step 2:
[1355] The terminal receives the text entered by the user, converts the text data into JSON format, and sends the converted JSON data to the server via the API over the Internet (input: user text, output: JSON data sent to the server).
[1356] Step 3:
[1357] The server receives user queries, temporarily stores them in a database, and prepares the data to send to a natural language processing engine for analysis (input: JSON data, output: data ready for analysis).
[1358] Step 4:
[1359] The server sends the JSON data to a natural language processing engine. The natural language processing engine analyzes the query and identifies the user's intent. In this case, the intent is identified as "Tell me how to set up Wi-Fi." (Input: JSON data, Output: Analysis results with intent identified).
[1360] Step 5:
[1361] The server receives the analysis results and sends the data to the emotion analysis means, which analyzes emotions from the text data and identifies emotions such as "confusion" or "anger" (input: analysis results, output: emotion analysis results).
[1362] Step 6:
[1363] The server integrates the analysis results from the natural language processing engine and the sentiment analysis method. The server then sends the integrated results to the generative AI model. This transmission is also done via API (input: integrated analysis results, output: data sent to the generative AI model).
[1364] Step 7:
[1365] The generative AI model generates an appropriate answer based on the analysis results it receives. It takes into account the results of sentiment analysis to generate an answer with a tone and content that takes the user's emotions into consideration. For example, it generates an answer like, "Open the Settings app, tap the 'Wi-Fi' tab, select an available network, and enter your password. We'll help you further if you need help." (Input: Analysis results, Output: Generated answer).
[1366] Step 8:
[1367] The server receives the answer received from the generative AI model, converts the data into JSON format, and sends the converted data to the user's information processing device (input: generated answer, output: JSON data sent to the information processing device).
[1368] Step 9:
[1369] The user's information processing device parses the JSON-formatted answer data received from the server and displays the answer text in the chat box. The user can view the displayed answer and perform operations according to the instructions (input: JSON-formatted answer data, output: answer displayed in the chat box).
[1370] (Application example 2)
[1371] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1372] Conventional systems often fail to provide appropriate responses to users when they make inquiries because they do not take their emotions into consideration. Furthermore, because responses are not generated based on emotions, users' frustration and confusion are often left unresolved, leading to a decrease in satisfaction. Furthermore, this problem is particularly pronounced on online shopping sites, where responses that do not take users' emotions into consideration ultimately result in a decrease in sales and an increase in the burden on customer support.
[1373] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input an inquiry via a chat-style interface on the terminal, server means for receiving the inquiry content, means for analyzing the received inquiry content with a natural language processing engine, means for generating an answer based on the analysis results using a generative AI model, means for combining with an emotion engine that recognizes the user's emotions in order to adjust the generated answer, and means for transmitting the generated answer to the user's terminal. This makes it possible to quickly provide an appropriate answer that takes the user's emotions into consideration.
[1374] "Terminal" refers to an electronic device operated by a user, and specifically includes smartphones, tablets, personal computers, etc.
[1375] A "chat-style interface" is a format in which users communicate with each other by entering text.
[1376] A "server" is a computer system that processes and stores data on a network.
[1377] A "natural language processing engine" is a technology that analyzes input text data and understands its meaning and intent.
[1378] A "generative AI model" is an artificial intelligence model that generates appropriate answers based on analysis results.
[1379] An "emotion engine" is a technology that recognizes emotions from a user's text and adjusts responses based on those emotions.
[1380] "Means for generating an answer" refers to the process of using a generative AI model to create an appropriate answer to a user's inquiry.
[1381] The "means for transmitting the answer to the user's terminal" is a function for transmitting the generated answer to the user's terminal via a network.
[1382] This invention is a system that, when a user inputs an inquiry in chat format, analyzes the inquiry content using a natural language processing engine and an emotion engine, and generates an answer using a generative AI model. This system includes the user's terminal, a server, a natural language processing engine, a generative AI model, and an emotion engine as its main components.
[1383] System Program
[1384] In this system, users input inquiries using a chat-style interface on their device. The text entered by the user is sent to a server via the Internet. The server receives the inquiry, sends it to a natural language processing engine to analyze its meaning, and simultaneously recognizes the user's emotions using an emotion engine. Based on the results of these analyses, a generative AI model generates an appropriate response, which is then sent back to the user's device via the server.
[1385] Processing Description
[1386] The server receives the query text entered by the user and then analyzes it using a natural language processing engine, such as the open-source Natural Language Toolkit (NLTK) or Google's Cloud Natural Language API. The emotion engine then recognizes the emotion in the user's text. Examples of emotion engines that can be used include Microsoft's Azure Text Analytics and IBM's Watson Tone Analyzer.
[1387] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. Generative AI models such as OpenAI's GPT-3 and GPT-4 can be used. The generative AI model generates a response based on the analysis results, adjusting the tone and content to take the user's emotions into consideration. The server then sends the generated response to the user's device, which displays the response in a chat box.
[1388] Specific examples
[1389] For example, if a user types "My order hasn't arrived yet, what should I do?" into the chat box, this text is sent over the internet to a server. The server then sends this text to a natural language processing engine and an emotion engine to analyze intent and emotion. The analysis results determine that the intent of the inquiry is about "delayed delivery of my order" and that the user is "confused." The generative AI model uses this analysis to generate a specific, emotion-sensitive response, providing the following answer: "We're sorry for the concern. To track your order, please go to the 'Order History' page, select the order, and view the tracking information. Please let us know if we need further assistance."
[1390] Prompt Sentence Examples
[1391] User Question: "My order hasn't arrived yet, what should I do?"
[1392] User Emotion: Confused
[1393] Response: "We're sorry for your concern. To track your order, please go to your 'Order History' page and select the order to view tracking information. Please let us know if we need further assistance."
[1394] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1395] Step 1:
[1396] A user launches the official app and enters an inquiry into the chat-style interface. For example, the user might enter, "My order hasn't arrived yet. What should I do?" This input is temporarily stored on the device as text data.
[1397] Input: User query text
[1398] Output: Save to the terminal as text data
[1399] Step 2:
[1400] The device sends the text entered by the user to the server via the API, which is transmitted over the Internet.
[1401] Input: User query text
[1402] Output: Data sent to the server
[1403] Step 3:
[1404] The server receives a user's inquiry and temporarily stores the received data. It then sends the received inquiry to the natural language processing engine. Specifically, it sends the text data to the API of the natural language processing engine.
[1405] Input: Query text
[1406] Output: Data sent to the natural language processing engine
[1407] Step 4:
[1408] The natural language processing engine analyzes the user's inquiry and understands its intent. For example, it identifies the inquiry as being about a "delayed delivery of an order." The analysis results are then sent back to the server.
[1409] Input: Query text
[1410] Output: Intention information as a result of analysis
[1411] Step 5:
[1412] The server receives the intent analysis results from the natural language processing engine and sends them to the emotion engine. The emotion engine analyzes the emotion from the query text and identifies emotions such as "confusion." The analysis results are then returned to the server.
[1413] Input: Query text and intent analysis results
[1414] Output: Emotion analysis results
[1415] Step 6:
[1416] The server receives the analysis results from the natural language processing engine and emotion engine and sends them to the generative AI model. This transmission is also done via API. The generative AI model generates an appropriate answer based on the analysis results.
[1417] Input: Results of natural language processing and sentiment analysis
[1418] Output: The generated answer
[1419] Step 7:
[1420] The generative AI model returns the generated content to the server, which then sends the generated answer to the user's device via the Internet.
[1421] Input: Generated answer
[1422] Output: Data sent to the user's device
[1423] Step 8:
[1424] The terminal displays the answer received from the server in the chat box, and the user can view the generated answer and perform operations according to the instructions.
[1425] Input: Generated answer from the server
[1426] Output: Displayed in the user's chat box
[1427] Through the above steps, the user can quickly receive an appropriate response that takes into consideration their feelings.
[1428] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1429] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1430] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1431] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1432] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1433] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1434] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1435] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1436] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1437] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1438] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1439] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1440] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1441] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1442] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1443] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1444] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1445] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1446] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1447] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1448] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1449] The following is further disclosed regarding the above embodiment.
[1450] (Claim 1)
[1451] a means for a user to input an inquiry through a chat-style interface on the device;
[1452] A server means for receiving the inquiry content;
[1453] A means for analyzing the received inquiry content using a natural language processing engine;
[1454] A means for generating answers based on the analysis results using a generative AI model;
[1455] means for transmitting the generated answer to the user's terminal;
[1456] A system including:
[1457] (Claim 2)
[1458] 10. The system of claim 1, further comprising means for displaying the generated answers.
[1459] (Claim 3)
[1460] 10. The system of claim 1, wherein the generated answers are based on the contents of the instruction manual.
[1461] "Example 1"
[1462] (Claim 1)
[1463] means for a user to input a query in a chat-style interface on the electronic device;
[1464] A server means for receiving the inquiry content;
[1465] A means for analyzing the received inquiry content using a natural language processing engine;
[1466] A means for generating answers based on the analysis results using a generative AI model;
[1467] means for transmitting the generated answer to the user's electronic device;
[1468] A system including:
[1469] (Claim 2)
[1470] 10. The system of claim 1, further comprising means for displaying the generated answers.
[1471] (Claim 3)
[1472] 10. The system of claim 1, wherein the generated answer is based on the content of the usage guide.
[1473] "Application Example 1"
[1474] (Claim 1)
[1475] a means for a user to input an inquiry through a chat-style interface on the device;
[1476] A server means for receiving the inquiry content;
[1477] A means for analyzing the received inquiry content using a natural language processing engine;
[1478] A means for generating answers based on the analysis results using a generative AI model;
[1479] means for transmitting the generated answer to the user's terminal;
[1480] A means for analyzing inquiries specifically regarding the operation of a robot in a factory and generating appropriate instructions;
[1481] A system including:
[1482] (Claim 2)
[1483] 10. The system of claim 1, further comprising means for displaying the generated answers.
[1484] (Claim 3)
[1485] 10. The system of claim 1, wherein the generated answers are based on the contents of the instruction manual.
[1486] "Example 2: Combining Emotion Engines"
[1487] (Claim 1)
[1488] A means for a user to input an inquiry through a chat-style interface of an information processing device;
[1489] an information processing server means for receiving the inquiry content;
[1490] A means for analyzing the received inquiry content using a natural language processing engine;
[1491] A sentiment analysis means for analyzing sentiment from text input by a user;
[1492] A means for generating an answer based on the analysis results and the emotion analysis results using a generative AI model;
[1493] means for transmitting the generated answer to the user's information processing device;
[1494] A system including:
[1495] (Claim 2)
[1496] 10. The system of claim 1, further comprising means for displaying the generated answers.
[1497] (Claim 3)
[1498] 10. The system of claim 1, wherein the generated answer is sensitive to the user's feelings.
[1499] "Application example 2 when combining emotion engines"
[1500] (Claim 1)
[1501] a means for a user to input an inquiry through a chat-style interface on the device;
[1502] A server means for receiving the inquiry content;
[1503] A means for analyzing the received inquiry content using a natural language processing engine;
[1504] A means for generating answers based on the analysis results using a generative AI model;
[1505] means for combining an emotion engine that recognizes the user's emotions to adjust the generated answers;
[1506] means for transmitting the generated answer to the user's terminal;
[1507] A system including:
[1508] (Claim 2)
[1509] 10. The system of claim 1, further comprising means for displaying the generated answers.
[1510] (Claim 3)
[1511] 10. The system of claim 1, wherein the generated answers are based on the contents of the instruction manual and adjust the tone and content to take into account the user's emotions. [Explanation of symbols]
[1512] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for a user to input an inquiry through a chat-style interface on the device; A server means for receiving the inquiry content; A means for analyzing the received inquiry content using a natural language processing engine; A means for generating answers based on the analysis results using a generative AI model; means for transmitting the generated answer to the user's terminal; A system including:
2. The system of claim 1 further comprising means for displaying the generated answers.
3. The system of claim 1 , wherein the generated answers are based on the contents of the instruction manual.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A