system
A system that processes natural language inputs for route guidance through a conversational agent on a server provides quick and easy access to route information, addressing the inefficiencies of conventional methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Conventional systems for obtaining route guidance require multiple operations and are time-consuming, making them difficult for beginners and the elderly to use effectively.
A system that allows users to input questions in natural language via a communication terminal, which are processed by a server using a conversational agent to generate and display route information quickly.
Enables users to easily and quickly obtain necessary route information without the need for complex operations.
Smart Images

Figure 2026063858000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, when a user uses route guidance, there has been a problem that it takes time to obtain search results using a browser or an application. Also, since a plurality of operations are required, it has been difficult to use especially for beginners and the elderly. The object of the present invention is to solve these problems and provide a system that enables a user to easily and quickly obtain necessary route information.
Means for Solving the Problems
[0005] The present invention provides a system comprising: means for a user to input a question in natural language via a communication terminal; means for sending the question to a server; means for the server to receive the question and generate a response using a dialogue agent; and means for sending the response back to the communication terminal and displaying it to the user. This allows the user to quickly obtain route information simply by asking a question in natural language.
[0006] A "user" is a person who uses the system to input questions in natural language.
[0007] A "communication terminal" is a device used by users to input questions and communicate with a server. Examples include smartphones and personal computers.
[0008] "Natural language" refers to the language that users use on a daily basis, not a specific programming language or command set, but a language used for communication between people.
[0009] A "server" is a computer system that receives requests sent from communication terminals, generates responses through conversational agents, and sends them back to the communication terminals.
[0010] A "question" is the content of an inquiry entered by a user into a communication terminal, expressed in natural language to obtain appropriate routing information and other relevant data.
[0011] "Transmission" refers to the act of transferring data from a communication terminal to a server.
[0012] "Receiving" refers to the act of a server receiving data transmitted from a communication terminal.
[0013] A "conversational agent" is software that runs on a server, analyzes user questions, and generates appropriate responses.
[0014] "Response" refers to the content of the answer generated by the dialogue agent and sent back from the server in response to the user's question.
[0015] "Return" refers to the act of the server sending back the generated response to the communication terminal.
[0016] "Display" refers to the act of the communication terminal visually presenting the response to the user.
Brief Description of the Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Embodiments for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] The present invention provides a system that allows users to input questions in natural language and quickly obtain route information. The embodiments thereof are described in detail below.
[0039] System-wide configuration
[0040] This system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0041] Communication terminal
[0042] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0043] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0044] server
[0045] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0046] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs.
[0047] Dialogue Agent
[0048] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0049] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0050] Sending and displaying responses
[0051] The server sends the response generated by the conversational agent back to the communication terminal. The communication terminal receives this response and displays it in the chat window. The user can then review the displayed response.
[0052] Specific examples
[0053] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station" into the chat window of their communication terminal and presses the send button.
[0054] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent.
[0055] The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please take the JR Yamanote Line. The journey takes approximately 10 minutes."
[0056] This response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0057] The above describes an embodiment of the system based on the present invention. By applying this system, users will be able to obtain route information easily and quickly.
[0058] The following describes the processing flow.
[0059] Step 1:
[0060] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0061] Step 2:
[0062] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0063] Step 3:
[0064] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0065] Step 4:
[0066] The server receives the request. The received data is converted into an appropriate format for analysis.
[0067] Step 5:
[0068] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0069] Step 6:
[0070] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0071] Step 7:
[0072] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0073] Step 8:
[0074] The generated response is returned to the server. The server converts this response into JSON data.
[0075] Step 9:
[0076] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0077] Step 10:
[0078] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0079] Step 11:
[0080] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0081] The above outlines the specific processing steps of this system's program.
[0082] (Example 1)
[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] In modern society, it is crucial for users to quickly and accurately obtain appropriate route information based on their location and destination. However, conventional systems often take a long time to process questions entered in natural language, resulting in inadequate responses. To solve this problem, a system is needed that can properly analyze user questions and quickly provide optimal route information.
[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0086] In this invention, the server includes means for a user to input a question in natural language via a communication device, means for transmitting the question to a data processing device, means for the data processing device to receive the question and generate a response using a conversational agent, means for sending the response back to the communication device and displaying it to the user, and includes question input at the communication device, question analysis at the data processing device, retrieval of an information database, and response generation. This makes it possible for the user to obtain route information easily and quickly.
[0087] A "communication device" is a device used by a user to input questions via an interface, and includes smartphones and personal computers.
[0088] A "data processing device" is a device that receives data sent by a user, analyzes it, and generates a response; it generally refers to a server.
[0089] A "conversational agent" is a program or system that analyzes a user's questions and generates appropriate responses, using natural language processing technology.
[0090] A "response" refers to the information generated by the conversational agent in response to a user's question, and it is presented in natural language.
[0091] An "information database" is a system for storing and managing necessary data, including route information and map information.
[0092] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.
[0093] The present invention provides a system that allows a user to input a question in natural language via a communication device and quickly obtain route information based on that question. Specific embodiments thereof are described in detail below.
[0094] Communication device
[0095] Users open a dedicated chat window using a communication device such as a smartphone or personal computer. In this chat window, users can input questions about their current location and destination using natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends the entered question to a data processing device (server).
[0096] Data processing device (server)
[0097] The data processing unit receives requests sent from communication devices and has the function of analyzing these requests. The received question content is passed to a conversation agent running on the server. The data processing unit automatically handles everything from receiving requests to generating and sending responses, and also has the function of saving received messages as logs.
[0098] Conversation Agent
[0099] The conversational agent's role is to analyze the content of questions sent by users using natural language processing technology. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the information database to obtain the optimal route information. The obtained route information is converted into a response to be sent back to the user in natural language. For example, it might say, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes."
[0100] Sending and displaying responses
[0101] The data processing unit (server) sends the response generated by the conversation agent back to the communication device. The communication device receives this response and displays it in the chat window. The user can review the displayed response and obtain the necessary routing information.
[0102] Specific examples
[0103] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station" into the chat window of the communication device and clicks the send button. The communication device sends this question to the data processing device. The data processing device receives the request and passes the question to the conversational agent. The conversational agent searches the information database and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the Y Line of the railway. The journey takes approximately 10 minutes." This response is returned to the data processing device and then sent back to the communication device. The communication device displays the received response in the chat window, allowing the user to confirm the answer.
[0104] Example of a prompt
[0105] Examples of prompts to input into a generative AI model include the following:
[0106] "Describe the sequence of events in which the system generates responses to questions submitted by users. Clearly define the roles of the server, communication device, and conversational agent."
[0107] The above describes the embodiments for carrying out the present invention. With this system, users can easily and quickly obtain route information.
[0108] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0109] Step 1:
[0110] The user opens a chat window via a communication device. This device could be a smartphone or a personal computer. Once the chat window appears, the user enters a question in natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" This question is then entered into the communication device.
[0111] Step 2:
[0112] When the user clicks the "Submit" button, the communication device sends the entered question content to the data processing device (server). Specifically, the communication device sends the question text to the server via the network. At this time, the input data is the user's question, and the output data is the question sent as a request to the server.
[0113] Step 3:
[0114] The server receives requests sent from communication devices. First, the server analyzes the received question. This analysis involves inputting the user's question as received data into a text analysis engine, which outputs analysis results that help understand the format and structure of the question. These analysis results are then passed to the conversational agent in the next step.
[0115] Step 4:
[0116] The conversational agent running on the server receives the analysis results and uses natural language processing techniques to further analyze the question. At this stage, the specific data processing involves identifying the origin and destination from the question; the input data is the analysis results, and the output data is the identified origin and destination. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent identifies "Shibuya" and "Shinjuku".
[0117] Step 5:
[0118] The conversational agent searches an information database based on the identified origin and destination. This database stores route information, and a search process is performed to obtain the optimal route information. The input data is the origin and destination, and the output data is the optimal route information. For example, information such as, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," might be output.
[0119] Step 6:
[0120] Based on the acquired route information, the conversational agent generates a response to send back to the user in natural language. At this stage, the input data is the optimal route information, and the output data is a natural language response. For example, a response in the form of, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," is generated.
[0121] Step 7:
[0122] The server sends the generated response back to the communication device. Specifically, the process involves sending the response to the communication device over the network. The input data is the response in natural language, and the output data is the transmission request to the communication device.
[0123] Step 8:
[0124] The communication device receives the response message sent back from the server and displays it in the chat window. The user can then review this response message and easily obtain the necessary routing information. The input data is the response message from the server, and the output data is the response displayed in the chat window.
[0125] The above outlines the specific processing steps of this system and the details of the inputs and outputs at each step.
[0126] (Application Example 1)
[0127] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0128] In food delivery operations, a challenge exists in that delivery personnel have difficulty quickly and efficiently obtaining the optimal route. Traditional systems require considerable effort from delivery personnel to acquire route information, resulting in delivery delays and inefficiencies. Furthermore, their ability to provide accurate route information instantly in response to natural language inquiries is limited.
[0129] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0130] In this invention, the server includes means for a user to input a question in natural language via a communication terminal, means for sending the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and means configured to obtain optimal route information when the user is performing food delivery work. This enables delivery personnel to quickly obtain optimal route information in natural language.
[0131] (Term definition)
[0132] A "user" is someone who uses a communication terminal to input questions in natural language and performs food delivery services.
[0133] A "communication terminal" is a device used by a user, such as a smartphone or smart glasses, that has a means of inputting questions in natural language and sending them to a server.
[0134] "Natural language" refers to the language that humans use in everyday life, and is a format that allows for easy input of questions and instructions, rather than specialized programming languages or commands.
[0135] A "server" is a computing system that receives questions from communication terminals, analyzes their content, generates appropriate responses, and sends them back to the communication terminals.
[0136] A "conversational agent" is a software program that runs within a server, analyzes user inquiries, and generates responses based on appropriate routing information.
[0137] A "response" is information generated by a conversational agent based on a user's question, and is a natural language response that includes routing information.
[0138] "Means of returning a response to a communication terminal and displaying it to the user" refers to a function that returns a response generated by the server to a communication terminal and provides that response to the user via screen or audio.
[0139] "Food delivery services" refer to the business activity of taking orders for food and beverages and delivering them quickly to the customer's designated location.
[0140] "Optimal route information" refers to information about the most efficient route from the user's specified starting point to the destination, taking into account factors such as time, distance, and traffic conditions.
[0141] A "route guidance system" is a system that calculates routes based on maps and traffic information and provides users with the most suitable route information.
[0142] Modes for carrying out the invention
[0143] System Configuration
[0144] This invention provides a system for food delivery services that allows delivery personnel to input questions in natural language and quickly obtain optimal route information. The system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0145] Hardware and software to be used
[0146] Hardware: Communication devices such as smartphones and smart glasses.
[0147] Software: Applications for Android® or iOS devices, conversational agent servers (e.g., IBM Watson®, Google® Dialogflow), and route guidance system servers.
[0148] Detailed explanation of the process
[0149] 1. User input:
[0150] Users can launch an application on their communication terminal and input route-related questions in natural language. For example, they might input, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input can be done via voice or text.
[0151] 2. Sending input to the server:
[0152] The communication terminal sends the entered question content to the server. The server receives this data and passes it on to the conversational agent.
[0153] 3. Dialogue Agent:
[0154] The conversational agent runs on a server and is responsible for analyzing the content of natural language questions it receives. For example, generative AI models such as IBM Watson or Google Dialogflow can be used. The conversational agent identifies the origin and destination from the user's questions and searches the navigation system's database to obtain the best route information.
[0155] 4. Response generation and return:
[0156] The conversational agent generates a response in natural language based on the acquired route information. The generated response is sent back to the communication terminal via the server. For example, a response such as "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes" is generated.
[0157] 5. Display of response:
[0158] The communication terminal receives the returned response and displays it in the chat window. In addition to text display, audio output is also available as a display method.
[0159] Specific examples and prompt statements
[0160] Specific example
[0161] Scenario: A pizza delivery driver checks the shortest route.
[0162] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0163] Analysis results from the conversational agent: Origin: "Pizza house in Shibuya", Destination: "Office in Ebisu"
[0164] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0165] Example prompts for generative AI models
[0166] Prompt: The user typed, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" Provide the best route information for this question in natural language.
[0167] Example data:
[0168] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0169] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0170] This allows delivery drivers in the food delivery industry to quickly and efficiently obtain optimal route information.
[0171] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0172] Detailed program processing steps
[0173] Step 1:
[0174] The user launches an application on their communication terminal and enters a question in natural language. For example, the user might enter, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input is sent to the communication terminal in either text or voice format.
[0175] Input: The user enters a question into the application using natural language.
[0176] Output: User's question (text or audio data)
[0177] Step 2:
[0178] The communication terminal sends the entered question content to the server. During this process, natural language text data and audio data are sent to the server.
[0179] Input: User's question (text or audio data)
[0180] Output: Question content sent to the server
[0181] Step 3:
[0182] The server receives the question and passes it to the conversational agent. The conversational agent analyzes the user's question. Here, a generative AI model (e.g., IBM Watson, Google Dialogflow) is used to identify the origin and destination from the question. In this analysis process, natural language processing techniques are used to extract the origin and destination and convert them into an appropriate format.
[0183] Input: Question content sent to the server
[0184] Output: Analyzed origin and destination
[0185] Step 4:
[0186] The conversational agent searches the route guidance system's database to obtain the optimal route information based on the analyzed origin ("pizza house in Shibuya") and destination ("office in Ebisu"). This database search process uses traffic information and map information to identify the best route, taking into account factors such as travel time and distance.
[0187] Input: Analyzed origin and destination
[0188] Output: Optimal route information
[0189] Step 5:
[0190] The conversational agent generates a response in natural language based on the acquired route information. This response is generated in a format such as, "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes." Here, the natural language generation function of the generative AI model is used to create an appropriate sentence.
[0191] Input: Optimal route information
[0192] Output: Response with route information generated in natural language
[0193] Step 6:
[0194] The server sends the generated response back to the communication terminal. The response is sent to the communication terminal in either text or audio format.
[0195] Input: Generated natural language response
[0196] Output: Response sent to communication terminal
[0197] Step 7:
[0198] The communication terminal receives a response and displays it to the user. Display methods include text-based chat window display and audio output. The user can review this response and obtain optimal routing information.
[0199] Input: Response sent from the server
[0200] Output: Route information displayed to the user
[0201] This series of processes allows users to quickly obtain optimal route information in response to questions entered in natural language, enabling them to efficiently carry out food delivery operations.
[0202] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0203] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0204] System-wide configuration
[0205] This system consists of a communication terminal used by the user, a server that receives requests, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotions.
[0206] Communication terminal
[0207] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0208] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0209] server
[0210] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0211] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs. Furthermore, the server adjusts the content of its responses based on sentiment data analyzed by the sentiment engine.
[0212] Dialogue Agent
[0213] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0214] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0215] Emotional Engine
[0216] The emotion engine analyzes the user's input and recognizes their emotional state. For example, if the user uses language that indicates frustration, the emotion engine will recognize that emotion as "anger." This emotional information is passed to the conversational agent and reflected in the generated response.
[0217] Response adjustment and return
[0218] The server adjusts the responses generated by the conversational agent based on the analysis results from the emotion engine. For example, if the user is irritated, the server may adjust the tone of the response to be calmer.
[0219] The adjusted response is sent back from the server to the communication terminal. The communication terminal receives this response and displays it in the chat window.
[0220] Specific examples
[0221] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication device and presses the send button.
[0222] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey takes approximately 10 minutes."
[0223] Simultaneously, the emotion engine within the server recognizes the user's emotion of urgency from their expression, "I'm in a hurry!" Based on this emotion information, a response tone that is more responsive to the urgent situation is set.
[0224] This adjusted response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0225] The above describes an embodiment of a system incorporating an emotion engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[0226] The following describes the processing flow.
[0227] Step 1:
[0228] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0229] Step 2:
[0230] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0231] Step 3:
[0232] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0233] Step 4:
[0234] The server receives the request. The received data is converted into an appropriate format for analysis.
[0235] Step 5:
[0236] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0237] Step 6:
[0238] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0239] Step 7:
[0240] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0241] Step 8:
[0242] The server passes the response from the conversational agent to the emotion engine. At the same time, the user's input is also passed.
[0243] Step 9:
[0244] The emotion engine on the server analyzes the user's input and recognizes their emotional state. For example, it recognizes the emotion of urgency from the expression "I'm in a hurry!"
[0245] Step 10:
[0246] The emotion engine adjusts its response based on the emotional information it recognizes. For example, it might change the tone to one that indicates a more urgent response.
[0247] Step 11:
[0248] The adjusted response is returned to the server. The server converts this response into JSON data.
[0249] Step 12:
[0250] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0251] Step 13:
[0252] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0253] Step 14:
[0254] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0255] The above outlines the specific processing steps of this system.
[0256] (Example 2)
[0257] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0258] Currently, systems that allow users to obtain route information via communication terminals often fail to consider the user's emotional state, resulting in decreased user satisfaction. Furthermore, the inability to respond flexibly to user emotions makes it difficult to provide appropriate service, especially when the user is in a hurry or irritated. Improving this situation and enhancing user satisfaction is crucial.
[0259] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language via a communication terminal, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, means for using an emotion analysis engine to recognize the user's emotions, and means for adjusting the response content based on the emotion information recognized by the emotion analysis engine. This makes it possible to provide a detailed response that corresponds to the user's emotions.
[0260] A "communication terminal" is a device used by users to input questions and communicate with a server, and specifically refers to devices such as smartphones and personal computers.
[0261] A "server" is a computer device that receives questions from users, analyzes the messages, generates responses, and sends them back to the users.
[0262] A "question" refers to an inquiry about routing information or other related matters that a user inputs in natural language through a communication terminal.
[0263] A "dialogue agent" is a software module that analyzes the content of a user's question and generates an appropriate response, using natural language processing technology.
[0264] A "response" is a natural language response generated by a conversational agent and sent back to the user via the server.
[0265] An "emotion analysis engine" is a software module that recognizes the user's emotional state from their input and assigns emotional labels such as "joy," "anger," and "sadness."
[0266] "Emotional information" refers to data representing the user's emotional state, which is recognized by the emotion analysis engine and used to adjust responses by the dialogue agent.
[0267] A "route information provision system" is a database and software module that provides optimal route information based on a specific origin and destination.
[0268] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion analysis engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0269] System-wide configuration
[0270] This system consists of the following hardware and software:
[0271] Communication devices used by the user (smartphones and personal computers)
[0272] Server that receives and analyzes messages
[0273] A conversational agent that generates responses to user questions.
[0274] A sentiment analysis engine that recognizes user emotions.
[0275] Communication terminal
[0276] Users can open a chat window using their communication device and enter questions in natural language. For example, they can type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends this question to the server.
[0277] server
[0278] The server receives and analyzes the content of questions sent from communication terminals. The received messages are passed to the conversational agent. The server also uses an emotion analysis engine to analyze the user's emotions and passes the results to the conversational agent. The server automatically processes everything from receiving requests to generating and sending responses, and also has a function to save received messages as logs.
[0279] Dialogue Agent
[0280] The conversational agent runs on the server and analyzes the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question, then searches the route information system (database) to obtain the optimal route information. The obtained route information is then converted into an appropriate natural language response by the conversational agent.
[0281] Emotion analysis engine
[0282] The emotion analysis engine analyzes the user's input and recognizes their emotional state. For example, if the user uses the expression "I'm in a hurry!", the emotion analysis engine recognizes that the user is feeling anxious. The analyzed emotional information is passed to the conversational agent and reflected in the generated response.
[0283] Response adjustment and return
[0284] The server adjusts the response generated by the dialogue agent based on the analysis results of the sentiment analysis engine. For example, when the user is impatient, adjustments such as calming down the tone of the response are made. The adjusted response is returned from the server to the communication terminal, and the communication terminal receives this response and displays it in the chat window.
[0285] Specific examples
[0286] Here, specific examples are given. The user enters "Please tell me the route from Tokyo Station to Shinagawa Station. I'm in a hurry!" in the chat window of the communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the dialogue agent. The dialogue agent searches the route information providing system and generates a response such as "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The required time is about 10 minutes." At the same time, the sentiment analysis engine recognizes the user's anxiety from the expression "I'm in a hurry!" and adjusts the response tone based on this sentiment information. This adjusted response is returned to the server and then sent to the communication terminal. The communication terminal displays the received response in the chat window, and the user can confirm the answer.
[0287] The above is an embodiment of the system combined with the sentiment analysis engine based on the present invention. By applying this system, the user can not only easily and quickly obtain route information, but also get a response according to their own feelings
[0288] The flow of specific processing in Example 2 will be described using FIG. 13.
[0289] Step 1: The user enters a question
[0290] The user enters a question into the chat window of their communication device. This question is in natural language. For example, they might enter, "What is the shortest route from Shibuya to Shinjuku?" The input data is in text format. Specifically, the user uses a smartphone or computer keyboard to type the text and clicks the send button. The input content is stored as text data on the communication device.
[0291] Step 2: Send the question to the server
[0292] The terminal sends the user's input to the server. The input data is in text format, and the output is a data packet sent to the server. Specifically, the terminal converts the input text into a data packet format and sends it to the server over the internet. The data is transmitted according to the TCP / IP protocol.
[0293] Step 3: The server receives the question and analyzes it.
[0294] The server receives the question content sent from the terminal and parses the message content. The input data is the text data received from the terminal, and the output is the parsed question content. Specifically, the server receives the data packet and converts it to text format. Next, it parses the question content using natural language processing technology and saves it as a log.
[0295] Step 4: Pass the question to the conversational agent.
[0296] The server passes the analyzed question content to the conversational agent. The input data is the analyzed question content, and the output is data for the conversational agent to generate a response. Specifically, the server makes an API call to the conversational agent and passes the analyzed question content. The conversational agent then starts generating a response based on this content.
[0297] Step 5: Analyze emotions with an emotion analysis engine
[0298] The server sends the user's question to the sentiment analysis engine, which then analyzes the emotional state. The input data is the user's question, and the output is data with emotional labels. Specifically, the server makes an API call to the sentiment analysis engine and sends the input text. The sentiment analysis engine analyzes the emotions from the text and assigns emotional labels such as "joy," "anger," and "sadness."
[0299] Step 6: The conversational agent generates a response.
[0300] The conversational agent identifies the origin and destination from the question, searches a route information system, and generates the optimal route information. The input data is the question, and the output is a response in natural language. Specifically, the conversational agent queries the database, retrieves appropriate route information, and converts that information into natural language.
[0301] Step 7: The server adjusts its response based on the sentiment analysis results.
[0302] The server combines the responses received from the conversational agent with the results of the emotion analysis engine to adjust the final response. The input data is the generated response and emotion information, and the output is the adjusted response. Specifically, the server applies emotion-specific adjustment logic to adjust the tone and content of the response. For example, if the user is irritated, the tone of the response will be softened.
[0303] Step 8: Send the adjusted response back to the terminal.
[0304] The server sends a pre-arranged response to the communication terminal. The input data is the pre-arranged response, and the output is the data packet sent to the terminal. Specifically, the server converts the response into a packet format and sends it to the terminal over the internet.
[0305] Step 9: The terminal displays a response
[0306] The terminal displays the response received from the server in the chat window. The input data is the text data received from the server, and the output is the response display that the user can view. As a specific operation, the terminal converts the received data packet into text format and displays it in the chat window. The user can check this to obtain the route information.
[0307] (Application Example 2)
[0308] Next, Application Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0309] In modern logistics centers, a large number of packages and goods need to be processed quickly and accurately. However, it is not easy for administrators to efficiently check the positions and delivery routes of packages and make optimal decisions. In addition, since the emotions and urgency of administrators are ignored, the user experience and efficiency may decline. Due to such situations, there are problems such as work delays and mistakes, resulting in a decline in overall business efficiency.
[0310] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for a user to input a question in natural language, means for transmitting the question to the server, means for the server to receive the question and generate a response by an interactive agent, means for returning the response to the communication terminal and displaying it to the user, and an emotion recognition engine for recognizing the user's emotional state and adjusting the response content based on the emotional state. Thereby, the administrator of the logistics center can quickly obtain information on the position of packages and the optimal delivery route using natural language, and it becomes possible to provide an optimal response according to the emotional state of the administrator.
[0311] The "communication terminal" is a device for a user to input a question in natural language and transmit it to the server.
[0312] A "server" is a computer system that receives questions from users, generates responses using dialogue agents, and sends them back to the communication terminal.
[0313] A "conversational agent" is a program that runs within a server, analyzes user questions, and generates appropriate responses.
[0314] An "emotion recognition engine" is a program that analyzes the user's emotional state from their natural language input and adjusts its response based on the results.
[0315] "Adjusting the response content" means changing the tone and details of the generated response based on the user's emotional state.
[0316] "Logistics" is a general term for the business processes involved in the storage, delivery, and management of packages and goods.
[0317] "Natural language" refers to text and audio composed of the language that users use on a daily basis.
[0318] "Means of sending questions" refers to the function of sending user questions to the server via a communication terminal.
[0319] "Means of sending and displaying a response" refers to a function that receives a response from a server on a communication terminal and displays it to the user.
[0320] "Route information" refers to data about the optimal route between a specified origin and destination.
[0321] This invention relates to a system that allows managers in logistics centers to input questions in natural language using a smartphone and receive prompt and appropriate route information in response. This system recognizes the user's emotional state and adjusts its responses accordingly to provide more effective support. Specific embodiments for carrying out this invention are described in detail below.
[0322] System-wide configuration
[0323] This system consists of a communication terminal used by the user, a server that receives questions, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotional state.
[0324] Communication terminal
[0325] The user launches the logistics management app using a communication device (e.g., a smartphone). In the app's chat window, the user can input questions about the package's location and delivery route using natural language. For example, they might input, "I need to send this package to point A as soon as possible, which route is best?" When the user clicks the "Send" button, the communication device sends the entered question to the server.
[0326] server
[0327] The server receives questions sent from communication terminals and forwards requests to conversational agents. The server also uses an emotion recognition engine to analyze the user's emotional state and adjusts the conversational agent's responses based on the results. Software used within the server includes Flask (web framework) and OpenAI® (natural language processing engine).
[0328] Dialogue Agent
[0329] The conversational agent runs on the server and is responsible for analyzing the user's questions. Specifically, it uses the OpenAI API to generate optimal route information from the question. For example, if a user asks, "I urgently need to send this package to point A, which route is best?", the conversational agent searches a logistics information database and obtains the optimal route information. This route information is then converted into a natural language response by the conversational agent.
[0330] Emotional Engine
[0331] The emotion engine analyzes the user's input and recognizes their emotional state. For example, it recognizes the user's anxiety from expressions like "urgent" or "in a hurry." This emotional information is passed to the conversational agent and used to adjust the response content and tone.
[0332] Response adjustment and return
[0333] The server adjusts the responses generated by the conversational agent based on the analysis results of the emotion engine. For example, if the user is in a hurry, a message in a gentler tone, such as "The best route is as follows. Thank you for your understanding," may be added. The adjusted response is sent back from the server to the communication terminal and displayed to the user.
[0334] Specific examples
[0335] The user types "I urgently need to send this package to point A, what's the best route?" into the chat window of the communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the logistics information database and generates a response: "The best route is as follows. Thank you for your understanding." At the same time, the emotion engine in the server recognizes the user's urgency from the expression "urgent." Based on this emotion information, the tone of the response is adjusted. This adjusted response is returned to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0336] Examples of prompts to input into a generative AI model:
[0337] Provide the best route for the query: "I urgently need to send this package to point A. What is the best route?"
[0338] In this way, the present invention can improve work efficiency in logistics centers and reduce stress for managers.
[0339] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0340] Step 1:
[0341] The user uses the chat window on the communication terminal to input a question in natural language. For example, a question like, "I urgently need to send this package to point A, what is the best route?" might be entered. This input is converted into digital data within the communication terminal.
[0342] Step 2:
[0343] When the user clicks the "Send" button, the communication device sends this digital data to the server. The transmitted data includes the user's question.
[0344] Step 3:
[0345] The server receives a question from the user and passes the request to the conversational agent. The conversational agent analyzes this data to identify the question (the location and destination of the package). Natural language processing techniques are used for this analysis.
[0346] Step 4:
[0347] The conversational agent searches the logistics information database to obtain the optimal route information. An example of a prompt for this database search is: "Provide the best route for the query: 'I urgently need to send this package to point A, which route is best?'" The generated route information is returned to the conversational agent.
[0348] Step 5:
[0349] The emotion recognition engine on the server analyzes the user's emotional state from their input. In this example, the expression "urgent" is used to recognize that the user is in a hurry. The emotion recognition engine then generates this emotional data.
[0350] Step 6:
[0351] The server adjusts the response generated by the dialogue agent based on emotion data from the emotion recognition engine. A gentler tone or additional sentences depending on the urgency are added to the response. For example, a response like, "Thank you for your understanding. The optimal route is as follows..." might be created.
[0352] Step 7:
[0353] The server sends a coordinated response back to the communication terminal. The returned data includes coordinated routing information.
[0354] Step 8:
[0355] The communication terminal receives a response from the server and displays it in the chat window. This allows the user to quickly confirm the optimal route information.
[0356] Through the steps described above, the present invention constructs a system that supports the efficient execution of tasks by managers in logistics centers and provides rapid and appropriate route information.
[0357] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0358] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0359] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0360] [Second Embodiment]
[0361] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0362] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0363] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0364] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0365] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0366] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0367] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0368] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0369] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0370] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0371] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0372] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0373] The present invention provides a system that allows users to input questions in natural language and quickly obtain route information. The embodiments thereof are described in detail below.
[0374] System-wide configuration
[0375] This system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0376] Communication terminal
[0377] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0378] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0379] server
[0380] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0381] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs.
[0382] Dialogue Agent
[0383] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0384] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0385] Sending and displaying responses
[0386] The server sends the response generated by the conversational agent back to the communication terminal. The communication terminal receives this response and displays it in the chat window. The user can then review the displayed response.
[0387] Specific examples
[0388] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station" into the chat window of their communication terminal and presses the send button.
[0389] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent.
[0390] The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please take the JR Yamanote Line. The journey takes approximately 10 minutes."
[0391] This response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0392] The above describes an embodiment of the system based on the present invention. By applying this system, users will be able to obtain route information easily and quickly.
[0393] The following describes the processing flow.
[0394] Step 1:
[0395] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0396] Step 2:
[0397] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0398] Step 3:
[0399] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0400] Step 4:
[0401] The server receives the request. The received data is converted into an appropriate format for analysis.
[0402] Step 5:
[0403] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0404] Step 6:
[0405] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0406] Step 7:
[0407] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0408] Step 8:
[0409] The generated response is returned to the server. The server converts this response into JSON data.
[0410] Step 9:
[0411] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0412] Step 10:
[0413] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0414] Step 11:
[0415] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0416] The above outlines the specific processing steps of this system's program.
[0417] (Example 1)
[0418] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0419] In modern society, it is crucial for users to quickly and accurately obtain appropriate route information based on their location and destination. However, conventional systems often take a long time to process questions entered in natural language, resulting in inadequate responses. To solve this problem, a system is needed that can properly analyze user questions and quickly provide optimal route information.
[0420] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0421] In this invention, the server includes means for a user to input a question in natural language via a communication device, means for transmitting the question to a data processing device, means for the data processing device to receive the question and generate a response using a conversational agent, means for sending the response back to the communication device and displaying it to the user, and includes question input at the communication device, question analysis at the data processing device, retrieval of an information database, and response generation. This makes it possible for the user to obtain route information easily and quickly.
[0422] A "communication device" is a device used by a user to input questions via an interface, and includes smartphones and personal computers.
[0423] A "data processing device" is a device that receives data sent by a user, analyzes it, and generates a response; it generally refers to a server.
[0424] A "conversational agent" is a program or system that analyzes a user's questions and generates appropriate responses, using natural language processing technology.
[0425] A "response" refers to the information generated by the conversational agent in response to a user's question, and it is presented in natural language.
[0426] An "information database" is a system for storing and managing necessary data, including route information and map information.
[0427] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.
[0428] The present invention provides a system that allows a user to input a question in natural language via a communication device and quickly obtain route information based on that question. Specific embodiments thereof are described in detail below.
[0429] Communication device
[0430] Users open a dedicated chat window using a communication device such as a smartphone or personal computer. In this chat window, users can input questions about their current location and destination using natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends the entered question to a data processing device (server).
[0431] Data processing device (server)
[0432] The data processing unit receives requests sent from communication devices and has the function of analyzing these requests. The received question content is passed to a conversation agent running on the server. The data processing unit automatically handles everything from receiving requests to generating and sending responses, and also has the function of saving received messages as logs.
[0433] Conversation Agent
[0434] The conversational agent's role is to analyze the content of questions sent by users using natural language processing technology. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the information database to obtain the optimal route information. The obtained route information is converted into a response to be sent back to the user in natural language. For example, it might say, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes."
[0435] Sending and displaying responses
[0436] The data processing unit (server) sends the response generated by the conversation agent back to the communication device. The communication device receives this response and displays it in the chat window. The user can review the displayed response and obtain the necessary routing information.
[0437] Specific examples
[0438] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station" into the chat window of the communication device and clicks the send button. The communication device sends this question to the data processing device. The data processing device receives the request and passes the question to the conversational agent. The conversational agent searches the information database and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the Y Line of the railway. The journey takes approximately 10 minutes." This response is returned to the data processing device and then sent back to the communication device. The communication device displays the received response in the chat window, allowing the user to confirm the answer.
[0439] Example of a prompt
[0440] Examples of prompts to input into a generative AI model include the following:
[0441] "Describe the sequence of events in which the system generates responses to questions submitted by users. Clearly define the roles of the server, communication device, and conversational agent."
[0442] The above describes the embodiments for carrying out the present invention. With this system, users can easily and quickly obtain route information.
[0443] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0444] Step 1:
[0445] The user opens a chat window via a communication device. This device could be a smartphone or a personal computer. Once the chat window appears, the user enters a question in natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" This question is then entered into the communication device.
[0446] Step 2:
[0447] When the user clicks the "Submit" button, the communication device sends the entered question content to the data processing device (server). Specifically, the communication device sends the question text to the server via the network. At this time, the input data is the user's question, and the output data is the question sent as a request to the server.
[0448] Step 3:
[0449] The server receives requests sent from communication devices. First, the server analyzes the received question. This analysis involves inputting the user's question as received data into a text analysis engine, which outputs analysis results that help understand the format and structure of the question. These analysis results are then passed to the conversational agent in the next step.
[0450] Step 4:
[0451] The conversational agent running on the server receives the analysis results and uses natural language processing techniques to further analyze the question. At this stage, the specific data processing involves identifying the origin and destination from the question; the input data is the analysis results, and the output data is the identified origin and destination. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent identifies "Shibuya" and "Shinjuku".
[0452] Step 5:
[0453] The conversational agent searches an information database based on the identified origin and destination. This database stores route information, and a search process is performed to obtain the optimal route information. The input data is the origin and destination, and the output data is the optimal route information. For example, information such as, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," might be output.
[0454] Step 6:
[0455] Based on the acquired route information, the conversational agent generates a response to send back to the user in natural language. At this stage, the input data is the optimal route information, and the output data is a natural language response. For example, a response in the form of, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," is generated.
[0456] Step 7:
[0457] The server sends the generated response back to the communication device. Specifically, the process involves sending the response to the communication device over the network. The input data is the response in natural language, and the output data is the transmission request to the communication device.
[0458] Step 8:
[0459] The communication device receives the response message sent back from the server and displays it in the chat window. The user can then review this response message and easily obtain the necessary routing information. The input data is the response message from the server, and the output data is the response displayed in the chat window.
[0460] The above outlines the specific processing steps of this system and the details of the inputs and outputs at each step.
[0461] (Application Example 1)
[0462] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0463] In food delivery operations, a challenge exists in that delivery personnel have difficulty quickly and efficiently obtaining the optimal route. Traditional systems require considerable effort from delivery personnel to acquire route information, resulting in delivery delays and inefficiencies. Furthermore, their ability to provide accurate route information instantly in response to natural language inquiries is limited.
[0464] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0465] In this invention, the server includes means for a user to input a question in natural language via a communication terminal, means for sending the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and means configured to obtain optimal route information when the user is performing food delivery work. This enables delivery personnel to quickly obtain optimal route information in natural language.
[0466] (Term definition)
[0467] A "user" is someone who uses a communication terminal to input questions in natural language and performs food delivery services.
[0468] A "communication terminal" is a device used by a user, such as a smartphone or smart glasses, that has a means of inputting questions in natural language and sending them to a server.
[0469] "Natural language" refers to the language that humans use in everyday life, and is a format that allows for easy input of questions and instructions, rather than specialized programming languages or commands.
[0470] A "server" is a computing system that receives questions from communication terminals, analyzes their content, generates appropriate responses, and sends them back to the communication terminals.
[0471] A "conversational agent" is a software program that runs within a server, analyzes user inquiries, and generates responses based on appropriate routing information.
[0472] A "response" is information generated by a conversational agent based on a user's question, and is a natural language response that includes routing information.
[0473] "Means of returning a response to a communication terminal and displaying it to the user" refers to a function that returns a response generated by the server to a communication terminal and provides that response to the user via screen or audio.
[0474] "Food delivery services" refer to the business activity of taking orders for food and beverages and delivering them quickly to the customer's designated location.
[0475] "Optimal route information" refers to information about the most efficient route from the user's specified starting point to the destination, taking into account factors such as time, distance, and traffic conditions.
[0476] A "route guidance system" is a system that calculates routes based on maps and traffic information and provides users with the most suitable route information.
[0477] Modes for carrying out the invention
[0478] System Configuration
[0479] This invention provides a system for food delivery services that allows delivery personnel to input questions in natural language and quickly obtain optimal route information. The system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0480] Hardware and software to be used
[0481] Hardware: Communication devices such as smartphones and smart glasses.
[0482] Software: Applications for Android or iOS devices, conversational agent servers (e.g., IBM Watson, Google Dialogflow), and route guidance system servers.
[0483] Detailed explanation of the process
[0484] 1. User input:
[0485] Users can launch an application on their communication terminal and input route-related questions in natural language. For example, they might input, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input can be done via voice or text.
[0486] 2. Sending input to the server:
[0487] The communication terminal sends the entered question content to the server. The server receives this data and passes it on to the conversational agent.
[0488] 3. Dialogue Agent:
[0489] The conversational agent runs on a server and is responsible for analyzing the content of natural language questions it receives. For example, generative AI models such as IBM Watson or Google Dialogflow can be used. The conversational agent identifies the origin and destination from the user's questions and searches the navigation system's database to obtain the best route information.
[0490] 4. Response generation and return:
[0491] The conversational agent generates a response in natural language based on the acquired route information. The generated response is sent back to the communication terminal via the server. For example, a response such as "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes" is generated.
[0492] 5. Display of response:
[0493] The communication terminal receives the returned response and displays it in the chat window. In addition to text display, audio output is also available as a display method.
[0494] Specific examples and prompt statements
[0495] Specific example
[0496] Scenario: A pizza delivery driver checks the shortest route.
[0497] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0498] Analysis results from the conversational agent: Origin: "Pizza house in Shibuya", Destination: "Office in Ebisu"
[0499] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0500] Example prompts for generative AI models
[0501] Prompt: The user typed, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" Provide the best route information for this question in natural language.
[0502] Example data:
[0503] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0504] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0505] This allows delivery drivers in the food delivery industry to quickly and efficiently obtain optimal route information.
[0506] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0507] Detailed program processing steps
[0508] Step 1:
[0509] The user launches an application on their communication terminal and enters a question in natural language. For example, the user might enter, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input is sent to the communication terminal in either text or voice format.
[0510] Input: The user enters a question into the application using natural language.
[0511] Output: User's question (text or audio data)
[0512] Step 2:
[0513] The communication terminal sends the entered question content to the server. During this process, natural language text data and audio data are sent to the server.
[0514] Input: User's question (text or audio data)
[0515] Output: Question content sent to the server
[0516] Step 3:
[0517] The server receives the question and passes it to the conversational agent. The conversational agent analyzes the user's question. Here, a generative AI model (e.g., IBM Watson, Google Dialogflow) is used to identify the origin and destination from the question. In this analysis process, natural language processing techniques are used to extract the origin and destination and convert them into an appropriate format.
[0518] Input: Question content sent to the server
[0519] Output: Analyzed origin and destination
[0520] Step 4:
[0521] The conversational agent searches the route guidance system's database to obtain the optimal route information based on the analyzed origin ("pizza house in Shibuya") and destination ("office in Ebisu"). This database search process uses traffic information and map information to identify the best route, taking into account factors such as travel time and distance.
[0522] Input: Analyzed origin and destination
[0523] Output: Optimal route information
[0524] Step 5:
[0525] The conversational agent generates a response in natural language based on the acquired route information. This response is generated in a format such as, "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes." Here, the natural language generation function of the generative AI model is used to create an appropriate sentence.
[0526] Input: Optimal route information
[0527] Output: Response with route information generated in natural language
[0528] Step 6:
[0529] The server sends the generated response back to the communication terminal. The response is sent to the communication terminal in either text or audio format.
[0530] Input: Generated natural language response
[0531] Output: Response sent to communication terminal
[0532] Step 7:
[0533] The communication terminal receives a response and displays it to the user. Display methods include text-based chat window display and audio output. The user can review this response and obtain optimal routing information.
[0534] Input: Response sent from the server
[0535] Output: Route information displayed to the user
[0536] This series of processes allows users to quickly obtain optimal route information in response to questions entered in natural language, enabling them to efficiently carry out food delivery operations.
[0537] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0538] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0539] System-wide configuration
[0540] This system consists of a communication terminal used by the user, a server that receives requests, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotions.
[0541] Communication terminal
[0542] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0543] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0544] server
[0545] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0546] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs. Furthermore, the server adjusts the content of its responses based on sentiment data analyzed by the sentiment engine.
[0547] Dialogue Agent
[0548] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0549] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0550] Emotional Engine
[0551] The emotion engine analyzes the user's input and recognizes their emotional state. For example, if the user uses language that indicates frustration, the emotion engine will recognize that emotion as "anger." This emotional information is passed to the conversational agent and reflected in the generated response.
[0552] Response adjustment and return
[0553] The server adjusts the responses generated by the conversational agent based on the analysis results from the emotion engine. For example, if the user is irritated, the server may adjust the tone of the response to be calmer.
[0554] The adjusted response is sent back from the server to the communication terminal. The communication terminal receives this response and displays it in the chat window.
[0555] Specific examples
[0556] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication device and presses the send button.
[0557] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey takes approximately 10 minutes."
[0558] Simultaneously, the emotion engine within the server recognizes the user's emotion of urgency from their expression, "I'm in a hurry!" Based on this emotion information, a response tone that is more responsive to the urgent situation is set.
[0559] This adjusted response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0560] The above describes an embodiment of a system incorporating an emotion engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[0561] The following describes the processing flow.
[0562] Step 1:
[0563] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0564] Step 2:
[0565] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0566] Step 3:
[0567] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0568] Step 4:
[0569] The server receives the request. The received data is converted into an appropriate format for analysis.
[0570] Step 5:
[0571] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0572] Step 6:
[0573] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0574] Step 7:
[0575] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0576] Step 8:
[0577] The server passes the response from the conversational agent to the emotion engine. At the same time, the user's input is also passed.
[0578] Step 9:
[0579] The emotion engine on the server analyzes the user's input and recognizes their emotional state. For example, it recognizes the emotion of urgency from the expression "I'm in a hurry!"
[0580] Step 10:
[0581] The emotion engine adjusts its response based on the emotional information it recognizes. For example, it might change the tone to one that indicates a more urgent response.
[0582] Step 11:
[0583] The adjusted response is returned to the server. The server converts this response into JSON data.
[0584] Step 12:
[0585] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0586] Step 13:
[0587] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0588] Step 14:
[0589] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0590] The above outlines the specific processing steps of this system.
[0591] (Example 2)
[0592] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0593] Currently, systems that allow users to obtain route information via communication terminals often fail to consider the user's emotional state, resulting in decreased user satisfaction. Furthermore, the inability to respond flexibly to user emotions makes it difficult to provide appropriate service, especially when the user is in a hurry or irritated. Improving this situation and enhancing user satisfaction is crucial.
[0594] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language via a communication terminal, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, means for using an emotion analysis engine to recognize the user's emotions, and means for adjusting the response content based on the emotion information recognized by the emotion analysis engine. This makes it possible to provide a detailed response that corresponds to the user's emotions.
[0595] A "communication terminal" is a device used by users to input questions and communicate with a server, and specifically refers to devices such as smartphones and personal computers.
[0596] A "server" is a computer device that receives questions from users, analyzes the messages, generates responses, and sends them back to the users.
[0597] A "question" refers to an inquiry about routing information or other related matters that a user inputs in natural language through a communication terminal.
[0598] A "dialogue agent" is a software module that analyzes the content of a user's question and generates an appropriate response, using natural language processing technology.
[0599] A "response" is a natural language response generated by a conversational agent and sent back to the user via the server.
[0600] An "emotion analysis engine" is a software module that recognizes the user's emotional state from their input and assigns emotional labels such as "joy," "anger," and "sadness."
[0601] "Emotional information" refers to data representing the user's emotional state, which is recognized by the emotion analysis engine and used to adjust responses by the dialogue agent.
[0602] A "route information provision system" is a database and software module that provides optimal route information based on a specific origin and destination.
[0603] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion analysis engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0604] System-wide configuration
[0605] This system consists of the following hardware and software:
[0606] Communication devices used by the user (smartphones and personal computers)
[0607] Server that receives and analyzes messages
[0608] A conversational agent that generates responses to user questions.
[0609] A sentiment analysis engine that recognizes user emotions.
[0610] Communication terminal
[0611] Users can open a chat window using their communication device and enter questions in natural language. For example, they can type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends this question to the server.
[0612] server
[0613] The server receives and analyzes the content of questions sent from communication terminals. The received messages are passed to the conversational agent. The server also uses an emotion analysis engine to analyze the user's emotions and passes the results to the conversational agent. The server automatically processes everything from receiving requests to generating and sending responses, and also has a function to save received messages as logs.
[0614] Dialogue Agent
[0615] The conversational agent runs on the server and analyzes the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question, then searches the route information system (database) to obtain the optimal route information. The obtained route information is then converted into an appropriate natural language response by the conversational agent.
[0616] Emotion analysis engine
[0617] The emotion analysis engine analyzes the user's input and recognizes their emotional state. For example, if the user uses the expression "I'm in a hurry!", the emotion analysis engine recognizes that the user is feeling anxious. The analyzed emotional information is passed to the conversational agent and reflected in the generated response.
[0618] Response adjustment and return
[0619] The server adjusts the responses generated by the conversational agent based on the results of the emotion analysis engine. For example, if the user is irritated, the server adjusts the tone of the response to be calmer. The adjusted response is sent back from the server to the communication terminal, which receives the response and displays it in the chat window.
[0620] Specific examples
[0621] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the route information system and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey will take approximately 10 minutes." At the same time, the sentiment analysis engine recognizes the user's urgency from the expression "I'm in a hurry!" and adjusts the response tone based on this sentiment information. This adjusted response is returned to the server and sent back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0622] The above describes an embodiment of a system incorporating an emotion analysis engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[0623] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0624] Step 1: The user enters the question.
[0625] The user enters a question into the chat window of their communication device. This question is in natural language. For example, they might enter, "What is the shortest route from Shibuya to Shinjuku?" The input data is in text format. Specifically, the user uses a smartphone or computer keyboard to type the text and clicks the send button. The input content is stored as text data on the communication device.
[0626] Step 2: Send the question to the server
[0627] The terminal sends the user's input to the server. The input data is in text format, and the output is a data packet sent to the server. Specifically, the terminal converts the input text into a data packet format and sends it to the server over the internet. The data is transmitted according to the TCP / IP protocol.
[0628] Step 3: The server receives the question and analyzes it.
[0629] The server receives the question content sent from the terminal and parses the message content. The input data is the text data received from the terminal, and the output is the parsed question content. Specifically, the server receives the data packet and converts it to text format. Next, it parses the question content using natural language processing technology and saves it as a log.
[0630] Step 4: Pass the question to the conversational agent.
[0631] The server passes the analyzed question content to the conversational agent. The input data is the analyzed question content, and the output is data for the conversational agent to generate a response. Specifically, the server makes an API call to the conversational agent and passes the analyzed question content. The conversational agent then starts generating a response based on this content.
[0632] Step 5: Analyze emotions with an emotion analysis engine
[0633] The server sends the user's question to the sentiment analysis engine, which then analyzes the emotional state. The input data is the user's question, and the output is data with emotional labels. Specifically, the server makes an API call to the sentiment analysis engine and sends the input text. The sentiment analysis engine analyzes the emotions from the text and assigns emotional labels such as "joy," "anger," and "sadness."
[0634] Step 6: The conversational agent generates a response.
[0635] The conversational agent identifies the origin and destination from the question, searches a route information system, and generates the optimal route information. The input data is the question, and the output is a response in natural language. Specifically, the conversational agent queries the database, retrieves appropriate route information, and converts that information into natural language.
[0636] Step 7: The server adjusts its response based on the sentiment analysis results.
[0637] The server combines the responses received from the conversational agent with the results of the emotion analysis engine to adjust the final response. The input data is the generated response and emotion information, and the output is the adjusted response. Specifically, the server applies emotion-specific adjustment logic to adjust the tone and content of the response. For example, if the user is irritated, the tone of the response will be softened.
[0638] Step 8: Send the adjusted response back to the terminal.
[0639] The server sends a pre-arranged response to the communication terminal. The input data is the pre-arranged response, and the output is the data packet sent to the terminal. Specifically, the server converts the response into a packet format and sends it to the terminal over the internet.
[0640] Step 9: The terminal displays a response
[0641] The terminal displays the response received from the server in the chat window. The input data is text data received from the server, and the output is the response display that the user can see. Specifically, the terminal converts the received data packet into text format and displays it in the chat window. The user can then verify this and obtain routing information.
[0642] (Application Example 2)
[0643] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0644] Modern logistics centers require the rapid and accurate processing of a large number of packages and shipments. However, it is not easy for managers to efficiently track package locations and delivery routes and make optimal decisions. Furthermore, disregarding managers' emotions and urgency can lead to a poor user experience and reduced efficiency. This situation results in delays and errors, and a decline in overall operational efficiency.
[0645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and an emotion recognition engine that recognizes the user's emotional state and adjusts the response content based on the emotional state. As a result, the manager of the logistics center can quickly obtain information on the location of packages and the optimal delivery route using natural language, and an optimal response can be provided according to the manager's emotional state.
[0646] A "communication terminal" is a device that allows users to input questions in natural language and send them to a server.
[0647] A "server" is a computer system that receives questions from users, generates responses using dialogue agents, and sends them back to the communication terminal.
[0648] A "conversational agent" is a program that runs within a server, analyzes user questions, and generates appropriate responses.
[0649] An "emotion recognition engine" is a program that analyzes the user's emotional state from their natural language input and adjusts its response based on the results.
[0650] "Adjusting the response content" means changing the tone and details of the generated response based on the user's emotional state.
[0651] "Logistics" is a general term for the business processes involved in the storage, delivery, and management of packages and goods.
[0652] "Natural language" refers to text and audio composed of the language that users use on a daily basis.
[0653] "Means of sending questions" refers to the function of sending user questions to the server via a communication terminal.
[0654] "Means of sending and displaying a response" refers to a function that receives a response from a server on a communication terminal and displays it to the user.
[0655] "Route information" refers to data about the optimal route between a specified origin and destination.
[0656] This invention relates to a system that allows managers in logistics centers to input questions in natural language using a smartphone and receive prompt and appropriate route information in response. This system recognizes the user's emotional state and adjusts its responses accordingly to provide more effective support. Specific embodiments for carrying out this invention are described in detail below.
[0657] System-wide configuration
[0658] This system consists of a communication terminal used by the user, a server that receives questions, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotional state.
[0659] Communication terminal
[0660] The user launches the logistics management app using a communication device (e.g., a smartphone). In the app's chat window, the user can input questions about the package's location and delivery route using natural language. For example, they might input, "I need to send this package to point A as soon as possible, which route is best?" When the user clicks the "Send" button, the communication device sends the entered question to the server.
[0661] server
[0662] The server receives questions sent from communication terminals and forwards requests to conversational agents. The server also uses an emotion recognition engine to analyze the user's emotional state and adjusts the conversational agent's responses based on the results. Software used within the server includes Flask (a web framework) and OpenAI (a natural language processing engine).
[0663] Dialogue Agent
[0664] The conversational agent runs on the server and is responsible for analyzing the user's questions. Specifically, it uses the OpenAI API to generate optimal route information from the question. For example, if a user asks, "I urgently need to send this package to point A, which route is best?", the conversational agent searches a logistics information database and obtains the optimal route information. This route information is then converted into a natural language response by the conversational agent.
[0665] Emotional Engine
[0666] The emotion engine analyzes the user's input and recognizes their emotional state. For example, it recognizes the user's anxiety from expressions like "urgent" or "in a hurry." This emotional information is passed to the conversational agent and used to adjust the response content and tone.
[0667] Response adjustment and return
[0668] The server adjusts the responses generated by the conversational agent based on the analysis results of the emotion engine. For example, if the user is in a hurry, a message in a gentler tone, such as "The best route is as follows. Thank you for your understanding," may be added. The adjusted response is sent back from the server to the communication terminal and displayed to the user.
[0669] Specific examples
[0670] The user types "I urgently need to send this package to point A, what's the best route?" into the chat window of the communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the logistics information database and generates a response: "The best route is as follows. Thank you for your understanding." At the same time, the emotion engine in the server recognizes the user's urgency from the expression "urgent." Based on this emotion information, the tone of the response is adjusted. This adjusted response is returned to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0671] Examples of prompts to input into a generative AI model:
[0672] Provide the best route for the query: "I urgently need to send this package to point A. What is the best route?"
[0673] In this way, the present invention can improve work efficiency in logistics centers and reduce stress for managers.
[0674] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0675] Step 1:
[0676] The user uses the chat window on the communication terminal to input a question in natural language. For example, a question like, "I urgently need to send this package to point A, what is the best route?" might be entered. This input is converted into digital data within the communication terminal.
[0677] Step 2:
[0678] When the user clicks the "Send" button, the communication device sends this digital data to the server. The transmitted data includes the user's question.
[0679] Step 3:
[0680] The server receives a question from the user and passes the request to the conversational agent. The conversational agent analyzes this data to identify the question (the location and destination of the package). Natural language processing techniques are used for this analysis.
[0681] Step 4:
[0682] The conversational agent searches the logistics information database to obtain the optimal route information. An example of a prompt for this database search is: "Provide the best route for the query: 'I urgently need to send this package to point A, which route is best?'" The generated route information is returned to the conversational agent.
[0683] Step 5:
[0684] The emotion recognition engine on the server analyzes the user's emotional state from their input. In this example, the expression "urgent" is used to recognize that the user is in a hurry. The emotion recognition engine then generates this emotional data.
[0685] Step 6:
[0686] The server adjusts the response generated by the dialogue agent based on emotion data from the emotion recognition engine. A gentler tone or additional sentences depending on the urgency are added to the response. For example, a response like, "Thank you for your understanding. The optimal route is as follows..." might be created.
[0687] Step 7:
[0688] The server sends a coordinated response back to the communication terminal. The returned data includes coordinated routing information.
[0689] Step 8:
[0690] The communication terminal receives a response from the server and displays it in the chat window. This allows the user to quickly confirm the optimal route information.
[0691] Through the steps described above, the present invention constructs a system that supports the efficient execution of tasks by managers in logistics centers and provides rapid and appropriate route information.
[0692] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0693] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0694] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0695] [Third Embodiment]
[0696] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0697] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0698] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0699] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0700] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0701] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0702] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0703] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0704] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0705] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0706] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0707] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0708] The present invention provides a system that allows users to input questions in natural language and quickly obtain route information. The embodiments thereof are described in detail below.
[0709] System-wide configuration
[0710] This system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0711] Communication terminal
[0712] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0713] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0714] server
[0715] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0716] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs.
[0717] Dialogue Agent
[0718] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0719] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0720] Sending and displaying responses
[0721] The server sends the response generated by the conversational agent back to the communication terminal. The communication terminal receives this response and displays it in the chat window. The user can then review the displayed response.
[0722] Specific examples
[0723] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station" into the chat window of their communication terminal and presses the send button.
[0724] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent.
[0725] The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please take the JR Yamanote Line. The journey takes approximately 10 minutes."
[0726] This response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0727] The above describes an embodiment of the system based on the present invention. By applying this system, users will be able to obtain route information easily and quickly.
[0728] The following describes the processing flow.
[0729] Step 1:
[0730] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0731] Step 2:
[0732] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0733] Step 3:
[0734] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0735] Step 4:
[0736] The server receives the request. The received data is converted into an appropriate format for analysis.
[0737] Step 5:
[0738] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0739] Step 6:
[0740] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0741] Step 7:
[0742] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0743] Step 8:
[0744] The generated response is returned to the server. The server converts this response into JSON data.
[0745] Step 9:
[0746] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0747] Step 10:
[0748] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0749] Step 11:
[0750] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0751] The above outlines the specific processing steps of this system's program.
[0752] (Example 1)
[0753] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0754] In modern society, it is crucial for users to quickly and accurately obtain appropriate route information based on their location and destination. However, conventional systems often take a long time to process questions entered in natural language, resulting in inadequate responses. To solve this problem, a system is needed that can properly analyze user questions and quickly provide optimal route information.
[0755] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0756] In this invention, the server includes means for a user to input a question in natural language via a communication device, means for transmitting the question to a data processing device, means for the data processing device to receive the question and generate a response using a conversational agent, means for sending the response back to the communication device and displaying it to the user, and includes question input at the communication device, question analysis at the data processing device, retrieval of an information database, and response generation. This makes it possible for the user to obtain route information easily and quickly.
[0757] A "communication device" is a device used by a user to input questions via an interface, and includes smartphones and personal computers.
[0758] A "data processing device" is a device that receives data sent by a user, analyzes it, and generates a response; it generally refers to a server.
[0759] A "conversational agent" is a program or system that analyzes a user's questions and generates appropriate responses, using natural language processing technology.
[0760] A "response" refers to the information generated by the conversational agent in response to a user's question, and it is presented in natural language.
[0761] An "information database" is a system for storing and managing necessary data, including route information and map information.
[0762] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.
[0763] The present invention provides a system that allows a user to input a question in natural language via a communication device and quickly obtain route information based on that question. Specific embodiments thereof are described in detail below.
[0764] Communication device
[0765] Users open a dedicated chat window using a communication device such as a smartphone or personal computer. In this chat window, users can input questions about their current location and destination using natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends the entered question to a data processing device (server).
[0766] Data processing device (server)
[0767] The data processing unit receives requests sent from communication devices and has the function of analyzing these requests. The received question content is passed to a conversation agent running on the server. The data processing unit automatically handles everything from receiving requests to generating and sending responses, and also has the function of saving received messages as logs.
[0768] Conversation Agent
[0769] The conversational agent's role is to analyze the content of questions sent by users using natural language processing technology. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the information database to obtain the optimal route information. The obtained route information is converted into a response to be sent back to the user in natural language. For example, it might say, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes."
[0770] Sending and displaying responses
[0771] The data processing unit (server) sends the response generated by the conversation agent back to the communication device. The communication device receives this response and displays it in the chat window. The user can review the displayed response and obtain the necessary routing information.
[0772] Specific examples
[0773] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station" into the chat window of the communication device and clicks the send button. The communication device sends this question to the data processing device. The data processing device receives the request and passes the question to the conversational agent. The conversational agent searches the information database and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the Y Line of the railway. The journey takes approximately 10 minutes." This response is returned to the data processing device and then sent back to the communication device. The communication device displays the received response in the chat window, allowing the user to confirm the answer.
[0774] Example of a prompt
[0775] Examples of prompts to input into a generative AI model include the following:
[0776] "Describe the sequence of events in which the system generates responses to questions submitted by users. Clearly define the roles of the server, communication device, and conversational agent."
[0777] The above describes the embodiments for carrying out the present invention. With this system, users can easily and quickly obtain route information.
[0778] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0779] Step 1:
[0780] The user opens a chat window via a communication device. This device could be a smartphone or a personal computer. Once the chat window appears, the user enters a question in natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" This question is then entered into the communication device.
[0781] Step 2:
[0782] When the user clicks the "Submit" button, the communication device sends the entered question content to the data processing device (server). Specifically, the communication device sends the question text to the server via the network. At this time, the input data is the user's question, and the output data is the question sent as a request to the server.
[0783] Step 3:
[0784] The server receives requests sent from communication devices. First, the server analyzes the received question. This analysis involves inputting the user's question as received data into a text analysis engine, which outputs analysis results that help understand the format and structure of the question. These analysis results are then passed to the conversational agent in the next step.
[0785] Step 4:
[0786] The conversational agent running on the server receives the analysis results and uses natural language processing techniques to further analyze the question. At this stage, the specific data processing involves identifying the origin and destination from the question; the input data is the analysis results, and the output data is the identified origin and destination. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent identifies "Shibuya" and "Shinjuku".
[0787] Step 5:
[0788] The conversational agent searches an information database based on the identified origin and destination. This database stores route information, and a search process is performed to obtain the optimal route information. The input data is the origin and destination, and the output data is the optimal route information. For example, information such as, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," might be output.
[0789] Step 6:
[0790] Based on the acquired route information, the conversational agent generates a response to send back to the user in natural language. At this stage, the input data is the optimal route information, and the output data is a natural language response. For example, a response in the form of, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," is generated.
[0791] Step 7:
[0792] The server sends the generated response back to the communication device. Specifically, the process involves sending the response to the communication device over the network. The input data is the response in natural language, and the output data is the transmission request to the communication device.
[0793] Step 8:
[0794] The communication device receives the response message sent back from the server and displays it in the chat window. The user can then review this response message and easily obtain the necessary routing information. The input data is the response message from the server, and the output data is the response displayed in the chat window.
[0795] The above outlines the specific processing steps of this system and the details of the inputs and outputs at each step.
[0796] (Application Example 1)
[0797] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0798] In food delivery operations, a challenge exists in that delivery personnel have difficulty quickly and efficiently obtaining the optimal route. Traditional systems require considerable effort from delivery personnel to acquire route information, resulting in delivery delays and inefficiencies. Furthermore, their ability to provide accurate route information instantly in response to natural language inquiries is limited.
[0799] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0800] In this invention, the server includes means for a user to input a question in natural language via a communication terminal, means for sending the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and means configured to obtain optimal route information when the user is performing food delivery work. This enables delivery personnel to quickly obtain optimal route information in natural language.
[0801] (Term definition)
[0802] A "user" is someone who uses a communication terminal to input questions in natural language and performs food delivery services.
[0803] A "communication terminal" is a device used by a user, such as a smartphone or smart glasses, that has a means of inputting questions in natural language and sending them to a server.
[0804] "Natural language" refers to the language that humans use in everyday life, and is a format that allows for easy input of questions and instructions, rather than specialized programming languages or commands.
[0805] A "server" is a computing system that receives questions from communication terminals, analyzes their content, generates appropriate responses, and sends them back to the communication terminals.
[0806] A "conversational agent" is a software program that runs within a server, analyzes user inquiries, and generates responses based on appropriate routing information.
[0807] A "response" is information generated by a conversational agent based on a user's question, and is a natural language response that includes routing information.
[0808] "Means of returning a response to a communication terminal and displaying it to the user" refers to a function that returns a response generated by the server to a communication terminal and provides that response to the user via screen or audio.
[0809] "Food delivery services" refer to the business activity of taking orders for food and beverages and delivering them quickly to the customer's designated location.
[0810] "Optimal route information" refers to information about the most efficient route from the user's specified starting point to the destination, taking into account factors such as time, distance, and traffic conditions.
[0811] A "route guidance system" is a system that calculates routes based on maps and traffic information and provides users with the most suitable route information.
[0812] Modes for carrying out the invention
[0813] System Configuration
[0814] This invention provides a system for food delivery services that allows delivery personnel to input questions in natural language and quickly obtain optimal route information. The system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[0815] Hardware and software to be used
[0816] Hardware: Communication devices such as smartphones and smart glasses.
[0817] Software: Applications for Android or iOS devices, conversational agent servers (e.g., IBM Watson, Google Dialogflow), and route guidance system servers.
[0818] Detailed explanation of the process
[0819] 1. User input:
[0820] Users can launch an application on their communication terminal and input route-related questions in natural language. For example, they might input, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input can be done via voice or text.
[0821] 2. Sending input to the server:
[0822] The communication terminal sends the entered question content to the server. The server receives this data and passes it on to the conversational agent.
[0823] 3. Dialogue Agent:
[0824] The conversational agent runs on a server and is responsible for analyzing the content of natural language questions it receives. For example, generative AI models such as IBM Watson or Google Dialogflow can be used. The conversational agent identifies the origin and destination from the user's questions and searches the navigation system's database to obtain the best route information.
[0825] 4. Response generation and return:
[0826] The conversational agent generates a response in natural language based on the acquired route information. The generated response is sent back to the communication terminal via the server. For example, a response such as "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes" is generated.
[0827] 5. Display of response:
[0828] The communication terminal receives the returned response and displays it in the chat window. In addition to text display, audio output is also available as a display method.
[0829] Specific examples and prompt statements
[0830] Specific example
[0831] Scenario: A pizza delivery driver checks the shortest route.
[0832] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0833] Analysis results from the conversational agent: Origin: "Pizza house in Shibuya", Destination: "Office in Ebisu"
[0834] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0835] Example prompts for generative AI models
[0836] Prompt: The user typed, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" Provide the best route information for this question in natural language.
[0837] Example data:
[0838] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[0839] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[0840] This allows delivery drivers in the food delivery industry to quickly and efficiently obtain optimal route information.
[0841] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0842] Detailed program processing steps
[0843] Step 1:
[0844] The user launches an application on their communication terminal and enters a question in natural language. For example, the user might enter, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input is sent to the communication terminal in either text or voice format.
[0845] Input: The user enters a question into the application using natural language.
[0846] Output: User's question (text or audio data)
[0847] Step 2:
[0848] The communication terminal sends the entered question content to the server. During this process, natural language text data and audio data are sent to the server.
[0849] Input: User's question (text or audio data)
[0850] Output: Question content sent to the server
[0851] Step 3:
[0852] The server receives the question and passes it to the conversational agent. The conversational agent analyzes the user's question. Here, a generative AI model (e.g., IBM Watson, Google Dialogflow) is used to identify the origin and destination from the question. In this analysis process, natural language processing techniques are used to extract the origin and destination and convert them into an appropriate format.
[0853] Input: Question content sent to the server
[0854] Output: Analyzed origin and destination
[0855] Step 4:
[0856] The conversational agent searches the route guidance system's database to obtain the optimal route information based on the analyzed origin ("pizza house in Shibuya") and destination ("office in Ebisu"). This database search process uses traffic information and map information to identify the best route, taking into account factors such as travel time and distance.
[0857] Input: Analyzed origin and destination
[0858] Output: Optimal route information
[0859] Step 5:
[0860] The conversational agent generates a response in natural language based on the acquired route information. This response is generated in a format such as, "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes." Here, the natural language generation function of the generative AI model is used to create an appropriate sentence.
[0861] Input: Optimal route information
[0862] Output: Response with route information generated in natural language
[0863] Step 6:
[0864] The server sends the generated response back to the communication terminal. The response is sent to the communication terminal in either text or audio format.
[0865] Input: Generated natural language response
[0866] Output: Response sent to communication terminal
[0867] Step 7:
[0868] The communication terminal receives a response and displays it to the user. Display methods include text-based chat window display and audio output. The user can review this response and obtain optimal routing information.
[0869] Input: Response sent from the server
[0870] Output: Route information displayed to the user
[0871] This series of processes allows users to quickly obtain optimal route information in response to questions entered in natural language, enabling them to efficiently carry out food delivery operations.
[0872] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0873] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0874] System-wide configuration
[0875] This system consists of a communication terminal used by the user, a server that receives requests, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotions.
[0876] Communication terminal
[0877] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[0878] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[0879] server
[0880] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[0881] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs. Furthermore, the server adjusts the content of its responses based on sentiment data analyzed by the sentiment engine.
[0882] Dialogue Agent
[0883] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[0884] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[0885] Emotional Engine
[0886] The emotion engine analyzes the user's input and recognizes their emotional state. For example, if the user uses language that indicates frustration, the emotion engine will recognize that emotion as "anger." This emotional information is passed to the conversational agent and reflected in the generated response.
[0887] Response adjustment and return
[0888] The server adjusts the responses generated by the conversational agent based on the analysis results from the emotion engine. For example, if the user is irritated, the server may adjust the tone of the response to be calmer.
[0889] The adjusted response is sent back from the server to the communication terminal. The communication terminal receives this response and displays it in the chat window.
[0890] Specific examples
[0891] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication device and presses the send button.
[0892] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey takes approximately 10 minutes."
[0893] Simultaneously, the emotion engine within the server recognizes the user's emotion of urgency from their expression, "I'm in a hurry!" Based on this emotion information, a response tone that is more responsive to the urgent situation is set.
[0894] This adjusted response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0895] The above describes an embodiment of a system incorporating an emotion engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[0896] The following describes the processing flow.
[0897] Step 1:
[0898] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[0899] Step 2:
[0900] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[0901] Step 3:
[0902] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[0903] Step 4:
[0904] The server receives the request. The received data is converted into an appropriate format for analysis.
[0905] Step 5:
[0906] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[0907] Step 6:
[0908] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[0909] Step 7:
[0910] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0911] Step 8:
[0912] The server passes the response from the conversational agent to the emotion engine. At the same time, the user's input is also passed.
[0913] Step 9:
[0914] The emotion engine on the server analyzes the user's input and recognizes their emotional state. For example, it recognizes the emotion of urgency from the expression "I'm in a hurry!"
[0915] Step 10:
[0916] The emotion engine adjusts its response based on the emotional information it recognizes. For example, it might change the tone to one that indicates a more urgent response.
[0917] Step 11:
[0918] The adjusted response is returned to the server. The server converts this response into JSON data.
[0919] Step 12:
[0920] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[0921] Step 13:
[0922] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[0923] Step 14:
[0924] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[0925] The above outlines the specific processing steps of this system.
[0926] (Example 2)
[0927] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0928] Currently, systems that allow users to obtain route information via communication terminals often fail to consider the user's emotional state, resulting in decreased user satisfaction. Furthermore, the inability to respond flexibly to user emotions makes it difficult to provide appropriate service, especially when the user is in a hurry or irritated. Improving this situation and enhancing user satisfaction is crucial.
[0929] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language via a communication terminal, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, means for using an emotion analysis engine to recognize the user's emotions, and means for adjusting the response content based on the emotion information recognized by the emotion analysis engine. This makes it possible to provide a detailed response that corresponds to the user's emotions.
[0930] A "communication terminal" is a device used by users to input questions and communicate with a server, and specifically refers to devices such as smartphones and personal computers.
[0931] A "server" is a computer device that receives questions from users, analyzes the messages, generates responses, and sends them back to the users.
[0932] A "question" refers to an inquiry about routing information or other related matters that a user inputs in natural language through a communication terminal.
[0933] A "dialogue agent" is a software module that analyzes the content of a user's question and generates an appropriate response, using natural language processing technology.
[0934] A "response" is a natural language response generated by a conversational agent and sent back to the user via the server.
[0935] An "emotion analysis engine" is a software module that recognizes the user's emotional state from their input and assigns emotional labels such as "joy," "anger," and "sadness."
[0936] "Emotional information" refers to data representing the user's emotional state, which is recognized by the emotion analysis engine and used to adjust responses by the dialogue agent.
[0937] A "route information provision system" is a database and software module that provides optimal route information based on a specific origin and destination.
[0938] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion analysis engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[0939] System-wide configuration
[0940] This system consists of the following hardware and software:
[0941] Communication devices used by the user (smartphones and personal computers)
[0942] Server that receives and analyzes messages
[0943] A conversational agent that generates responses to user questions.
[0944] A sentiment analysis engine that recognizes user emotions.
[0945] Communication terminal
[0946] Users can open a chat window using their communication device and enter questions in natural language. For example, they can type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends this question to the server.
[0947] server
[0948] The server receives and analyzes the content of questions sent from communication terminals. The received messages are passed to the conversational agent. The server also uses an emotion analysis engine to analyze the user's emotions and passes the results to the conversational agent. The server automatically processes everything from receiving requests to generating and sending responses, and also has a function to save received messages as logs.
[0949] Dialogue Agent
[0950] The conversational agent runs on the server and analyzes the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question, then searches the route information system (database) to obtain the optimal route information. The obtained route information is then converted into an appropriate natural language response by the conversational agent.
[0951] Emotion analysis engine
[0952] The emotion analysis engine analyzes the user's input and recognizes their emotional state. For example, if the user uses the expression "I'm in a hurry!", the emotion analysis engine recognizes that the user is feeling anxious. The analyzed emotional information is passed to the conversational agent and reflected in the generated response.
[0953] Response adjustment and return
[0954] The server adjusts the responses generated by the conversational agent based on the results of the emotion analysis engine. For example, if the user is irritated, the server adjusts the tone of the response to be calmer. The adjusted response is sent back from the server to the communication terminal, which receives the response and displays it in the chat window.
[0955] Specific examples
[0956] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the route information system and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey will take approximately 10 minutes." At the same time, the sentiment analysis engine recognizes the user's urgency from the expression "I'm in a hurry!" and adjusts the response tone based on this sentiment information. This adjusted response is returned to the server and sent back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[0957] The above describes an embodiment of a system incorporating an emotion analysis engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[0958] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0959] Step 1: The user enters the question.
[0960] The user enters a question into the chat window of their communication device. This question is in natural language. For example, they might enter, "What is the shortest route from Shibuya to Shinjuku?" The input data is in text format. Specifically, the user uses a smartphone or computer keyboard to type the text and clicks the send button. The input content is stored as text data on the communication device.
[0961] Step 2: Send the question to the server
[0962] The terminal sends the user's input to the server. The input data is in text format, and the output is a data packet sent to the server. Specifically, the terminal converts the input text into a data packet format and sends it to the server over the internet. The data is transmitted according to the TCP / IP protocol.
[0963] Step 3: The server receives the question and analyzes it.
[0964] The server receives the question content sent from the terminal and parses the message content. The input data is the text data received from the terminal, and the output is the parsed question content. Specifically, the server receives the data packet and converts it to text format. Next, it parses the question content using natural language processing technology and saves it as a log.
[0965] Step 4: Pass the question to the conversational agent.
[0966] The server passes the analyzed question content to the conversational agent. The input data is the analyzed question content, and the output is data for the conversational agent to generate a response. Specifically, the server makes an API call to the conversational agent and passes the analyzed question content. The conversational agent then starts generating a response based on this content.
[0967] Step 5: Analyze emotions with an emotion analysis engine
[0968] The server sends the user's question to the sentiment analysis engine, which then analyzes the emotional state. The input data is the user's question, and the output is data with emotional labels. Specifically, the server makes an API call to the sentiment analysis engine and sends the input text. The sentiment analysis engine analyzes the emotions from the text and assigns emotional labels such as "joy," "anger," and "sadness."
[0969] Step 6: The conversational agent generates a response.
[0970] The conversational agent identifies the origin and destination from the question, searches a route information system, and generates the optimal route information. The input data is the question, and the output is a response in natural language. Specifically, the conversational agent queries the database, retrieves appropriate route information, and converts that information into natural language.
[0971] Step 7: The server adjusts its response based on the sentiment analysis results.
[0972] The server combines the responses received from the conversational agent with the results of the emotion analysis engine to adjust the final response. The input data is the generated response and emotion information, and the output is the adjusted response. Specifically, the server applies emotion-specific adjustment logic to adjust the tone and content of the response. For example, if the user is irritated, the tone of the response will be softened.
[0973] Step 8: Send the adjusted response back to the terminal.
[0974] The server sends a pre-arranged response to the communication terminal. The input data is the pre-arranged response, and the output is the data packet sent to the terminal. Specifically, the server converts the response into a packet format and sends it to the terminal over the internet.
[0975] Step 9: The terminal displays a response
[0976] The terminal displays the response received from the server in the chat window. The input data is text data received from the server, and the output is the response display that the user can see. Specifically, the terminal converts the received data packet into text format and displays it in the chat window. The user can then verify this and obtain routing information.
[0977] (Application Example 2)
[0978] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0979] Modern logistics centers require the rapid and accurate processing of a large number of packages and shipments. However, it is not easy for managers to efficiently track package locations and delivery routes and make optimal decisions. Furthermore, disregarding managers' emotions and urgency can lead to a poor user experience and reduced efficiency. This situation results in delays and errors, and a decline in overall operational efficiency.
[0980] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and an emotion recognition engine that recognizes the user's emotional state and adjusts the response content based on the emotional state. As a result, the manager of the logistics center can quickly obtain information on the location of packages and the optimal delivery route using natural language, and an optimal response can be provided according to the manager's emotional state.
[0981] A "communication terminal" is a device that allows users to input questions in natural language and send them to a server.
[0982] A "server" is a computer system that receives questions from users, generates responses using dialogue agents, and sends them back to the communication terminal.
[0983] A "conversational agent" is a program that runs within a server, analyzes user questions, and generates appropriate responses.
[0984] An "emotion recognition engine" is a program that analyzes the user's emotional state from their natural language input and adjusts its response based on the results.
[0985] "Adjusting the response content" means changing the tone and details of the generated response based on the user's emotional state.
[0986] "Logistics" is a general term for the business processes involved in the storage, delivery, and management of packages and goods.
[0987] "Natural language" refers to text and audio composed of the language that users use on a daily basis.
[0988] "Means of sending questions" refers to the function of sending user questions to the server via a communication terminal.
[0989] "Means of sending and displaying a response" refers to a function that receives a response from a server on a communication terminal and displays it to the user.
[0990] "Route information" refers to data about the optimal route between a specified origin and destination.
[0991] This invention relates to a system that allows managers in logistics centers to input questions in natural language using a smartphone and receive prompt and appropriate route information in response. This system recognizes the user's emotional state and adjusts its responses accordingly to provide more effective support. Specific embodiments for carrying out this invention are described in detail below.
[0992] System-wide configuration
[0993] This system consists of a communication terminal used by the user, a server that receives questions, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotional state.
[0994] Communication terminal
[0995] The user launches the logistics management app using a communication device (e.g., a smartphone). In the app's chat window, the user can input questions about the package's location and delivery route using natural language. For example, they might input, "I need to send this package to point A as soon as possible, which route is best?" When the user clicks the "Send" button, the communication device sends the entered question to the server.
[0996] server
[0997] The server receives questions sent from communication terminals and forwards requests to conversational agents. The server also uses an emotion recognition engine to analyze the user's emotional state and adjusts the conversational agent's responses based on the results. Software used within the server includes Flask (a web framework) and OpenAI (a natural language processing engine).
[0998] Dialogue Agent
[0999] The conversational agent runs on the server and is responsible for analyzing the user's questions. Specifically, it uses the OpenAI API to generate optimal route information from the question. For example, if a user asks, "I urgently need to send this package to point A, which route is best?", the conversational agent searches a logistics information database and obtains the optimal route information. This route information is then converted into a natural language response by the conversational agent.
[1000] Emotional Engine
[1001] The emotion engine analyzes the user's input and recognizes their emotional state. For example, it recognizes the user's anxiety from expressions like "urgent" or "in a hurry." This emotional information is passed to the conversational agent and used to adjust the response content and tone.
[1002] Response adjustment and return
[1003] The server adjusts the responses generated by the conversational agent based on the analysis results of the emotion engine. For example, if the user is in a hurry, a message in a gentler tone, such as "The best route is as follows. Thank you for your understanding," may be added. The adjusted response is sent back from the server to the communication terminal and displayed to the user.
[1004] Specific examples
[1005] The user types "I urgently need to send this package to point A, what's the best route?" into the chat window of the communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the logistics information database and generates a response: "The best route is as follows. Thank you for your understanding." At the same time, the emotion engine in the server recognizes the user's urgency from the expression "urgent." Based on this emotion information, the tone of the response is adjusted. This adjusted response is returned to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[1006] Examples of prompts to input into a generative AI model:
[1007] Provide the best route for the query: "I urgently need to send this package to point A. What is the best route?"
[1008] In this way, the present invention can improve work efficiency in logistics centers and reduce stress for managers.
[1009] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1010] Step 1:
[1011] The user uses the chat window on the communication terminal to input a question in natural language. For example, a question like, "I urgently need to send this package to point A, what is the best route?" might be entered. This input is converted into digital data within the communication terminal.
[1012] Step 2:
[1013] When the user clicks the "Send" button, the communication device sends this digital data to the server. The transmitted data includes the user's question.
[1014] Step 3:
[1015] The server receives a question from the user and passes the request to the conversational agent. The conversational agent analyzes this data to identify the question (the location and destination of the package). Natural language processing techniques are used for this analysis.
[1016] Step 4:
[1017] The conversational agent searches the logistics information database to obtain the optimal route information. An example of a prompt for this database search is: "Provide the best route for the query: 'I urgently need to send this package to point A, which route is best?'" The generated route information is returned to the conversational agent.
[1018] Step 5:
[1019] The emotion recognition engine on the server analyzes the user's emotional state from their input. In this example, the expression "urgent" is used to recognize that the user is in a hurry. The emotion recognition engine then generates this emotional data.
[1020] Step 6:
[1021] The server adjusts the response generated by the dialogue agent based on emotion data from the emotion recognition engine. A gentler tone or additional sentences depending on the urgency are added to the response. For example, a response like, "Thank you for your understanding. The optimal route is as follows..." might be created.
[1022] Step 7:
[1023] The server sends a coordinated response back to the communication terminal. The returned data includes coordinated routing information.
[1024] Step 8:
[1025] The communication terminal receives a response from the server and displays it in the chat window. This allows the user to quickly confirm the optimal route information.
[1026] Through the steps described above, the present invention constructs a system that supports the efficient execution of tasks by managers in logistics centers and provides rapid and appropriate route information.
[1027] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1028] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1029] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1030] [Fourth Embodiment]
[1031] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1032] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1033] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1034] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1035] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1036] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1037] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1038] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1039] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1040] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1041] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1042] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1043] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1044] The present invention provides a system that allows users to input questions in natural language and quickly obtain route information. The embodiments thereof are described in detail below.
[1045] System-wide configuration
[1046] This system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[1047] Communication terminal
[1048] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[1049] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[1050] server
[1051] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[1052] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs.
[1053] Dialogue Agent
[1054] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[1055] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[1056] Sending and displaying responses
[1057] The server sends the response generated by the conversational agent back to the communication terminal. The communication terminal receives this response and displays it in the chat window. The user can then review the displayed response.
[1058] Specific examples
[1059] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station" into the chat window of their communication terminal and presses the send button.
[1060] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent.
[1061] The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please take the JR Yamanote Line. The journey takes approximately 10 minutes."
[1062] This response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[1063] The above describes an embodiment of the system based on the present invention. By applying this system, users will be able to obtain route information easily and quickly.
[1064] The following describes the processing flow.
[1065] Step 1:
[1066] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[1067] Step 2:
[1068] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[1069] Step 3:
[1070] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[1071] Step 4:
[1072] The server receives the request. The received data is converted into an appropriate format for analysis.
[1073] Step 5:
[1074] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[1075] Step 6:
[1076] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[1077] Step 7:
[1078] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[1079] Step 8:
[1080] The generated response is returned to the server. The server converts this response into JSON data.
[1081] Step 9:
[1082] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[1083] Step 10:
[1084] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[1085] Step 11:
[1086] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[1087] The above outlines the specific processing steps of this system's program.
[1088] (Example 1)
[1089] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1090] In modern society, it is crucial for users to quickly and accurately obtain appropriate route information based on their location and destination. However, conventional systems often take a long time to process questions entered in natural language, resulting in inadequate responses. To solve this problem, a system is needed that can properly analyze user questions and quickly provide optimal route information.
[1091] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1092] In this invention, the server includes means for a user to input a question in natural language via a communication device, means for transmitting the question to a data processing device, means for the data processing device to receive the question and generate a response using a conversational agent, means for sending the response back to the communication device and displaying it to the user, and includes question input at the communication device, question analysis at the data processing device, retrieval of an information database, and response generation. This makes it possible for the user to obtain route information easily and quickly.
[1093] A "communication device" is a device used by a user to input questions via an interface, and includes smartphones and personal computers.
[1094] A "data processing device" is a device that receives data sent by a user, analyzes it, and generates a response; it generally refers to a server.
[1095] A "conversational agent" is a program or system that analyzes a user's questions and generates appropriate responses, using natural language processing technology.
[1096] A "response" refers to the information generated by the conversational agent in response to a user's question, and it is presented in natural language.
[1097] An "information database" is a system for storing and managing necessary data, including route information and map information.
[1098] "Natural language processing technology" refers to the technology that enables computers to understand, analyze, and generate human language.
[1099] The present invention provides a system that allows a user to input a question in natural language via a communication device and quickly obtain route information based on that question. Specific embodiments thereof are described in detail below.
[1100] Communication device
[1101] Users open a dedicated chat window using a communication device such as a smartphone or personal computer. In this chat window, users can input questions about their current location and destination using natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends the entered question to a data processing device (server).
[1102] Data processing device (server)
[1103] The data processing unit receives requests sent from communication devices and has the function of analyzing these requests. The received question content is passed to a conversation agent running on the server. The data processing unit automatically handles everything from receiving requests to generating and sending responses, and also has the function of saving received messages as logs.
[1104] Conversation Agent
[1105] The conversational agent's role is to analyze the content of questions sent by users using natural language processing technology. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the information database to obtain the optimal route information. The obtained route information is converted into a response to be sent back to the user in natural language. For example, it might say, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes."
[1106] Sending and displaying responses
[1107] The data processing unit (server) sends the response generated by the conversation agent back to the communication device. The communication device receives this response and displays it in the chat window. The user can review the displayed response and obtain the necessary routing information.
[1108] Specific examples
[1109] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station" into the chat window of the communication device and clicks the send button. The communication device sends this question to the data processing device. The data processing device receives the request and passes the question to the conversational agent. The conversational agent searches the information database and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the Y Line of the railway. The journey takes approximately 10 minutes." This response is returned to the data processing device and then sent back to the communication device. The communication device displays the received response in the chat window, allowing the user to confirm the answer.
[1110] Example of a prompt
[1111] Examples of prompts to input into a generative AI model include the following:
[1112] "Describe the sequence of events in which the system generates responses to questions submitted by users. Clearly define the roles of the server, communication device, and conversational agent."
[1113] The above describes the embodiments for carrying out the present invention. With this system, users can easily and quickly obtain route information.
[1114] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1115] Step 1:
[1116] The user opens a chat window via a communication device. This device could be a smartphone or a personal computer. Once the chat window appears, the user enters a question in natural language. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?" This question is then entered into the communication device.
[1117] Step 2:
[1118] When the user clicks the "Submit" button, the communication device sends the entered question content to the data processing device (server). Specifically, the communication device sends the question text to the server via the network. At this time, the input data is the user's question, and the output data is the question sent as a request to the server.
[1119] Step 3:
[1120] The server receives requests sent from communication devices. First, the server analyzes the received question. This analysis involves inputting the user's question as received data into a text analysis engine, which outputs analysis results that help understand the format and structure of the question. These analysis results are then passed to the conversational agent in the next step.
[1121] Step 4:
[1122] The conversational agent running on the server receives the analysis results and uses natural language processing techniques to further analyze the question. At this stage, the specific data processing involves identifying the origin and destination from the question; the input data is the analysis results, and the output data is the identified origin and destination. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent identifies "Shibuya" and "Shinjuku".
[1123] Step 5:
[1124] The conversational agent searches an information database based on the identified origin and destination. This database stores route information, and a search process is performed to obtain the optimal route information. The input data is the origin and destination, and the output data is the optimal route information. For example, information such as, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," might be output.
[1125] Step 6:
[1126] Based on the acquired route information, the conversational agent generates a response to send back to the user in natural language. At this stage, the input data is the optimal route information, and the output data is a natural language response. For example, a response in the form of, "The shortest route from Shibuya to Shinjuku is via the X-ray train. The journey takes approximately 15 minutes," is generated.
[1127] Step 7:
[1128] The server sends the generated response back to the communication device. Specifically, the process involves sending the response to the communication device over the network. The input data is the response in natural language, and the output data is the transmission request to the communication device.
[1129] Step 8:
[1130] The communication device receives the response message sent back from the server and displays it in the chat window. The user can then review this response message and easily obtain the necessary routing information. The input data is the response message from the server, and the output data is the response displayed in the chat window.
[1131] The above outlines the specific processing steps of this system and the details of the inputs and outputs at each step.
[1132] (Application Example 1)
[1133] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1134] In food delivery operations, a challenge exists in that delivery personnel have difficulty quickly and efficiently obtaining the optimal route. Traditional systems require considerable effort from delivery personnel to acquire route information, resulting in delivery delays and inefficiencies. Furthermore, their ability to provide accurate route information instantly in response to natural language inquiries is limited.
[1135] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1136] In this invention, the server includes means for a user to input a question in natural language via a communication terminal, means for sending the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and means configured to obtain optimal route information when the user is performing food delivery work. This enables delivery personnel to quickly obtain optimal route information in natural language.
[1137] (Term definition)
[1138] A "user" is someone who uses a communication terminal to input questions in natural language and performs food delivery services.
[1139] A "communication terminal" is a device used by a user, such as a smartphone or smart glasses, that has a means of inputting questions in natural language and sending them to a server.
[1140] "Natural language" refers to the language that humans use in everyday life, and is a format that allows for easy input of questions and instructions, rather than specialized programming languages or commands.
[1141] A "server" is a computing system that receives questions from communication terminals, analyzes their content, generates appropriate responses, and sends them back to the communication terminals.
[1142] A "conversational agent" is a software program that runs within a server, analyzes user inquiries, and generates responses based on appropriate routing information.
[1143] A "response" is information generated by a conversational agent based on a user's question, and is a natural language response that includes routing information.
[1144] "Means of returning a response to a communication terminal and displaying it to the user" refers to a function that returns a response generated by the server to a communication terminal and provides that response to the user via screen or audio.
[1145] "Food delivery services" refer to the business activity of taking orders for food and beverages and delivering them quickly to the customer's designated location.
[1146] "Optimal route information" refers to information about the most efficient route from the user's specified starting point to the destination, taking into account factors such as time, distance, and traffic conditions.
[1147] A "route guidance system" is a system that calculates routes based on maps and traffic information and provides users with the most suitable route information.
[1148] Modes for carrying out the invention
[1149] System Configuration
[1150] This invention provides a system for food delivery services that allows delivery personnel to input questions in natural language and quickly obtain optimal route information. The system consists of a communication terminal used by the user, a server that receives requests, and a dialogue agent that generates responses.
[1151] Hardware and software to be used
[1152] Hardware: Communication devices such as smartphones and smart glasses.
[1153] Software: Applications for Android or iOS devices, conversational agent servers (e.g., IBM Watson, Google Dialogflow), and route guidance system servers.
[1154] Detailed explanation of the process
[1155] 1. User input:
[1156] Users can launch an application on their communication terminal and input route-related questions in natural language. For example, they might input, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input can be done via voice or text.
[1157] 2. Sending input to the server:
[1158] The communication terminal sends the entered question content to the server. The server receives this data and passes it on to the conversational agent.
[1159] 3. Dialogue Agent:
[1160] The conversational agent runs on a server and is responsible for analyzing the content of natural language questions it receives. For example, generative AI models such as IBM Watson or Google Dialogflow can be used. The conversational agent identifies the origin and destination from the user's questions and searches the navigation system's database to obtain the best route information.
[1161] 4. Response generation and return:
[1162] The conversational agent generates a response in natural language based on the acquired route information. The generated response is sent back to the communication terminal via the server. For example, a response such as "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes" is generated.
[1163] 5. Display of response:
[1164] The communication terminal receives the returned response and displays it in the chat window. In addition to text display, audio output is also available as a display method.
[1165] Specific examples and prompt statements
[1166] Specific example
[1167] Scenario: A pizza delivery driver checks the shortest route.
[1168] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[1169] Analysis results from the conversational agent: Origin: "Pizza house in Shibuya", Destination: "Office in Ebisu"
[1170] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[1171] Example prompts for generative AI models
[1172] Prompt: The user typed, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" Provide the best route information for this question in natural language.
[1173] Example data:
[1174] User input: "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?"
[1175] AI response: "The shortest route is via the JR Yamanote Line, and the journey takes approximately 15 minutes."
[1176] This allows delivery drivers in the food delivery industry to quickly and efficiently obtain optimal route information.
[1177] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1178] Detailed program processing steps
[1179] Step 1:
[1180] The user launches an application on their communication terminal and enters a question in natural language. For example, the user might enter, "What is the shortest route from the pizza place in Shibuya to the office in Ebisu?" This input is sent to the communication terminal in either text or voice format.
[1181] Input: The user enters a question into the application using natural language.
[1182] Output: User's question (text or audio data)
[1183] Step 2:
[1184] The communication terminal sends the entered question content to the server. During this process, natural language text data and audio data are sent to the server.
[1185] Input: User's question (text or audio data)
[1186] Output: Question content sent to the server
[1187] Step 3:
[1188] The server receives the question and passes it to the conversational agent. The conversational agent analyzes the user's question. Here, a generative AI model (e.g., IBM Watson, Google Dialogflow) is used to identify the origin and destination from the question. In this analysis process, natural language processing techniques are used to extract the origin and destination and convert them into an appropriate format.
[1189] Input: Question content sent to the server
[1190] Output: Analyzed origin and destination
[1191] Step 4:
[1192] The conversational agent searches the route guidance system's database to obtain the optimal route information based on the analyzed origin ("pizza house in Shibuya") and destination ("office in Ebisu"). This database search process uses traffic information and map information to identify the best route, taking into account factors such as travel time and distance.
[1193] Input: Analyzed origin and destination
[1194] Output: Optimal route information
[1195] Step 5:
[1196] The conversational agent generates a response in natural language based on the acquired route information. This response is generated in a format such as, "The shortest route is via the JR Yamanote Line, and the travel time is approximately 15 minutes." Here, the natural language generation function of the generative AI model is used to create an appropriate sentence.
[1197] Input: Optimal route information
[1198] Output: Response with route information generated in natural language
[1199] Step 6:
[1200] The server sends the generated response back to the communication terminal. The response is sent to the communication terminal in either text or audio format.
[1201] Input: Generated natural language response
[1202] Output: Response sent to communication terminal
[1203] Step 7:
[1204] The communication terminal receives a response and displays it to the user. Display methods include text-based chat window display and audio output. The user can review this response and obtain optimal routing information.
[1205] Input: Response sent from the server
[1206] Output: Route information displayed to the user
[1207] This series of processes allows users to quickly obtain optimal route information in response to questions entered in natural language, enabling them to efficiently carry out food delivery operations.
[1208] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1209] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[1210] System-wide configuration
[1211] This system consists of a communication terminal used by the user, a server that receives requests, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotions.
[1212] Communication terminal
[1213] Users open the CrewNavi chat window using a communication device (e.g., a smartphone or personal computer). In the chat window, users can input questions about their current location and destination using natural language. For example, they can input, "What is the shortest route from Shibuya to Shinjuku?"
[1214] When the user clicks the "Submit" button, the communication terminal sends the entered question content to the server.
[1215] server
[1216] The server receives requests sent from communication terminals and parses the message content. The received message content is then passed to the conversational agent.
[1217] The server automatically handles everything from receiving requests to generating and sending responses. It also has a function to save received messages as logs. Furthermore, the server adjusts the content of its responses based on sentiment data analyzed by the sentiment engine.
[1218] Dialogue Agent
[1219] The conversational agent runs within the server and is responsible for analyzing the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question. Then, it searches the CrewNavi information database to obtain the optimal route information.
[1220] The acquired route information is converted into an appropriate natural language response by the conversational agent. For example, a response such as, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes," is generated.
[1221] Emotional Engine
[1222] The emotion engine analyzes the user's input and recognizes their emotional state. For example, if the user uses language that indicates frustration, the emotion engine will recognize that emotion as "anger." This emotional information is passed to the conversational agent and reflected in the generated response.
[1223] Response adjustment and return
[1224] The server adjusts the responses generated by the conversational agent based on the analysis results from the emotion engine. For example, if the user is irritated, the server may adjust the tone of the response to be calmer.
[1225] The adjusted response is sent back from the server to the communication terminal. The communication terminal receives this response and displays it in the chat window.
[1226] Specific examples
[1227] Here is a specific example. A user types "Please tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication device and presses the send button.
[1228] The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the CrewNavi information database and generates the response, "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey takes approximately 10 minutes."
[1229] Simultaneously, the emotion engine within the server recognizes the user's emotion of urgency from their expression, "I'm in a hurry!" Based on this emotion information, a response tone that is more responsive to the urgent situation is set.
[1230] This adjusted response is sent back to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[1231] The above describes an embodiment of a system incorporating an emotion engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[1232] The following describes the processing flow.
[1233] Step 1:
[1234] The user enters a question into the chat window of their communication device. For example, they might type, "What is the shortest route from Shibuya to Shinjuku?"
[1235] Step 2:
[1236] The user clicks the "Submit" button. This sends the question to the server via the communication device.
[1237] Step 3:
[1238] The terminal converts the question content into JSON data and sends it to the server using HTTP or WebSocket.
[1239] Step 4:
[1240] The server receives the request. The received data is converted into an appropriate format for analysis.
[1241] Step 5:
[1242] The server extracts the question content from the received request and passes it to the conversational agent. At this time, the question content is stored in natural language.
[1243] Step 6:
[1244] The conversational agent on the server analyzes the question. For example, it identifies keywords such as "Shibuya" or "Shinjuku."
[1245] Step 7:
[1246] The conversational agent searches CrewNavi's information database and generates the best possible response to a question. For example, it might generate a response like, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[1247] Step 8:
[1248] The server passes the response from the conversational agent to the emotion engine. At the same time, the user's input is also passed.
[1249] Step 9:
[1250] The emotion engine on the server analyzes the user's input and recognizes their emotional state. For example, it recognizes the emotion of urgency from the expression "I'm in a hurry!"
[1251] Step 10:
[1252] The emotion engine adjusts its response based on the emotional information it recognizes. For example, it might change the tone to one that indicates a more urgent response.
[1253] Step 11:
[1254] The adjusted response is returned to the server. The server converts this response into JSON data.
[1255] Step 12:
[1256] The server sends response data in JSON format back to the communication terminal. Communication is conducted via HTTP response or WebSocket message.
[1257] Step 13:
[1258] The communication terminal receives the response data. The received data is analyzed within the terminal and displayed in the chat window.
[1259] Step 14:
[1260] Users can see the responses displayed in the chat window of their communication device. For example, a message might appear stating, "The shortest route from Shibuya to Shinjuku is via the JR Yamanote Line. The journey takes approximately 15 minutes."
[1261] The above outlines the specific processing steps of this system.
[1262] (Example 2)
[1263] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1264] Currently, systems that allow users to obtain route information via communication terminals often fail to consider the user's emotional state, resulting in decreased user satisfaction. Furthermore, the inability to respond flexibly to user emotions makes it difficult to provide appropriate service, especially when the user is in a hurry or irritated. Improving this situation and enhancing user satisfaction is crucial.
[1265] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language via a communication terminal, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, means for using an emotion analysis engine to recognize the user's emotions, and means for adjusting the response content based on the emotion information recognized by the emotion analysis engine. This makes it possible to provide a detailed response that corresponds to the user's emotions.
[1266] A "communication terminal" is a device used by users to input questions and communicate with a server, and specifically refers to devices such as smartphones and personal computers.
[1267] A "server" is a computer device that receives questions from users, analyzes the messages, generates responses, and sends them back to the users.
[1268] A "question" refers to an inquiry about routing information or other related matters that a user inputs in natural language through a communication terminal.
[1269] A "dialogue agent" is a software module that analyzes the content of a user's question and generates an appropriate response, using natural language processing technology.
[1270] A "response" is a natural language response generated by a conversational agent and sent back to the user via the server.
[1271] An "emotion analysis engine" is a software module that recognizes the user's emotional state from their input and assigns emotional labels such as "joy," "anger," and "sadness."
[1272] "Emotional information" refers to data representing the user's emotional state, which is recognized by the emotion analysis engine and used to adjust responses by the dialogue agent.
[1273] A "route information provision system" is a database and software module that provides optimal route information based on a specific origin and destination.
[1274] This invention combines a system that allows users to input questions in natural language and provides quick and appropriate route information in response to those questions with an emotion analysis engine that recognizes the user's emotions. The embodiments thereof are described in detail below.
[1275] System-wide configuration
[1276] This system consists of the following hardware and software:
[1277] Communication devices used by the user (smartphones and personal computers)
[1278] Server that receives and analyzes messages
[1279] A conversational agent that generates responses to user questions.
[1280] A sentiment analysis engine that recognizes user emotions.
[1281] Communication terminal
[1282] Users can open a chat window using their communication device and enter questions in natural language. For example, they can type, "What is the shortest route from Shibuya to Shinjuku?" When the user clicks the "Send" button, the communication device sends this question to the server.
[1283] server
[1284] The server receives and analyzes the content of questions sent from communication terminals. The received messages are passed to the conversational agent. The server also uses an emotion analysis engine to analyze the user's emotions and passes the results to the conversational agent. The server automatically processes everything from receiving requests to generating and sending responses, and also has a function to save received messages as logs.
[1285] Dialogue Agent
[1286] The conversational agent runs on the server and analyzes the user's questions. For example, if a user asks, "What is the shortest route from Shibuya to Shinjuku?", the conversational agent first identifies the starting point and destination from the question, then searches the route information system (database) to obtain the optimal route information. The obtained route information is then converted into an appropriate natural language response by the conversational agent.
[1287] Emotion analysis engine
[1288] The emotion analysis engine analyzes the user's input and recognizes their emotional state. For example, if the user uses the expression "I'm in a hurry!", the emotion analysis engine recognizes that the user is feeling anxious. The analyzed emotional information is passed to the conversational agent and reflected in the generated response.
[1289] Response adjustment and return
[1290] The server adjusts the responses generated by the conversational agent based on the results of the emotion analysis engine. For example, if the user is irritated, the server adjusts the tone of the response to be calmer. The adjusted response is sent back from the server to the communication terminal, which receives the response and displays it in the chat window.
[1291] Specific examples
[1292] Here is a specific example. A user types "Tell me the route from Tokyo Station to Shinagawa Station, I'm in a hurry!" into the chat window of their communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the route information system and generates a response: "To get from Tokyo Station to Shinagawa Station, please use the JR Yamanote Line. The journey will take approximately 10 minutes." At the same time, the sentiment analysis engine recognizes the user's urgency from the expression "I'm in a hurry!" and adjusts the response tone based on this sentiment information. This adjusted response is returned to the server and sent back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[1293] The above describes an embodiment of a system incorporating an emotion analysis engine based on the present invention. By applying this system, users can not only easily and quickly obtain route information but also receive responses that correspond to their emotions.
[1294] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1295] Step 1: The user enters the question.
[1296] The user enters a question into the chat window of their communication device. This question is in natural language. For example, they might enter, "What is the shortest route from Shibuya to Shinjuku?" The input data is in text format. Specifically, the user uses a smartphone or computer keyboard to type the text and clicks the send button. The input content is stored as text data on the communication device.
[1297] Step 2: Send the question to the server
[1298] The terminal sends the user's input to the server. The input data is in text format, and the output is a data packet sent to the server. Specifically, the terminal converts the input text into a data packet format and sends it to the server over the internet. The data is transmitted according to the TCP / IP protocol.
[1299] Step 3: The server receives the question and analyzes it.
[1300] The server receives the question content sent from the terminal and parses the message content. The input data is the text data received from the terminal, and the output is the parsed question content. Specifically, the server receives the data packet and converts it to text format. Next, it parses the question content using natural language processing technology and saves it as a log.
[1301] Step 4: Pass the question to the conversational agent.
[1302] The server passes the analyzed question content to the conversational agent. The input data is the analyzed question content, and the output is data for the conversational agent to generate a response. Specifically, the server makes an API call to the conversational agent and passes the analyzed question content. The conversational agent then starts generating a response based on this content.
[1303] Step 5: Analyze emotions with an emotion analysis engine
[1304] The server sends the user's question to the sentiment analysis engine, which then analyzes the emotional state. The input data is the user's question, and the output is data with emotional labels. Specifically, the server makes an API call to the sentiment analysis engine and sends the input text. The sentiment analysis engine analyzes the emotions from the text and assigns emotional labels such as "joy," "anger," and "sadness."
[1305] Step 6: The conversational agent generates a response.
[1306] The conversational agent identifies the origin and destination from the question, searches a route information system, and generates the optimal route information. The input data is the question, and the output is a response in natural language. Specifically, the conversational agent queries the database, retrieves appropriate route information, and converts that information into natural language.
[1307] Step 7: The server adjusts its response based on the sentiment analysis results.
[1308] The server combines the responses received from the conversational agent with the results of the emotion analysis engine to adjust the final response. The input data is the generated response and emotion information, and the output is the adjusted response. Specifically, the server applies emotion-specific adjustment logic to adjust the tone and content of the response. For example, if the user is irritated, the tone of the response will be softened.
[1309] Step 8: Send the adjusted response back to the terminal.
[1310] The server sends a pre-arranged response to the communication terminal. The input data is the pre-arranged response, and the output is the data packet sent to the terminal. Specifically, the server converts the response into a packet format and sends it to the terminal over the internet.
[1311] Step 9: The terminal displays a response
[1312] The terminal displays the response received from the server in the chat window. The input data is text data received from the server, and the output is the response display that the user can see. Specifically, the terminal converts the received data packet into text format and displays it in the chat window. The user can then verify this and obtain routing information.
[1313] (Application Example 2)
[1314] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1315] Modern logistics centers require the rapid and accurate processing of a large number of packages and shipments. However, it is not easy for managers to efficiently track package locations and delivery routes and make optimal decisions. Furthermore, disregarding managers' emotions and urgency can lead to a poor user experience and reduced efficiency. This situation results in delays and errors, and a decline in overall operational efficiency.
[1316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input a question in natural language, means for transmitting the question to the server, means for the server to receive the question and generate a response using a dialogue agent, means for sending the response back to the communication terminal and displaying it to the user, and an emotion recognition engine that recognizes the user's emotional state and adjusts the response content based on the emotional state. As a result, the manager of the logistics center can quickly obtain information on the location of packages and the optimal delivery route using natural language, and an optimal response can be provided according to the manager's emotional state.
[1317] A "communication terminal" is a device that allows users to input questions in natural language and send them to a server.
[1318] A "server" is a computer system that receives questions from users, generates responses using dialogue agents, and sends them back to the communication terminal.
[1319] A "conversational agent" is a program that runs within a server, analyzes user questions, and generates appropriate responses.
[1320] An "emotion recognition engine" is a program that analyzes the user's emotional state from their natural language input and adjusts its response based on the results.
[1321] "Adjusting the response content" means changing the tone and details of the generated response based on the user's emotional state.
[1322] "Logistics" is a general term for the business processes involved in the storage, delivery, and management of packages and goods.
[1323] "Natural language" refers to text and audio composed of the language that users use on a daily basis.
[1324] "Means of sending questions" refers to the function of sending user questions to the server via a communication terminal.
[1325] "Means of sending and displaying a response" refers to a function that receives a response from a server on a communication terminal and displays it to the user.
[1326] "Route information" refers to data about the optimal route between a specified origin and destination.
[1327] This invention relates to a system that allows managers in logistics centers to input questions in natural language using a smartphone and receive prompt and appropriate route information in response. This system recognizes the user's emotional state and adjusts its responses accordingly to provide more effective support. Specific embodiments for carrying out this invention are described in detail below.
[1328] System-wide configuration
[1329] This system consists of a communication terminal used by the user, a server that receives questions, a dialogue agent that generates responses, and an emotion engine that recognizes the user's emotional state.
[1330] Communication terminal
[1331] The user launches the logistics management app using a communication device (e.g., a smartphone). In the app's chat window, the user can input questions about the package's location and delivery route using natural language. For example, they might input, "I need to send this package to point A as soon as possible, which route is best?" When the user clicks the "Send" button, the communication device sends the entered question to the server.
[1332] server
[1333] The server receives questions sent from communication terminals and forwards requests to conversational agents. The server also uses an emotion recognition engine to analyze the user's emotional state and adjusts the conversational agent's responses based on the results. Software used within the server includes Flask (a web framework) and OpenAI (a natural language processing engine).
[1334] Dialogue Agent
[1335] The conversational agent runs on the server and is responsible for analyzing the user's questions. Specifically, it uses the OpenAI API to generate optimal route information from the question. For example, if a user asks, "I urgently need to send this package to point A, which route is best?", the conversational agent searches a logistics information database and obtains the optimal route information. This route information is then converted into a natural language response by the conversational agent.
[1336] Emotional Engine
[1337] The emotion engine analyzes the user's input and recognizes their emotional state. For example, it recognizes the user's anxiety from expressions like "urgent" or "in a hurry." This emotional information is passed to the conversational agent and used to adjust the response content and tone.
[1338] Response adjustment and return
[1339] The server adjusts the responses generated by the conversational agent based on the analysis results of the emotion engine. For example, if the user is in a hurry, a message in a gentler tone, such as "The best route is as follows. Thank you for your understanding," may be added. The adjusted response is sent back from the server to the communication terminal and displayed to the user.
[1340] Specific examples
[1341] The user types "I urgently need to send this package to point A, what's the best route?" into the chat window of the communication terminal and presses the send button. The communication terminal sends this question to the server. The server receives the request and passes the question to the conversational agent. The conversational agent searches the logistics information database and generates a response: "The best route is as follows. Thank you for your understanding." At the same time, the emotion engine in the server recognizes the user's urgency from the expression "urgent." Based on this emotion information, the tone of the response is adjusted. This adjusted response is returned to the server, which then sends the response back to the communication terminal. The communication terminal displays the received response in the chat window, allowing the user to confirm the answer.
[1342] Examples of prompts to input into a generative AI model:
[1343] Provide the best route for the query: "I urgently need to send this package to point A. What is the best route?"
[1344] In this way, the present invention can improve work efficiency in logistics centers and reduce stress for managers.
[1345] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1346] Step 1:
[1347] The user uses the chat window on the communication terminal to input a question in natural language. For example, a question like, "I urgently need to send this package to point A, what is the best route?" might be entered. This input is converted into digital data within the communication terminal.
[1348] Step 2:
[1349] When the user clicks the "Send" button, the communication device sends this digital data to the server. The transmitted data includes the user's question.
[1350] Step 3:
[1351] The server receives a question from the user and passes the request to the conversational agent. The conversational agent analyzes this data to identify the question (the location and destination of the package). Natural language processing techniques are used for this analysis.
[1352] Step 4:
[1353] The conversational agent searches the logistics information database to obtain the optimal route information. An example of a prompt for this database search is: "Provide the best route for the query: 'I urgently need to send this package to point A, which route is best?'" The generated route information is returned to the conversational agent.
[1354] Step 5:
[1355] The emotion recognition engine on the server analyzes the user's emotional state from their input. In this example, the expression "urgent" is used to recognize that the user is in a hurry. The emotion recognition engine then generates this emotional data.
[1356] Step 6:
[1357] The server adjusts the response generated by the dialogue agent based on emotion data from the emotion recognition engine. A gentler tone or additional sentences depending on the urgency are added to the response. For example, a response like, "Thank you for your understanding. The optimal route is as follows..." might be created.
[1358] Step 7:
[1359] The server sends a coordinated response back to the communication terminal. The returned data includes coordinated routing information.
[1360] Step 8:
[1361] The communication terminal receives a response from the server and displays it in the chat window. This allows the user to quickly confirm the optimal route information.
[1362] Through the steps described above, the present invention constructs a system that supports the efficient execution of tasks by managers in logistics centers and provides rapid and appropriate route information.
[1363] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1364] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1365] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1366] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1367] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1368] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1369] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1370] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1371] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1372] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1373] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1374] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1375] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1376] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1377] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1378] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1379] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1380] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1381] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1382] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1383] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1384] The following is further disclosed regarding the embodiments described above.
[1385] (Claim 1)
[1386] A means by which the user inputs a question in natural language via a communication terminal,
[1387] Means for sending the aforementioned question to the server,
[1388] The server receives the question and a dialogue agent generates a response,
[1389] means for sending the aforementioned response back to the communication terminal and displaying it to the user,
[1390] A system that includes this.
[1391] (Claim 2)
[1392] The system according to claim 1, wherein the dialogue agent is configured to analyze the content of the user's question and generate an appropriate response.
[1393] (Claim 3)
[1394] The system according to claim 1, wherein the response is generated based on information from CrewNavi.
[1395] "Example 1"
[1396] (Claim 1)
[1397] A means by which the user inputs a question in natural language via a communication device,
[1398] Means for transmitting the aforementioned question to a data processing device,
[1399] The data processing device includes means for receiving the question and generating a response by a conversational agent,
[1400] means for sending the aforementioned response back to the communication device and displaying it to the user,
[1401] Means including inputting a question in a communication device, analyzing a question in a data processing device, searching an information database, and generating a response,
[1402] A system that includes this.
[1403] (Claim 2)
[1404] The system according to claim 1, wherein the conversational agent is configured to analyze the content of the user's question, identify the origin and destination using natural language processing technology, and generate an appropriate response.
[1405] (Claim 3)
[1406] The system according to claim 1, wherein the response is generated based on a database of routing information.
[1407] "Application Example 1"
[1408] Claims for a New Invention
[1409] (Claim 1)
[1410] A means by which the user inputs a question in natural language via a communication terminal,
[1411] Means for sending the aforementioned question to the server,
[1412] The server receives the question and a dialogue agent generates a response,
[1413] means for sending the aforementioned response back to the communication terminal and displaying it to the user,
[1414] A means configured to obtain optimal route information when a user performs food delivery work,
[1415] A system that includes this.
[1416] (Claim 2)
[1417] The system according to claim 1, wherein the dialogue agent is configured to analyze the content of the user's question and generate an appropriate response.
[1418] (Claim 3)
[1419] The system according to claim 1, wherein the response is generated based on information from the route guidance system.
[1420] "Example 2 of combining an emotion engine"
[1421] (Claim 1)
[1422] A means by which the user inputs a question in natural language via a communication terminal,
[1423] Means for sending the aforementioned question to the server,
[1424] The server receives the question and a dialogue agent generates a response,
[1425] means for sending the aforementioned response back to the communication terminal and displaying it to the user,
[1426] A method of using an emotion analysis engine to recognize the user's emotions,
[1427] Means for adjusting the response content based on emotional information recognized by the aforementioned emotion analysis engine,
[1428] A system that includes this.
[1429] (Claim 2)
[1430] The system according to claim 1, wherein the dialogue agent is configured to analyze the content of the user's question and generate an appropriate response.
[1431] (Claim 3)
[1432] The system according to claim 1, wherein the response is generated based on data from a routing information provision system.
[1433] "Application example 2 of combining emotional engines"
[1434] Claims for a New Invention
[1435] (Claim 1)
[1436] A means by which the user inputs a question in natural language via a communication terminal,
[1437] Means for sending the aforementioned question to the server,
[1438] The server receives the question and a dialogue agent generates a response,
[1439] means for sending the aforementioned response back to the communication terminal and displaying it to the user,
[1440] Means including an emotion recognition engine that recognizes the user's emotional state and adjusts the response content based on the said emotional state,
[1441] A system that includes this.
[1442] (Claim 2)
[1443] The system according to claim 1, wherein the dialogue agent is configured to analyze the content of the user's question and generate an appropriate response.
[1444] (Claim 3)
[1445] The system according to claim 1, wherein the response is generated based on logistics information. [Explanation of symbols]
[1446] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means by which the user inputs a question in natural language via a communication terminal, Means for sending the aforementioned question to the server, The server receives the question and a dialogue agent generates a response, means for sending the aforementioned response back to the communication terminal and displaying it to the user, A system that includes this.
2. The system according to claim 1, wherein the dialogue agent is configured to analyze the content of the user's question and generate an appropriate response.
3. The system according to claim 1, wherein the response is generated based on information from CrewNavi.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A