System

The system addresses user interaction challenges with AI by analyzing input, providing feedback, and refining responses, enhancing interaction quality and personalization through continuous learning and emotion recognition.

JP2026021181APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122863
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Users face challenges in effectively interacting with AI systems due to difficulty in asking appropriate questions, leading to decreased accuracy and usefulness of information provided, and a lack of personalized feedback and interaction data utilization.

Method used

A system that receives user input, analyzes it using natural language processing, provides feedback to clarify questions, learns from interaction data to improve response models, and generates refined responses, incorporating an emotion engine to enhance user experience.

Benefits of technology

The system clarifies user queries, improves interaction quality, and enhances the accuracy and personalization of AI responses by refining user questioning methods and continuously updating response models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021181000001_ABST
    Figure 2026021181000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving input data from a user; means for analyzing the input data; means for providing feedback to the user based on the analysis; means for improving how the user interacts based on the feedback; means for performing learning using the input data and historical interaction data to update a response model; and means for generating a refined response based on the updated response model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, with the evolution of generative AI models, there has been a growing demand for improving the quality of interactions between AI and users. However, there is a challenge in that users have difficulty learning how to ask appropriate questions. As a result, the accuracy and usefulness of the information provided by AI can decrease, and the user experience can be impaired. Furthermore, there is a lack of effective ways to utilize interaction data and provide personalized feedback to users. [Means for solving the problem]

[0005] The present invention provides a system that receives input data from a user, analyzes it, and provides feedback. Specifically, the system includes a means for receiving input data from a user, a means for analyzing the input data, a means for providing feedback to the user based on the analysis, a means for improving the user's interaction method based on the feedback, a means for learning using the input data and past interaction data to update a response model, and a means for generating a refined response based on the updated response model. This improves the user's questioning method and makes interactions with AI more effective. The system also includes a means for retrieving interaction data from a database and sharing the analysis results, enabling more highly personalized interactions and improving the user experience. Furthermore, the system analyzes the input data using a natural language processing model to identify ambiguous parts, helping the user ask more specific and useful questions.

[0006] "Input data" is text information such as questions or instructions sent by a user to the system.

[0007] The "means of analysis" refers to natural language processing algorithms and models that understand the content of input data and interpret its meaning and intent.

[0008] "Feedback" refers to a response such as advice or instructions for revision that the system provides to the user based on the analysis results.

[0009] The "interaction method" refers to the format and content of questions asked when a user interacts with a system.

[0010] "Learning means" refers to machine learning algorithms that allow the system to improve its response model using past interaction data.

[0011] An "answer model" is a model that allows the system to generate appropriate answers to user questions.

[0012] A "refinement response" is an answer from a system whose accuracy and usefulness have been improved through learning and analysis.

[0013] A "database" is a storage device that stores dialogue data, analysis results, etc.

[0014] A "natural language processing model" is an algorithm or technology for analyzing human language and understanding its meaning.

[0015] "Ambiguous parts" are parts of the input data that are difficult to interpret or understand, especially ambiguous words or contexts. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. This system includes elements such as "input data," "analysis means," "feedback," "dialogue method," "learning means," "response model," "refined response," "database," "natural language processing model," and "ambiguous parts."

[0038] Server Roles

[0039] The server performs the following main functions:

[0040] 1. Receiving and storing data: The server receives the input data sent by the user via the terminal and stores it in a database.

[0041] 2. Data analysis: The server analyzes the received input data using a natural language processing model to understand the meaning and intent of the input.

[0042] 3. Feedback generation: Based on the analysis results, feedback is generated to help users improve their questioning methods.

[0043] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0044] 5. Run learning: Using collected input data and past interaction data, the machine learning algorithm runs and continuously updates and improves the response model.

[0045] 6. Response Generation: Based on the updated response model, a refined response to the user's question is generated.

[0046] Device Role

[0047] The terminal acts as a relay for the interaction between the user and the server:

[0048] 1. Sending input data: Questions and instructions entered by the user are sent to the server in real time.

[0049] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0050] User Roles

[0051] The user is the direct interlocutor of the system and goes through the following process:

[0052] 1. Entering a question: Enter a question or instruction through the terminal.

[0053] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0054] Specific examples

[0055] As an example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server receives this data and analyzes it using a natural language processing model. As a result, it detects that the question is ambiguous and generates feedback such as "Please specify a specific time." The server sends this feedback to the device, which then displays it to the user.

[0056] The user checks this feedback and resubmits the question, revising it to "What's the weather like in Shinjuku tomorrow at 3 PM?" The server analyzes the data again, this time including a specific time, generating a refined answer and sending it to the device. The device then displays this answer to the user, completing the dialogue.

[0057] In this way, the present invention provides a system that can refine the content of user questions and improve the quality of dialogue with AI. Furthermore, by utilizing the database for continuous learning, the system can improve the response model and enhance the user experience.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[0061] Step 2:

[0062] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[0063] Step 3:

[0064] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[0065] Step 4:

[0066] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[0067] Step 5:

[0068] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[0069] Step 6:

[0070] The server generates feedback to the user based on the analysis results, such as a message saying "Please specify a specific time."

[0071] Step 7:

[0072] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[0073] Step 8:

[0074] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[0075] Step 9:

[0076] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[0077] Step 10:

[0078] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[0079] Step 11:

[0080] The server verifies that the revised question is specific and generates a refined response, such as "The weather in Shinjuku tomorrow at 3 PM will be sunny."

[0081] Step 12:

[0082] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[0083] Step 13:

[0084] The user confirms the refined response received through the terminal, thereby completing the interaction.

[0085] Step 14:

[0086] The server runs a machine learning algorithm based on the current dialogue data and updates the response model, enabling even more accurate responses the next time the dialogue occurs.

[0087] By following the above steps, the system of the present invention effectively manages and improves the interaction between the user and the server, achieving higher quality communication.

[0088] Example 1

[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0090] In conventional dialogue systems, if a user's question or instruction is ambiguous, the question itself cannot be accurately analyzed, resulting in an inability to provide an appropriate response. Furthermore, due to a lack of means to provide clear feedback, users are forced to go through a lot of trial and error. Furthermore, the response model is not continuously improved, which can lead to a decline in the quality of the dialogue.

[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0092] In this invention, the server includes means for receiving input data from a user via a terminal, means for analyzing the input data using a natural language processing model to understand the meaning and intent of the input data, means for providing feedback to the user based on the analysis, means for correcting the user's question based on the feedback, means for learning using the input data and past dialogue data and updating a response model, means for generating a refined response based on the updated response model, and means for transmitting the generated refined response to the terminal and displaying it. This makes it possible to specify the content of the user's question and improve the quality of the response provided by AI.

[0093] A "terminal" is a device through which a user inputs questions and instructions and transmits and receives data to and from a server.

[0094] "Input data" refers to data including questions and instructions sent by a user via a terminal.

[0095] A "server" is a device or system that receives input data sent from a terminal and performs processing such as analysis, feedback generation, and response generation.

[0096] A "natural language processing model" is a general term for algorithms and models that analyze input data and understand its meaning and intent.

[0097] "Feedback" refers to answers or instructions that the server provides to the user based on the analysis results, and is used to clarify the content of the user's question.

[0098] "Analysis" is the process of using a natural language processing model to understand the meaning and intent of input data and make that content concrete.

[0099] "Learning" is the process by which the server runs machine learning algorithms based on input data and past interaction data to continuously update and improve the response model.

[0100] A "response model" is a model for generating answers to user questions, and is updated based on learning by the server.

[0101] An "elaborate response" is a specific and clear answer that is generated to provide appropriate information to a user's question.

[0102] A "database" is a storage device that stores input data and past dialogue data and can be accessed by a server for analysis and learning.

[0103] An "ambiguous part" is a part of a user's question or instruction that is unclear or difficult to interpret, and feedback is generated by identifying this.

[0104] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. In this system, the server receives data entered by the user via the terminal, and performs processing such as analysis, feedback generation, and response generation to concretize the content of the user's question and provide a high-quality response using AI. Specific embodiments of the present invention are described below.

[0105] The server mainly uses the following hardware and software:

[0106] Server hardware: A server machine with a powerful CPU and sufficient memory

[0107] Database: A database system (e.g., MySQL, PostgreSQL) for storing input data and past dialogue data.

[0108] Natural language processing model: A natural language processing model (e.g., GPT-4) to analyze the input data.

[0109] Machine learning algorithms: Algorithms for learning and updating response models using past data (e.g., TensorFlow, PyTorch)

[0110] The terminal acts as a relay for data between the user and the server. The terminal can be a device such as a PC, smartphone, or tablet that runs an application or web browser that communicates with the server via an internet connection.

[0111] The user operates the terminal to input questions and instructions to the system. The operation of the system will be explained below using a concrete example.

[0112] As a concrete example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server analyzes the received data using a natural language processing model (GPT-4) to understand the intent of the question. If the server detects that the question is ambiguous, it generates feedback such as "Please specify a specific time." The server then sends this feedback to the device, which displays it to the user.

[0113] The user checks the feedback from the server, modifies the question, and re-enters it. For example, the user enters a specific question such as, "What is the weather like in Shinjuku tomorrow at 3 PM?" The device then sends the modified data to the server again. The server again analyzes the data and generates a more specific answer this time. It generates a refined response, "The weather will be cloudy in Shinjuku tomorrow at 3 PM," and sends it to the device. The device then displays the received response to the user.

[0114] In this way, the system can clarify the user's ambiguous questions and provide appropriate feedback, improving the quality of the AI ​​responses. Furthermore, the server uses input data and past dialogue data to perform machine learning and continuously update and improve the response model, improving the user experience.

[0115] Prompt Sentence Examples

[0116] First input: "What's the weather like in Shinjuku tomorrow?"

[0117] Corrected input: "What's the weather like in Shinjuku tomorrow at 3pm?"

[0118] In this way, the present invention provides a system that clarifies the content of a user's question and improves the quality of dialogue with AI.

[0119] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0120] Step 1: User enters question

[0121] The user inputs a question or instruction into the terminal. For example, the user inputs "Please tell me what the weather will be like in Shinjuku tomorrow." The input text becomes the input data for the terminal.

[0122] Step 2: Submitting input data

[0123] The device sends the user-entered question data to the server in real time via an HTTP POST request, with the input data included in the body of the API request.

[0124] Step 3: Receiving and storing data

[0125] The server receives the input data sent from the terminal. The received data is saved in a database on the server. The database stores the user's question ("What is the weather in Shinjuku tomorrow?").

[0126] Step 4: Analysis using natural language processing models

[0127] The server analyzes the received input data using a natural language processing model (GPT-4). The input data is input into the model, and an analysis result is generated. The analysis result includes the intent of the question and any ambiguities.

[0128] Step 5: Generate feedback

[0129] The server generates feedback based on the analysis results of the natural language processing model. For example, if the question is ambiguous, the server generates feedback such as "Please specify a specific time." The generated feedback is in text format.

[0130] Step 6: Submit your feedback

[0131] The server sends the generated feedback to the terminal. A feedback message ("Please specify a specific time") is sent as an HTTP response.

[0132] Step 7: View your feedback

[0133] The terminal displays the feedback received from the server to the user, and the feedback message is displayed on the screen of the terminal to notify the user.

[0134] Step 8: Modifying the Question

[0135] The user modifies the question based on the received feedback. For example, the user inputs a specific question such as, "Please tell me the weather in Shinjuku tomorrow at 3:00 PM." The modified question becomes the new input data.

[0136] Step 9: Resend the corrected data

[0137] The terminal sends the corrected input data to the server again via an HTTP POST request. The corrected input data is included in the body of the API request.

[0138] Step 10: Reanalyze the data

[0139] The server then analyzes the revised data using a natural language processing model. The revised data is input into the model, and an analysis result is generated. The analysis result for the specific question is obtained.

[0140] Step 11: Generate a response

[0141] The server generates a refined response based on the analysis results. For example, it may generate an answer such as "The weather in Shinjuku tomorrow at 3:00 PM will be cloudy." The generated response is in text format.

[0142] Step 12: Sending a Response

[0143] The server sends the generated response to the terminal. The response message ("The weather in Shinjuku tomorrow at 3 PM will be cloudy") is sent as an HTTP response.

[0144] Step 13: View the response

[0145] The terminal displays the response received from the server to the user. The response message is displayed on the terminal screen to notify the user. This display completes the interaction.

[0146] (Application example 1)

[0147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0148] In today's retail industry, when customers ask questions about products in physical stores, staff are expected to respond quickly and accurately. However, many stores struggle to immediately provide customers with the specific product and inventory information they require, resulting in lower customer satisfaction and lost sales opportunities. Furthermore, the quality of the information provided varies depending on the staff's knowledge and experience, resulting in an inconsistent customer experience. Therefore, there is a growing need for a system that enables staff in physical stores to efficiently respond to customers and provide accurate information.

[0149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0150] In this invention, the server includes means for receiving input data from a user, means for analyzing the input data, means for providing feedback to the user based on the analysis, means for improving the user's interaction method based on the feedback, means for learning using the input data and past interaction data and updating a response model, means for generating refined responses based on the updated response model, means for supporting interaction with the user in real time using a smart device, and means for providing product information and inventory information based on the support, thereby enabling staff to provide product information and inventory information to customers quickly and accurately using the smart device.

[0151] "Means for receiving input data from a user" refers to the function of collecting data through an interface for users to input questions or instructions and sending it to the system.

[0152] The "means for analyzing the input data" refers to a function that analyzes the received data using natural language processing technology, etc., and performs processing to understand its meaning and intent.

[0153] The "means for providing feedback to the user based on the analysis" refers to a function that provides advice and correction instructions to the user to ask more specific questions based on the analysis results.

[0154] "Means for improving the user's interaction method based on the feedback" refers to a process by which the user understands the feedback and improves the content of the questions and instructions.

[0155] "Means for learning using the input data and past dialogue data and updating the response model" refers to a function that performs machine learning based on collected data, continuously improves the system's response model, and enables it to generate more accurate responses.

[0156] "Means for generating a refined response based on the updated response model" refers to a function that uses an improved response model to generate a specific and accurate answer to a user's question.

[0157] "Means for supporting real-time interaction with users using smart devices" refers to a function for using devices such as smart glasses or smartphones to interact with users in real time and provide specific information.

[0158] "Means for providing product information and inventory information based on the support" refers to a function that retrieves product information and inventory status from a database based on a user's question and presents that information to the user.

[0159] The present invention provides a system for supporting interactions with customers in a physical store. A specific embodiment of this system will be described below.

[0160] System Program

[0161] 1. Hardware configuration:

[0162] It uses servers, smart devices (such as smart glasses and smartphones), and communication networks (Wi-Fi or 5G). It uses AWS EC2 instances for the servers, Amazon RDS for the database, and AWS Sagemaker for natural language processing.

[0163] 2. Software configuration:

[0164] An Android-based application is installed on the smart device, and a Node.js server is operated on the server side. Python is used for data analysis, and BERT is used as the natural language processing model.

[0165] System Operation Overview

[0166] 1. Getting and sending user input data:

[0167] The user (staff member) uses a smart device to input customer questions by voice, and the smart device converts the voice data into text using the Google Speech-to-Text API and sends the text data to the server.

[0168] 2. Data Analysis:

[0169] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker), which identifies the intent of the question and any ambiguities, and generates appropriate feedback.

[0170] 3. Feedback and response generation:

[0171] Based on the analysis results, the server generates feedback to help the user refine their question. If necessary, it sends additional questions or correction instructions to the smart device. If the user re-enters the necessary information, the server analyzes the data again and generates a refined response.

[0172] 4. Information provision:

[0173] The refined response includes product and inventory information retrieved from a database (Amazon RDS), allowing users to provide accurate and prompt information to customers.

[0174] Specific examples

[0175] Example 1

[0176] Scenario: A customer asks, "Tell me about this product."

[0177] Prompt statement:

[0178] Provide specific product information. You need to tell customers the description, price, and features of this product.

[0179] Example response:

[0180] The smart device will display the message, "This product is the latest model and costs 5,000 yen. Its features include waterproof functionality and a long battery life."

[0181] Example 2

[0182] Scenario: A customer asks, "Is this item in stock?"

[0183] Prompt statement:

[0184] Please check the inventory information and provide it to your customers. We need to check the current availability of this item.

[0185] Example response:

[0186] The smart device will display the message, "There are currently 5 of this item in stock in the store."

[0187] In this way, the present invention provides a specific system configuration and processing procedure for efficiently and accurately handling customers in a brick-and-mortar store.

[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0189] Step 1: Getting and sending input data from the user

[0190] A user (staff member) uses a smart device to voice-input a customer question. The smart device then converts the voice data into text using the Google Speech-to-Text API. The input is voice data, and the output is text data. The smart device then sends the converted text data to the server.

[0191] Step 2: Analyze the data

[0192] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker). Specifically, it processes and calculates the data to identify the intent and ambiguity of the text data. The input is the text data, and the output is the analysis results (intent, ambiguity).

[0193] Step 3: Generate feedback

[0194] Based on the analysis results, the server generates feedback to help the user refine their query. The feedback includes additional questions and correction instructions as needed. The input is the analysis results, and the output is a feedback message. This is sent to the smart device and displayed to the user.

[0195] Step 4: User Modification of Question

[0196] The user re-enters a question based on the feedback and sends the revised text data to the server. The input is the feedback message and the revised question, and the output is the text data of the revised question.

[0197] Step 5: Reparsing and generating a response

[0198] The server then analyzes the revised text data again using a natural language processing model. Based on the analysis results, it generates a refined response. The input is the text data of the revised question, and the output is the refined response.

[0199] Step 6: Provide information

[0200] Based on the refined response, the server retrieves product information and inventory information from the database (Amazon RDS) and generates information to be provided to the user. The input is the refined response and information from the database, and the output is the final information provided to the user.

[0201] Step 7: Display to the user

[0202] The final information is sent to the smart device and displayed to the user, allowing the user to provide accurate product and inventory information to customers. The input is the final information from the server, and the output is the information displayed on the smart device.

[0203] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0204] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. It also incorporates an emotion engine that recognizes the user's emotions and adjusts or enhances responses based on those emotions. This system includes elements such as input data, analysis means, feedback, a dialogue method, learning means, response model, refined responses, a database, a natural language processing model, ambiguity, and an emotion engine.

[0205] Server Roles

[0206] The server performs the following functions:

[0207] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[0208] 2. Data analysis: The received input data is analyzed using a natural language processing model and an emotion engine to understand the meaning, intent, and emotional state of the input.

[0209] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[0210] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0211] 5. Run Learning: Using collected input data and past interaction data, machine learning algorithms are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[0212] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[0213] Device Role

[0214] The terminal acts as a relay for the interaction between the user and the server:

[0215] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[0216] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0217] User Roles

[0218] The user is the direct interlocutor of the system and goes through the following process:

[0219] 1. Entering a question: Enter a question or instruction through the terminal.

[0220] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0221] Specific examples

[0222] As an example, consider a question about the weather. The user enters "What will the weather be like in Shinjuku tomorrow?" into the device and sends it. The device then sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server then sends this feedback to the device, which then displays it to the user.

[0223] The user checks the feedback, modifies the question to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[0224] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[0225] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, the response model and emotion engine can be improved, resulting in a better user experience.

[0226] The processing flow will be explained below.

[0227] Step 1:

[0228] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[0229] Step 2:

[0230] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[0231] Step 3:

[0232] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[0233] Step 4:

[0234] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[0235] Step 5:

[0236] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[0237] Step 6:

[0238] The server uses an emotion engine to analyze the user's emotional state, for example, by inferring the user's emotions from text and detecting stress, joy, anxiety, etc.

[0239] Step 7:

[0240] The server generates feedback to the user based on the analysis results, such as "Please specify a specific time," and adds expressions that take into account the user's emotional state.

[0241] Step 8:

[0242] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[0243] Step 9:

[0244] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[0245] Step 10:

[0246] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[0247] Step 11:

[0248] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[0249] Step 12:

[0250] The server then verifies that the revised question is specific and generates a refined response, adding expressions that take the user's emotional state into consideration. For example, the server generates a response such as, "The weather in Shinjuku tomorrow at 3 p.m. will be sunny. Have a nice afternoon."

[0251] Step 13:

[0252] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[0253] Step 14:

[0254] The user confirms the refined response received through the terminal, thereby completing the interaction.

[0255] Step 15:

[0256] The server runs a machine learning algorithm based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses in the next conversation.

[0257] By following the above steps, the system of the present invention effectively manages and improves the dialogue between the user and the server, realizing higher quality and more emotionally sensitive communication.

[0258] Example 2

[0259] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0260] Conventional dialogue systems lack natural and rich communication because they do not consider the user's emotional state. They also fail to properly handle ambiguity in user questions and instructions, resulting in insufficient feedback and an inability to quickly respond to user needs. Furthermore, they lack a means to effectively utilize past dialogue data to continuously improve the system's response model, limiting the user experience.

[0261] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data from a user and storing it in a database, means for analyzing the input data using a natural language processing model and an emotion engine to understand the meaning, intention, and emotional state of the input data, means for generating feedback corresponding to the user's question method and emotional state based on the analysis and transmitting the feedback to the terminal, means for improving the user's dialogue method based on the feedback, means for executing a machine learning algorithm using the input data and past dialogue data to continuously update a response model and emotion engine, and means for generating refined responses based on the updated response model and adjusting them based on emotion understanding. This makes it possible to realize natural and effective dialogue that takes the user's emotional state into consideration, reduce ambiguity in the user's questions and instructions, and continuously improve the system's response model.

[0262] A "user" is an entity that provides input data to and receives feedback from a system.

[0263] "Input Data" refers to questions, instructions, or other information provided by a user to a system.

[0264] A "database" is a data management system that stores input data and past dialogue data and retrieves them as needed.

[0265] A "natural language processing model" is a machine learning model that analyzes human language and understands its meaning and intent.

[0266] An "emotion engine" is an analytical system for detecting a user's emotional state from input data and adjusting responses accordingly.

[0267] "Analysis" refers to processing input data using a natural language processing model and emotion engine to understand its meaning, intent, and emotional state.

[0268] "Feedback" refers to the response or instructions provided to the user based on the analysis results.

[0269] "Terminal" refers to a device used by a user to provide input data to a system.

[0270] A "machine learning algorithm" is a computational method for training a model based on data and continuously improving its performance.

[0271] A "response model" is a model for generating appropriate responses to user questions and instructions.

[0272] "Refined responses" refer to detailed responses generated based on an updated response model and tailored based on emotion understanding.

[0273] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. The present invention incorporates an emotion engine that recognizes the user's emotions and adjusts and enhances responses based on those emotions. The system includes the following elements:

[0274] Server Roles

[0275] Data Receipt and Storage:

[0276] The server receives input data sent by the user via the terminal and stores the data in a database. The database used is, for example, MySQL. When a user enters "What is the weather in Shinjuku tomorrow?", the server receives the data via an HTTP POST request and stores it in the database.

[0277] Data analysis:

[0278] The server analyzes the received input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). This allows it to understand the meaning, intent, and emotional state of the input. The server analyzes the input data, "What's the weather in Shinjuku tomorrow?" and extracts the keywords "Shinjuku," "tomorrow," and "weather," as well as the user's emotional state (curiosity).

[0279] Generate feedback:

[0280] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. This feedback is generated using a response model (e.g., GPT-3). It detects ambiguous time specifications and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration.

[0281] Send feedback:

[0282] The generated feedback is sent to the terminal and displayed to the user. The server sends feedback to the terminal saying "Please specify a specific time," which is displayed to the user.

[0283] Run the training:

[0284] The server uses the collected input data and past dialogue data to run machine learning algorithms (e.g., TensorFlow or PyTorch) and continuously update and improve the response model and emotion engine. The server inputs the dialogue history from the database into TensorFlow to retrain the response model and emotion engine.

[0285] Response generation:

[0286] Based on the updated response model, the server generates a refined response to the user's question and adjusts it based on sentiment understanding. In response to the question, "What is the weather like in Shinjuku tomorrow at 3 PM?", the server generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." and sends it to the device.

[0287] Device Role

[0288] Sending input data:

[0289] The device sends questions, instructions, and emotion-related data entered by the user to the server in real time. For example, if a user enters "What's the weather in Shinjuku tomorrow?" and presses the send button, the data is sent to the server.

[0290] Receiving and viewing feedback:

[0291] The terminal receives the feedback from the server and displays it in an easy-to-understand manner for the user. The terminal displays the feedback received from the server, "Please specify a specific time," in the chat window.

[0292] User Roles

[0293] Enter your question:

[0294] The user inputs questions or instructions through the terminal. The user types "Please tell me the weather in Shinjuku tomorrow" into the text box on the terminal and clicks the send button.

[0295] Review and correct feedback:

[0296] The user checks the feedback sent from the server, modifies the question if necessary, and resubmits it. The user sees the feedback "Please specify a specific time," modifies the question to "What is the weather like in Shinjuku at 3 PM tomorrow," and submits it again.

[0297] Specific examples

[0298] For example, consider the case where a user inputs and sends "What's the weather in Shinjuku tomorrow?" into a terminal. The terminal sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server sends this feedback to the terminal, which then displays it to the user.

[0299] The user checks the feedback, amends it to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[0300] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[0301] Prompt Sentence Examples

[0302] Here are some examples of prompts for generative AI models:

[0303] Prompt statement:

[0304] Use an emotion engine to generate a response to the user's question, "What will the weather be like in Shinjuku tomorrow?" If no specific time is specified, provide feedback, and generate a specific response when the question is entered again.

[0305] Example 1: User's first question

[0306] User: What's the weather like in Shinjuku tomorrow?

[0307] System: What time of day would you like to know the weather in Shinjuku tomorrow? Please specify a specific time.

[0308] Example 2: User question after modification

[0309] User: What's the weather like in Shinjuku tomorrow at 3pm?

[0310] System: The weather in Shinjuku will be sunny tomorrow at 3 PM. Have a nice afternoon.

[0311] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, it is possible to improve the response model and emotion engine, thereby improving the user experience.

[0312] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0313] Step 1:

[0314] Sending input data

[0315] The user types "Please tell me the weather in Shinjuku tomorrow" into the terminal and clicks the send button.

[0316] The terminal sends this input data to the server as an HTTP POST request.

[0317] Input: User question: "What's the weather like in Shinjuku tomorrow?"

[0318] Output: Data sent to the server as an HTTP POST request

[0319] Step 2:

[0320] Receiving and storing data

[0321] The server receives the input data submitted by the user. This data is stored in a database, such as MySQL. The server extracts the data from the POST request and stores it in the "inputs" table in the database.

[0322] Input: HTTP POST request

[0323] Output: Input data stored in a database

[0324] Step 3:

[0325] Data analysis

[0326] The server analyzes the stored input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). Through this analysis, the meaning and intent of the text and the user's emotional state are extracted. Specifically, keywords are extracted from the text and the emotional state is detected.

[0327] Input: Input data stored in the database

[0328] Output: Keywords and emotional state (e.g., "Shinjuku," "tomorrow," "weather," where the emotion is curiosity)

[0329] Step 4:

[0330] Generate feedback

[0331] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. It uses a response model (e.g., GPT-3) to detect ambiguous time specifications and generates feedback such as "Please specify a specific time." The feedback includes wording that takes the user's emotions into consideration.

[0332] Input: Keywords and emotional states

[0333] Output: Feedback: "Please specify a specific time."

[0334] Step 5:

[0335] Send Feedback

[0336] The server sends the generated feedback to the terminal as an HTTP response. The terminal receives this feedback and displays it in an easy-to-understand manner for the user. Specifically, the feedback is displayed in a chat window or similar.

[0337] Input: "Please specify a specific time" feedback

[0338] Output: Feedback displayed on the terminal

[0339] Step 6:

[0340] Review and correct feedback

[0341] The user checks the feedback from the server and modifies the question, re-entering "What is the weather like in Shinjuku tomorrow at 3 PM?" and clicking the resend button. The device then resends the modified question to the server.

[0342] Input: User's revised question "What's the weather like in Shinjuku tomorrow at 3pm?"

[0343] Output: Resent data

[0344] Step 7:

[0345] Reparsing and generating a response

[0346] The server receives the data again and analyzes it using the natural language processing model and emotion engine. Based on the results, it generates a specific response. For example, it generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon."

[0347] Input: Corrected input data

[0348] Output: The specific response generated

[0349] Step 8:

[0350] Sending a Response

[0351] The server generates a specific response and sends it to the terminal, which receives it and displays it to the user.

[0352] Input: Specific response

[0353] Output: Response displayed on the terminal

[0354] Step 9:

[0355] Execution of training

[0356] The server uses the collected input data and past dialogue data to run machine learning algorithms and continuously update and improve the response model and emotion engine. The dialogue history from the database is input into TensorFlow to retrain the model.

[0357] Input: Collected input data and past interaction data

[0358] Output: Updated response model and emotion engine

[0359] (Application example 2)

[0360] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0361] In recent years, improving the customer experience in brick-and-mortar stores has become increasingly important, but conventional systems have struggled to respond appropriately to customer questions and requests in real time. Furthermore, there is a lack of technology to provide feedback that takes into account the customer's emotional state, limiting the improvement of customer satisfaction. In particular, there is a need for an effective concierge service that can respond to ambiguous questions and emotional changes.

[0362] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving input data from a user; means for analyzing the input data; means for providing feedback to the user based on the analysis; means for improving the user's interaction method based on the feedback; means for learning using the input data and past interaction data and updating a response model; means for generating a refined response based on the updated response model; means for analyzing the input data using a natural language processing model and identifying ambiguous parts; and means for providing a concierge service in response to customer questions in a physical store, including an emotion engine that recognizes the user's emotional state and adjusts responses based on the user's emotional state. This makes it possible to respond to customer questions in real time and provide feedback that takes the user's emotional state into consideration, thereby significantly improving customer satisfaction.

[0363] "Input data" refers to questions, instructions, and other text information that a user sends to the system via a terminal.

[0364] "Means for analyzing" are the technical means for processing received input data and understanding its meaning, intent, and emotional state.

[0365] "Feedback" refers to responses or advice provided to the user based on the analysis results, and is information used to respond to the user's questions or requests.

[0366] "Means for improving the interaction method" are technical measures that use user feedback and correction data to enable the system to respond more accurately the next time the interaction is performed.

[0367] The "learning means" is a machine learning algorithm that uses past interaction data and newly collected input data to continuously improve the system's response model.

[0368] A "response model" is a model within the system that generates optimal responses to user questions and instructions.

[0369] A "refined response" is generated based on the updated response model and is a very specific and relevant answer to the user's question.

[0370] A "natural language processing model" is an artificial intelligence technology that analyzes input data from a user as language and understands its meaning and structure.

[0371] An "ambiguous part" is a part of the user's input data where the intent of the question or request is unclear.

[0372] An "emotion engine" is a technology that analyzes a user's emotional state and adjusts feedback and responses according to that emotional state.

[0373] A "physical store" is a store or shop that exists in a physical location and provides products and services in a face-to-face manner to customers.

[0374] A "concierge service" is a service provided in a physical store that responds to customer questions and requests and provides information about and offers products and services.

[0375] A "database" is a storage device and its management system for centrally storing input data and dialogue data collected by the system.

[0376] This invention provides a system for realizing a concierge service that provides highly accurate and emotionally sensitive feedback in real time when a customer inputs a question or request in a store. Hereinafter, an embodiment of the invention will be described in detail.

[0377] System Configuration

[0378] server

[0379] The server performs the following functions:

[0380] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[0381] 2. Data analysis: The received input data is analyzed using a natural language processing model (e.g., BERT or GPT) and an emotion engine (e.g., Microsoft Azure Emotion API) to understand the meaning, intent, and emotional state of the input.

[0382] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[0383] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0384] 5. Run learning: Using the collected input data and past interaction data, machine learning algorithms (e.g., TensorFlow) are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[0385] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[0386] Terminal

[0387] The terminal acts as a relay for the interaction between the user and the server:

[0388] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[0389] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0390] User

[0391] The user is the direct interlocutor of the system and goes through the following process:

[0392] 1. Entering a question: Enter a question or instruction through the terminal.

[0393] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0394] Specific examples

[0395] For example, a user might enter "I'm looking for a special gift today. Which one would be good?" into a terminal in a physical store and send it. The terminal sends this input data in real time to a server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it recognizes high expectations and generates feedback such as "I see you're looking for a special gift today. Here are some recommended products suitable for your special occasion. We'd be happy to help you make a great choice." This feedback is displayed to the user through the terminal.

[0396] Prompt Sentence Examples

[0397] By inputting the following prompt into the generative AI model, we generate a response based on high expectations:

[0398] "Generate a response that reflects high expectations for a customer looking for a special gift."

[0399] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0400] Step 1:

[0401] The user inputs and sends questions and instructions through the terminal.

[0402] Input: Text data of user questions and instructions

[0403] Output: The input data sent

[0404] What happens: Using a smartphone or smart glasses, the user types, "I'm looking for a special gift today. What would be good?" and submits that data.

[0405] Step 2:

[0406] The terminal transmits input data to the server in real time.

[0407] Input: Input data submitted by the user

[0408] Output: The request containing the input data sent

[0409] Specific operation: The terminal receives the user's input data and sends it to the server via the Internet.

[0410] Step 3:

[0411] The server receives the input data and stores it in a database.

[0412] Input: Input data sent from the terminal

[0413] Output: Input data stored in a database

[0414] Specific operation: The server's data receiving system receives the input data and stores it in a database, including the user's ID and timestamp.

[0415] Step 4:

[0416] The server analyzes the received input data using a natural language processing model (e.g., BERT or GPT).

[0417] Input: Saved input data

[0418] Output: Analysis results (question intent, content, related keywords)

[0419] What it does: The model on the server analyzes the input data and extracts the meaning and intent of the text.

[0420] Step 5:

[0421] The server uses the analysis results to understand the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API).

[0422] Input: Analysis results

[0423] Output: Analysis results including emotional state

[0424] Specific operation: The server inputs the analysis results into the emotion engine to detect high expectations and other emotions.

[0425] Step 6:

[0426] The server generates feedback based on the analysis results and emotional state.

[0427] Input: Analysis results including emotional state

[0428] Output: Feedback text data

[0429] What it does: Based on the analysis results and emotional state, the server generates feedback such as, "You're looking for a special gift today. Here are some recommended products for your special occasion. We'd be happy to help you make a great choice."

[0430] Step 7:

[0431] The server generates feedback and sends it to the device.

[0432] Input: Text data of generated feedback

[0433] Output: Feedback data sent

[0434] Specific operations: The server generates and sends a request to send feedback data to the terminal.

[0435] Step 8:

[0436] The terminal receives the feedback and displays it to the user.

[0437] Input: Feedback data sent from the server

[0438] Output: Displayed feedback

[0439] Specific operation: The device displays the received feedback data on the screen, and the user checks it and decides the next action.

[0440] Step 9:

[0441] The server uses collected input and feedback data to learn and update the response model.

[0442] Input: Input and feedback data

[0443] Output: Updated response model

[0444] What it does: Machine learning algorithms on the server continuously train and improve response models based on collected interaction data.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0446] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0448] [Second embodiment]

[0449] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0450] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0451] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0453] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0455] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0456] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0457] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0460] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0461] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. This system includes elements such as "input data," "analysis means," "feedback," "dialogue method," "learning means," "response model," "refined response," "database," "natural language processing model," and "ambiguous parts."

[0462] Server Roles

[0463] The server performs the following main functions:

[0464] 1. Receiving and storing data: The server receives the input data sent by the user via the terminal and stores it in a database.

[0465] 2. Data analysis: The server analyzes the received input data using a natural language processing model to understand the meaning and intent of the input.

[0466] 3. Feedback generation: Based on the analysis results, feedback is generated to help users improve their questioning methods.

[0467] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0468] 5. Run learning: Using collected input data and past interaction data, the machine learning algorithm runs and continuously updates and improves the response model.

[0469] 6. Response Generation: Based on the updated response model, a refined response to the user's question is generated.

[0470] Device Role

[0471] The terminal acts as a relay for the interaction between the user and the server:

[0472] 1. Sending input data: Questions and instructions entered by the user are sent to the server in real time.

[0473] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0474] User Roles

[0475] The user is the direct interlocutor of the system and goes through the following process:

[0476] 1. Entering a question: Enter a question or instruction through the terminal.

[0477] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0478] Specific examples

[0479] As an example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server receives this data and analyzes it using a natural language processing model. As a result, it detects that the question is ambiguous and generates feedback such as "Please specify a specific time." The server sends this feedback to the device, which then displays it to the user.

[0480] The user checks this feedback and resubmits the question, revising it to "What's the weather like in Shinjuku tomorrow at 3 PM?" The server analyzes the data again, this time including a specific time, generating a refined answer and sending it to the device. The device then displays this answer to the user, completing the dialogue.

[0481] In this way, the present invention provides a system that can refine the content of user questions and improve the quality of dialogue with AI. Furthermore, by utilizing the database for continuous learning, the system can improve the response model and enhance the user experience.

[0482] The processing flow will be explained below.

[0483] Step 1:

[0484] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[0485] Step 2:

[0486] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[0487] Step 3:

[0488] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[0489] Step 4:

[0490] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[0491] Step 5:

[0492] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[0493] Step 6:

[0494] The server generates feedback to the user based on the analysis results, such as a message saying "Please specify a specific time."

[0495] Step 7:

[0496] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[0497] Step 8:

[0498] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[0499] Step 9:

[0500] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[0501] Step 10:

[0502] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[0503] Step 11:

[0504] The server verifies that the revised question is specific and generates a refined response, such as "The weather in Shinjuku tomorrow at 3 PM will be sunny."

[0505] Step 12:

[0506] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[0507] Step 13:

[0508] The user confirms the refined response received through the terminal, thereby completing the interaction.

[0509] Step 14:

[0510] The server runs a machine learning algorithm based on the current dialogue data and updates the response model, enabling even more accurate responses the next time the dialogue occurs.

[0511] By following the above steps, the system of the present invention effectively manages and improves the interaction between the user and the server, achieving higher quality communication.

[0512] Example 1

[0513] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0514] In conventional dialogue systems, if a user's question or instruction is ambiguous, the question itself cannot be accurately analyzed, resulting in an inability to provide an appropriate response. Furthermore, due to a lack of means to provide clear feedback, users are forced to go through a lot of trial and error. Furthermore, the response model is not continuously improved, which can lead to a decline in the quality of the dialogue.

[0515] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0516] In this invention, the server includes means for receiving input data from a user via a terminal, means for analyzing the input data using a natural language processing model to understand the meaning and intent of the input data, means for providing feedback to the user based on the analysis, means for correcting the user's question based on the feedback, means for learning using the input data and past dialogue data and updating a response model, means for generating a refined response based on the updated response model, and means for transmitting the generated refined response to the terminal and displaying it. This makes it possible to specify the content of the user's question and improve the quality of the response provided by AI.

[0517] A "terminal" is a device through which a user inputs questions and instructions and transmits and receives data to and from a server.

[0518] "Input data" refers to data including questions and instructions sent by a user via a terminal.

[0519] A "server" is a device or system that receives input data sent from a terminal and performs processing such as analysis, feedback generation, and response generation.

[0520] A "natural language processing model" is a general term for algorithms and models that analyze input data and understand its meaning and intent.

[0521] "Feedback" refers to answers or instructions that the server provides to the user based on the analysis results, and is used to clarify the content of the user's question.

[0522] "Analysis" is the process of using a natural language processing model to understand the meaning and intent of input data and make that content concrete.

[0523] "Learning" is the process by which the server runs machine learning algorithms based on input data and past interaction data to continuously update and improve the response model.

[0524] A "response model" is a model for generating answers to user questions, and is updated based on learning by the server.

[0525] An "elaborate response" is a specific and clear answer that is generated to provide appropriate information to a user's question.

[0526] A "database" is a storage device that stores input data and past dialogue data and can be accessed by a server for analysis and learning.

[0527] An "ambiguous part" is a part of a user's question or instruction that is unclear or difficult to interpret, and feedback is generated by identifying this.

[0528] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. In this system, the server receives data entered by the user via the terminal, and performs processing such as analysis, feedback generation, and response generation to concretize the content of the user's question and provide a high-quality response using AI. Specific embodiments of the present invention are described below.

[0529] The server mainly uses the following hardware and software:

[0530] Server hardware: A server machine with a powerful CPU and sufficient memory

[0531] Database: A database system (e.g., MySQL, PostgreSQL) for storing input data and past dialogue data.

[0532] Natural language processing model: A natural language processing model (e.g., GPT-4) to analyze the input data.

[0533] Machine learning algorithms: Algorithms for learning and updating response models using past data (e.g., TensorFlow, PyTorch)

[0534] The terminal acts as a relay for data between the user and the server. The terminal can be a device such as a PC, smartphone, or tablet that runs an application or web browser that communicates with the server via an internet connection.

[0535] The user operates the terminal to input questions and instructions to the system. The operation of the system will be explained below using a concrete example.

[0536] As a concrete example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server analyzes the received data using a natural language processing model (GPT-4) to understand the intent of the question. If the server detects that the question is ambiguous, it generates feedback such as "Please specify a specific time." The server then sends this feedback to the device, which displays it to the user.

[0537] The user checks the feedback from the server, modifies the question, and re-enters it. For example, the user enters a specific question such as, "What is the weather like in Shinjuku tomorrow at 3 PM?" The device then sends the modified data to the server again. The server again analyzes the data and generates a more specific answer this time. It generates a refined response, "The weather will be cloudy in Shinjuku tomorrow at 3 PM," and sends it to the device. The device then displays the received response to the user.

[0538] In this way, the system can clarify the user's ambiguous questions and provide appropriate feedback, improving the quality of the AI ​​responses. Furthermore, the server uses input data and past dialogue data to perform machine learning and continuously update and improve the response model, improving the user experience.

[0539] Prompt Sentence Examples

[0540] First input: "What's the weather like in Shinjuku tomorrow?"

[0541] Corrected input: "What's the weather like in Shinjuku tomorrow at 3pm?"

[0542] In this way, the present invention provides a system that clarifies the content of a user's question and improves the quality of dialogue with AI.

[0543] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0544] Step 1: User enters question

[0545] The user inputs a question or instruction into the terminal. For example, the user inputs "Please tell me what the weather will be like in Shinjuku tomorrow." The input text becomes the input data for the terminal.

[0546] Step 2: Submitting input data

[0547] The device sends the user-entered question data to the server in real time via an HTTP POST request, with the input data included in the body of the API request.

[0548] Step 3: Receiving and storing data

[0549] The server receives the input data sent from the terminal. The received data is saved in a database on the server. The database stores the user's question ("What is the weather in Shinjuku tomorrow?").

[0550] Step 4: Analysis using natural language processing models

[0551] The server analyzes the received input data using a natural language processing model (GPT-4). The input data is input into the model, and an analysis result is generated. The analysis result includes the intent of the question and any ambiguities.

[0552] Step 5: Generate feedback

[0553] The server generates feedback based on the analysis results of the natural language processing model. For example, if the question is ambiguous, the server generates feedback such as "Please specify a specific time." The generated feedback is in text format.

[0554] Step 6: Submit your feedback

[0555] The server sends the generated feedback to the terminal. A feedback message ("Please specify a specific time") is sent as an HTTP response.

[0556] Step 7: View your feedback

[0557] The terminal displays the feedback received from the server to the user, and the feedback message is displayed on the screen of the terminal to notify the user.

[0558] Step 8: Modifying the Question

[0559] The user modifies the question based on the received feedback. For example, the user inputs a specific question such as, "Please tell me the weather in Shinjuku tomorrow at 3:00 PM." The modified question becomes the new input data.

[0560] Step 9: Resend the corrected data

[0561] The terminal sends the corrected input data to the server again via an HTTP POST request. The corrected input data is included in the body of the API request.

[0562] Step 10: Reanalyze the data

[0563] The server then analyzes the revised data using a natural language processing model. The revised data is input into the model, and an analysis result is generated. The analysis result for the specific question is obtained.

[0564] Step 11: Generate a response

[0565] The server generates a refined response based on the analysis results. For example, it may generate an answer such as "The weather in Shinjuku tomorrow at 3:00 PM will be cloudy." The generated response is in text format.

[0566] Step 12: Sending a Response

[0567] The server sends the generated response to the terminal. The response message ("The weather in Shinjuku tomorrow at 3 PM will be cloudy") is sent as an HTTP response.

[0568] Step 13: View the response

[0569] The terminal displays the response received from the server to the user. The response message is displayed on the terminal screen to notify the user. This display completes the interaction.

[0570] (Application example 1)

[0571] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0572] In today's retail industry, when customers ask questions about products in physical stores, staff are expected to respond quickly and accurately. However, many stores struggle to immediately provide customers with the specific product and inventory information they require, resulting in lower customer satisfaction and lost sales opportunities. Furthermore, the quality of the information provided varies depending on the staff's knowledge and experience, resulting in an inconsistent customer experience. Therefore, there is a growing need for a system that enables staff in physical stores to efficiently respond to customers and provide accurate information.

[0573] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0574] In this invention, the server includes means for receiving input data from a user, means for analyzing the input data, means for providing feedback to the user based on the analysis, means for improving the user's interaction method based on the feedback, means for learning using the input data and past interaction data and updating a response model, means for generating refined responses based on the updated response model, means for supporting interaction with the user in real time using a smart device, and means for providing product information and inventory information based on the support, thereby enabling staff to provide product information and inventory information to customers quickly and accurately using the smart device.

[0575] "Means for receiving input data from a user" refers to the function of collecting data through an interface for users to input questions or instructions and sending it to the system.

[0576] The "means for analyzing the input data" refers to a function that analyzes the received data using natural language processing technology, etc., and performs processing to understand its meaning and intent.

[0577] The "means for providing feedback to the user based on the analysis" refers to a function that provides advice and correction instructions to the user to ask more specific questions based on the analysis results.

[0578] "Means for improving the user's interaction method based on the feedback" refers to a process by which the user understands the feedback and improves the content of the questions and instructions.

[0579] "Means for learning using the input data and past dialogue data and updating the response model" refers to a function that performs machine learning based on collected data, continuously improves the system's response model, and enables it to generate more accurate responses.

[0580] "Means for generating a refined response based on the updated response model" refers to a function that uses an improved response model to generate a specific and accurate answer to a user's question.

[0581] "Means for supporting real-time interaction with users using smart devices" refers to a function for using devices such as smart glasses or smartphones to interact with users in real time and provide specific information.

[0582] "Means for providing product information and inventory information based on the support" refers to a function that retrieves product information and inventory status from a database based on a user's question and presents that information to the user.

[0583] The present invention provides a system for supporting interactions with customers in a physical store. A specific embodiment of this system will be described below.

[0584] System Program

[0585] 1. Hardware configuration:

[0586] It uses servers, smart devices (such as smart glasses and smartphones), and communication networks (Wi-Fi or 5G). It uses AWS EC2 instances for the servers, Amazon RDS for the database, and AWS Sagemaker for natural language processing.

[0587] 2. Software configuration:

[0588] An Android-based application is installed on the smart device, and a Node.js server is operated on the server side. Python is used for data analysis, and BERT is used as the natural language processing model.

[0589] System Operation Overview

[0590] 1. Getting and sending user input data:

[0591] The user (staff member) uses a smart device to input customer questions by voice, and the smart device converts the voice data into text using the Google Speech-to-Text API and sends the text data to the server.

[0592] 2. Data Analysis:

[0593] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker), which identifies the intent of the question and any ambiguities, and generates appropriate feedback.

[0594] 3. Feedback and response generation:

[0595] Based on the analysis results, the server generates feedback to help the user refine their question. If necessary, it sends additional questions or correction instructions to the smart device. If the user re-enters the necessary information, the server analyzes the data again and generates a refined response.

[0596] 4. Information provision:

[0597] The refined response includes product and inventory information retrieved from a database (Amazon RDS), allowing users to provide accurate and prompt information to customers.

[0598] Specific examples

[0599] Example 1

[0600] Scenario: A customer asks, "Tell me about this product."

[0601] Prompt statement:

[0602] Provide specific product information. You need to tell customers the description, price, and features of this product.

[0603] Example response:

[0604] The smart device will display the message, "This product is the latest model and costs 5,000 yen. Its features include waterproof functionality and a long battery life."

[0605] Example 2

[0606] Scenario: A customer asks, "Is this item in stock?"

[0607] Prompt statement:

[0608] Please check the inventory information and provide it to your customers. We need to check the current availability of this item.

[0609] Example response:

[0610] The smart device will display the message, "There are currently 5 of this item in stock in the store."

[0611] In this way, the present invention provides a specific system configuration and processing procedure for efficiently and accurately handling customers in a brick-and-mortar store.

[0612] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0613] Step 1: Getting and sending input data from the user

[0614] A user (staff member) uses a smart device to voice-input a customer question. The smart device then converts the voice data into text using the Google Speech-to-Text API. The input is voice data, and the output is text data. The smart device then sends the converted text data to the server.

[0615] Step 2: Analyze the data

[0616] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker). Specifically, it processes and calculates the data to identify the intent and ambiguity of the text data. The input is the text data, and the output is the analysis results (intent, ambiguity).

[0617] Step 3: Generate feedback

[0618] Based on the analysis results, the server generates feedback to help the user refine their query. The feedback includes additional questions and correction instructions as needed. The input is the analysis results, and the output is a feedback message. This is sent to the smart device and displayed to the user.

[0619] Step 4: User Modification of Question

[0620] The user re-enters a question based on the feedback and sends the revised text data to the server. The input is the feedback message and the revised question, and the output is the text data of the revised question.

[0621] Step 5: Reparsing and generating a response

[0622] The server then analyzes the revised text data again using a natural language processing model. Based on the analysis results, it generates a refined response. The input is the text data of the revised question, and the output is the refined response.

[0623] Step 6: Provide information

[0624] Based on the refined response, the server retrieves product information and inventory information from the database (Amazon RDS) and generates information to be provided to the user. The input is the refined response and information from the database, and the output is the final information provided to the user.

[0625] Step 7: Display to the user

[0626] The final information is sent to the smart device and displayed to the user, allowing the user to provide accurate product and inventory information to customers. The input is the final information from the server, and the output is the information displayed on the smart device.

[0627] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0628] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. It also incorporates an emotion engine that recognizes the user's emotions and adjusts or enhances responses based on those emotions. This system includes elements such as input data, analysis means, feedback, a dialogue method, learning means, response model, refined responses, a database, a natural language processing model, ambiguity, and an emotion engine.

[0629] Server Roles

[0630] The server performs the following functions:

[0631] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[0632] 2. Data analysis: The received input data is analyzed using a natural language processing model and an emotion engine to understand the meaning, intent, and emotional state of the input.

[0633] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[0634] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0635] 5. Run Learning: Using collected input data and past interaction data, machine learning algorithms are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[0636] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[0637] Device Role

[0638] The terminal acts as a relay for the interaction between the user and the server:

[0639] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[0640] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0641] User Roles

[0642] The user is the direct interlocutor of the system and goes through the following process:

[0643] 1. Entering a question: Enter a question or instruction through the terminal.

[0644] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0645] Specific examples

[0646] As an example, consider a question about the weather. The user enters "What will the weather be like in Shinjuku tomorrow?" into the device and sends it. The device then sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server then sends this feedback to the device, which then displays it to the user.

[0647] The user checks the feedback, modifies the question to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[0648] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[0649] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, the response model and emotion engine can be improved, resulting in a better user experience.

[0650] The processing flow will be explained below.

[0651] Step 1:

[0652] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[0653] Step 2:

[0654] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[0655] Step 3:

[0656] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[0657] Step 4:

[0658] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[0659] Step 5:

[0660] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[0661] Step 6:

[0662] The server uses an emotion engine to analyze the user's emotional state, for example, by inferring the user's emotions from text and detecting stress, joy, anxiety, etc.

[0663] Step 7:

[0664] The server generates feedback to the user based on the analysis results, such as "Please specify a specific time," and adds expressions that take into account the user's emotional state.

[0665] Step 8:

[0666] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[0667] Step 9:

[0668] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[0669] Step 10:

[0670] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[0671] Step 11:

[0672] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[0673] Step 12:

[0674] The server then verifies that the revised question is specific and generates a refined response, adding expressions that take the user's emotional state into consideration. For example, the server generates a response such as, "The weather in Shinjuku tomorrow at 3 p.m. will be sunny. Have a nice afternoon."

[0675] Step 13:

[0676] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[0677] Step 14:

[0678] The user confirms the refined response received through the terminal, thereby completing the interaction.

[0679] Step 15:

[0680] The server runs a machine learning algorithm based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses in the next conversation.

[0681] By following the above steps, the system of the present invention effectively manages and improves the dialogue between the user and the server, realizing higher quality and more emotionally sensitive communication.

[0682] Example 2

[0683] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0684] Conventional dialogue systems lack natural and rich communication because they do not consider the user's emotional state. They also fail to properly handle ambiguity in user questions and instructions, resulting in insufficient feedback and an inability to quickly respond to user needs. Furthermore, they lack a means to effectively utilize past dialogue data to continuously improve the system's response model, limiting the user experience.

[0685] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data from a user and storing it in a database, means for analyzing the input data using a natural language processing model and an emotion engine to understand the meaning, intention, and emotional state of the input data, means for generating feedback corresponding to the user's question method and emotional state based on the analysis and transmitting the feedback to the terminal, means for improving the user's dialogue method based on the feedback, means for executing a machine learning algorithm using the input data and past dialogue data to continuously update a response model and emotion engine, and means for generating refined responses based on the updated response model and adjusting them based on emotion understanding. This makes it possible to realize natural and effective dialogue that takes the user's emotional state into consideration, reduce ambiguity in the user's questions and instructions, and continuously improve the system's response model.

[0686] A "user" is an entity that provides input data to and receives feedback from a system.

[0687] "Input Data" refers to questions, instructions, or other information provided by a user to a system.

[0688] A "database" is a data management system that stores input data and past dialogue data and retrieves them as needed.

[0689] A "natural language processing model" is a machine learning model that analyzes human language and understands its meaning and intent.

[0690] An "emotion engine" is an analytical system for detecting a user's emotional state from input data and adjusting responses accordingly.

[0691] "Analysis" refers to processing input data using a natural language processing model and emotion engine to understand its meaning, intent, and emotional state.

[0692] "Feedback" refers to the response or instructions provided to the user based on the analysis results.

[0693] "Terminal" refers to a device used by a user to provide input data to a system.

[0694] A "machine learning algorithm" is a computational method for training a model based on data and continuously improving its performance.

[0695] A "response model" is a model for generating appropriate responses to user questions and instructions.

[0696] "Refined responses" refer to detailed responses generated based on an updated response model and tailored based on emotion understanding.

[0697] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. The present invention incorporates an emotion engine that recognizes the user's emotions and adjusts and enhances responses based on those emotions. The system includes the following elements:

[0698] Server Roles

[0699] Data Receipt and Storage:

[0700] The server receives input data sent by the user via the terminal and stores the data in a database. The database used is, for example, MySQL. When a user enters "What is the weather in Shinjuku tomorrow?", the server receives the data via an HTTP POST request and stores it in the database.

[0701] Data analysis:

[0702] The server analyzes the received input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). This allows it to understand the meaning, intent, and emotional state of the input. The server analyzes the input data, "What's the weather in Shinjuku tomorrow?" and extracts the keywords "Shinjuku," "tomorrow," and "weather," as well as the user's emotional state (curiosity).

[0703] Generate feedback:

[0704] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. This feedback is generated using a response model (e.g., GPT-3). It detects ambiguous time specifications and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration.

[0705] Send feedback:

[0706] The generated feedback is sent to the terminal and displayed to the user. The server sends feedback to the terminal saying "Please specify a specific time," which is displayed to the user.

[0707] Run the training:

[0708] The server uses the collected input data and past dialogue data to run machine learning algorithms (e.g., TensorFlow or PyTorch) and continuously update and improve the response model and emotion engine. The server inputs the dialogue history from the database into TensorFlow to retrain the response model and emotion engine.

[0709] Response generation:

[0710] Based on the updated response model, the server generates a refined response to the user's question and adjusts it based on sentiment understanding. In response to the question, "What is the weather like in Shinjuku tomorrow at 3 PM?", the server generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." and sends it to the device.

[0711] Device Role

[0712] Sending input data:

[0713] The device sends questions, instructions, and emotion-related data entered by the user to the server in real time. For example, if a user enters "What's the weather in Shinjuku tomorrow?" and presses the send button, the data is sent to the server.

[0714] Receiving and viewing feedback:

[0715] The terminal receives the feedback from the server and displays it in an easy-to-understand manner for the user. The terminal displays the feedback received from the server, "Please specify a specific time," in the chat window.

[0716] User Roles

[0717] Enter your question:

[0718] The user inputs questions or instructions through the terminal. The user types "Please tell me the weather in Shinjuku tomorrow" into the text box on the terminal and clicks the send button.

[0719] Review and correct feedback:

[0720] The user checks the feedback sent from the server, modifies the question if necessary, and resubmits it. The user sees the feedback "Please specify a specific time," modifies the question to "What is the weather like in Shinjuku at 3 PM tomorrow," and submits it again.

[0721] Specific examples

[0722] For example, consider the case where a user inputs and sends "What's the weather in Shinjuku tomorrow?" into a terminal. The terminal sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server sends this feedback to the terminal, which then displays it to the user.

[0723] The user checks the feedback, amends it to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[0724] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[0725] Prompt Sentence Examples

[0726] Here are some examples of prompts for generative AI models:

[0727] Prompt statement:

[0728] Use an emotion engine to generate a response to the user's question, "What will the weather be like in Shinjuku tomorrow?" If no specific time is specified, provide feedback, and generate a specific response when the question is entered again.

[0729] Example 1: User's first question

[0730] User: What's the weather like in Shinjuku tomorrow?

[0731] System: What time of day would you like to know the weather in Shinjuku tomorrow? Please specify a specific time.

[0732] Example 2: User question after modification

[0733] User: What's the weather like in Shinjuku tomorrow at 3pm?

[0734] System: The weather in Shinjuku will be sunny tomorrow at 3 PM. Have a nice afternoon.

[0735] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, it is possible to improve the response model and emotion engine, thereby improving the user experience.

[0736] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0737] Step 1:

[0738] Sending input data

[0739] The user types "Please tell me the weather in Shinjuku tomorrow" into the terminal and clicks the send button.

[0740] The terminal sends this input data to the server as an HTTP POST request.

[0741] Input: User question: "What's the weather like in Shinjuku tomorrow?"

[0742] Output: Data sent to the server as an HTTP POST request

[0743] Step 2:

[0744] Receiving and storing data

[0745] The server receives the input data submitted by the user. This data is stored in a database, such as MySQL. The server extracts the data from the POST request and stores it in the "inputs" table in the database.

[0746] Input: HTTP POST request

[0747] Output: Input data stored in a database

[0748] Step 3:

[0749] Data analysis

[0750] The server analyzes the stored input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). Through this analysis, the meaning and intent of the text and the user's emotional state are extracted. Specifically, keywords are extracted from the text and the emotional state is detected.

[0751] Input: Input data stored in the database

[0752] Output: Keywords and emotional state (e.g., "Shinjuku," "tomorrow," "weather," where the emotion is curiosity)

[0753] Step 4:

[0754] Generate feedback

[0755] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. It uses a response model (e.g., GPT-3) to detect ambiguous time specifications and generates feedback such as "Please specify a specific time." The feedback includes wording that takes the user's emotions into consideration.

[0756] Input: Keywords and emotional states

[0757] Output: Feedback: "Please specify a specific time."

[0758] Step 5:

[0759] Send Feedback

[0760] The server sends the generated feedback to the terminal as an HTTP response. The terminal receives this feedback and displays it in an easy-to-understand manner for the user. Specifically, the feedback is displayed in a chat window or similar.

[0761] Input: "Please specify a specific time" feedback

[0762] Output: Feedback displayed on the terminal

[0763] Step 6:

[0764] Review and correct feedback

[0765] The user checks the feedback from the server and modifies the question, re-entering "What is the weather like in Shinjuku tomorrow at 3 PM?" and clicking the resend button. The device then resends the modified question to the server.

[0766] Input: User's revised question "What's the weather like in Shinjuku tomorrow at 3pm?"

[0767] Output: Resent data

[0768] Step 7:

[0769] Reparsing and generating a response

[0770] The server receives the data again and analyzes it using the natural language processing model and emotion engine. Based on the results, it generates a specific response. For example, it generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon."

[0771] Input: Corrected input data

[0772] Output: The specific response generated

[0773] Step 8:

[0774] Sending a Response

[0775] The server generates a specific response and sends it to the terminal, which receives it and displays it to the user.

[0776] Input: Specific response

[0777] Output: Response displayed on the terminal

[0778] Step 9:

[0779] Execution of training

[0780] The server uses the collected input data and past dialogue data to run machine learning algorithms and continuously update and improve the response model and emotion engine. The dialogue history from the database is input into TensorFlow to retrain the model.

[0781] Input: Collected input data and past interaction data

[0782] Output: Updated response model and emotion engine

[0783] (Application example 2)

[0784] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0785] In recent years, improving the customer experience in brick-and-mortar stores has become increasingly important, but conventional systems have struggled to respond appropriately to customer questions and requests in real time. Furthermore, there is a lack of technology to provide feedback that takes into account the customer's emotional state, limiting the improvement of customer satisfaction. In particular, there is a need for an effective concierge service that can respond to ambiguous questions and emotional changes.

[0786] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving input data from a user; means for analyzing the input data; means for providing feedback to the user based on the analysis; means for improving the user's interaction method based on the feedback; means for learning using the input data and past interaction data and updating a response model; means for generating a refined response based on the updated response model; means for analyzing the input data using a natural language processing model and identifying ambiguous parts; and means for providing a concierge service in response to customer questions in a physical store, including an emotion engine that recognizes the user's emotional state and adjusts responses based on the user's emotional state. This makes it possible to respond to customer questions in real time and provide feedback that takes the user's emotional state into consideration, thereby significantly improving customer satisfaction.

[0787] "Input data" refers to questions, instructions, and other text information that a user sends to the system via a terminal.

[0788] "Means for analyzing" are the technical means for processing received input data and understanding its meaning, intent, and emotional state.

[0789] "Feedback" refers to responses or advice provided to the user based on the analysis results, and is information used to respond to the user's questions or requests.

[0790] "Means for improving the interaction method" are technical measures that use user feedback and correction data to enable the system to respond more accurately the next time the interaction is performed.

[0791] The "learning means" is a machine learning algorithm that uses past interaction data and newly collected input data to continuously improve the system's response model.

[0792] A "response model" is a model within the system that generates optimal responses to user questions and instructions.

[0793] A "refined response" is generated based on the updated response model and is a very specific and relevant answer to the user's question.

[0794] A "natural language processing model" is an artificial intelligence technology that analyzes input data from a user as language and understands its meaning and structure.

[0795] An "ambiguous part" is a part of the user's input data where the intent of the question or request is unclear.

[0796] An "emotion engine" is a technology that analyzes a user's emotional state and adjusts feedback and responses according to that emotional state.

[0797] A "physical store" is a store or shop that exists in a physical location and provides products and services in a face-to-face manner to customers.

[0798] A "concierge service" is a service provided in a physical store that responds to customer questions and requests and provides information about and offers products and services.

[0799] A "database" is a storage device and its management system for centrally storing input data and dialogue data collected by the system.

[0800] This invention provides a system for realizing a concierge service that provides highly accurate and emotionally sensitive feedback in real time when a customer inputs a question or request in a store. Hereinafter, an embodiment of the invention will be described in detail.

[0801] System Configuration

[0802] server

[0803] The server performs the following functions:

[0804] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[0805] 2. Data analysis: The received input data is analyzed using a natural language processing model (e.g., BERT or GPT) and an emotion engine (e.g., Microsoft Azure Emotion API) to understand the meaning, intent, and emotional state of the input.

[0806] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[0807] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0808] 5. Run learning: Using the collected input data and past interaction data, machine learning algorithms (e.g., TensorFlow) are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[0809] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[0810] Terminal

[0811] The terminal acts as a relay for the interaction between the user and the server:

[0812] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[0813] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0814] User

[0815] The user is the direct interlocutor of the system and goes through the following process:

[0816] 1. Entering a question: Enter a question or instruction through the terminal.

[0817] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0818] Specific examples

[0819] For example, a user might enter "I'm looking for a special gift today. Which one would be good?" into a terminal in a physical store and send it. The terminal sends this input data in real time to a server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it recognizes high expectations and generates feedback such as "I see you're looking for a special gift today. Here are some recommended products suitable for your special occasion. We'd be happy to help you make a great choice." This feedback is displayed to the user through the terminal.

[0820] Prompt Sentence Examples

[0821] By inputting the following prompt into the generative AI model, we generate a response based on high expectations:

[0822] "Generate a response that reflects high expectations for a customer looking for a special gift."

[0823] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0824] Step 1:

[0825] The user inputs and sends questions and instructions through the terminal.

[0826] Input: Text data of user questions and instructions

[0827] Output: The input data sent

[0828] What happens: Using a smartphone or smart glasses, the user types, "I'm looking for a special gift today. What would be good?" and submits that data.

[0829] Step 2:

[0830] The terminal transmits input data to the server in real time.

[0831] Input: Input data submitted by the user

[0832] Output: The request containing the input data sent

[0833] Specific operation: The terminal receives the user's input data and sends it to the server via the Internet.

[0834] Step 3:

[0835] The server receives the input data and stores it in a database.

[0836] Input: Input data sent from the terminal

[0837] Output: Input data stored in a database

[0838] Specific operation: The server's data receiving system receives the input data and stores it in a database, including the user's ID and timestamp.

[0839] Step 4:

[0840] The server analyzes the received input data using a natural language processing model (e.g., BERT or GPT).

[0841] Input: Saved input data

[0842] Output: Analysis results (question intent, content, related keywords)

[0843] What it does: The model on the server analyzes the input data and extracts the meaning and intent of the text.

[0844] Step 5:

[0845] The server uses the analysis results to understand the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API).

[0846] Input: Analysis results

[0847] Output: Analysis results including emotional state

[0848] Specific operation: The server inputs the analysis results into the emotion engine to detect high expectations and other emotions.

[0849] Step 6:

[0850] The server generates feedback based on the analysis results and emotional state.

[0851] Input: Analysis results including emotional state

[0852] Output: Feedback text data

[0853] What it does: Based on the analysis results and emotional state, the server generates feedback such as, "You're looking for a special gift today. Here are some recommended products for your special occasion. We'd be happy to help you make a great choice."

[0854] Step 7:

[0855] The server generates feedback and sends it to the device.

[0856] Input: Text data of generated feedback

[0857] Output: Feedback data sent

[0858] Specific operations: The server generates and sends a request to send feedback data to the terminal.

[0859] Step 8:

[0860] The terminal receives the feedback and displays it to the user.

[0861] Input: Feedback data sent from the server

[0862] Output: Displayed feedback

[0863] Specific operation: The device displays the received feedback data on the screen, and the user checks it and decides the next action.

[0864] Step 9:

[0865] The server uses collected input and feedback data to learn and update the response model.

[0866] Input: Input and feedback data

[0867] Output: Updated response model

[0868] What it does: Machine learning algorithms on the server continuously train and improve response models based on collected interaction data.

[0869] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0870] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0871] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0872] [Third embodiment]

[0873] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0874] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0875] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0876] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0877] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0878] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0879] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0880] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0881] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0882] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0883] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0884] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0885] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. This system includes elements such as "input data," "analysis means," "feedback," "dialogue method," "learning means," "response model," "refined response," "database," "natural language processing model," and "ambiguous parts."

[0886] Server Roles

[0887] The server performs the following main functions:

[0888] 1. Receiving and storing data: The server receives the input data sent by the user via the terminal and stores it in a database.

[0889] 2. Data analysis: The server analyzes the received input data using a natural language processing model to understand the meaning and intent of the input.

[0890] 3. Feedback generation: Based on the analysis results, feedback is generated to help users improve their questioning methods.

[0891] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[0892] 5. Run learning: Using collected input data and past interaction data, the machine learning algorithm runs and continuously updates and improves the response model.

[0893] 6. Response Generation: Based on the updated response model, a refined response to the user's question is generated.

[0894] Device Role

[0895] The terminal acts as a relay for the interaction between the user and the server:

[0896] 1. Sending input data: Questions and instructions entered by the user are sent to the server in real time.

[0897] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[0898] User Roles

[0899] The user is the direct interlocutor of the system and goes through the following process:

[0900] 1. Entering a question: Enter a question or instruction through the terminal.

[0901] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[0902] Specific examples

[0903] As an example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server receives this data and analyzes it using a natural language processing model. As a result, it detects that the question is ambiguous and generates feedback such as "Please specify a specific time." The server sends this feedback to the device, which then displays it to the user.

[0904] The user checks this feedback and resubmits the question, revising it to "What's the weather like in Shinjuku tomorrow at 3 PM?" The server analyzes the data again, this time including a specific time, generating a refined answer and sending it to the device. The device then displays this answer to the user, completing the dialogue.

[0905] In this way, the present invention provides a system that can refine the content of user questions and improve the quality of dialogue with AI. Furthermore, by utilizing the database for continuous learning, the system can improve the response model and enhance the user experience.

[0906] The processing flow will be explained below.

[0907] Step 1:

[0908] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[0909] Step 2:

[0910] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[0911] Step 3:

[0912] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[0913] Step 4:

[0914] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[0915] Step 5:

[0916] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[0917] Step 6:

[0918] The server generates feedback to the user based on the analysis results, such as a message saying "Please specify a specific time."

[0919] Step 7:

[0920] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[0921] Step 8:

[0922] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[0923] Step 9:

[0924] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[0925] Step 10:

[0926] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[0927] Step 11:

[0928] The server verifies that the revised question is specific and generates a refined response, such as "The weather in Shinjuku tomorrow at 3 PM will be sunny."

[0929] Step 12:

[0930] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[0931] Step 13:

[0932] The user confirms the refined response received through the terminal, thereby completing the interaction.

[0933] Step 14:

[0934] The server runs a machine learning algorithm based on the current dialogue data and updates the response model, enabling even more accurate responses the next time the dialogue occurs.

[0935] By following the above steps, the system of the present invention effectively manages and improves the interaction between the user and the server, achieving higher quality communication.

[0936] Example 1

[0937] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0938] In conventional dialogue systems, if a user's question or instruction is ambiguous, the question itself cannot be accurately analyzed, resulting in an inability to provide an appropriate response. Furthermore, due to a lack of means to provide clear feedback, users are forced to go through a lot of trial and error. Furthermore, the response model is not continuously improved, which can lead to a decline in the quality of the dialogue.

[0939] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0940] In this invention, the server includes means for receiving input data from a user via a terminal, means for analyzing the input data using a natural language processing model to understand the meaning and intent of the input data, means for providing feedback to the user based on the analysis, means for correcting the user's question based on the feedback, means for learning using the input data and past dialogue data and updating a response model, means for generating a refined response based on the updated response model, and means for transmitting the generated refined response to the terminal and displaying it. This makes it possible to specify the content of the user's question and improve the quality of the response provided by AI.

[0941] A "terminal" is a device through which a user inputs questions and instructions and transmits and receives data to and from a server.

[0942] "Input data" refers to data including questions and instructions sent by a user via a terminal.

[0943] A "server" is a device or system that receives input data sent from a terminal and performs processing such as analysis, feedback generation, and response generation.

[0944] A "natural language processing model" is a general term for algorithms and models that analyze input data and understand its meaning and intent.

[0945] "Feedback" refers to answers or instructions that the server provides to the user based on the analysis results, and is used to clarify the content of the user's question.

[0946] "Analysis" is the process of using a natural language processing model to understand the meaning and intent of input data and make that content concrete.

[0947] "Learning" is the process by which the server runs machine learning algorithms based on input data and past interaction data to continuously update and improve the response model.

[0948] A "response model" is a model for generating answers to user questions, and is updated based on learning by the server.

[0949] An "elaborate response" is a specific and clear answer that is generated to provide appropriate information to a user's question.

[0950] A "database" is a storage device that stores input data and past dialogue data and can be accessed by a server for analysis and learning.

[0951] An "ambiguous part" is a part of a user's question or instruction that is unclear or difficult to interpret, and feedback is generated by identifying this.

[0952] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. In this system, the server receives data entered by the user via the terminal, and performs processing such as analysis, feedback generation, and response generation to concretize the content of the user's question and provide a high-quality response using AI. Specific embodiments of the present invention are described below.

[0953] The server mainly uses the following hardware and software:

[0954] Server hardware: A server machine with a powerful CPU and sufficient memory

[0955] Database: A database system (e.g., MySQL, PostgreSQL) for storing input data and past dialogue data.

[0956] Natural language processing model: A natural language processing model (e.g., GPT-4) to analyze the input data.

[0957] Machine learning algorithms: Algorithms for learning and updating response models using past data (e.g., TensorFlow, PyTorch)

[0958] The terminal acts as a relay for data between the user and the server. The terminal can be a device such as a PC, smartphone, or tablet that runs an application or web browser that communicates with the server via an internet connection.

[0959] The user operates the terminal to input questions and instructions to the system. The operation of the system will be explained below using a concrete example.

[0960] As a concrete example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server analyzes the received data using a natural language processing model (GPT-4) to understand the intent of the question. If the server detects that the question is ambiguous, it generates feedback such as "Please specify a specific time." The server then sends this feedback to the device, which displays it to the user.

[0961] The user checks the feedback from the server, modifies the question, and re-enters it. For example, the user enters a specific question such as, "What is the weather like in Shinjuku tomorrow at 3 PM?" The device then sends the modified data to the server again. The server again analyzes the data and generates a more specific answer this time. It generates a refined response, "The weather will be cloudy in Shinjuku tomorrow at 3 PM," and sends it to the device. The device then displays the received response to the user.

[0962] In this way, the system can clarify the user's ambiguous questions and provide appropriate feedback, improving the quality of the AI ​​responses. Furthermore, the server uses input data and past dialogue data to perform machine learning and continuously update and improve the response model, improving the user experience.

[0963] Prompt Sentence Examples

[0964] First input: "What's the weather like in Shinjuku tomorrow?"

[0965] Corrected input: "What's the weather like in Shinjuku tomorrow at 3pm?"

[0966] In this way, the present invention provides a system that clarifies the content of a user's question and improves the quality of dialogue with AI.

[0967] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0968] Step 1: User enters question

[0969] The user inputs a question or instruction into the terminal. For example, the user inputs "Please tell me what the weather will be like in Shinjuku tomorrow." The input text becomes the input data for the terminal.

[0970] Step 2: Submitting input data

[0971] The device sends the user-entered question data to the server in real time via an HTTP POST request, with the input data included in the body of the API request.

[0972] Step 3: Receiving and storing data

[0973] The server receives the input data sent from the terminal. The received data is saved in a database on the server. The database stores the user's question ("What is the weather in Shinjuku tomorrow?").

[0974] Step 4: Analysis using natural language processing models

[0975] The server analyzes the received input data using a natural language processing model (GPT-4). The input data is input into the model, and an analysis result is generated. The analysis result includes the intent of the question and any ambiguities.

[0976] Step 5: Generate feedback

[0977] The server generates feedback based on the analysis results of the natural language processing model. For example, if the question is ambiguous, the server generates feedback such as "Please specify a specific time." The generated feedback is in text format.

[0978] Step 6: Submit your feedback

[0979] The server sends the generated feedback to the terminal. A feedback message ("Please specify a specific time") is sent as an HTTP response.

[0980] Step 7: View your feedback

[0981] The terminal displays the feedback received from the server to the user, and the feedback message is displayed on the screen of the terminal to notify the user.

[0982] Step 8: Modifying the Question

[0983] The user modifies the question based on the received feedback. For example, the user inputs a specific question such as, "Please tell me the weather in Shinjuku tomorrow at 3:00 PM." The modified question becomes the new input data.

[0984] Step 9: Resend the corrected data

[0985] The terminal sends the corrected input data to the server again via an HTTP POST request. The corrected input data is included in the body of the API request.

[0986] Step 10: Reanalyze the data

[0987] The server then analyzes the revised data using a natural language processing model. The revised data is input into the model, and an analysis result is generated. The analysis result for the specific question is obtained.

[0988] Step 11: Generate a response

[0989] The server generates a refined response based on the analysis results. For example, it may generate an answer such as "The weather in Shinjuku tomorrow at 3:00 PM will be cloudy." The generated response is in text format.

[0990] Step 12: Sending a Response

[0991] The server sends the generated response to the terminal. The response message ("The weather in Shinjuku tomorrow at 3 PM will be cloudy") is sent as an HTTP response.

[0992] Step 13: View the response

[0993] The terminal displays the response received from the server to the user. The response message is displayed on the terminal screen to notify the user. This display completes the interaction.

[0994] (Application example 1)

[0995] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0996] In today's retail industry, when customers ask questions about products in physical stores, staff are expected to respond quickly and accurately. However, many stores struggle to immediately provide customers with the specific product and inventory information they require, resulting in lower customer satisfaction and lost sales opportunities. Furthermore, the quality of the information provided varies depending on the staff's knowledge and experience, resulting in an inconsistent customer experience. Therefore, there is a growing need for a system that enables staff in physical stores to efficiently respond to customers and provide accurate information.

[0997] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0998] In this invention, the server includes means for receiving input data from a user, means for analyzing the input data, means for providing feedback to the user based on the analysis, means for improving the user's interaction method based on the feedback, means for learning using the input data and past interaction data and updating a response model, means for generating refined responses based on the updated response model, means for supporting interaction with the user in real time using a smart device, and means for providing product information and inventory information based on the support, thereby enabling staff to provide product information and inventory information to customers quickly and accurately using the smart device.

[0999] "Means for receiving input data from a user" refers to the function of collecting data through an interface for users to input questions or instructions and sending it to the system.

[1000] The "means for analyzing the input data" refers to a function that analyzes the received data using natural language processing technology, etc., and performs processing to understand its meaning and intent.

[1001] The "means for providing feedback to the user based on the analysis" refers to a function that provides advice and correction instructions to the user to ask more specific questions based on the analysis results.

[1002] "Means for improving the user's interaction method based on the feedback" refers to a process by which the user understands the feedback and improves the content of the questions and instructions.

[1003] "Means for learning using the input data and past dialogue data and updating the response model" refers to a function that performs machine learning based on collected data, continuously improves the system's response model, and enables it to generate more accurate responses.

[1004] "Means for generating a refined response based on the updated response model" refers to a function that uses an improved response model to generate a specific and accurate answer to a user's question.

[1005] "Means for supporting real-time interaction with users using smart devices" refers to a function for using devices such as smart glasses or smartphones to interact with users in real time and provide specific information.

[1006] "Means for providing product information and inventory information based on the support" refers to a function that retrieves product information and inventory status from a database based on a user's question and presents that information to the user.

[1007] The present invention provides a system for supporting interactions with customers in a physical store. A specific embodiment of this system will be described below.

[1008] System Program

[1009] 1. Hardware configuration:

[1010] It uses servers, smart devices (such as smart glasses and smartphones), and communication networks (Wi-Fi or 5G). It uses AWS EC2 instances for the servers, Amazon RDS for the database, and AWS Sagemaker for natural language processing.

[1011] 2. Software configuration:

[1012] An Android-based application is installed on the smart device, and a Node.js server is operated on the server side. Python is used for data analysis, and BERT is used as the natural language processing model.

[1013] System Operation Overview

[1014] 1. Getting and sending user input data:

[1015] The user (staff member) uses a smart device to input customer questions by voice, and the smart device converts the voice data into text using the Google Speech-to-Text API and sends the text data to the server.

[1016] 2. Data Analysis:

[1017] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker), which identifies the intent of the question and any ambiguities, and generates appropriate feedback.

[1018] 3. Feedback and response generation:

[1019] Based on the analysis results, the server generates feedback to help the user refine their question. If necessary, it sends additional questions or correction instructions to the smart device. If the user re-enters the necessary information, the server analyzes the data again and generates a refined response.

[1020] 4. Information provision:

[1021] The refined response includes product and inventory information retrieved from a database (Amazon RDS), allowing users to provide accurate and prompt information to customers.

[1022] Specific examples

[1023] Example 1

[1024] Scenario: A customer asks, "Tell me about this product."

[1025] Prompt statement:

[1026] Provide specific product information. You need to tell customers the description, price, and features of this product.

[1027] Example response:

[1028] The smart device will display the message, "This product is the latest model and costs 5,000 yen. Its features include waterproof functionality and a long battery life."

[1029] Example 2

[1030] Scenario: A customer asks, "Is this item in stock?"

[1031] Prompt statement:

[1032] Please check the inventory information and provide it to your customers. We need to check the current availability of this item.

[1033] Example response:

[1034] The smart device will display the message, "There are currently 5 of this item in stock in the store."

[1035] In this way, the present invention provides a specific system configuration and processing procedure for efficiently and accurately handling customers in a brick-and-mortar store.

[1036] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1037] Step 1: Getting and sending input data from the user

[1038] A user (staff member) uses a smart device to voice-input a customer question. The smart device then converts the voice data into text using the Google Speech-to-Text API. The input is voice data, and the output is text data. The smart device then sends the converted text data to the server.

[1039] Step 2: Analyze the data

[1040] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker). Specifically, it processes and calculates the data to identify the intent and ambiguity of the text data. The input is the text data, and the output is the analysis results (intent, ambiguity).

[1041] Step 3: Generate feedback

[1042] Based on the analysis results, the server generates feedback to help the user refine their query. The feedback includes additional questions and correction instructions as needed. The input is the analysis results, and the output is a feedback message. This is sent to the smart device and displayed to the user.

[1043] Step 4: User Modification of Question

[1044] The user re-enters a question based on the feedback and sends the revised text data to the server. The input is the feedback message and the revised question, and the output is the text data of the revised question.

[1045] Step 5: Reparsing and generating a response

[1046] The server then analyzes the revised text data again using a natural language processing model. Based on the analysis results, it generates a refined response. The input is the text data of the revised question, and the output is the refined response.

[1047] Step 6: Provide information

[1048] Based on the refined response, the server retrieves product information and inventory information from the database (Amazon RDS) and generates information to be provided to the user. The input is the refined response and information from the database, and the output is the final information provided to the user.

[1049] Step 7: Display to the user

[1050] The final information is sent to the smart device and displayed to the user, allowing the user to provide accurate product and inventory information to customers. The input is the final information from the server, and the output is the information displayed on the smart device.

[1051] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1052] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. It also incorporates an emotion engine that recognizes the user's emotions and adjusts or enhances responses based on those emotions. This system includes elements such as input data, analysis means, feedback, a dialogue method, learning means, response model, refined responses, a database, a natural language processing model, ambiguity, and an emotion engine.

[1053] Server Roles

[1054] The server performs the following functions:

[1055] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[1056] 2. Data analysis: The received input data is analyzed using a natural language processing model and an emotion engine to understand the meaning, intent, and emotional state of the input.

[1057] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[1058] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[1059] 5. Run Learning: Using collected input data and past interaction data, machine learning algorithms are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[1060] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[1061] Device Role

[1062] The terminal acts as a relay for the interaction between the user and the server:

[1063] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[1064] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[1065] User Roles

[1066] The user is the direct interlocutor of the system and goes through the following process:

[1067] 1. Entering a question: Enter a question or instruction through the terminal.

[1068] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[1069] Specific examples

[1070] As an example, consider a question about the weather. The user enters "What will the weather be like in Shinjuku tomorrow?" into the device and sends it. The device then sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server then sends this feedback to the device, which then displays it to the user.

[1071] The user checks the feedback, modifies the question to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[1072] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[1073] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, the response model and emotion engine can be improved, resulting in a better user experience.

[1074] The processing flow will be explained below.

[1075] Step 1:

[1076] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[1077] Step 2:

[1078] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[1079] Step 3:

[1080] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[1081] Step 4:

[1082] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[1083] Step 5:

[1084] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[1085] Step 6:

[1086] The server uses an emotion engine to analyze the user's emotional state, for example, by inferring the user's emotions from text and detecting stress, joy, anxiety, etc.

[1087] Step 7:

[1088] The server generates feedback to the user based on the analysis results, such as "Please specify a specific time," and adds expressions that take into account the user's emotional state.

[1089] Step 8:

[1090] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[1091] Step 9:

[1092] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[1093] Step 10:

[1094] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[1095] Step 11:

[1096] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[1097] Step 12:

[1098] The server then verifies that the revised question is specific and generates a refined response, adding expressions that take the user's emotional state into consideration. For example, the server generates a response such as, "The weather in Shinjuku tomorrow at 3 p.m. will be sunny. Have a nice afternoon."

[1099] Step 13:

[1100] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[1101] Step 14:

[1102] The user confirms the refined response received through the terminal, thereby completing the interaction.

[1103] Step 15:

[1104] The server runs a machine learning algorithm based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses in the next conversation.

[1105] By following the above steps, the system of the present invention effectively manages and improves the dialogue between the user and the server, realizing higher quality and more emotionally sensitive communication.

[1106] Example 2

[1107] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1108] Conventional dialogue systems lack natural and rich communication because they do not consider the user's emotional state. They also fail to properly handle ambiguity in user questions and instructions, resulting in insufficient feedback and an inability to quickly respond to user needs. Furthermore, they lack a means to effectively utilize past dialogue data to continuously improve the system's response model, limiting the user experience.

[1109] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data from a user and storing it in a database, means for analyzing the input data using a natural language processing model and an emotion engine to understand the meaning, intention, and emotional state of the input data, means for generating feedback corresponding to the user's question method and emotional state based on the analysis and transmitting the feedback to the terminal, means for improving the user's dialogue method based on the feedback, means for executing a machine learning algorithm using the input data and past dialogue data to continuously update a response model and emotion engine, and means for generating refined responses based on the updated response model and adjusting them based on emotion understanding. This makes it possible to realize natural and effective dialogue that takes the user's emotional state into consideration, reduce ambiguity in the user's questions and instructions, and continuously improve the system's response model.

[1110] A "user" is an entity that provides input data to and receives feedback from a system.

[1111] "Input Data" refers to questions, instructions, or other information provided by a user to a system.

[1112] A "database" is a data management system that stores input data and past dialogue data and retrieves them as needed.

[1113] A "natural language processing model" is a machine learning model that analyzes human language and understands its meaning and intent.

[1114] An "emotion engine" is an analytical system for detecting a user's emotional state from input data and adjusting responses accordingly.

[1115] "Analysis" refers to processing input data using a natural language processing model and emotion engine to understand its meaning, intent, and emotional state.

[1116] "Feedback" refers to the response or instructions provided to the user based on the analysis results.

[1117] "Terminal" refers to a device used by a user to provide input data to a system.

[1118] A "machine learning algorithm" is a computational method for training a model based on data and continuously improving its performance.

[1119] A "response model" is a model for generating appropriate responses to user questions and instructions.

[1120] "Refined responses" refer to detailed responses generated based on an updated response model and tailored based on emotion understanding.

[1121] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. The present invention incorporates an emotion engine that recognizes the user's emotions and adjusts and enhances responses based on those emotions. The system includes the following elements:

[1122] Server Roles

[1123] Data Receipt and Storage:

[1124] The server receives input data sent by the user via the terminal and stores the data in a database. The database used is, for example, MySQL. When a user enters "What is the weather in Shinjuku tomorrow?", the server receives the data via an HTTP POST request and stores it in the database.

[1125] Data analysis:

[1126] The server analyzes the received input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). This allows it to understand the meaning, intent, and emotional state of the input. The server analyzes the input data, "What's the weather in Shinjuku tomorrow?" and extracts the keywords "Shinjuku," "tomorrow," and "weather," as well as the user's emotional state (curiosity).

[1127] Generate feedback:

[1128] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. This feedback is generated using a response model (e.g., GPT-3). It detects ambiguous time specifications and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration.

[1129] Send feedback:

[1130] The generated feedback is sent to the terminal and displayed to the user. The server sends feedback to the terminal saying "Please specify a specific time," which is displayed to the user.

[1131] Run the training:

[1132] The server uses the collected input data and past dialogue data to run machine learning algorithms (e.g., TensorFlow or PyTorch) and continuously update and improve the response model and emotion engine. The server inputs the dialogue history from the database into TensorFlow to retrain the response model and emotion engine.

[1133] Response generation:

[1134] Based on the updated response model, the server generates a refined response to the user's question and adjusts it based on sentiment understanding. In response to the question, "What is the weather like in Shinjuku tomorrow at 3 PM?", the server generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." and sends it to the device.

[1135] Device Role

[1136] Sending input data:

[1137] The device sends questions, instructions, and emotion-related data entered by the user to the server in real time. For example, if a user enters "What's the weather in Shinjuku tomorrow?" and presses the send button, the data is sent to the server.

[1138] Receiving and viewing feedback:

[1139] The terminal receives the feedback from the server and displays it in an easy-to-understand manner for the user. The terminal displays the feedback received from the server, "Please specify a specific time," in the chat window.

[1140] User Roles

[1141] Enter your question:

[1142] The user inputs questions or instructions through the terminal. The user types "Please tell me the weather in Shinjuku tomorrow" into the text box on the terminal and clicks the send button.

[1143] Review and correct feedback:

[1144] The user checks the feedback sent from the server, modifies the question if necessary, and resubmits it. The user sees the feedback "Please specify a specific time," modifies the question to "What is the weather like in Shinjuku at 3 PM tomorrow," and submits it again.

[1145] Specific examples

[1146] For example, consider the case where a user inputs and sends "What's the weather in Shinjuku tomorrow?" into a terminal. The terminal sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server sends this feedback to the terminal, which then displays it to the user.

[1147] The user checks the feedback, amends it to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[1148] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[1149] Prompt Sentence Examples

[1150] Here are some examples of prompts for generative AI models:

[1151] Prompt statement:

[1152] Use an emotion engine to generate a response to the user's question, "What will the weather be like in Shinjuku tomorrow?" If no specific time is specified, provide feedback, and generate a specific response when the question is entered again.

[1153] Example 1: User's first question

[1154] User: What's the weather like in Shinjuku tomorrow?

[1155] System: What time of day would you like to know the weather in Shinjuku tomorrow? Please specify a specific time.

[1156] Example 2: User question after modification

[1157] User: What's the weather like in Shinjuku tomorrow at 3pm?

[1158] System: The weather in Shinjuku will be sunny tomorrow at 3 PM. Have a nice afternoon.

[1159] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, it is possible to improve the response model and emotion engine, thereby improving the user experience.

[1160] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1161] Step 1:

[1162] Sending input data

[1163] The user types "Please tell me the weather in Shinjuku tomorrow" into the terminal and clicks the send button.

[1164] The terminal sends this input data to the server as an HTTP POST request.

[1165] Input: User question: "What's the weather like in Shinjuku tomorrow?"

[1166] Output: Data sent to the server as an HTTP POST request

[1167] Step 2:

[1168] Receiving and storing data

[1169] The server receives the input data submitted by the user. This data is stored in a database, such as MySQL. The server extracts the data from the POST request and stores it in the "inputs" table in the database.

[1170] Input: HTTP POST request

[1171] Output: Input data stored in a database

[1172] Step 3:

[1173] Data analysis

[1174] The server analyzes the stored input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). Through this analysis, the meaning and intent of the text and the user's emotional state are extracted. Specifically, keywords are extracted from the text and the emotional state is detected.

[1175] Input: Input data stored in the database

[1176] Output: Keywords and emotional state (e.g., "Shinjuku," "tomorrow," "weather," where the emotion is curiosity)

[1177] Step 4:

[1178] Generate feedback

[1179] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. It uses a response model (e.g., GPT-3) to detect ambiguous time specifications and generates feedback such as "Please specify a specific time." The feedback includes wording that takes the user's emotions into consideration.

[1180] Input: Keywords and emotional states

[1181] Output: Feedback: "Please specify a specific time."

[1182] Step 5:

[1183] Send Feedback

[1184] The server sends the generated feedback to the terminal as an HTTP response. The terminal receives this feedback and displays it in an easy-to-understand manner for the user. Specifically, the feedback is displayed in a chat window or similar.

[1185] Input: "Please specify a specific time" feedback

[1186] Output: Feedback displayed on the terminal

[1187] Step 6:

[1188] Review and correct feedback

[1189] The user checks the feedback from the server and modifies the question, re-entering "What is the weather like in Shinjuku tomorrow at 3 PM?" and clicking the resend button. The device then resends the modified question to the server.

[1190] Input: User's revised question "What's the weather like in Shinjuku tomorrow at 3pm?"

[1191] Output: Resent data

[1192] Step 7:

[1193] Reparsing and generating a response

[1194] The server receives the data again and analyzes it using the natural language processing model and emotion engine. Based on the results, it generates a specific response. For example, it generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon."

[1195] Input: Corrected input data

[1196] Output: The specific response generated

[1197] Step 8:

[1198] Sending a Response

[1199] The server generates a specific response and sends it to the terminal, which receives it and displays it to the user.

[1200] Input: Specific response

[1201] Output: Response displayed on the terminal

[1202] Step 9:

[1203] Execution of training

[1204] The server uses the collected input data and past dialogue data to run machine learning algorithms and continuously update and improve the response model and emotion engine. The dialogue history from the database is input into TensorFlow to retrain the model.

[1205] Input: Collected input data and past interaction data

[1206] Output: Updated response model and emotion engine

[1207] (Application example 2)

[1208] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1209] In recent years, improving the customer experience in brick-and-mortar stores has become increasingly important, but conventional systems have struggled to respond appropriately to customer questions and requests in real time. Furthermore, there is a lack of technology to provide feedback that takes into account the customer's emotional state, limiting the improvement of customer satisfaction. In particular, there is a need for an effective concierge service that can respond to ambiguous questions and emotional changes.

[1210] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving input data from a user; means for analyzing the input data; means for providing feedback to the user based on the analysis; means for improving the user's interaction method based on the feedback; means for learning using the input data and past interaction data and updating a response model; means for generating a refined response based on the updated response model; means for analyzing the input data using a natural language processing model and identifying ambiguous parts; and means for providing a concierge service in response to customer questions in a physical store, including an emotion engine that recognizes the user's emotional state and adjusts responses based on the user's emotional state. This makes it possible to respond to customer questions in real time and provide feedback that takes the user's emotional state into consideration, thereby significantly improving customer satisfaction.

[1211] "Input data" refers to questions, instructions, and other text information that a user sends to the system via a terminal.

[1212] "Means for analyzing" are the technical means for processing received input data and understanding its meaning, intent, and emotional state.

[1213] "Feedback" refers to responses or advice provided to the user based on the analysis results, and is information used to respond to the user's questions or requests.

[1214] "Means for improving the interaction method" are technical measures that use user feedback and correction data to enable the system to respond more accurately the next time the interaction is performed.

[1215] The "learning means" is a machine learning algorithm that uses past interaction data and newly collected input data to continuously improve the system's response model.

[1216] A "response model" is a model within the system that generates optimal responses to user questions and instructions.

[1217] A "refined response" is generated based on the updated response model and is a very specific and relevant answer to the user's question.

[1218] A "natural language processing model" is an artificial intelligence technology that analyzes input data from a user as language and understands its meaning and structure.

[1219] An "ambiguous part" is a part of the user's input data where the intent of the question or request is unclear.

[1220] An "emotion engine" is a technology that analyzes a user's emotional state and adjusts feedback and responses according to that emotional state.

[1221] A "physical store" is a store or shop that exists in a physical location and provides products and services in a face-to-face manner to customers.

[1222] A "concierge service" is a service provided in a physical store that responds to customer questions and requests and provides information about and offers products and services.

[1223] A "database" is a storage device and its management system for centrally storing input data and dialogue data collected by the system.

[1224] This invention provides a system for realizing a concierge service that provides highly accurate and emotionally sensitive feedback in real time when a customer inputs a question or request in a store. Hereinafter, an embodiment of the invention will be described in detail.

[1225] System Configuration

[1226] server

[1227] The server performs the following functions:

[1228] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[1229] 2. Data analysis: The received input data is analyzed using a natural language processing model (e.g., BERT or GPT) and an emotion engine (e.g., Microsoft Azure Emotion API) to understand the meaning, intent, and emotional state of the input.

[1230] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[1231] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[1232] 5. Run learning: Using the collected input data and past interaction data, machine learning algorithms (e.g., TensorFlow) are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[1233] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[1234] Terminal

[1235] The terminal acts as a relay for the interaction between the user and the server:

[1236] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[1237] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[1238] User

[1239] The user is the direct interlocutor of the system and goes through the following process:

[1240] 1. Entering a question: Enter a question or instruction through the terminal.

[1241] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[1242] Specific examples

[1243] For example, a user might enter "I'm looking for a special gift today. Which one would be good?" into a terminal in a physical store and send it. The terminal sends this input data in real time to a server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it recognizes high expectations and generates feedback such as "I see you're looking for a special gift today. Here are some recommended products suitable for your special occasion. We'd be happy to help you make a great choice." This feedback is displayed to the user through the terminal.

[1244] Prompt Sentence Examples

[1245] By inputting the following prompt into the generative AI model, we generate a response based on high expectations:

[1246] "Generate a response that reflects high expectations for a customer looking for a special gift."

[1247] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1248] Step 1:

[1249] The user inputs and sends questions and instructions through the terminal.

[1250] Input: Text data of user questions and instructions

[1251] Output: The input data sent

[1252] What happens: Using a smartphone or smart glasses, the user types, "I'm looking for a special gift today. What would be good?" and submits that data.

[1253] Step 2:

[1254] The terminal transmits input data to the server in real time.

[1255] Input: Input data submitted by the user

[1256] Output: The request containing the input data sent

[1257] Specific operation: The terminal receives the user's input data and sends it to the server via the Internet.

[1258] Step 3:

[1259] The server receives the input data and stores it in a database.

[1260] Input: Input data sent from the terminal

[1261] Output: Input data stored in a database

[1262] Specific operation: The server's data receiving system receives the input data and stores it in a database, including the user's ID and timestamp.

[1263] Step 4:

[1264] The server analyzes the received input data using a natural language processing model (e.g., BERT or GPT).

[1265] Input: Saved input data

[1266] Output: Analysis results (question intent, content, related keywords)

[1267] What it does: The model on the server analyzes the input data and extracts the meaning and intent of the text.

[1268] Step 5:

[1269] The server uses the analysis results to understand the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API).

[1270] Input: Analysis results

[1271] Output: Analysis results including emotional state

[1272] Specific operation: The server inputs the analysis results into the emotion engine to detect high expectations and other emotions.

[1273] Step 6:

[1274] The server generates feedback based on the analysis results and emotional state.

[1275] Input: Analysis results including emotional state

[1276] Output: Feedback text data

[1277] What it does: Based on the analysis results and emotional state, the server generates feedback such as, "You're looking for a special gift today. Here are some recommended products for your special occasion. We'd be happy to help you make a great choice."

[1278] Step 7:

[1279] The server generates feedback and sends it to the device.

[1280] Input: Text data of generated feedback

[1281] Output: Feedback data sent

[1282] Specific operations: The server generates and sends a request to send feedback data to the terminal.

[1283] Step 8:

[1284] The terminal receives the feedback and displays it to the user.

[1285] Input: Feedback data sent from the server

[1286] Output: Displayed feedback

[1287] Specific operation: The device displays the received feedback data on the screen, and the user checks it and decides the next action.

[1288] Step 9:

[1289] The server uses collected input and feedback data to learn and update the response model.

[1290] Input: Input and feedback data

[1291] Output: Updated response model

[1292] What it does: Machine learning algorithms on the server continuously train and improve response models based on collected interaction data.

[1293] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1294] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1295] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1296] [Fourth embodiment]

[1297] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1298] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1299] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1300] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1301] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1302] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1303] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1304] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1305] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1306] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1307] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1308] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1309] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1310] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. This system includes elements such as "input data," "analysis means," "feedback," "dialogue method," "learning means," "response model," "refined response," "database," "natural language processing model," and "ambiguous parts."

[1311] Server Roles

[1312] The server performs the following main functions:

[1313] 1. Receiving and storing data: The server receives the input data sent by the user via the terminal and stores it in a database.

[1314] 2. Data analysis: The server analyzes the received input data using a natural language processing model to understand the meaning and intent of the input.

[1315] 3. Feedback generation: Based on the analysis results, feedback is generated to help users improve their questioning methods.

[1316] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[1317] 5. Run learning: Using collected input data and past interaction data, the machine learning algorithm runs and continuously updates and improves the response model.

[1318] 6. Response Generation: Based on the updated response model, a refined response to the user's question is generated.

[1319] Device Role

[1320] The terminal acts as a relay for the interaction between the user and the server:

[1321] 1. Sending input data: Questions and instructions entered by the user are sent to the server in real time.

[1322] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[1323] User Roles

[1324] The user is the direct interlocutor of the system and goes through the following process:

[1325] 1. Entering a question: Enter a question or instruction through the terminal.

[1326] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[1327] Specific examples

[1328] As an example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server receives this data and analyzes it using a natural language processing model. As a result, it detects that the question is ambiguous and generates feedback such as "Please specify a specific time." The server sends this feedback to the device, which then displays it to the user.

[1329] The user checks this feedback and resubmits the question, revising it to "What's the weather like in Shinjuku tomorrow at 3 PM?" The server analyzes the data again, this time including a specific time, generating a refined answer and sending it to the device. The device then displays this answer to the user, completing the dialogue.

[1330] In this way, the present invention provides a system that can refine the content of user questions and improve the quality of dialogue with AI. Furthermore, by utilizing the database for continuous learning, the system can improve the response model and enhance the user experience.

[1331] The processing flow will be explained below.

[1332] Step 1:

[1333] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[1334] Step 2:

[1335] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[1336] Step 3:

[1337] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[1338] Step 4:

[1339] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[1340] Step 5:

[1341] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[1342] Step 6:

[1343] The server generates feedback to the user based on the analysis results, such as a message saying "Please specify a specific time."

[1344] Step 7:

[1345] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[1346] Step 8:

[1347] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[1348] Step 9:

[1349] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[1350] Step 10:

[1351] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[1352] Step 11:

[1353] The server verifies that the revised question is specific and generates a refined response, such as "The weather in Shinjuku tomorrow at 3 PM will be sunny."

[1354] Step 12:

[1355] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[1356] Step 13:

[1357] The user confirms the refined response received through the terminal, thereby completing the interaction.

[1358] Step 14:

[1359] The server runs a machine learning algorithm based on the current dialogue data and updates the response model, enabling even more accurate responses the next time the dialogue occurs.

[1360] By following the above steps, the system of the present invention effectively manages and improves the interaction between the user and the server, achieving higher quality communication.

[1361] Example 1

[1362] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1363] In conventional dialogue systems, if a user's question or instruction is ambiguous, the question itself cannot be accurately analyzed, resulting in an inability to provide an appropriate response. Furthermore, due to a lack of means to provide clear feedback, users are forced to go through a lot of trial and error. Furthermore, the response model is not continuously improved, which can lead to a decline in the quality of the dialogue.

[1364] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1365] In this invention, the server includes means for receiving input data from a user via a terminal, means for analyzing the input data using a natural language processing model to understand the meaning and intent of the input data, means for providing feedback to the user based on the analysis, means for correcting the user's question based on the feedback, means for learning using the input data and past dialogue data and updating a response model, means for generating a refined response based on the updated response model, and means for transmitting the generated refined response to the terminal and displaying it. This makes it possible to specify the content of the user's question and improve the quality of the response provided by AI.

[1366] A "terminal" is a device through which a user inputs questions and instructions and transmits and receives data to and from a server.

[1367] "Input data" refers to data including questions and instructions sent by a user via a terminal.

[1368] A "server" is a device or system that receives input data sent from a terminal and performs processing such as analysis, feedback generation, and response generation.

[1369] A "natural language processing model" is a general term for algorithms and models that analyze input data and understand its meaning and intent.

[1370] "Feedback" refers to answers or instructions that the server provides to the user based on the analysis results, and is used to clarify the content of the user's question.

[1371] "Analysis" is the process of using a natural language processing model to understand the meaning and intent of input data and make that content concrete.

[1372] "Learning" is the process by which the server runs machine learning algorithms based on input data and past interaction data to continuously update and improve the response model.

[1373] A "response model" is a model for generating answers to user questions, and is updated based on learning by the server.

[1374] An "elaborate response" is a specific and clear answer that is generated to provide appropriate information to a user's question.

[1375] A "database" is a storage device that stores input data and past dialogue data and can be accessed by a server for analysis and learning.

[1376] An "ambiguous part" is a part of a user's question or instruction that is unclear or difficult to interpret, and feedback is generated by identifying this.

[1377] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. In this system, the server receives data entered by the user via the terminal, and performs processing such as analysis, feedback generation, and response generation to concretize the content of the user's question and provide a high-quality response using AI. Specific embodiments of the present invention are described below.

[1378] The server mainly uses the following hardware and software:

[1379] Server hardware: A server machine with a powerful CPU and sufficient memory

[1380] Database: A database system (e.g., MySQL, PostgreSQL) for storing input data and past dialogue data.

[1381] Natural language processing model: A natural language processing model (e.g., GPT-4) to analyze the input data.

[1382] Machine learning algorithms: Algorithms for learning and updating response models using past data (e.g., TensorFlow, PyTorch)

[1383] The terminal acts as a relay for data between the user and the server. The terminal can be a device such as a PC, smartphone, or tablet that runs an application or web browser that communicates with the server via an internet connection.

[1384] The user operates the terminal to input questions and instructions to the system. The operation of the system will be explained below using a concrete example.

[1385] As a concrete example, consider the case where a user asks a question about the weather. The user inputs "What will the weather be like in Shinjuku tomorrow?" into the device. The device sends this input data to the server. The server analyzes the received data using a natural language processing model (GPT-4) to understand the intent of the question. If the server detects that the question is ambiguous, it generates feedback such as "Please specify a specific time." The server then sends this feedback to the device, which displays it to the user.

[1386] The user checks the feedback from the server, modifies the question, and re-enters it. For example, the user enters a specific question such as, "What is the weather like in Shinjuku tomorrow at 3 PM?" The device then sends the modified data to the server again. The server again analyzes the data and generates a more specific answer this time. It generates a refined response, "The weather will be cloudy in Shinjuku tomorrow at 3 PM," and sends it to the device. The device then displays the received response to the user.

[1387] In this way, the system can clarify the user's ambiguous questions and provide appropriate feedback, improving the quality of the AI ​​responses. Furthermore, the server uses input data and past dialogue data to perform machine learning and continuously update and improve the response model, improving the user experience.

[1388] Prompt Sentence Examples

[1389] First input: "What's the weather like in Shinjuku tomorrow?"

[1390] Corrected input: "What's the weather like in Shinjuku tomorrow at 3pm?"

[1391] In this way, the present invention provides a system that clarifies the content of a user's question and improves the quality of dialogue with AI.

[1392] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1393] Step 1: User enters question

[1394] The user inputs a question or instruction into the terminal. For example, the user inputs "Please tell me what the weather will be like in Shinjuku tomorrow." The input text becomes the input data for the terminal.

[1395] Step 2: Submitting input data

[1396] The device sends the user-entered question data to the server in real time via an HTTP POST request, with the input data included in the body of the API request.

[1397] Step 3: Receiving and storing data

[1398] The server receives the input data sent from the terminal. The received data is saved in a database on the server. The database stores the user's question ("What is the weather in Shinjuku tomorrow?").

[1399] Step 4: Analysis using natural language processing models

[1400] The server analyzes the received input data using a natural language processing model (GPT-4). The input data is input into the model, and an analysis result is generated. The analysis result includes the intent of the question and any ambiguities.

[1401] Step 5: Generate feedback

[1402] The server generates feedback based on the analysis results of the natural language processing model. For example, if the question is ambiguous, the server generates feedback such as "Please specify a specific time." The generated feedback is in text format.

[1403] Step 6: Submit your feedback

[1404] The server sends the generated feedback to the terminal. A feedback message ("Please specify a specific time") is sent as an HTTP response.

[1405] Step 7: View your feedback

[1406] The terminal displays the feedback received from the server to the user, and the feedback message is displayed on the screen of the terminal to notify the user.

[1407] Step 8: Modifying the Question

[1408] The user modifies the question based on the received feedback. For example, the user inputs a specific question such as, "Please tell me the weather in Shinjuku tomorrow at 3:00 PM." The modified question becomes the new input data.

[1409] Step 9: Resend the corrected data

[1410] The terminal sends the corrected input data to the server again via an HTTP POST request. The corrected input data is included in the body of the API request.

[1411] Step 10: Reanalyze the data

[1412] The server then analyzes the revised data using a natural language processing model. The revised data is input into the model, and an analysis result is generated. The analysis result for the specific question is obtained.

[1413] Step 11: Generate a response

[1414] The server generates a refined response based on the analysis results. For example, it may generate an answer such as "The weather in Shinjuku tomorrow at 3:00 PM will be cloudy." The generated response is in text format.

[1415] Step 12: Sending a Response

[1416] The server sends the generated response to the terminal. The response message ("The weather in Shinjuku tomorrow at 3 PM will be cloudy") is sent as an HTTP response.

[1417] Step 13: View the response

[1418] The terminal displays the response received from the server to the user. The response message is displayed on the terminal screen to notify the user. This display completes the interaction.

[1419] (Application example 1)

[1420] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1421] In today's retail industry, when customers ask questions about products in physical stores, staff are expected to respond quickly and accurately. However, many stores struggle to immediately provide customers with the specific product and inventory information they require, resulting in lower customer satisfaction and lost sales opportunities. Furthermore, the quality of the information provided varies depending on the staff's knowledge and experience, resulting in an inconsistent customer experience. Therefore, there is a growing need for a system that enables staff in physical stores to efficiently respond to customers and provide accurate information.

[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1423] In this invention, the server includes means for receiving input data from a user, means for analyzing the input data, means for providing feedback to the user based on the analysis, means for improving the user's interaction method based on the feedback, means for learning using the input data and past interaction data and updating a response model, means for generating refined responses based on the updated response model, means for supporting interaction with the user in real time using a smart device, and means for providing product information and inventory information based on the support, thereby enabling staff to provide product information and inventory information to customers quickly and accurately using the smart device.

[1424] "Means for receiving input data from a user" refers to the function of collecting data through an interface for users to input questions or instructions and sending it to the system.

[1425] The "means for analyzing the input data" refers to a function that analyzes the received data using natural language processing technology, etc., and performs processing to understand its meaning and intent.

[1426] The "means for providing feedback to the user based on the analysis" refers to a function that provides advice and correction instructions to the user to ask more specific questions based on the analysis results.

[1427] "Means for improving the user's interaction method based on the feedback" refers to a process by which the user understands the feedback and improves the content of the questions and instructions.

[1428] "Means for learning using the input data and past dialogue data and updating the response model" refers to a function that performs machine learning based on collected data, continuously improves the system's response model, and enables it to generate more accurate responses.

[1429] "Means for generating a refined response based on the updated response model" refers to a function that uses an improved response model to generate a specific and accurate answer to a user's question.

[1430] "Means for supporting real-time interaction with users using smart devices" refers to a function for using devices such as smart glasses or smartphones to interact with users in real time and provide specific information.

[1431] "Means for providing product information and inventory information based on the support" refers to a function that retrieves product information and inventory status from a database based on a user's question and presents that information to the user.

[1432] The present invention provides a system for supporting interactions with customers in a physical store. A specific embodiment of this system will be described below.

[1433] System Program

[1434] 1. Hardware configuration:

[1435] It uses servers, smart devices (such as smart glasses and smartphones), and communication networks (Wi-Fi or 5G). It uses AWS EC2 instances for the servers, Amazon RDS for the database, and AWS Sagemaker for natural language processing.

[1436] 2. Software configuration:

[1437] An Android-based application is installed on the smart device, and a Node.js server is operated on the server side. Python is used for data analysis, and BERT is used as the natural language processing model.

[1438] System Operation Overview

[1439] 1. Getting and sending user input data:

[1440] The user (staff member) uses a smart device to input customer questions by voice, and the smart device converts the voice data into text using the Google Speech-to-Text API and sends the text data to the server.

[1441] 2. Data Analysis:

[1442] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker), which identifies the intent of the question and any ambiguities, and generates appropriate feedback.

[1443] 3. Feedback and response generation:

[1444] Based on the analysis results, the server generates feedback to help the user refine their question. If necessary, it sends additional questions or correction instructions to the smart device. If the user re-enters the necessary information, the server analyzes the data again and generates a refined response.

[1445] 4. Information provision:

[1446] The refined response includes product and inventory information retrieved from a database (Amazon RDS), allowing users to provide accurate and prompt information to customers.

[1447] Specific examples

[1448] Example 1

[1449] Scenario: A customer asks, "Tell me about this product."

[1450] Prompt statement:

[1451] Provide specific product information. You need to tell customers the description, price, and features of this product.

[1452] Example response:

[1453] The smart device will display the message, "This product is the latest model and costs 5,000 yen. Its features include waterproof functionality and a long battery life."

[1454] Example 2

[1455] Scenario: A customer asks, "Is this item in stock?"

[1456] Prompt statement:

[1457] Please check the inventory information and provide it to your customers. We need to check the current availability of this item.

[1458] Example response:

[1459] The smart device will display the message, "There are currently 5 of this item in stock in the store."

[1460] In this way, the present invention provides a specific system configuration and processing procedure for efficiently and accurately handling customers in a brick-and-mortar store.

[1461] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1462] Step 1: Getting and sending input data from the user

[1463] A user (staff member) uses a smart device to voice-input a customer question. The smart device then converts the voice data into text using the Google Speech-to-Text API. The input is voice data, and the output is text data. The smart device then sends the converted text data to the server.

[1464] Step 2: Analyze the data

[1465] The server analyzes the received text data using a natural language processing model (the BERT model on AWS Sagemaker). Specifically, it processes and calculates the data to identify the intent and ambiguity of the text data. The input is the text data, and the output is the analysis results (intent, ambiguity).

[1466] Step 3: Generate feedback

[1467] Based on the analysis results, the server generates feedback to help the user refine their query. The feedback includes additional questions and correction instructions as needed. The input is the analysis results, and the output is a feedback message. This is sent to the smart device and displayed to the user.

[1468] Step 4: User Modification of Question

[1469] The user re-enters a question based on the feedback and sends the revised text data to the server. The input is the feedback message and the revised question, and the output is the text data of the revised question.

[1470] Step 5: Reparsing and generating a response

[1471] The server then analyzes the revised text data again using a natural language processing model. Based on the analysis results, it generates a refined response. The input is the text data of the revised question, and the output is the refined response.

[1472] Step 6: Provide information

[1473] Based on the refined response, the server retrieves product information and inventory information from the database (Amazon RDS) and generates information to be provided to the user. The input is the refined response and information from the database, and the output is the final information provided to the user.

[1474] Step 7: Display to the user

[1475] The final information is sent to the smart device and displayed to the user, allowing the user to provide accurate product and inventory information to customers. The input is the final information from the server, and the output is the information displayed on the smart device.

[1476] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1477] The system of the present invention provides a means for effective dialogue between a server, a terminal, and a user. It also incorporates an emotion engine that recognizes the user's emotions and adjusts or enhances responses based on those emotions. This system includes elements such as input data, analysis means, feedback, a dialogue method, learning means, response model, refined responses, a database, a natural language processing model, ambiguity, and an emotion engine.

[1478] Server Roles

[1479] The server performs the following functions:

[1480] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[1481] 2. Data analysis: The received input data is analyzed using a natural language processing model and an emotion engine to understand the meaning, intent, and emotional state of the input.

[1482] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[1483] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[1484] 5. Run Learning: Using collected input data and past interaction data, machine learning algorithms are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[1485] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[1486] Device Role

[1487] The terminal acts as a relay for the interaction between the user and the server:

[1488] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[1489] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[1490] User Roles

[1491] The user is the direct interlocutor of the system and goes through the following process:

[1492] 1. Entering a question: Enter a question or instruction through the terminal.

[1493] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[1494] Specific examples

[1495] As an example, consider a question about the weather. The user enters "What will the weather be like in Shinjuku tomorrow?" into the device and sends it. The device then sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server then sends this feedback to the device, which then displays it to the user.

[1496] The user checks the feedback, modifies the question to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[1497] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[1498] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, the response model and emotion engine can be improved, resulting in a better user experience.

[1499] The processing flow will be explained below.

[1500] Step 1:

[1501] The user inputs a question using the input interface of the terminal. For example, the user inputs "What is the weather in Shinjuku tomorrow?" into the text box.

[1502] Step 2:

[1503] After completing the input, the user clicks the send button, and the terminal transmits the input question data to the server in real time.

[1504] Step 3:

[1505] The server receives the question data sent from the terminal. At this point, the text data "Please tell me what the weather will be like in Shinjuku tomorrow" arrives at the server.

[1506] Step 4:

[1507] The server stores the received question data in a database, thereby accumulating a record of user interactions.

[1508] Step 5:

[1509] The server analyzes the received question data using a natural language processing model, specifically detecting the ambiguous time specification "tomorrow."

[1510] Step 6:

[1511] The server uses an emotion engine to analyze the user's emotional state, for example, by inferring the user's emotions from text and detecting stress, joy, anxiety, etc.

[1512] Step 7:

[1513] The server generates feedback to the user based on the analysis results, such as "Please specify a specific time," and adds expressions that take into account the user's emotional state.

[1514] Step 8:

[1515] The server sends the generated feedback message to the terminal, so that the feedback is delivered to the user.

[1516] Step 9:

[1517] The terminal displays the feedback message received from the server to the user, and the user confirms the message "Please specify a specific time."

[1518] Step 10:

[1519] The user checks the feedback and modifies the question, for example, by entering "What is the weather like in Shinjuku at 3 PM tomorrow?" and clicking the submit button again.

[1520] Step 11:

[1521] The terminal again transmits the revised question data to the server, which receives the revised data and analyzes it again using the natural language processing model.

[1522] Step 12:

[1523] The server then verifies that the revised question is specific and generates a refined response, adding expressions that take the user's emotional state into consideration. For example, the server generates a response such as, "The weather in Shinjuku tomorrow at 3 p.m. will be sunny. Have a nice afternoon."

[1524] Step 13:

[1525] The server generates a response and sends it to the terminal, which receives it and displays it to the user.

[1526] Step 14:

[1527] The user confirms the refined response received through the terminal, thereby completing the interaction.

[1528] Step 15:

[1529] The server runs a machine learning algorithm based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses in the next conversation.

[1530] By following the above steps, the system of the present invention effectively manages and improves the dialogue between the user and the server, realizing higher quality and more emotionally sensitive communication.

[1531] Example 2

[1532] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1533] Conventional dialogue systems lack natural and rich communication because they do not consider the user's emotional state. They also fail to properly handle ambiguity in user questions and instructions, resulting in insufficient feedback and an inability to quickly respond to user needs. Furthermore, they lack a means to effectively utilize past dialogue data to continuously improve the system's response model, limiting the user experience.

[1534] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving input data from a user and storing it in a database, means for analyzing the input data using a natural language processing model and an emotion engine to understand the meaning, intention, and emotional state of the input data, means for generating feedback corresponding to the user's question method and emotional state based on the analysis and transmitting the feedback to the terminal, means for improving the user's dialogue method based on the feedback, means for executing a machine learning algorithm using the input data and past dialogue data to continuously update a response model and emotion engine, and means for generating refined responses based on the updated response model and adjusting them based on emotion understanding. This makes it possible to realize natural and effective dialogue that takes the user's emotional state into consideration, reduce ambiguity in the user's questions and instructions, and continuously improve the system's response model.

[1535] A "user" is an entity that provides input data to and receives feedback from a system.

[1536] "Input Data" refers to questions, instructions, or other information provided by a user to a system.

[1537] A "database" is a data management system that stores input data and past dialogue data and retrieves them as needed.

[1538] A "natural language processing model" is a machine learning model that analyzes human language and understands its meaning and intent.

[1539] An "emotion engine" is an analytical system for detecting a user's emotional state from input data and adjusting responses accordingly.

[1540] "Analysis" refers to processing input data using a natural language processing model and emotion engine to understand its meaning, intent, and emotional state.

[1541] "Feedback" refers to the response or instructions provided to the user based on the analysis results.

[1542] "Terminal" refers to a device used by a user to provide input data to a system.

[1543] A "machine learning algorithm" is a computational method for training a model based on data and continuously improving its performance.

[1544] A "response model" is a model for generating appropriate responses to user questions and instructions.

[1545] "Refined responses" refer to detailed responses generated based on an updated response model and tailored based on emotion understanding.

[1546] The present invention is a system for effectively conducting dialogue between a server, a terminal, and a user. The present invention incorporates an emotion engine that recognizes the user's emotions and adjusts and enhances responses based on those emotions. The system includes the following elements:

[1547] Server Roles

[1548] Data Receipt and Storage:

[1549] The server receives input data sent by the user via the terminal and stores the data in a database. The database used is, for example, MySQL. When a user enters "What is the weather in Shinjuku tomorrow?", the server receives the data via an HTTP POST request and stores it in the database.

[1550] Data analysis:

[1551] The server analyzes the received input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). This allows it to understand the meaning, intent, and emotional state of the input. The server analyzes the input data, "What's the weather in Shinjuku tomorrow?" and extracts the keywords "Shinjuku," "tomorrow," and "weather," as well as the user's emotional state (curiosity).

[1552] Generate feedback:

[1553] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. This feedback is generated using a response model (e.g., GPT-3). It detects ambiguous time specifications and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration.

[1554] Send feedback:

[1555] The generated feedback is sent to the terminal and displayed to the user. The server sends feedback to the terminal saying "Please specify a specific time," which is displayed to the user.

[1556] Run the training:

[1557] The server uses the collected input data and past dialogue data to run machine learning algorithms (e.g., TensorFlow or PyTorch) and continuously update and improve the response model and emotion engine. The server inputs the dialogue history from the database into TensorFlow to retrain the response model and emotion engine.

[1558] Response generation:

[1559] Based on the updated response model, the server generates a refined response to the user's question and adjusts it based on sentiment understanding. In response to the question, "What is the weather like in Shinjuku tomorrow at 3 PM?", the server generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." and sends it to the device.

[1560] Device Role

[1561] Sending input data:

[1562] The device sends questions, instructions, and emotion-related data entered by the user to the server in real time. For example, if a user enters "What's the weather in Shinjuku tomorrow?" and presses the send button, the data is sent to the server.

[1563] Receiving and viewing feedback:

[1564] The terminal receives the feedback from the server and displays it in an easy-to-understand manner for the user. The terminal displays the feedback received from the server, "Please specify a specific time," in the chat window.

[1565] User Roles

[1566] Enter your question:

[1567] The user inputs questions or instructions through the terminal. The user types "Please tell me the weather in Shinjuku tomorrow" into the text box on the terminal and clicks the send button.

[1568] Review and correct feedback:

[1569] The user checks the feedback sent from the server, modifies the question if necessary, and resubmits it. The user sees the feedback "Please specify a specific time," modifies the question to "What is the weather like in Shinjuku at 3 PM tomorrow," and submits it again.

[1570] Specific examples

[1571] For example, consider the case where a user inputs and sends "What's the weather in Shinjuku tomorrow?" into a terminal. The terminal sends this input data to the server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it detects an ambiguous time specification and the user's emotional state, and generates feedback such as "Please specify a specific time." This feedback is expressed in language that takes the user's emotions into consideration. The server sends this feedback to the terminal, which then displays it to the user.

[1572] The user checks the feedback, amends it to "What's the weather like in Shinjuku tomorrow at 3 PM?" and resubmits it. The server analyzes the data again and generates a specific response. Using an emotion engine, a response is generated that takes into account the user's emotional state, such as "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon." This response is displayed to the user on their device, completing the dialogue.

[1573] The server runs machine learning algorithms based on the conversation data and updates the response model and emotion engine, enabling more accurate and emotion-sensitive responses the next time the conversation occurs.

[1574] Prompt Sentence Examples

[1575] Here are some examples of prompts for generative AI models:

[1576] Prompt statement:

[1577] Use an emotion engine to generate a response to the user's question, "What will the weather be like in Shinjuku tomorrow?" If no specific time is specified, provide feedback, and generate a specific response when the question is entered again.

[1578] Example 1: User's first question

[1579] User: What's the weather like in Shinjuku tomorrow?

[1580] System: What time of day would you like to know the weather in Shinjuku tomorrow? Please specify a specific time.

[1581] Example 2: User question after modification

[1582] User: What's the weather like in Shinjuku tomorrow at 3pm?

[1583] System: The weather in Shinjuku will be sunny tomorrow at 3 PM. Have a nice afternoon.

[1584] As described above, the system of the present invention can improve the quality of dialogue by taking into account the user's question content and emotional state. Furthermore, by utilizing the database and continuously learning, it is possible to improve the response model and emotion engine, thereby improving the user experience.

[1585] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1586] Step 1:

[1587] Sending input data

[1588] The user types "Please tell me the weather in Shinjuku tomorrow" into the terminal and clicks the send button.

[1589] The terminal sends this input data to the server as an HTTP POST request.

[1590] Input: User question: "What's the weather like in Shinjuku tomorrow?"

[1591] Output: Data sent to the server as an HTTP POST request

[1592] Step 2:

[1593] Receiving and storing data

[1594] The server receives the input data submitted by the user. This data is stored in a database, such as MySQL. The server extracts the data from the POST request and stores it in the "inputs" table in the database.

[1595] Input: HTTP POST request

[1596] Output: Input data stored in a database

[1597] Step 3:

[1598] Data analysis

[1599] The server analyzes the stored input data using a natural language processing model (e.g., SpaCy or BERT) and an emotion engine (e.g., Affectiva). Through this analysis, the meaning and intent of the text and the user's emotional state are extracted. Specifically, keywords are extracted from the text and the emotional state is detected.

[1600] Input: Input data stored in the database

[1601] Output: Keywords and emotional state (e.g., "Shinjuku," "tomorrow," "weather," where the emotion is curiosity)

[1602] Step 4:

[1603] Generate feedback

[1604] Based on the analysis results, the server generates feedback according to the user's question style and emotional state. It uses a response model (e.g., GPT-3) to detect ambiguous time specifications and generates feedback such as "Please specify a specific time." The feedback includes wording that takes the user's emotions into consideration.

[1605] Input: Keywords and emotional states

[1606] Output: Feedback: "Please specify a specific time."

[1607] Step 5:

[1608] Send Feedback

[1609] The server sends the generated feedback to the terminal as an HTTP response. The terminal receives this feedback and displays it in an easy-to-understand manner for the user. Specifically, the feedback is displayed in a chat window or similar.

[1610] Input: "Please specify a specific time" feedback

[1611] Output: Feedback displayed on the terminal

[1612] Step 6:

[1613] Review and correct feedback

[1614] The user checks the feedback from the server and modifies the question, re-entering "What is the weather like in Shinjuku tomorrow at 3 PM?" and clicking the resend button. The device then resends the modified question to the server.

[1615] Input: User's revised question "What's the weather like in Shinjuku tomorrow at 3pm?"

[1616] Output: Resent data

[1617] Step 7:

[1618] Reparsing and generating a response

[1619] The server receives the data again and analyzes it using the natural language processing model and emotion engine. Based on the results, it generates a specific response. For example, it generates a response such as, "The weather in Shinjuku tomorrow at 3 PM is sunny. Have a nice afternoon."

[1620] Input: Corrected input data

[1621] Output: The specific response generated

[1622] Step 8:

[1623] Sending a Response

[1624] The server generates a specific response and sends it to the terminal, which receives it and displays it to the user.

[1625] Input: Specific response

[1626] Output: Response displayed on the terminal

[1627] Step 9:

[1628] Execution of training

[1629] The server uses the collected input data and past dialogue data to run machine learning algorithms and continuously update and improve the response model and emotion engine. The dialogue history from the database is input into TensorFlow to retrain the model.

[1630] Input: Collected input data and past interaction data

[1631] Output: Updated response model and emotion engine

[1632] (Application example 2)

[1633] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1634] In recent years, improving the customer experience in brick-and-mortar stores has become increasingly important, but conventional systems have struggled to respond appropriately to customer questions and requests in real time. Furthermore, there is a lack of technology to provide feedback that takes into account the customer's emotional state, limiting the improvement of customer satisfaction. In particular, there is a need for an effective concierge service that can respond to ambiguous questions and emotional changes.

[1635] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving input data from a user; means for analyzing the input data; means for providing feedback to the user based on the analysis; means for improving the user's interaction method based on the feedback; means for learning using the input data and past interaction data and updating a response model; means for generating a refined response based on the updated response model; means for analyzing the input data using a natural language processing model and identifying ambiguous parts; and means for providing a concierge service in response to customer questions in a physical store, including an emotion engine that recognizes the user's emotional state and adjusts responses based on the user's emotional state. This makes it possible to respond to customer questions in real time and provide feedback that takes the user's emotional state into consideration, thereby significantly improving customer satisfaction.

[1636] "Input data" refers to questions, instructions, and other text information that a user sends to the system via a terminal.

[1637] "Means for analyzing" are the technical means for processing received input data and understanding its meaning, intent, and emotional state.

[1638] "Feedback" refers to responses or advice provided to the user based on the analysis results, and is information used to respond to the user's questions or requests.

[1639] "Means for improving the interaction method" are technical measures that use user feedback and correction data to enable the system to respond more accurately the next time the interaction is performed.

[1640] The "learning means" is a machine learning algorithm that uses past interaction data and newly collected input data to continuously improve the system's response model.

[1641] A "response model" is a model within the system that generates optimal responses to user questions and instructions.

[1642] A "refined response" is generated based on the updated response model and is a very specific and relevant answer to the user's question.

[1643] A "natural language processing model" is an artificial intelligence technology that analyzes input data from a user as language and understands its meaning and structure.

[1644] An "ambiguous part" is a part of the user's input data where the intent of the question or request is unclear.

[1645] An "emotion engine" is a technology that analyzes a user's emotional state and adjusts feedback and responses according to that emotional state.

[1646] A "physical store" is a store or shop that exists in a physical location and provides products and services in a face-to-face manner to customers.

[1647] A "concierge service" is a service provided in a physical store that responds to customer questions and requests and provides information about and offers products and services.

[1648] A "database" is a storage device and its management system for centrally storing input data and dialogue data collected by the system.

[1649] This invention provides a system for realizing a concierge service that provides highly accurate and emotionally sensitive feedback in real time when a customer inputs a question or request in a store. Hereinafter, an embodiment of the invention will be described in detail.

[1650] System Configuration

[1651] server

[1652] The server performs the following functions:

[1653] 1. Receiving and storing data: Receives input data sent by the user via the terminal and stores it in a database.

[1654] 2. Data analysis: The received input data is analyzed using a natural language processing model (e.g., BERT or GPT) and an emotion engine (e.g., Microsoft Azure Emotion API) to understand the meaning, intent, and emotional state of the input.

[1655] 3. Feedback generation: Based on the analysis results, feedback is generated according to the user's question method and emotional state.

[1656] 4. Sending feedback: The generated feedback is sent to the device and displayed to the user.

[1657] 5. Run learning: Using the collected input data and past interaction data, machine learning algorithms (e.g., TensorFlow) are run to continuously update and improve the response model, which in turn enhances the emotion engine.

[1658] 6. Response Generation: Based on the updated response model, generate refined responses to the user's questions and adjust them based on sentiment understanding.

[1659] Terminal

[1660] The terminal acts as a relay for the interaction between the user and the server:

[1661] 1. Sending input data: Questions, instructions, and emotion-related data entered by the user are sent to the server in real time.

[1662] 2. Receiving and displaying feedback: Receives feedback from the server and displays it in a way that is easy for the user to understand.

[1663] User

[1664] The user is the direct interlocutor of the system and goes through the following process:

[1665] 1. Entering a question: Enter a question or instruction through the terminal.

[1666] 2. Review and correct feedback: Review the feedback sent by the server, correct your question if necessary, and resubmit.

[1667] Specific examples

[1668] For example, a user might enter "I'm looking for a special gift today. Which one would be good?" into a terminal in a physical store and send it. The terminal sends this input data in real time to a server. The server receives the data and analyzes it using a natural language processing model and an emotion engine. As a result, it recognizes high expectations and generates feedback such as "I see you're looking for a special gift today. Here are some recommended products suitable for your special occasion. We'd be happy to help you make a great choice." This feedback is displayed to the user through the terminal.

[1669] Prompt Sentence Examples

[1670] By inputting the following prompt into the generative AI model, we generate a response based on high expectations:

[1671] "Generate a response that reflects high expectations for a customer looking for a special gift."

[1672] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1673] Step 1:

[1674] The user inputs and sends questions and instructions through the terminal.

[1675] Input: Text data of user questions and instructions

[1676] Output: The input data sent

[1677] What happens: Using a smartphone or smart glasses, the user types, "I'm looking for a special gift today. What would be good?" and submits that data.

[1678] Step 2:

[1679] The terminal transmits input data to the server in real time.

[1680] Input: Input data submitted by the user

[1681] Output: The request containing the input data sent

[1682] Specific operation: The terminal receives the user's input data and sends it to the server via the Internet.

[1683] Step 3:

[1684] The server receives the input data and stores it in a database.

[1685] Input: Input data sent from the terminal

[1686] Output: Input data stored in a database

[1687] Specific operation: The server's data receiving system receives the input data and stores it in a database, including the user's ID and timestamp.

[1688] Step 4:

[1689] The server analyzes the received input data using a natural language processing model (e.g., BERT or GPT).

[1690] Input: Saved input data

[1691] Output: Analysis results (question intent, content, related keywords)

[1692] What it does: The model on the server analyzes the input data and extracts the meaning and intent of the text.

[1693] Step 5:

[1694] The server uses the analysis results to understand the emotional state using an emotion engine (e.g., Microsoft Azure Emotion API).

[1695] Input: Analysis results

[1696] Output: Analysis results including emotional state

[1697] Specific operation: The server inputs the analysis results into the emotion engine to detect high expectations and other emotions.

[1698] Step 6:

[1699] The server generates feedback based on the analysis results and emotional state.

[1700] Input: Analysis results including emotional state

[1701] Output: Feedback text data

[1702] What it does: Based on the analysis results and emotional state, the server generates feedback such as, "You're looking for a special gift today. Here are some recommended products for your special occasion. We'd be happy to help you make a great choice."

[1703] Step 7:

[1704] The server generates feedback and sends it to the device.

[1705] Input: Text data of generated feedback

[1706] Output: Feedback data sent

[1707] Specific operations: The server generates and sends a request to send feedback data to the terminal.

[1708] Step 8:

[1709] The terminal receives the feedback and displays it to the user.

[1710] Input: Feedback data sent from the server

[1711] Output: Displayed feedback

[1712] Specific operation: The device displays the received feedback data on the screen, and the user checks it and decides the next action.

[1713] Step 9:

[1714] The server uses collected input and feedback data to learn and update the response model.

[1715] Input: Input and feedback data

[1716] Output: Updated response model

[1717] What it does: Machine learning algorithms on the server continuously train and improve response models based on collected interaction data.

[1718] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1719] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1720] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1721] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1722] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1723] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1724] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1725] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1726] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1727] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1728] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1729] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1730] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1731] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1732] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1733] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1734] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1735] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1736] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1737] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1738] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1739] The following is further disclosed regarding the above embodiment.

[1740] (Claim 1)

[1741] means for receiving input data from a user;

[1742] means for analyzing the input data;

[1743] means for providing feedback to a user based on said analysis;

[1744] means for improving the user's interaction method based on said feedback;

[1745] means for performing learning using the input data and past dialogue data and updating a response model;

[1746] means for generating a refined response based on the updated response model;

[1747] A system including:

[1748] (Claim 2)

[1749] a database for storing the input data and past dialogue data;

[1750] a means for acquiring dialogue data from the database and sharing the analysis results;

[1751] The system of claim 1 further comprising:

[1752] (Claim 3)

[1753] 2. The system according to claim 1, wherein the analyzing means includes means for analyzing input data using a natural language processing model and identifying ambiguous portions.

[1754] "Example 1"

[1755] (Claim 1)

[1756] means for receiving input data from a user via a terminal;

[1757] means for analyzing the input data using a natural language processing model to understand meaning and intent;

[1758] means for providing feedback to a user based on said analysis;

[1759] a means for modifying the content of the user's question based on the feedback;

[1760] means for performing learning using the input data and past dialogue data and updating a response model;

[1761] means for generating a refined response based on the updated response model;

[1762] means for transmitting the generated refined response to a terminal for display;

[1763] A system including:

[1764] (Claim 2)

[1765] a database for storing the input data and past dialogue data;

[1766] a means for acquiring dialogue data from the database and sharing the analysis results;

[1767] and means for continuously learning based on the shared analysis results;

[1768] The system of claim 1 further comprising:

[1769] (Claim 3)

[1770] 2. The system of claim 1, wherein the analyzing means includes means for analyzing input data using a natural language processing model, identifying ambiguities, and generating specific feedback.

[1771] "Application Example 1"

[1772] (Claim 1)

[1773] means for receiving input data from a user;

[1774] means for analyzing the input data;

[1775] means for providing feedback to a user based on said analysis;

[1776] means for improving the user's interaction method based on said feedback;

[1777] means for performing learning using the input data and past dialogue data and updating a response model;

[1778] means for generating a refined response based on the updated response model;

[1779] A means for supporting real-time interaction with a user using a smart device;

[1780] means for providing product information and inventory information based on the support;

[1781] A system including:

[1782] (Claim 2)

[1783] a database for storing the input data and past dialogue data;

[1784] a means for acquiring dialogue data from the database and sharing the analysis results;

[1785] The system of claim 1 further comprising:

[1786] (Claim 3)

[1787] 2. The system according to claim 1, wherein the analyzing means includes means for analyzing input data using a natural language processing model and identifying ambiguous portions.

[1788] "Example 2: Combining Emotion Engines"

[1789] (Claim 1)

[1790] means for receiving input data from a user and storing it in a database;

[1791] means for analyzing the input data using a natural language processing model and an emotion engine to grasp the meaning, intent, and emotional state of the input data;

[1792] means for generating feedback according to the user's questioning method and emotional state based on the analysis and transmitting the feedback to the terminal;

[1793] means for improving the user's interaction method based on said feedback;

[1794] means for executing a machine learning algorithm using the input data and past interaction data to continuously update a response model and an emotion engine;

[1795] means for generating a refined response based on the updated response model and adjusting the refined response based on emotion understanding;

[1796] A system including:

[1797] (Claim 2)

[1798] a database for storing the input data and past dialogue data;

[1799] a means for acquiring dialogue data from the database and sharing the analysis results;

[1800] The system of claim 1 further comprising:

[1801] (Claim 3)

[1802] 2. The system according to claim 1, wherein the analyzing means includes means for analyzing input data using a natural language processing model, detecting ambiguous parts, and grasping an emotional state using an emotion engine.

[1803] "Application example 2 when combining emotion engines"

[1804] (Claim 1)

[1805] means for receiving input data from a user;

[1806] means for analyzing the input data;

[1807] means for providing feedback to a user based on said analysis;

[1808] means for improving the user's interaction method based on said feedback;

[1809] means for performing learning using the input data and past dialogue data and updating a response model;

[1810] means for generating a refined response based on the updated response model;

[1811] the analyzing means analyzes input data using a natural language processing model and identifies ambiguous parts;

[1812] an emotion engine that recognizes the user's emotional state and adjusts responses accordingly;

[1813] A means for providing concierge services to customers in physical stores to answer their questions;

[1814] A system including:

[1815] (Claim 2)

[1816] a database for storing the input data and past dialogue data;

[1817] a means for acquiring dialogue data from the database and sharing the analysis results;

[1818] an analysis engine for recognizing said emotional state;

[1819] a means for generating feedback to guide the customer to products and services in the store after identifying the ambiguous portion;

[1820] The system of claim 1 further comprising:

[1821] (Claim 3)

[1822] The system according to claim 1, wherein the analysis means includes means for analyzing input data using a natural language processing model, identifying ambiguous parts and the customer's emotional state, and providing guidance on products and services suitable for the customer based on the results. [Explanation of symbols]

[1823] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving input data from a user; means for analyzing the input data; means for providing feedback to a user based on said analysis; means for improving the user's interaction method based on said feedback; means for performing learning using the input data and past dialogue data and updating a response model; means for generating a refined response based on the updated response model; A system including:

2. a database for storing the input data and past dialogue data; a means for acquiring dialogue data from the database and sharing the analysis results; The system of claim 1 further comprising:

3. 2. The system according to claim 1, wherein the analyzing means includes means for analyzing the input data using a natural language processing model and identifying ambiguous portions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A