system
The system facilitates realistic customer service training through role selection and AI-driven dialogue simulations, addressing the lack of effective training methods by providing real-time feedback for skill improvement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Existing customer service training methods lack effective and practical ways for new crew members to simulate real customer service scenarios, often requiring external assistance and lacking real-time feedback, which hinders skill improvement.
A system that allows users to select roles and engage in realistic dialogue simulations using artificial intelligence models, providing real-time feedback and evaluation.
Enables users to improve their customer service skills efficiently through realistic simulations without external assistance, offering timely feedback for skill enhancement.
Smart Images

Figure 2026064667000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005] , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the customer service industry, there is a problem that there is a lack of effective and practical training methods. In particular, it is a problem that new crew members have limited opportunities to simulate actual customer service scenes. In addition, crew members often do not receive real-time feedback or evaluation for improving their dialogue skills with customers. As a result, the skill improvement of the customer service crew may be delayed. With the current training method, other manpower is required to play the roles of customers and crew members, which is not efficient.
Means for Solving the Problems
[0005] This invention provides a system that generates an interface for users to select a role and generates and displays initial questions and topics using an artificial intelligence model according to the selected role. Specifically, it includes means for the user to select a role, means for selecting an artificial intelligence model based on the selection result, means for displaying the generated questions on a client terminal, means for receiving the user's response and generating the next question or comment, and means for evaluating the response and generating feedback when the dialogue ends. This allows users to improve their skills through realistic customer service simulations without the need for external assistance. Furthermore, real-time feedback can efficiently promote user growth.
[0006] A "user" refers to an individual who uses this system to perform role selection and dialogue simulations.
[0007] "Role" refers to a role within the system, specifically, whether it's a "crew member" or a "customer."
[0008] "Interface" refers to the screen or input method that a user uses to interact with a system.
[0009] An "artificial intelligence model" refers to a program that uses machine learning algorithms to generate and analyze interactions with users.
[0010] A "client terminal" refers to an electronic device used by a user to access and interact with a system.
[0011] A "question" refers to an inquiry presented to the user during a conversation.
[0012] "Topic" refers to the theme or subject of a conversation.
[0013] "Response" refers to the answer that a user enters in response to a question or topic presented to them.
[0014] "Feedback" refers to the evaluation and advice provided to the user after the conversation ends.
[0015] "Means of selection" refers to the mechanism used by the user to select a specific option (e.g., role or answer).
Brief Description of the Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory where information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides a system for using AI to perform realistic dialogue simulations when users conduct customer service training. Specific embodiments and their processes are described below.
[0038] System Configuration
[0039] This system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting AI models, generating dialogues, and evaluation, while the terminal accepts user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[0040] Program processing flow
[0041] User role settings
[0042] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[0043] AI Model Selection
[0044] The server selects the appropriate AI model based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer" model. Conversely, if the user selects "Customer," the server will select the "Crew" model.
[0045] Start of dialogue
[0046] The server prompts the selected AI model to generate initial questions and topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[0047] Continuing the dialogue
[0048] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an AI model to generate the next question or comment. For example, if the user responds, "Hello! Our recommendation is the cappuccino. Let me show you around," the AI (the customer model) will generate the next question, "Do you have any seats available?" This process is repeated until the dialogue reaches its goal.
[0049] Evaluation and Feedback
[0050] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, and accuracy. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[0051] Specific example
[0052] User acting as crew member and AI acting as customer.
[0053] Let's say a user uses the system to improve their customer service skills.
[0054] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[0055] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[0056] 3. The server selects a "customer role model" based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[0057] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0058] 5. The server receives the response and generates the next question, "Are there any seats available?", and the conversation continues.
[0059] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[0060] Thus, by using the present invention, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[0061] The following describes the processing flow.
[0062] Step 1:
[0063] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[0064] Step 2:
[0065] The terminal displays the generated role selection screen to the user.
[0066] Step 3:
[0067] Users can choose to play either the role of a "crew member" or a "customer."
[0068] Step 4:
[0069] The device sends the user's selection results to the server.
[0070] Step 5:
[0071] The server checks the received role selection results and selects an appropriate artificial intelligence model according to the selected role.
[0072] Step 6:
[0073] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[0074] Step 7:
[0075] The server sends the initial generated questions and topics to the terminal.
[0076] Step 8:
[0077] The device displays received questions and topics to the user.
[0078] Step 9:
[0079] The user enters their response to the displayed question.
[0080] Step 10:
[0081] The terminal sends the user's response to the server.
[0082] Step 11:
[0083] The server analyzes the user's response and uses an artificial intelligence model to generate the next question or comment.
[0084] Step 12:
[0085] The server sends the next generated question or comment to the terminal.
[0086] Step 13:
[0087] The device displays the next question or comment received by the user.
[0088] Step 14:
[0089] Repeat the process from Step 9 to Step 13 until the dialogue reaches its goal.
[0090] Step 15:
[0091] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[0092] Step 16:
[0093] The server generates feedback based on the evaluation results.
[0094] Step 17:
[0095] The server sends the generated feedback to the terminal.
[0096] Step 18:
[0097] The device displays feedback to the user.
[0098] (Example 1)
[0099] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] In today's service industry, customer service skills are a crucial element. However, traditional customer service training is time-consuming and expensive, limiting opportunities for implementation. Furthermore, it is difficult to accurately replicate real-world situations, resulting in insufficient skill improvement among employees. Therefore, there is a need for a realistic dialogue simulation system that enables employees to efficiently and effectively improve their customer service skills.
[0101] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0102] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the terminal according to the selected role; means for displaying the generated questions and topics on the terminal; means for receiving responses from the user and using a natural language processing algorithm to generate subsequent questions and comments; and means for evaluating the user's responses when the dialogue ends and generating feedback based on the evaluation. This allows the user to have a realistic dialogue simulation while saving time and money.
[0103] A "user" refers to a person who operates the system, selects a role, and participates in dialogue simulations.
[0104] A "role" refers to the role that a user chooses in a dialogue simulation, and includes roles such as staff member or visitor.
[0105] "Interface" refers to the screens and functions that users use to operate a system, and specifically includes screens for selecting roles.
[0106] "Terminal" refers to a device operated by a user, and includes personal computers, smartphones, tablets, and other similar devices.
[0107] "Initial questions or topics" refer to the inquiries or topics that the system initially generates when starting a dialogue simulation.
[0108] An "artificial intelligence model" is a model that operates based on machine learning algorithms and is used to generate dialogues.
[0109] A "natural language processing algorithm" refers to a technology that analyzes user input and generates an appropriate response based on that analysis.
[0110] "Evaluation" refers to the process of analyzing the user's responses after the dialogue simulation is completed and measuring performance based on items such as politeness, fluency, and accuracy.
[0111] "Feedback" refers to the improvement suggestions and performance evaluations provided to users based on the evaluation results.
[0112] Modes for carrying out the invention
[0113] This invention is a system that uses artificial intelligence to conduct realistic dialogue simulations when users undergo customer service training. The system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting an AI model, generating dialogues, and evaluation, while the terminal receives user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[0114] When a user accesses the system, the server first generates a role selection screen. This uses web technologies such as HTML, CSS, and JavaScript (registered trademark). The generated role selection screen is sent to the terminal, which displays it on the user's screen. The user selects either "Crew" or "Customer" on the role selection screen. The selection result is sent from the terminal to the server.
[0115] Next, the server selects an appropriate AI model based on the user's chosen role. Specifically, for example, if the user selects "Crew," it will select the "Customer model," and conversely, if the user selects "Customer," it will select the "Crew model." This utilizes AI models that have been pre-trained using machine learning frameworks such as TENSORFLOW® or PyTorch.
[0116] The server initializes the selected AI model and generates the first question or topic. For example, the "customer model" would generate the question, "Hello, what coffee do you recommend here?" This information is sent to the device in JSON format, and the device displays it to the user. The user enters a response to the question, and that response is sent from the device to the server.
[0117] The server analyzes the user's responses and uses natural language processing algorithms to generate subsequent questions and comments. This process is repeated until the dialogue ends. After the dialogue ends, the server analyzes the user's entire dialogue log and generates feedback based on evaluation criteria such as politeness, fluency, and accuracy. The feedback is sent to the terminal in JSON format and displayed to the user.
[0118] Specific example
[0119] Specifically, the system operates in the following steps:
[0120] 1. When a user accesses the system, the server generates a role selection screen and sends it to the terminal.
[0121] 2. The terminal displays a role selection screen to the user.
[0122] 3. When a user selects a "crew member" role, the selection is sent to the server.
[0123] 4. The server selects the "customer role model" and generates the first question, "Hello, what coffee do you recommend here?" and sends it to the terminal.
[0124] 5. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0125] 6. The server receives the response, generates the next question "Are there any seats available?", and sends it back to the terminal.
[0126] In this way, users can efficiently and effectively improve their customer service skills through realistic dialogue simulations.
[0127] Example of a prompt
[0128] Examples of specific prompt messages include the following:
[0129] Prompt for users who chose the crew role: "As a crew member, please recommend a coffee to a customer who has come into the store."
[0130] Prompt for users who chose to play the customer role: "As a customer, act out a scenario where you visit a cafe and ask the staff for their coffee recommendation."
[0131] This system allows users to conduct realistic customer service training while saving time and money.
[0132] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0133] System processing flow
[0134] Step 1: Generate and display the role selection screen.
[0135] The server generates a role selection screen using HTML, CSS, and JavaScript when a user accesses the system.
[0136] Output: Data from the role selection screen (HTML, CSS, JavaScript) sent from the server to the terminal.
[0137] The terminal displays the received role selection screen data on the user's display.
[0138] • Users can choose to play either the role of a "crew member" or a "customer."
[0139] • Input: Select user role
[0140] Output: POST request from terminal to server containing the user's role selection results.
[0141] Step 2: Select and initialize the AI model
[0142] The server receives the user's selection results, parses them in JSON format, and selects the appropriate artificial intelligence model based on the selected role.
[0143] • Input: User role selection results (JSON format)
[0144] Output: Selected AI model
[0145] For example, if the user selects the "Crew Role," select the "Customer Role Model."
[0146] The server initializes the selected AI model and prepares it for the start of the interaction.
[0147] • Specific example: Loading a pre-trained model from TensorFlow or PyTorch.
[0148] Step 3: Generate and display the first question.
[0149] The server generates initial questions and topics. For example, for a selected "customer role model," it generates the initial question, "Hi, what coffee do you recommend here?"
[0150] • Input: Selected AI model
[0151] Output: Initial Question
[0152] The server sends the generated initial questions to the terminal in JSON format.
[0153] The terminal displays the initial question received to the user.
[0154] Step 4: Receiving and analyzing user responses
[0155] • Users enter appropriate responses to the displayed questions. For example, they might answer, "Our recommendation is the cappuccino."
[0156] • Input: User response
[0157] Output: User response data (JSON format) sent from the terminal to the server.
[0158] The server analyzes the user's response and uses natural language processing algorithms to generate the next question or comment.
[0159] • Input: User response data
[0160] Output: Next question or comment
[0161] Step 5: Continue the dialogue
[0162] The server sends the generated next question or comment to the terminal in JSON format.
[0163] The device displays received questions or comments to the user.
[0164] This process is repeated until the dialogue is finished.
[0165] • Input: User responses and generated questions or comments
[0166] • Output: Generate and display the next question or comment.
[0167] Step 6: Evaluate the dialogue and generate feedback
[0168] • Once the conversation ends, the server analyzes the entire conversation log and evaluates the user's performance based on evaluation criteria such as politeness, fluency, and accuracy.
[0169] • Input: Overall dialogue log
[0170] • Output: Evaluation results
[0171] The server generates feedback based on the evaluation results and sends it to the terminal in JSON format.
[0172] The device displays the generated feedback to the user, allowing them to self-evaluate and understand areas for improvement next time.
[0173] • Specific example: The evaluation is displayed in a score format, such as "Courtesy: 80 points, Fluency: 70 points, Accuracy: 90 points."
[0174] This allows users to improve their customer service skills through realistic dialogue simulations.
[0175] (Application Example 1)
[0176] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0177] Traditional customer service training methods have been difficult to implement effectively due to limited opportunities for direct interaction with real customers. Furthermore, the consistency and timeliness of feedback were also problematic, hindering efficient training. In addition, in real-world store environments, securing training time during busy periods was difficult, leading to delays in skill development.
[0178] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0179] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on a client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using the artificial intelligence model to generate subsequent questions and comments; means for evaluating the user's responses and generating feedback when the interaction ends; and means for the user to select a role as a crew member or a customer, obtain an artificial intelligence model corresponding to the selected role on the server, and run it on a smart device. This enables the user to conduct customer service training in a realistic manner and efficiently improve their skills regardless of time or location.
[0180] A "role" refers to the role a user chooses within a system.
[0181] "Interface" refers to the screens and input methods that users use to interact with a system.
[0182] A "client terminal" refers to a smart device operated by a user.
[0183] An "artificial intelligence model" refers to a machine learning algorithm used to generate and analyze dialogue through natural language processing.
[0184] A "server" refers to a remote computer that is responsible for the main processing of a system.
[0185] A "question or topic" refers to the starting point of a dialogue that is generated by the system and presented to the user.
[0186] "Response" refers to the reply that a user enters into the system.
[0187] "Feedback" refers to the improvement suggestions and evaluation information provided after evaluating the user's interaction results.
[0188] "Crew role" refers to a user who takes on the role of providing customer service during customer service training.
[0189] "Customer role" refers to the user who takes on the role of the person being served in customer service training.
[0190] "Smart devices" refer to mobile information terminals connected to the internet, such as smartphones and tablet devices.
[0191] "Retrieving on a server" refers to downloading specific information or models from a remote computer and using them.
[0192] "Execution" refers to a system carrying out a specific process.
[0193] The system in this invention provides the necessary functions for users to conduct customer service training. The system mainly consists of a server, a client terminal (smart device), and the user. The server is responsible for major data processing and execution of AI models, while the client terminal functions as the user interface.
[0194] Hardware and software to be used
[0195] hardware
[0196] Smart devices (smartphones and tablet devices)
[0197] Server (using a cloud server, e.g., AWS®)
[0198] software
[0199] Mobile applications (developed in Swift for iOS and Kotlin for Android®)
[0200] Server-side platform (combination of Node.js and Python)
[0201] Artificial intelligence model (using GPT-4® registered trademark)
[0202] System program processing flow
[0203] User Interface
[0204] The terminal generates and displays an interface for the user to select a role (crew member or customer). This allows the user to choose which role to simulate.
[0205] Server-side processing
[0206] When a user selects a role, that information is sent from the terminal to the server. The server then selects the corresponding artificial intelligence model based on the selected role. For example, if the user chooses the "crew" role, the server will select the "customer" model.
[0207] Initial question generation and display
[0208] The server uses the selected AI model to generate initial questions and topics. These questions are sent to the client terminal, which then displays them to the user. For example, the question "Hi, what coffee do you recommend here?" might be displayed.
[0209] Continuing the dialogue
[0210] When the user responds to a question, the device sends that response to the server. The server uses an AI model to generate the next question or comment and sends it back to the device. The system repeats this process, continuing the dialogue.
[0211] Evaluation and Feedback
[0212] Once the conversation ends, the server evaluates the entire interaction and generates feedback. This feedback evaluates the user's courtesy, fluency, and accuracy, among other things. This feedback is sent to the terminal and displayed to the user.
[0213] Adding specific examples
[0214] When a user begins customer service training using their smartphone, the following steps are taken:
[0215] 1. The user launches the app and logs in.
[0216] 2. On the role selection screen, select either "Crew Member" or "Customer."
[0217] 3. The server selects either a "customer model" or a "crew model" and generates an initial question (e.g., "Hello, what coffee do you recommend here?").
[0218] 4. When the user responds, the server generates the next question (e.g., User: "Our recommendation is the cappuccino" → Server: "Do you have any seats available?").
[0219] 5. After the interaction ends, the server evaluates the user's response and generates and returns feedback.
[0220] Example of a prompt
[0221] You will act as an AI for customer service training. When a user responds as a crew member, generate the following questions and comments as needed. For example, if the user answers "Our recommendation is the cappuccino," then ask "Do you have any seats available?"
[0222] In this way, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[0223] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0224] Step 1:
[0225] User login
[0226] The user launches the app on their smartphone and enters their email address and password on the login screen. The device sends the entered authentication information to the server, which then verifies it against its database. If authentication is successful, the server returns an authentication success message to the device.
[0227] Enter: Email address, password
[0228] Output: Authentication success message
[0229] Specific operation: The terminal sends the entered email address and password to the server as an HTTP request, and the server checks the database and returns the authentication result.
[0230] Step 2:
[0231] Role Selection
[0232] Once the user successfully logs in, the terminal displays a role selection screen. The user chooses either "Crew Member" or "Customer." The terminal then sends the selected role information to the server.
[0233] Input: Role selection (Crew member or customer)
[0234] Output: Role information
[0235] Specific operation: The terminal sends the user's role selection information to the server as an HTTP request. The server receives the role information and starts the process of selecting the appropriate AI model.
[0236] Step 3:
[0237] Selection of an artificial intelligence model
[0238] The server selects the corresponding artificial intelligence model based on the role information chosen by the user. If the user chooses the "crew" role, the server selects the "customer" model; otherwise, it selects the "crew" model.
[0239] Input: Role information
[0240] Output: Corresponding AI model
[0241] Specific operation: The server selects the appropriate AI model (e.g., a GPT-4 based model) based on its internal logic and loads the model.
[0242] Step 4:
[0243] Generating initial questions
[0244] The server uses the selected AI model to generate initial questions and topics. The generated questions are then sent to the terminal.
[0245] Input: AI model, initial prompt
[0246] Output: Initial Question
[0247] Specific operation: The server inputs a prompt message into the AI model and sends the initial question generated by the AI to the terminal (e.g., "Hello, what coffee do you recommend here?").
[0248] Step 5:
[0249] Receiving and analyzing user responses
[0250] The terminal receives the user's response and sends it to the server. The server analyzes the user's response and generates the next question or comment.
[0251] Input: User response
[0252] Output: The following questions and comments
[0253] Specific operation: The terminal sends the user response as text data to the server, and the server uses an AI model to generate the next question (e.g., User: "Our recommendation is the cappuccino" → AI: "Do you have any seats available?").
[0254] Step 6:
[0255] Continuing the dialogue
[0256] The server sends the next question or comment to the terminal, which then displays it to the user. This process repeats until the interaction is complete.
[0257] Input: Next question or comment
[0258] Output: User response
[0259] Specific operation: The terminal displays the next question received from the server to the user, and the user responds and sends the response back to the server.
[0260] Step 7:
[0261] End of dialogue and evaluation
[0262] When the conversation ends, the server analyzes the entire conversation and evaluates the user's performance. This evaluation includes aspects such as politeness, fluency, and accuracy. Based on the evaluation results, the server generates feedback and sends it to the terminal.
[0263] Input: Entire dialogue
[0264] Output: Evaluation results, feedback
[0265] Specific operation: The server analyzes the entire conversation using an AI algorithm, calculates an evaluation score based on that analysis, generates a feedback message, and sends it to the terminal.
[0266] In this way, users can receive customer service training in a realistic setting and efficiently improve their skills.
[0267] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0268] This invention is a system that supports customer service training, providing a realistic customer service experience through user role selection and dialogue simulation. Furthermore, by combining it with an emotion engine that recognizes user emotions, it achieves a more responsive and realistic interaction.
[0269] System Configuration
[0270] This system consists of three main elements: a server, a terminal, and a user. The server integrates an artificial intelligence model and an emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[0271] Program processing flow
[0272] User role settings
[0273] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[0274] AI Model and Emotion Engine Selection
[0275] The server selects the appropriate artificial intelligence model and emotion engine based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer model" and emotion engine. Conversely, if the user selects "Customer," the server will select the "Crew model" and emotion engine.
[0276] Start of dialogue
[0277] The server causes the selected artificial intelligence model to generate initial questions or topics. For example, in the case of the "customer role model", it generates a question such as "Hello, what is the recommended coffee here?". The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[0278] Continuation of the conversation
[0279] The user inputs an appropriate response to the displayed question, and the terminal sends the response to the server. The server analyzes the user's response and uses the emotion engine to recognize the user's emotion. For example, if the user has a worried expression, the emotion engine identifies this as "worried". Then, it uses the artificial intelligence model to generate the next question or comment, and the generated next question or comment is sent to the terminal. The terminal displays it to the user again. This process is repeated until the conversation reaches the goal.
[0280] Evaluation and feedback
[0281] When the conversation ends, the server executes an algorithm to evaluate the entire conversation. This evaluation includes the user's etiquette, fluency, accuracy, and the results of emotion recognition. The server generates feedback based on the evaluation results and sends the feedback content to the terminal. The terminal displays the feedback to the user, and the user can check their performance and understand the areas for improvement next time.
[0282] Specific example
[0283] User in the role of crew and customer service AI
[0284] Suppose a user uses the system to improve their customer service skills.
[0285] 1. When the user operates the terminal to access the system, the server generates a role selection screen and sends it to the terminal.
[0286] 2. When the user selects the "crew member" role, the terminal sends that selection to the server.
[0287] 3. Based on the user's role, the server selects the "customer role model" and the emotion engine, and generates the initial question "Hello, what is the recommended coffee here?"
[0288] 4. The terminal displays this question to the user, and the user responds with "Our recommendation is a cappuccino."
[0289] 5. The server receives the response, uses the emotion engine to recognize the user's emotion (e.g., detects a smile and identifies it as "happiness"), and then generates the next question "Is there an empty seat?" and the conversation continues.
[0290] 6. After the conversation ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user through the terminal.
[0291] In this way, by using the present invention, the user can conduct customer service training in a realistic manner and efficiently improve their skills. By introducing the emotion engine, it becomes possible to respond according to the user's emotion, and practical training can be carried out.
[0292] The following explains the processing flow.
[0293] Step 1:
[0294] When the user accesses, the server generates a role selection screen and sends it to the terminal.
[0295] Step 2:
[0296] The terminal displays the generated role selection screen to the user.
[0297] <00 The user selects either the "crew member role" or the "customer role".
[0299] Step 4:
[0300] The terminal sends the user's selection result to the server.
[0301] Step 5:
[0302] The server checks the received selection result and selects an artificial intelligence model and an emotion engine according to the selected role.
[0303] Step 6:
[0304] The server instructs the selected artificial intelligence model to generate initial questions or topics.
[0305] Step 7:
[0306] The server sends the generated initial questions or topics to the terminal.
[0307] Step 8:
[0308] The terminal displays the received questions or topics to the user.
[0309] Step 9:
[0310] The user inputs a response to the displayed question.
[0311] Step 10:
[0312] The terminal records the user's response data and emotion expression data and sends it to the server.
[0313] Step 11:
[0314] The server analyzes the received user response data and uses the emotion engine to recognize the user's emotion.
[0315] Step 12:
[0316] Based on the recognized sentiment information, the server uses an artificial intelligence model to generate the next questions and comments.
[0317] Step 13:
[0318] The server sends the next generated question or comment to the terminal.
[0319] Step 14:
[0320] The device displays the next question or comment received by the user.
[0321] Step 15:
[0322] The device supplies the user's voice data and facial expression data to the emotion engine in real time and records the emotion recognition results.
[0323] Step 16:
[0324] Repeat the process from Step 9 to Step 15 until the dialogue reaches its goal.
[0325] Step 17:
[0326] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[0327] Step 18:
[0328] The server generates detailed feedback based on the evaluation results.
[0329] Step 19:
[0330] The server sends the generated feedback to the terminal.
[0331] Step 20:
[0332] The device displays feedback to the user.
[0333] (Example 2)
[0334] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0335] The present invention aims to provide a system for efficiently conducting customer service training. Specifically, it solves the problem of supplementing aspects that conventional training methods could not fully cover by providing a system that can perform realistic dialogue simulations according to the role selected by the user, recognize the user's emotions, and generate appropriate feedback.
[0336] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a display screen for the user to select a role; means for selecting an artificial intelligence engine that generates initial questions and topics to initiate a dialogue with the client device according to the selected role; means for displaying the generated questions and topics on the client device; means for receiving responses from the user and using an emotion analysis engine to recognize the user's emotions; means for using an artificial intelligence engine that generates the next questions and comments based on the emotion analysis results; and means for analyzing and evaluating the user's responses and generating feedback when the dialogue ends. This makes it possible for the user to receive practical training through realistic dialogue simulations.
[0337] A "display screen" is a screen that provides an interface for users to select roles.
[0338] An "artificial intelligence engine" is an engine that has algorithms for generating initial questions and topics based on the user's choices.
[0339] A "client device" is a terminal operated by a user, which communicates with a server to display information and receive input.
[0340] "Initial questions or topics to start a conversation" refer to questions or topics used in the introductory part of a conversation, designed to allow the user to smoothly begin the customer service simulation.
[0341] An "emotion analysis engine" is an engine that analyzes and identifies emotions from user responses and facial expression data.
[0342] "Means for evaluation and generating feedback" refers to a means of comprehensively evaluating the user's performance at the end of an interaction and informing the user of areas for improvement and weaknesses based on the results.
[0343] "Means for generating initial questions and topics" refer to algorithms or programs used to automatically generate appropriate questions and topics at the start of a dialogue.
[0344] This invention relates to a customer service training system and aims to provide realistic dialogue simulations based on roles selected by the user. The system consists of three main elements: a server, a terminal, and the user.
[0345] The server primarily integrates an artificial intelligence engine (generative AI model) and an emotion analysis engine, and is responsible for generating, analyzing, and evaluating dialogue. Specific software used includes large-scale language models such as GPT-3 (registered trademark) and emotion recognition tools such as the Affectiva SDK. The terminal is a device operated by the user, receiving user selections and inputs, and displaying information received from the server. Users interact with the system as either a "crew member" or a "customer," and receive training based on the content of those interactions.
[0346] 1. User role settings
[0347] When a user accesses the system from their device, the server generates a role selection screen and sends it to the device. The device displays the role selection screen using HTML and JavaScript. The user selects either "Crew" or "Customer" and sends the selection result from the device to the server. The server analyzes this selection result and selects the appropriate dialogue model and sentiment analysis engine.
[0348] 2. Selection of AI Model and Emotion Engine
[0349] The server selects and initializes the appropriate artificial intelligence engine (generative AI model) and sentiment analysis engine based on the user's selection. For example, if the user selects "Crew Role," the server initializes the "Customer Role Model" and sentiment analysis engine.
[0350] 3. Starting the dialogue
[0351] The server inputs prompt text into the selected artificial intelligence engine, generating an initial question or topic. For example, it might generate the question, "Hello, what coffee do you recommend here?" The generated question is sent from the server to the terminal and displayed to the user on the terminal.
[0352] 4. Continue the dialogue
[0353] The user enters a response to a displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion analysis engine to recognize the user's emotion. For example, if the user displays "smile," the emotion analysis engine identifies this as "joy." The server then generates the next question or comment and sends it back to the device. This process is repeated until the dialogue reaches its goal.
[0354] 5. Evaluation and Feedback
[0355] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, accuracy, and the outcome of sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[0356] Specific example
[0357] When a user uses the system to improve their customer service skills, the system operates as follows:
[0358] 1. The user accesses the system by operating a terminal. The server generates a role selection screen and sends it to the terminal.
[0359] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[0360] 3. The server selects a "customer role model" and an emotion analysis engine, and generates the initial question, "Hi, what coffee do you recommend here?"
[0361] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0362] 5. The server receives the response, performs emotion recognition, and for example, detects a smile and identifies it as "joy." Based on this, it generates the next question, "Are there any seats available?", and the dialogue continues.
[0363] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. The feedback is sent to the terminal and displayed to the user.
[0364] Example prompt: "Generate an initial question for when the user chooses to play the role of a crew member. For example, 'Hi, what coffee do you recommend here?'"
[0365] Thus, by utilizing this invention, users can conduct customer service training in a realistic setting and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for more practical training.
[0366] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0367] Step 1: Generate the user role selection screen.
[0368] Input: The user accesses the system via the terminal's web browser.
[0369] Processing: The server uses HTML and JavaScript to generate a role selection screen, which includes the options "Crew Role" and "Customer Role".
[0370] Output: The server sends the generated role selection screen to the terminal as an HTTP response.
[0371] Specific operation: When a user accesses a URL, the server sends a response to the client containing the appropriate HTML and JavaScript code, which the terminal then renders and displays to the user.
[0372] Step 2: Submit the role selection results
[0373] Input: The user selects either "Crew Role" or "Customer Role" on the role selection screen.
[0374] Processing: The terminal retrieves the user's selection results using JavaScript and sends the selection data to the server as a JSON payload.
[0375] Output: The server receives the selected data.
[0376] Specific operation: When the user clicks a button, a JavaScript event handler retrieves the selection result and sends an AJAX request to the server.
[0377] Step 3: Select and initialize the AI model and emotion engine.
[0378] Input: The server receives the user's role selection results.
[0379] Processing: Based on the selected role, the server selects an appropriate AI generative model (e.g., GPT-3) and sentiment analysis engine (e.g., Affectiva SDK), and initializes them.
[0380] Output: The server generates an instance of the pre-configured AI model and emotion engine.
[0381] Specific operation: The server analyzes the received selection data, and if "Crew Role" is selected, it initializes the "Customer Role Model" and emotion engine, and loads the respective libraries and APIs.
[0382] Step 4: Generate and submit initial questions
[0383] Input: An AI model initialized on the server.
[0384] Processing: The server inputs a prompt sentence into the AI model for initial question generation. For example, it generates the question, "Hello, what coffee do you recommend here?"
[0385] Output: The server retrieves the generated question in text format and sends it to the terminal.
[0386] Specific operation: The server provides prompt text to the AI model, retrieves the generated question, and sends it to the client as an HTTP response.
[0387] Step 5: Display the initial question
[0388] Input: The initial question received by the user's device from the server.
[0389] Processing: The terminal displays the initial question it received on the screen.
[0390] Output: The screen displaying the initial question.
[0391] Specific operation: The device's browser receives the server response and displays the question text to the user as an HTML element.
[0392] Step 6: Sending the User Response
[0393] Input: The user enters their response to the initial question displayed on the screen.
[0394] Processing: The terminal receives the input response and sends it to the server as a JSON payload.
[0395] Output: The server receives the user's response data.
[0396] Specific operation: When the user enters a response in the text box and presses the submit button, the device sends an AJAX request to the server.
[0397] Step 7: Emotion recognition and next question generation
[0398] Input: Response data from users who reached the server.
[0399] Processing: The server uses an emotion analysis engine to recognize the user's emotions and, based on the results, has the AI model generate the next question or comment.
[0400] Output: Sentiment recognition results and generated questions or comments.
[0401] Specific operation: The server analyzes the response text and identifies emotions such as "joy" using emotion data, for example, "smile." Then, it generates the next question and sends it to the terminal.
[0402] Step 8: Continuing the Dialogue
[0403] Input: The following questions and comments generated from the server.
[0404] Processing: The terminal receives the next question or comment and displays it on the screen.
[0405] Output: The screen showing the following questions and comments.
[0406] Specific operation: The device's browser displays the new question text to the user again as an HTML element. This process is repeated until the interaction reaches its goal.
[0407] Step 9: Evaluation and feedback on the entire dialogue
[0408] Input: All dialogue data stored on the server when the dialogue ends.
[0409] Processing: The server runs an algorithm that evaluates the entire interaction, comprehensively assessing the user's politeness, fluency, accuracy, and sentiment recognition.
[0410] Output: Evaluation results and feedback messages.
[0411] Specific operation: The server analyzes the dialogue data, evaluates the user's performance using an evaluation algorithm, generates feedback including what to do next and areas for improvement, and sends it to the terminal.
[0412] Step 10: Displaying Feedback
[0413] Input: Feedback message received from the server.
[0414] Processing: The device displays the received feedback message to the user.
[0415] Output: Screen showing feedback.
[0416] Specific operation: The device's browser displays feedback text to the user as an HTML element, allowing the user to check their performance.
[0417] This detailed, step-by-step flow allows users to effectively conduct realistic and responsive customer service training. The introduction of an emotion analysis engine enables appropriate responses based on the user's emotions, resulting in more practical training.
[0418] (Application Example 2)
[0419] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0420] In real-world customer service training, it is crucial for trainees to practice in truly realistic scenarios and receive feedback based on their own emotional responses. However, conventional systems have limitations in training effectiveness due to insufficient emotional recognition and interactions that are not always real-time and responsive. Furthermore, for users to improve their conversational skills as store staff, training in an environment closer to actual customer service situations is required.
[0421] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0422] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using an artificial intelligence model to generate subsequent questions and comments using an emotion engine that recognizes emotions based on the generated responses; and means for evaluating the user's responses and generating feedback when the interaction ends. This enables the user to receive realistic customer service training, is provided with emotion-based interaction and feedback, and enables practical skill improvement.
[0423] A "role" refers to the specific actions or functions that a user is responsible for within a system.
[0424] "Means for generating an interface" refers to means that have the function of creating screens and operating mechanisms for users to access and operate a system.
[0425] An "artificial intelligence model for generating initial questions and topics" is an artificial intelligence model that generates questions and topics presented at the start of a conversation, based on the role selected by the user.
[0426] A "client terminal" refers to a computer or digital device that a user directly operates.
[0427] "Means for displaying questions and topics" refers to means that have the functionality to display generated questions and topics on a client terminal.
[0428] An "emotion-recognizing emotion engine" is an engine that analyzes a user's emotions from their facial expressions and voice, and generates an appropriate response based on the analysis results.
[0429] An "artificial intelligence model for generating the next question or comment" is an artificial intelligence-powered model that automatically generates the next question or comment based on the user's response and the results of sentiment analysis.
[0430] "Means for evaluating user responses and generating feedback" refers to methods for analyzing user interactions, evaluating their performance, and providing feedback on areas for improvement and positive aspects.
[0431] This system was built to help users improve their customer service skills in physical stores. The embodiments for carrying out the invention are shown below.
[0432] System Configuration
[0433] This system consists of three main elements: a server, a terminal (such as a smartphone or tablet), and the user. The server integrates an artificial intelligence model and emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[0434] Hardware and software to use
[0435] Server: A high-performance computer system used to run AI models and emotion engines.
[0436] Terminal: A user device such as a smartphone or tablet, which is operated by the user through an interface.
[0437] Software: AI models written in Python, an emotion engine for emotion recognition, and an application for the user interface.
[0438] Data processing and data calculation
[0439] Role Selection: When a user operates a terminal to select either the role of crew member or customer, that information is sent to the server.
[0440] Initial Question Generation: Based on the user's selection, the server selects an appropriate AI model (generative AI model) and generates initial questions and topics.
[0441] Display and Response: Generated questions and topics are displayed on the terminal, and responses from the user are retrieved. These responses are then sent to the server.
[0442] Emotion Recognition: The server processes user responses using an emotion engine and analyzes the emotion data. Based on this analysis, the next questions and comments are generated.
[0443] Evaluation and Feedback: After the interaction ends, the server comprehensively evaluates the user's responses and generates performance feedback. This feedback is provided to the user via the terminal, suggesting areas for improvement in training.
[0444] Specific Scenario Examples
[0445] When new staff members at a cafe undergo customer service training, the following steps are taken:
[0446] 1. The user selects a crew role and accesses the system.
[0447] 2. The server selects the "customer role model" and generates the initial question, "Hello, what coffee do you recommend here?"
[0448] 3. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0449] 4. The server analyzes this response using an emotion engine and recognizes the user's emotion (e.g., "joy").
[0450] 5. Generate the following question or comment. For example, "Are there any seats available?"
[0451] 6. When the interaction ends, the server evaluates the entire interaction and generates feedback. This feedback is displayed to the user through the terminal.
[0452] Example of a prompt
[0453] The following are examples of prompt statements that can actually be used.
[0454] "Please generate questions that can be used for customer service training at a cafe. Please use native-level language, and keep them simple and polite."
[0455] Thus, by using the system of the present invention, users can efficiently improve their skills by conducting customer service training in a realistic manner and receiving feedback that takes emotions into account.
[0456] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0457] Step 1:
[0458] The user operates the terminal to access the system, and the role selection screen is displayed.
[0459] Specifically, the terminal sends a request to the server, asking for a role selection screen. Upon receiving this request, the server generates an interface and sends it to the terminal. The user selects either the "Crew Role" or the "Customer Role," and this selection is sent to the server via the terminal.
[0460] Input: User role selection request
[0461] Output: Role selection screen
[0462] Step 2:
[0463] The server selects an appropriate artificial intelligence model based on the user's choices and generates initial questions and topics.
[0464] Specifically, the server selects either a "crew role model" or a "customer role model" and an emotion engine based on the role selected by the user. Next, it uses a generative AI model to generate initial questions and topics based on the prompt text.
[0465] Input: Select user role
[0466] Output: Initial questions and topics
[0467] Step 3:
[0468] The terminal displays initial questions and topics received from the server to the user.
[0469] In terms of specific operations, the server sends the initial generated questions and topics to the terminal, which then displays them to the user. The user then responds to the displayed questions.
[0470] Input: Initial questions or topics
[0471] Output: User response
[0472] Step 4:
[0473] The server receives responses from users and uses an emotion engine to recognize those emotions.
[0474] In terms of specific operations, the terminal sends the user's response to the server, where an emotion engine analyzes the response and recognizes the emotion. For example, it performs facial expression analysis and voice analysis to determine the user's emotional state.
[0475] Input: User response
[0476] Output: Emotion recognition result
[0477] Step 5:
[0478] Based on the results of the sentiment engine, the server generates the following questions and comments.
[0479] In terms of specific operation, the server uses emotion recognition data as input and a generative AI model to generate the next question or comment. This allows the dialogue to continue.
[0480] Input: Sentiment recognition result
[0481] Output: The following questions and comments
[0482] Step 6:
[0483] The terminal displays the next question or comment received from the server to the user.
[0484] In terms of specific actions, this process is carried out similarly to step 3, with the terminal receiving data from the server and displaying it to the user. The user then responds again.
[0485] Input: Next question or comment
[0486] Output: User response
[0487] Step 7:
[0488] Once the interaction is complete, the server evaluates the entire interaction and generates feedback.
[0489] Specifically, the server analyzes the entire dialogue script and evaluates it based on the user's politeness, fluency, accuracy, and emotion recognition results. Based on this evaluation, it generates feedback and sends it to the terminal.
[0490] Input: Content of the entire dialogue
[0491] Output: User feedback
[0492] Step 8:
[0493] The terminal displays feedback from the server to the user.
[0494] In terms of specific actions, the device displays feedback received from the server to the user, allowing the user to check their performance and understand areas for improvement for the next training session.
[0495] Input: Feedback
[0496] Output: Feedback display
[0497] Through this series of processing steps, users can conduct customer service training in a realistic setting and receive emotion-based interaction and feedback.
[0498] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0499] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0500] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0501] [Second Embodiment]
[0502] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0503] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0504] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0505] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0506] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0507] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0508] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0509] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0510] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0511] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0512] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0513] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0514] This invention provides a system for using AI to perform realistic dialogue simulations when users conduct customer service training. Specific embodiments and their processes are described below.
[0515] System Configuration
[0516] This system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting AI models, generating dialogues, and evaluation, while the terminal accepts user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[0517] Program processing flow
[0518] User role settings
[0519] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[0520] AI Model Selection
[0521] The server selects the appropriate AI model based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer" model. Conversely, if the user selects "Customer," the server will select the "Crew" model.
[0522] Start of dialogue
[0523] The server prompts the selected AI model to generate initial questions and topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[0524] Continuing the dialogue
[0525] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an AI model to generate the next question or comment. For example, if the user responds, "Hello! Our recommendation is the cappuccino. Let me show you around," the AI (the customer model) will generate the next question, "Do you have any seats available?" This process is repeated until the dialogue reaches its goal.
[0526] Evaluation and Feedback
[0527] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, and accuracy. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[0528] Specific example
[0529] User acting as crew member and AI acting as customer.
[0530] Let's say a user uses the system to improve their customer service skills.
[0531] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[0532] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[0533] 3. The server selects a "customer role model" based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[0534] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0535] 5. The server receives the response and generates the next question, "Are there any seats available?", and the conversation continues.
[0536] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[0537] Thus, by using the present invention, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[0538] The following describes the processing flow.
[0539] Step 1:
[0540] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[0541] Step 2:
[0542] The terminal displays the generated role selection screen to the user.
[0543] Step 3:
[0544] Users can choose to play either the role of a "crew member" or a "customer."
[0545] Step 4:
[0546] The device sends the user's selection results to the server.
[0547] Step 5:
[0548] The server checks the received role selection results and selects an appropriate artificial intelligence model according to the selected role.
[0549] Step 6:
[0550] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[0551] Step 7:
[0552] The server sends the initial generated questions and topics to the terminal.
[0553] Step 8:
[0554] The device displays received questions and topics to the user.
[0555] Step 9:
[0556] The user enters their response to the displayed question.
[0557] Step 10:
[0558] The terminal sends the user's response to the server.
[0559] Step 11:
[0560] The server analyzes the user's response and uses an artificial intelligence model to generate the next question or comment.
[0561] Step 12:
[0562] The server sends the next generated question or comment to the terminal.
[0563] Step 13:
[0564] The device displays the next question or comment received by the user.
[0565] Step 14:
[0566] Repeat the process from Step 9 to Step 13 until the dialogue reaches its goal.
[0567] Step 15:
[0568] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[0569] Step 16:
[0570] The server generates feedback based on the evaluation results.
[0571] Step 17:
[0572] The server sends the generated feedback to the terminal.
[0573] Step 18:
[0574] The device displays feedback to the user.
[0575] (Example 1)
[0576] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0577] In today's service industry, customer service skills are a crucial element. However, traditional customer service training is time-consuming and expensive, limiting opportunities for implementation. Furthermore, it is difficult to accurately replicate real-world situations, resulting in insufficient skill improvement among employees. Therefore, there is a need for a realistic dialogue simulation system that enables employees to efficiently and effectively improve their customer service skills.
[0578] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0579] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the terminal according to the selected role; means for displaying the generated questions and topics on the terminal; means for receiving responses from the user and using a natural language processing algorithm to generate subsequent questions and comments; and means for evaluating the user's responses when the dialogue ends and generating feedback based on the evaluation. This allows the user to have a realistic dialogue simulation while saving time and money.
[0580] A "user" refers to a person who operates the system, selects a role, and participates in dialogue simulations.
[0581] A "role" refers to the role that a user chooses in a dialogue simulation, and includes roles such as staff member or visitor.
[0582] "Interface" refers to the screens and functions that users use to operate a system, and specifically includes screens for selecting roles.
[0583] "Terminal" refers to a device operated by a user, and includes personal computers, smartphones, tablets, and other similar devices.
[0584] "Initial questions or topics" refer to the inquiries or topics that the system initially generates when starting a dialogue simulation.
[0585] An "artificial intelligence model" is a model that operates based on machine learning algorithms and is used to generate dialogues.
[0586] A "natural language processing algorithm" refers to a technology that analyzes user input and generates an appropriate response based on that analysis.
[0587] "Evaluation" refers to the process of analyzing the user's responses after the dialogue simulation is completed and measuring performance based on items such as politeness, fluency, and accuracy.
[0588] "Feedback" refers to the improvement suggestions and performance evaluations provided to users based on the evaluation results.
[0589] Modes for carrying out the invention
[0590] This invention is a system that uses artificial intelligence to conduct realistic dialogue simulations when users undergo customer service training. The system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting an AI model, generating dialogues, and evaluation, while the terminal receives user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[0591] When a user accesses the system, the server first generates a role selection screen. This uses web technologies such as HTML, CSS, and JavaScript. The generated role selection screen is sent to the terminal, which displays it on the user's screen. The user selects either "Crew" or "Customer" on the role selection screen. The selection result is sent from the terminal to the server.
[0592] Next, the server selects an appropriate AI model based on the user's chosen role. Specifically, for example, if the user selects "Crew," it will select the "Customer model," and conversely, if the user selects "Customer," it will select the "Crew model." This utilizes AI models that have been pre-trained using machine learning frameworks such as TensorFlow or PyTorch.
[0593] The server initializes the selected AI model and generates the first question or topic. For example, the "customer model" would generate the question, "Hello, what coffee do you recommend here?" This information is sent to the device in JSON format, and the device displays it to the user. The user enters a response to the question, and that response is sent from the device to the server.
[0594] The server analyzes the user's responses and uses natural language processing algorithms to generate subsequent questions and comments. This process is repeated until the dialogue ends. After the dialogue ends, the server analyzes the user's entire dialogue log and generates feedback based on evaluation criteria such as politeness, fluency, and accuracy. The feedback is sent to the terminal in JSON format and displayed to the user.
[0595] Specific example
[0596] Specifically, the system operates in the following steps:
[0597] 1. When a user accesses the system, the server generates a role selection screen and sends it to the terminal.
[0598] 2. The terminal displays a role selection screen to the user.
[0599] 3. When a user selects a "crew member" role, the selection is sent to the server.
[0600] 4. The server selects the "customer role model" and generates the first question, "Hello, what coffee do you recommend here?" and sends it to the terminal.
[0601] 5. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0602] 6. The server receives the response, generates the next question "Are there any seats available?", and sends it back to the terminal.
[0603] In this way, users can efficiently and effectively improve their customer service skills through realistic dialogue simulations.
[0604] Example of a prompt
[0605] Examples of specific prompt messages include the following:
[0606] Prompt for users who chose the crew role: "As a crew member, please recommend a coffee to a customer who has come into the store."
[0607] Prompt for users who chose to play the customer role: "As a customer, act out a scenario where you visit a cafe and ask the staff for their coffee recommendation."
[0608] This system allows users to conduct realistic customer service training while saving time and money.
[0609] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0610] System processing flow
[0611] Step 1: Generate and display the role selection screen.
[0612] The server generates a role selection screen using HTML, CSS, and JavaScript when a user accesses the system.
[0613] Output: Data from the role selection screen (HTML, CSS, JavaScript) sent from the server to the terminal.
[0614] The terminal displays the received role selection screen data on the user's display.
[0615] • Users can choose to play either the role of a "crew member" or a "customer."
[0616] • Input: Select user role
[0617] Output: POST request from terminal to server containing the user's role selection results.
[0618] Step 2: Select and initialize the AI model
[0619] The server receives the user's selection results, parses them in JSON format, and selects the appropriate artificial intelligence model based on the selected role.
[0620] • Input: User role selection results (JSON format)
[0621] Output: Selected AI model
[0622] For example, if the user selects the "Crew Role," select the "Customer Role Model."
[0623] The server initializes the selected AI model and prepares it for the start of the interaction.
[0624] • Specific example: Loading a pre-trained model from TensorFlow or PyTorch.
[0625] Step 3: Generate and display the first question.
[0626] The server generates initial questions and topics. For example, for a selected "customer role model," it generates the initial question, "Hi, what coffee do you recommend here?"
[0627] • Input: Selected AI model
[0628] Output: Initial Question
[0629] The server sends the generated initial questions to the terminal in JSON format.
[0630] The terminal displays the initial question received to the user.
[0631] Step 4: Receiving and analyzing user responses
[0632] • Users enter appropriate responses to the displayed questions. For example, they might answer, "Our recommendation is the cappuccino."
[0633] • Input: User response
[0634] Output: User response data (JSON format) sent from the terminal to the server.
[0635] The server analyzes the user's response and uses natural language processing algorithms to generate the next question or comment.
[0636] • Input: User response data
[0637] Output: Next question or comment
[0638] Step 5: Continue the dialogue
[0639] The server sends the generated next question or comment to the terminal in JSON format.
[0640] The device displays received questions or comments to the user.
[0641] This process is repeated until the dialogue is finished.
[0642] • Input: User responses and generated questions or comments
[0643] • Output: Generate and display the next question or comment.
[0644] Step 6: Evaluate the dialogue and generate feedback
[0645] • Once the conversation ends, the server analyzes the entire conversation log and evaluates the user's performance based on evaluation criteria such as politeness, fluency, and accuracy.
[0646] • Input: Overall dialogue log
[0647] • Output: Evaluation results
[0648] The server generates feedback based on the evaluation results and sends it to the terminal in JSON format.
[0649] The device displays the generated feedback to the user, allowing them to self-evaluate and understand areas for improvement next time.
[0650] • Specific example: The evaluation is displayed in a score format, such as "Courtesy: 80 points, Fluency: 70 points, Accuracy: 90 points."
[0651] This allows users to improve their customer service skills through realistic dialogue simulations.
[0652] (Application Example 1)
[0653] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0654] Traditional customer service training methods have been difficult to implement effectively due to limited opportunities for direct interaction with real customers. Furthermore, the consistency and timeliness of feedback were also problematic, hindering efficient training. In addition, in real-world store environments, securing training time during busy periods was difficult, leading to delays in skill development.
[0655] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0656] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on a client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using the artificial intelligence model to generate subsequent questions and comments; means for evaluating the user's responses and generating feedback when the interaction ends; and means for the user to select a role as a crew member or a customer, obtain an artificial intelligence model corresponding to the selected role on the server, and run it on a smart device. This enables the user to conduct customer service training in a realistic manner and efficiently improve their skills regardless of time or location.
[0657] A "role" refers to the role a user chooses within a system.
[0658] "Interface" refers to the screens and input methods that users use to interact with a system.
[0659] A "client terminal" refers to a smart device operated by a user.
[0660] An "artificial intelligence model" refers to a machine learning algorithm used to generate and analyze dialogue through natural language processing.
[0661] A "server" refers to a remote computer that is responsible for the main processing of a system.
[0662] A "question or topic" refers to the starting point of a dialogue that is generated by the system and presented to the user.
[0663] "Response" refers to the reply that a user enters into the system.
[0664] "Feedback" refers to the improvement suggestions and evaluation information provided after evaluating the user's interaction results.
[0665] "Crew role" refers to a user who takes on the role of providing customer service during customer service training.
[0666] "Customer role" refers to the user who takes on the role of the person being served in customer service training.
[0667] "Smart devices" refer to mobile information terminals connected to the internet, such as smartphones and tablet devices.
[0668] "Retrieving on a server" refers to downloading specific information or models from a remote computer and using them.
[0669] "Execution" refers to a system carrying out a specific process.
[0670] The system in this invention provides the necessary functions for users to conduct customer service training. The system mainly consists of a server, a client terminal (smart device), and the user. The server is responsible for major data processing and execution of AI models, while the client terminal functions as the user interface.
[0671] Hardware and software to be used
[0672] hardware
[0673] Smart devices (smartphones and tablet devices)
[0674] Server (using a cloud server, e.g., AWS)
[0675] software
[0676] Mobile applications (developed in Swift for iOS and Kotlin for Android)
[0677] Server-side platform (combination of Node.js and Python)
[0678] Artificial intelligence model (using GPT-4)
[0679] System program processing flow
[0680] User Interface
[0681] The terminal generates and displays an interface for the user to select a role (crew member or customer). This allows the user to choose which role to simulate.
[0682] Server-side processing
[0683] When a user selects a role, that information is sent from the terminal to the server. The server then selects the corresponding artificial intelligence model based on the selected role. For example, if the user chooses the "crew" role, the server will select the "customer" model.
[0684] Initial question generation and display
[0685] The server uses the selected AI model to generate initial questions and topics. These questions are sent to the client terminal, which then displays them to the user. For example, the question "Hi, what coffee do you recommend here?" might be displayed.
[0686] Continuing the dialogue
[0687] When the user responds to a question, the device sends that response to the server. The server uses an AI model to generate the next question or comment and sends it back to the device. The system repeats this process, continuing the dialogue.
[0688] Evaluation and Feedback
[0689] Once the conversation ends, the server evaluates the entire interaction and generates feedback. This feedback evaluates the user's courtesy, fluency, and accuracy, among other things. This feedback is sent to the terminal and displayed to the user.
[0690] Adding specific examples
[0691] When a user begins customer service training using their smartphone, the following steps are taken:
[0692] 1. The user launches the app and logs in.
[0693] 2. On the role selection screen, select either "Crew Member" or "Customer."
[0694] 3. The server selects either a "customer model" or a "crew model" and generates an initial question (e.g., "Hello, what coffee do you recommend here?").
[0695] 4. When the user responds, the server generates the next question (e.g., User: "Our recommendation is the cappuccino" → Server: "Do you have any seats available?").
[0696] 5. After the interaction ends, the server evaluates the user's response and generates and returns feedback.
[0697] Example of a prompt
[0698] You will act as an AI for customer service training. When a user responds as a crew member, generate the following questions and comments as needed. For example, if the user answers "Our recommendation is the cappuccino," then ask "Do you have any seats available?"
[0699] In this way, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[0700] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0701] Step 1:
[0702] User login
[0703] The user launches the app on their smartphone and enters their email address and password on the login screen. The device sends the entered authentication information to the server, which then verifies it against its database. If authentication is successful, the server returns an authentication success message to the device.
[0704] Enter: Email address, password
[0705] Output: Authentication success message
[0706] Specific operation: The terminal sends the entered email address and password to the server as an HTTP request, and the server checks the database and returns the authentication result.
[0707] Step 2:
[0708] Role Selection
[0709] Once the user successfully logs in, the terminal displays a role selection screen. The user chooses either "Crew Member" or "Customer." The terminal then sends the selected role information to the server.
[0710] Input: Role selection (Crew member or customer)
[0711] Output: Role information
[0712] Specific operation: The terminal sends the user's role selection information to the server as an HTTP request. The server receives the role information and starts the process of selecting the appropriate AI model.
[0713] Step 3:
[0714] Selection of an artificial intelligence model
[0715] The server selects the corresponding artificial intelligence model based on the role information chosen by the user. If the user chooses the "crew" role, the server selects the "customer" model; otherwise, it selects the "crew" model.
[0716] Input: Role information
[0717] Output: Corresponding AI model
[0718] Specific operation: The server selects the appropriate AI model (e.g., a GPT-4 based model) based on its internal logic and loads the model.
[0719] Step 4:
[0720] Generating initial questions
[0721] The server uses the selected AI model to generate initial questions and topics. The generated questions are then sent to the terminal.
[0722] Input: AI model, initial prompt
[0723] Output: Initial Question
[0724] Specific operation: The server inputs a prompt message into the AI model and sends the initial question generated by the AI to the terminal (e.g., "Hello, what coffee do you recommend here?").
[0725] Step 5:
[0726] Receiving and analyzing user responses
[0727] The terminal receives the user's response and sends it to the server. The server analyzes the user's response and generates the next question or comment.
[0728] Input: User response
[0729] Output: The following questions and comments
[0730] Specific operation: The terminal sends the user response as text data to the server, and the server uses an AI model to generate the next question (e.g., User: "Our recommendation is the cappuccino" → AI: "Do you have any seats available?").
[0731] Step 6:
[0732] Continuing the dialogue
[0733] The server sends the next question or comment to the terminal, which then displays it to the user. This process repeats until the interaction is complete.
[0734] Input: Next question or comment
[0735] Output: User response
[0736] Specific operation: The terminal displays the next question received from the server to the user, and the user responds and sends the response back to the server.
[0737] Step 7:
[0738] End of dialogue and evaluation
[0739] When the conversation ends, the server analyzes the entire conversation and evaluates the user's performance. This evaluation includes aspects such as politeness, fluency, and accuracy. Based on the evaluation results, the server generates feedback and sends it to the terminal.
[0740] Input: Entire dialogue
[0741] Output: Evaluation results, feedback
[0742] Specific operation: The server analyzes the entire conversation using an AI algorithm, calculates an evaluation score based on that analysis, generates a feedback message, and sends it to the terminal.
[0743] In this way, users can receive customer service training in a realistic setting and efficiently improve their skills.
[0744] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0745] This invention is a system that supports customer service training, providing a realistic customer service experience through user role selection and dialogue simulation. Furthermore, by combining it with an emotion engine that recognizes user emotions, it achieves a more responsive and realistic interaction.
[0746] System Configuration
[0747] This system consists of three main elements: a server, a terminal, and a user. The server integrates an artificial intelligence model and an emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[0748] Program processing flow
[0749] User role settings
[0750] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[0751] AI Model and Emotion Engine Selection
[0752] The server selects the appropriate artificial intelligence model and emotion engine based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer model" and emotion engine. Conversely, if the user selects "Customer," the server will select the "Crew model" and emotion engine.
[0753] Start of dialogue
[0754] The server causes the selected artificial intelligence model to generate initial questions or topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[0755] Continuing the dialogue
[0756] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion engine to recognize the user's emotions. For example, if the user has an anxious expression, the emotion engine identifies this as "anxiety." The AI model then generates the next question or comment, and the generated next question or comment is sent to the device. The device displays this to the user again. This process is repeated until the dialogue reaches its goal.
[0757] Evaluation and Feedback
[0758] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes the user's courtesy, fluency, accuracy, and sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[0759] Specific example
[0760] User acting as crew member and AI acting as customer.
[0761] Let's say a user uses the system to improve their customer service skills.
[0762] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[0763] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[0764] 3. The server selects a "customer role model" and an emotion engine based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[0765] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0766] 5. The server receives the response and uses its emotion engine to recognize the user's emotion (for example, detecting a smile and identifying it as "joy"). It then generates the next question, "Are there any seats available?", and the conversation continues.
[0767] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[0768] Thus, by using this invention, users can conduct customer service training in a realistic manner and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for practical training.
[0769] The following describes the processing flow.
[0770] Step 1:
[0771] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[0772] Step 2:
[0773] The terminal displays the generated role selection screen to the user.
[0774] Step 3:
[0775] Users can choose to play either the role of a "crew member" or a "customer."
[0776] Step 4:
[0777] The device sends the user's selection results to the server.
[0778] Step 5:
[0779] The server reviews the received selection results and selects an artificial intelligence model and emotion engine appropriate to the chosen role.
[0780] Step 6:
[0781] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[0782] Step 7:
[0783] The server sends the initial generated questions and topics to the terminal.
[0784] Step 8:
[0785] The device displays received questions and topics to the user.
[0786] Step 9:
[0787] The user enters their response to the displayed question.
[0788] Step 10:
[0789] The device records user response data and emotional expression data and sends it to the server.
[0790] Step 11:
[0791] The server analyzes the received user response data and uses an emotion engine to recognize the user's emotions.
[0792] Step 12:
[0793] Based on the recognized sentiment information, the server uses an artificial intelligence model to generate the next questions and comments.
[0794] Step 13:
[0795] The server sends the next generated question or comment to the terminal.
[0796] Step 14:
[0797] The device displays the next question or comment received by the user.
[0798] Step 15:
[0799] The device supplies the user's voice data and facial expression data to the emotion engine in real time and records the emotion recognition results.
[0800] Step 16:
[0801] Repeat the process from Step 9 to Step 15 until the dialogue reaches its goal.
[0802] Step 17:
[0803] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[0804] Step 18:
[0805] The server generates detailed feedback based on the evaluation results.
[0806] Step 19:
[0807] The server sends the generated feedback to the terminal.
[0808] Step 20:
[0809] The device displays feedback to the user.
[0810] (Example 2)
[0811] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0812] The present invention aims to provide a system for efficiently conducting customer service training. Specifically, it solves the problem of supplementing aspects that conventional training methods could not fully cover by providing a system that can perform realistic dialogue simulations according to the role selected by the user, recognize the user's emotions, and generate appropriate feedback.
[0813] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a display screen for the user to select a role; means for selecting an artificial intelligence engine that generates initial questions and topics to initiate a dialogue with the client device according to the selected role; means for displaying the generated questions and topics on the client device; means for receiving responses from the user and using an emotion analysis engine to recognize the user's emotions; means for using an artificial intelligence engine that generates the next questions and comments based on the emotion analysis results; and means for analyzing and evaluating the user's responses and generating feedback when the dialogue ends. This makes it possible for the user to receive practical training through realistic dialogue simulations.
[0814] A "display screen" is a screen that provides an interface for users to select roles.
[0815] An "artificial intelligence engine" is an engine that has algorithms for generating initial questions and topics based on the user's choices.
[0816] A "client device" is a terminal operated by a user, which communicates with a server to display information and receive input.
[0817] "Initial questions or topics to start a conversation" refer to questions or topics used in the introductory part of a conversation, designed to allow the user to smoothly begin the customer service simulation.
[0818] An "emotion analysis engine" is an engine that analyzes and identifies emotions from user responses and facial expression data.
[0819] "Means for evaluation and generating feedback" refers to a means of comprehensively evaluating the user's performance at the end of an interaction and informing the user of areas for improvement and weaknesses based on the results.
[0820] "Means for generating initial questions and topics" refer to algorithms or programs used to automatically generate appropriate questions and topics at the start of a dialogue.
[0821] This invention relates to a customer service training system and aims to provide realistic dialogue simulations based on roles selected by the user. The system consists of three main elements: a server, a terminal, and the user.
[0822] The server primarily integrates an artificial intelligence engine (generative AI model) and an emotion analysis engine, and is responsible for generating, analyzing, and evaluating dialogue. Specific software used includes large-scale language models such as GPT-3 and emotion recognition tools such as the Affectiva SDK. The terminal is a device operated by the user, receiving user selections and inputs, and displaying information received from the server. Users interact with the system as either a "crew member" or a "customer," and receive training based on the content of those interactions.
[0823] 1. User role settings
[0824] When a user accesses the system from their device, the server generates a role selection screen and sends it to the device. The device displays the role selection screen using HTML and JavaScript. The user selects either "Crew" or "Customer" and sends the selection result from the device to the server. The server analyzes this selection result and selects the appropriate dialogue model and sentiment analysis engine.
[0825] 2. Selection of AI Model and Emotion Engine
[0826] The server selects and initializes the appropriate artificial intelligence engine (generative AI model) and sentiment analysis engine based on the user's selection. For example, if the user selects "Crew Role," the server initializes the "Customer Role Model" and sentiment analysis engine.
[0827] 3. Starting the dialogue
[0828] The server inputs prompt text into the selected artificial intelligence engine, generating an initial question or topic. For example, it might generate the question, "Hello, what coffee do you recommend here?" The generated question is sent from the server to the terminal and displayed to the user on the terminal.
[0829] 4. Continue the dialogue
[0830] The user enters a response to a displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion analysis engine to recognize the user's emotion. For example, if the user displays "smile," the emotion analysis engine identifies this as "joy." The server then generates the next question or comment and sends it back to the device. This process is repeated until the dialogue reaches its goal.
[0831] 5. Evaluation and Feedback
[0832] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, accuracy, and the outcome of sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[0833] Specific example
[0834] When a user uses the system to improve their customer service skills, the system operates as follows:
[0835] 1. The user accesses the system by operating a terminal. The server generates a role selection screen and sends it to the terminal.
[0836] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[0837] 3. The server selects a "customer role model" and an emotion analysis engine, and generates the initial question, "Hi, what coffee do you recommend here?"
[0838] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0839] 5. The server receives the response, performs emotion recognition, and for example, detects a smile and identifies it as "joy." Based on this, it generates the next question, "Are there any seats available?", and the dialogue continues.
[0840] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. The feedback is sent to the terminal and displayed to the user.
[0841] Example prompt: "Generate an initial question for when the user chooses to play the role of a crew member. For example, 'Hi, what coffee do you recommend here?'"
[0842] Thus, by utilizing this invention, users can conduct customer service training in a realistic setting and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for more practical training.
[0843] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0844] Step 1: Generate the user role selection screen.
[0845] Input: The user accesses the system via the terminal's web browser.
[0846] Processing: The server uses HTML and JavaScript to generate a role selection screen, which includes the options "Crew Role" and "Customer Role".
[0847] Output: The server sends the generated role selection screen to the terminal as an HTTP response.
[0848] Specific operation: When a user accesses a URL, the server sends a response to the client containing the appropriate HTML and JavaScript code, which the terminal then renders and displays to the user.
[0849] Step 2: Submit the role selection results
[0850] Input: The user selects either "Crew Role" or "Customer Role" on the role selection screen.
[0851] Processing: The terminal retrieves the user's selection results using JavaScript and sends the selection data to the server as a JSON payload.
[0852] Output: The server receives the selected data.
[0853] Specific operation: When the user clicks a button, a JavaScript event handler retrieves the selection result and sends an AJAX request to the server.
[0854] Step 3: Select and initialize the AI model and emotion engine.
[0855] Input: The server receives the user's role selection results.
[0856] Processing: Based on the selected role, the server selects an appropriate AI generative model (e.g., GPT-3) and sentiment analysis engine (e.g., Affectiva SDK), and initializes them.
[0857] Output: The server generates an instance of the pre-configured AI model and emotion engine.
[0858] Specific operation: The server analyzes the received selection data, and if "Crew Role" is selected, it initializes the "Customer Role Model" and emotion engine, and loads the respective libraries and APIs.
[0859] Step 4: Generate and submit initial questions
[0860] Input: An AI model initialized on the server.
[0861] Processing: The server inputs a prompt sentence into the AI model for initial question generation. For example, it generates the question, "Hello, what coffee do you recommend here?"
[0862] Output: The server retrieves the generated question in text format and sends it to the terminal.
[0863] Specific operation: The server provides prompt text to the AI model, retrieves the generated question, and sends it to the client as an HTTP response.
[0864] Step 5: Display the initial question
[0865] Input: The initial question received by the user's device from the server.
[0866] Processing: The terminal displays the initial question it received on the screen.
[0867] Output: The screen displaying the initial question.
[0868] Specific operation: The device's browser receives the server response and displays the question text to the user as an HTML element.
[0869] Step 6: Sending the User Response
[0870] Input: The user enters their response to the initial question displayed on the screen.
[0871] Processing: The terminal receives the input response and sends it to the server as a JSON payload.
[0872] Output: The server receives the user's response data.
[0873] Specific operation: When the user enters a response in the text box and presses the submit button, the device sends an AJAX request to the server.
[0874] Step 7: Emotion recognition and next question generation
[0875] Input: Response data from users who reached the server.
[0876] Processing: The server uses an emotion analysis engine to recognize the user's emotions and, based on the results, has the AI model generate the next question or comment.
[0877] Output: Sentiment recognition results and generated questions or comments.
[0878] Specific operation: The server analyzes the response text and identifies emotions such as "joy" using emotion data, for example, "smile." Then, it generates the next question and sends it to the terminal.
[0879] Step 8: Continuing the Dialogue
[0880] Input: The following questions and comments generated from the server.
[0881] Processing: The terminal receives the next question or comment and displays it on the screen.
[0882] Output: The screen showing the following questions and comments.
[0883] Specific operation: The device's browser displays the new question text to the user again as an HTML element. This process is repeated until the interaction reaches its goal.
[0884] Step 9: Evaluation and feedback on the entire dialogue
[0885] Input: All dialogue data stored on the server when the dialogue ends.
[0886] Processing: The server runs an algorithm that evaluates the entire interaction, comprehensively assessing the user's politeness, fluency, accuracy, and sentiment recognition.
[0887] Output: Evaluation results and feedback messages.
[0888] Specific operation: The server analyzes the dialogue data, evaluates the user's performance using an evaluation algorithm, generates feedback including what to do next and areas for improvement, and sends it to the terminal.
[0889] Step 10: Displaying Feedback
[0890] Input: Feedback message received from the server.
[0891] Processing: The device displays the received feedback message to the user.
[0892] Output: Screen showing feedback.
[0893] Specific operation: The device's browser displays feedback text to the user as an HTML element, allowing the user to check their performance.
[0894] This detailed, step-by-step flow allows users to effectively conduct realistic and responsive customer service training. The introduction of an emotion analysis engine enables appropriate responses based on the user's emotions, resulting in more practical training.
[0895] (Application Example 2)
[0896] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0897] In real-world customer service training, it is crucial for trainees to practice in truly realistic scenarios and receive feedback based on their own emotional responses. However, conventional systems have limitations in training effectiveness due to insufficient emotional recognition and interactions that are not always real-time and responsive. Furthermore, for users to improve their conversational skills as store staff, training in an environment closer to actual customer service situations is required.
[0898] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0899] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using an artificial intelligence model to generate subsequent questions and comments using an emotion engine that recognizes emotions based on the generated responses; and means for evaluating the user's responses and generating feedback when the interaction ends. This enables the user to receive realistic customer service training, is provided with emotion-based interaction and feedback, and enables practical skill improvement.
[0900] A "role" refers to the specific actions or functions that a user is responsible for within a system.
[0901] "Means for generating an interface" refers to means that have the function of creating screens and operating mechanisms for users to access and operate a system.
[0902] An "artificial intelligence model for generating initial questions and topics" is an artificial intelligence model that generates questions and topics presented at the start of a conversation, based on the role selected by the user.
[0903] A "client terminal" refers to a computer or digital device that a user directly operates.
[0904] "Means for displaying questions and topics" refers to means that have the functionality to display generated questions and topics on a client terminal.
[0905] An "emotion-recognizing emotion engine" is an engine that analyzes a user's emotions from their facial expressions and voice, and generates an appropriate response based on the analysis results.
[0906] An "artificial intelligence model for generating the next question or comment" is an artificial intelligence-powered model that automatically generates the next question or comment based on the user's response and the results of sentiment analysis.
[0907] "Means for evaluating user responses and generating feedback" refers to methods for analyzing user interactions, evaluating their performance, and providing feedback on areas for improvement and positive aspects.
[0908] This system was built to help users improve their customer service skills in physical stores. The embodiments for carrying out the invention are shown below.
[0909] System Configuration
[0910] This system consists of three main elements: a server, a terminal (such as a smartphone or tablet), and the user. The server integrates an artificial intelligence model and emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[0911] Hardware and software to use
[0912] Server: A high-performance computer system used to run AI models and emotion engines.
[0913] Terminal: A user device such as a smartphone or tablet, which is operated by the user through an interface.
[0914] Software: AI models written in Python, an emotion engine for emotion recognition, and an application for the user interface.
[0915] Data processing and data calculation
[0916] Role Selection: When a user operates a terminal to select either the role of crew member or customer, that information is sent to the server.
[0917] Initial Question Generation: Based on the user's selection, the server selects an appropriate AI model (generative AI model) and generates initial questions and topics.
[0918] Display and Response: Generated questions and topics are displayed on the terminal, and responses from the user are retrieved. These responses are then sent to the server.
[0919] Emotion Recognition: The server processes user responses using an emotion engine and analyzes the emotion data. Based on this analysis, the next questions and comments are generated.
[0920] Evaluation and Feedback: After the interaction ends, the server comprehensively evaluates the user's responses and generates performance feedback. This feedback is provided to the user via the terminal, suggesting areas for improvement in training.
[0921] Specific Scenario Examples
[0922] When new staff members at a cafe undergo customer service training, the following steps are taken:
[0923] 1. The user selects a crew role and accesses the system.
[0924] 2. The server selects the "customer role model" and generates the initial question, "Hello, what coffee do you recommend here?"
[0925] 3. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[0926] 4. The server analyzes this response using an emotion engine and recognizes the user's emotion (e.g., "joy").
[0927] 5. Generate the following question or comment. For example, "Are there any seats available?"
[0928] 6. When the interaction ends, the server evaluates the entire interaction and generates feedback. This feedback is displayed to the user through the terminal.
[0929] Example of a prompt
[0930] The following are examples of prompt statements that can actually be used.
[0931] "Please generate questions that can be used for customer service training at a cafe. Please use native-level language, and keep them simple and polite."
[0932] Thus, by using the system of the present invention, users can efficiently improve their skills by conducting customer service training in a realistic manner and receiving feedback that takes emotions into account.
[0933] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0934] Step 1:
[0935] The user operates the terminal to access the system, and the role selection screen is displayed.
[0936] Specifically, the terminal sends a request to the server, asking for a role selection screen. Upon receiving this request, the server generates an interface and sends it to the terminal. The user selects either the "Crew Role" or the "Customer Role," and this selection is sent to the server via the terminal.
[0937] Input: User role selection request
[0938] Output: Role selection screen
[0939] Step 2:
[0940] The server selects an appropriate artificial intelligence model based on the user's choices and generates initial questions and topics.
[0941] Specifically, the server selects either a "crew role model" or a "customer role model" and an emotion engine based on the role selected by the user. Next, it uses a generative AI model to generate initial questions and topics based on the prompt text.
[0942] Input: Select user role
[0943] Output: Initial questions and topics
[0944] Step 3:
[0945] The terminal displays initial questions and topics received from the server to the user.
[0946] In terms of specific operations, the server sends the initial generated questions and topics to the terminal, which then displays them to the user. The user then responds to the displayed questions.
[0947] Input: Initial questions or topics
[0948] Output: User response
[0949] Step 4:
[0950] The server receives responses from users and uses an emotion engine to recognize those emotions.
[0951] In terms of specific operations, the terminal sends the user's response to the server, where an emotion engine analyzes the response and recognizes the emotion. For example, it performs facial expression analysis and voice analysis to determine the user's emotional state.
[0952] Input: User response
[0953] Output: Emotion recognition result
[0954] Step 5:
[0955] Based on the results of the sentiment engine, the server generates the following questions and comments.
[0956] In terms of specific operation, the server uses emotion recognition data as input and a generative AI model to generate the next question or comment. This allows the dialogue to continue.
[0957] Input: Sentiment recognition result
[0958] Output: The following questions and comments
[0959] Step 6:
[0960] The terminal displays the next question or comment received from the server to the user.
[0961] In terms of specific actions, this process is carried out similarly to step 3, with the terminal receiving data from the server and displaying it to the user. The user then responds again.
[0962] Input: Next question or comment
[0963] Output: User response
[0964] Step 7:
[0965] Once the interaction is complete, the server evaluates the entire interaction and generates feedback.
[0966] Specifically, the server analyzes the entire dialogue script and evaluates it based on the user's politeness, fluency, accuracy, and emotion recognition results. Based on this evaluation, it generates feedback and sends it to the terminal.
[0967] Input: Content of the entire dialogue
[0968] Output: User feedback
[0969] Step 8:
[0970] The terminal displays feedback from the server to the user.
[0971] In terms of specific actions, the device displays feedback received from the server to the user, allowing the user to check their performance and understand areas for improvement for the next training session.
[0972] Input: Feedback
[0973] Output: Feedback display
[0974] Through this series of processing steps, users can conduct customer service training in a realistic setting and receive emotion-based interaction and feedback.
[0975] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0976] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0977] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0978] [Third Embodiment]
[0979] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0980] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0981] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0982] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0983] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0984] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0985] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0986] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0987] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0988] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0989] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0990] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0991] This invention provides a system for using AI to perform realistic dialogue simulations when users conduct customer service training. Specific embodiments and their processes are described below.
[0992] System Configuration
[0993] This system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting AI models, generating dialogues, and evaluation, while the terminal accepts user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[0994] Program processing flow
[0995] User role settings
[0996] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[0997] AI Model Selection
[0998] The server selects the appropriate AI model based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer" model. Conversely, if the user selects "Customer," the server will select the "Crew" model.
[0999] Start of dialogue
[1000] The server prompts the selected AI model to generate initial questions and topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[1001] Continuing the dialogue
[1002] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an AI model to generate the next question or comment. For example, if the user responds, "Hello! Our recommendation is the cappuccino. Let me show you around," the AI (the customer model) will generate the next question, "Do you have any seats available?" This process is repeated until the dialogue reaches its goal.
[1003] Evaluation and Feedback
[1004] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, and accuracy. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1005] Specific example
[1006] User acting as crew member and AI acting as customer.
[1007] Let's say a user uses the system to improve their customer service skills.
[1008] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[1009] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1010] 3. The server selects a "customer role model" based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[1011] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1012] 5. The server receives the response and generates the next question, "Are there any seats available?", and the conversation continues.
[1013] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[1014] Thus, by using the present invention, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[1015] The following describes the processing flow.
[1016] Step 1:
[1017] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[1018] Step 2:
[1019] The terminal displays the generated role selection screen to the user.
[1020] Step 3:
[1021] Users can choose to play either the role of a "crew member" or a "customer."
[1022] Step 4:
[1023] The device sends the user's selection results to the server.
[1024] Step 5:
[1025] The server checks the received role selection results and selects an appropriate artificial intelligence model according to the selected role.
[1026] Step 6:
[1027] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[1028] Step 7:
[1029] The server sends the initial generated questions and topics to the terminal.
[1030] Step 8:
[1031] The device displays received questions and topics to the user.
[1032] Step 9:
[1033] The user enters their response to the displayed question.
[1034] Step 10:
[1035] The terminal sends the user's response to the server.
[1036] Step 11:
[1037] The server analyzes the user's response and uses an artificial intelligence model to generate the next question or comment.
[1038] Step 12:
[1039] The server sends the next generated question or comment to the terminal.
[1040] Step 13:
[1041] The device displays the next question or comment received by the user.
[1042] Step 14:
[1043] Repeat the process from Step 9 to Step 13 until the dialogue reaches its goal.
[1044] Step 15:
[1045] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[1046] Step 16:
[1047] The server generates feedback based on the evaluation results.
[1048] Step 17:
[1049] The server sends the generated feedback to the terminal.
[1050] Step 18:
[1051] The device displays feedback to the user.
[1052] (Example 1)
[1053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1054] In today's service industry, customer service skills are a crucial element. However, traditional customer service training is time-consuming and expensive, limiting opportunities for implementation. Furthermore, it is difficult to accurately replicate real-world situations, resulting in insufficient skill improvement among employees. Therefore, there is a need for a realistic dialogue simulation system that enables employees to efficiently and effectively improve their customer service skills.
[1055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1056] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the terminal according to the selected role; means for displaying the generated questions and topics on the terminal; means for receiving responses from the user and using a natural language processing algorithm to generate subsequent questions and comments; and means for evaluating the user's responses when the dialogue ends and generating feedback based on the evaluation. This allows the user to have a realistic dialogue simulation while saving time and money.
[1057] A "user" refers to a person who operates the system, selects a role, and participates in dialogue simulations.
[1058] A "role" refers to the role that a user chooses in a dialogue simulation, and includes roles such as staff member or visitor.
[1059] "Interface" refers to the screens and functions that users use to operate a system, and specifically includes screens for selecting roles.
[1060] "Terminal" refers to a device operated by a user, and includes personal computers, smartphones, tablets, and other similar devices.
[1061] "Initial questions or topics" refer to the inquiries or topics that the system initially generates when starting a dialogue simulation.
[1062] An "artificial intelligence model" is a model that operates based on machine learning algorithms and is used to generate dialogues.
[1063] A "natural language processing algorithm" refers to a technology that analyzes user input and generates an appropriate response based on that analysis.
[1064] "Evaluation" refers to the process of analyzing the user's responses after the dialogue simulation is completed and measuring performance based on items such as politeness, fluency, and accuracy.
[1065] "Feedback" refers to the improvement suggestions and performance evaluations provided to users based on the evaluation results.
[1066] Modes for carrying out the invention
[1067] This invention is a system that uses artificial intelligence to conduct realistic dialogue simulations when users undergo customer service training. The system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting an AI model, generating dialogues, and evaluation, while the terminal receives user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[1068] When a user accesses the system, the server first generates a role selection screen. This uses web technologies such as HTML, CSS, and JavaScript. The generated role selection screen is sent to the terminal, which displays it on the user's screen. The user selects either "Crew" or "Customer" on the role selection screen. The selection result is sent from the terminal to the server.
[1069] Next, the server selects an appropriate AI model based on the user's chosen role. Specifically, for example, if the user selects "Crew," it will select the "Customer model," and conversely, if the user selects "Customer," it will select the "Crew model." This utilizes AI models that have been pre-trained using machine learning frameworks such as TensorFlow or PyTorch.
[1070] The server initializes the selected AI model and generates the first question or topic. For example, the "customer model" would generate the question, "Hello, what coffee do you recommend here?" This information is sent to the device in JSON format, and the device displays it to the user. The user enters a response to the question, and that response is sent from the device to the server.
[1071] The server analyzes the user's responses and uses natural language processing algorithms to generate subsequent questions and comments. This process is repeated until the dialogue ends. After the dialogue ends, the server analyzes the user's entire dialogue log and generates feedback based on evaluation criteria such as politeness, fluency, and accuracy. The feedback is sent to the terminal in JSON format and displayed to the user.
[1072] Specific example
[1073] Specifically, the system operates in the following steps:
[1074] 1. When a user accesses the system, the server generates a role selection screen and sends it to the terminal.
[1075] 2. The terminal displays a role selection screen to the user.
[1076] 3. When a user selects a "crew member" role, the selection is sent to the server.
[1077] 4. The server selects the "customer role model" and generates the first question, "Hello, what coffee do you recommend here?" and sends it to the terminal.
[1078] 5. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1079] 6. The server receives the response, generates the next question "Are there any seats available?", and sends it back to the terminal.
[1080] In this way, users can efficiently and effectively improve their customer service skills through realistic dialogue simulations.
[1081] Example of a prompt
[1082] Examples of specific prompt messages include the following:
[1083] Prompt for users who chose the crew role: "As a crew member, please recommend a coffee to a customer who has come into the store."
[1084] Prompt for users who chose to play the customer role: "As a customer, act out a scenario where you visit a cafe and ask the staff for their coffee recommendation."
[1085] This system allows users to conduct realistic customer service training while saving time and money.
[1086] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1087] System processing flow
[1088] Step 1: Generate and display the role selection screen.
[1089] The server generates a role selection screen using HTML, CSS, and JavaScript when a user accesses the system.
[1090] Output: Data from the role selection screen (HTML, CSS, JavaScript) sent from the server to the terminal.
[1091] The terminal displays the received role selection screen data on the user's display.
[1092] • Users can choose to play either the role of a "crew member" or a "customer."
[1093] • Input: Select user role
[1094] Output: POST request from terminal to server containing the user's role selection results.
[1095] Step 2: Select and initialize the AI model
[1096] The server receives the user's selection results, parses them in JSON format, and selects the appropriate artificial intelligence model based on the selected role.
[1097] • Input: User role selection results (JSON format)
[1098] Output: Selected AI model
[1099] For example, if the user selects the "Crew Role," select the "Customer Role Model."
[1100] The server initializes the selected AI model and prepares it for the start of the interaction.
[1101] • Specific example: Loading a pre-trained model from TensorFlow or PyTorch.
[1102] Step 3: Generate and display the first question.
[1103] The server generates initial questions and topics. For example, for a selected "customer role model," it generates the initial question, "Hi, what coffee do you recommend here?"
[1104] • Input: Selected AI model
[1105] Output: Initial Question
[1106] The server sends the generated initial questions to the terminal in JSON format.
[1107] The terminal displays the initial question received to the user.
[1108] Step 4: Receiving and analyzing user responses
[1109] • Users enter appropriate responses to the displayed questions. For example, they might answer, "Our recommendation is the cappuccino."
[1110] • Input: User response
[1111] Output: User response data (JSON format) sent from the terminal to the server.
[1112] The server analyzes the user's response and uses natural language processing algorithms to generate the next question or comment.
[1113] • Input: User response data
[1114] Output: Next question or comment
[1115] Step 5: Continue the dialogue
[1116] The server sends the generated next question or comment to the terminal in JSON format.
[1117] The device displays received questions or comments to the user.
[1118] This process is repeated until the dialogue is finished.
[1119] • Input: User responses and generated questions or comments
[1120] • Output: Generate and display the next question or comment.
[1121] Step 6: Evaluate the dialogue and generate feedback
[1122] • Once the conversation ends, the server analyzes the entire conversation log and evaluates the user's performance based on evaluation criteria such as politeness, fluency, and accuracy.
[1123] • Input: Overall dialogue log
[1124] • Output: Evaluation results
[1125] The server generates feedback based on the evaluation results and sends it to the terminal in JSON format.
[1126] The device displays the generated feedback to the user, allowing them to self-evaluate and understand areas for improvement next time.
[1127] • Specific example: The evaluation is displayed in a score format, such as "Courtesy: 80 points, Fluency: 70 points, Accuracy: 90 points."
[1128] This allows users to improve their customer service skills through realistic dialogue simulations.
[1129] (Application Example 1)
[1130] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1131] Traditional customer service training methods have been difficult to implement effectively due to limited opportunities for direct interaction with real customers. Furthermore, the consistency and timeliness of feedback were also problematic, hindering efficient training. In addition, in real-world store environments, securing training time during busy periods was difficult, leading to delays in skill development.
[1132] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1133] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on a client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using the artificial intelligence model to generate subsequent questions and comments; means for evaluating the user's responses and generating feedback when the interaction ends; and means for the user to select a role as a crew member or a customer, obtain an artificial intelligence model corresponding to the selected role on the server, and run it on a smart device. This enables the user to conduct customer service training in a realistic manner and efficiently improve their skills regardless of time or location.
[1134] A "role" refers to the role a user chooses within a system.
[1135] "Interface" refers to the screens and input methods that users use to interact with a system.
[1136] A "client terminal" refers to a smart device operated by a user.
[1137] An "artificial intelligence model" refers to a machine learning algorithm used to generate and analyze dialogue through natural language processing.
[1138] A "server" refers to a remote computer that is responsible for the main processing of a system.
[1139] A "question or topic" refers to the starting point of a dialogue that is generated by the system and presented to the user.
[1140] "Response" refers to the reply that a user enters into the system.
[1141] "Feedback" refers to the improvement suggestions and evaluation information provided after evaluating the user's interaction results.
[1142] "Crew role" refers to a user who takes on the role of providing customer service during customer service training.
[1143] "Customer role" refers to the user who takes on the role of the person being served in customer service training.
[1144] "Smart devices" refer to mobile information terminals connected to the internet, such as smartphones and tablet devices.
[1145] "Retrieving on a server" refers to downloading specific information or models from a remote computer and using them.
[1146] "Execution" refers to a system carrying out a specific process.
[1147] The system in this invention provides the necessary functions for users to conduct customer service training. The system mainly consists of a server, a client terminal (smart device), and the user. The server is responsible for major data processing and execution of AI models, while the client terminal functions as the user interface.
[1148] Hardware and software to be used
[1149] hardware
[1150] Smart devices (smartphones and tablet devices)
[1151] Server (using a cloud server, e.g., AWS)
[1152] software
[1153] Mobile applications (developed in Swift for iOS and Kotlin for Android)
[1154] Server-side platform (combination of Node.js and Python)
[1155] Artificial intelligence model (using GPT-4)
[1156] System program processing flow
[1157] User Interface
[1158] The terminal generates and displays an interface for the user to select a role (crew member or customer). This allows the user to choose which role to simulate.
[1159] Server-side processing
[1160] When a user selects a role, that information is sent from the terminal to the server. The server then selects the corresponding artificial intelligence model based on the selected role. For example, if the user chooses the "crew" role, the server will select the "customer" model.
[1161] Initial question generation and display
[1162] The server uses the selected AI model to generate initial questions and topics. These questions are sent to the client terminal, which then displays them to the user. For example, the question "Hi, what coffee do you recommend here?" might be displayed.
[1163] Continuing the dialogue
[1164] When the user responds to a question, the device sends that response to the server. The server uses an AI model to generate the next question or comment and sends it back to the device. The system repeats this process, continuing the dialogue.
[1165] Evaluation and Feedback
[1166] Once the conversation ends, the server evaluates the entire interaction and generates feedback. This feedback evaluates the user's courtesy, fluency, and accuracy, among other things. This feedback is sent to the terminal and displayed to the user.
[1167] Adding specific examples
[1168] When a user begins customer service training using their smartphone, the following steps are taken:
[1169] 1. The user launches the app and logs in.
[1170] 2. On the role selection screen, select either "Crew Member" or "Customer."
[1171] 3. The server selects either a "customer model" or a "crew model" and generates an initial question (e.g., "Hello, what coffee do you recommend here?").
[1172] 4. When the user responds, the server generates the next question (e.g., User: "Our recommendation is the cappuccino" → Server: "Do you have any seats available?").
[1173] 5. After the interaction ends, the server evaluates the user's response and generates and returns feedback.
[1174] Example of a prompt
[1175] You will act as an AI for customer service training. When a user responds as a crew member, generate the following questions and comments as needed. For example, if the user answers "Our recommendation is the cappuccino," then ask "Do you have any seats available?"
[1176] In this way, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[1177] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1178] Step 1:
[1179] User login
[1180] The user launches the app on their smartphone and enters their email address and password on the login screen. The device sends the entered authentication information to the server, which then verifies it against its database. If authentication is successful, the server returns an authentication success message to the device.
[1181] Enter: Email address, password
[1182] Output: Authentication success message
[1183] Specific operation: The terminal sends the entered email address and password to the server as an HTTP request, and the server checks the database and returns the authentication result.
[1184] Step 2:
[1185] Role Selection
[1186] Once the user successfully logs in, the terminal displays a role selection screen. The user chooses either "Crew Member" or "Customer." The terminal then sends the selected role information to the server.
[1187] Input: Role selection (Crew member or customer)
[1188] Output: Role information
[1189] Specific operation: The terminal sends the user's role selection information to the server as an HTTP request. The server receives the role information and starts the process of selecting the appropriate AI model.
[1190] Step 3:
[1191] Selection of an artificial intelligence model
[1192] The server selects the corresponding artificial intelligence model based on the role information chosen by the user. If the user chooses the "crew" role, the server selects the "customer" model; otherwise, it selects the "crew" model.
[1193] Input: Role information
[1194] Output: Corresponding AI model
[1195] Specific operation: The server selects the appropriate AI model (e.g., a GPT-4 based model) based on its internal logic and loads the model.
[1196] Step 4:
[1197] Generating initial questions
[1198] The server uses the selected AI model to generate initial questions and topics. The generated questions are then sent to the terminal.
[1199] Input: AI model, initial prompt
[1200] Output: Initial Question
[1201] Specific operation: The server inputs a prompt message into the AI model and sends the initial question generated by the AI to the terminal (e.g., "Hello, what coffee do you recommend here?").
[1202] Step 5:
[1203] Receiving and analyzing user responses
[1204] The terminal receives the user's response and sends it to the server. The server analyzes the user's response and generates the next question or comment.
[1205] Input: User response
[1206] Output: The following questions and comments
[1207] Specific operation: The terminal sends the user response as text data to the server, and the server uses an AI model to generate the next question (e.g., User: "Our recommendation is the cappuccino" → AI: "Do you have any seats available?").
[1208] Step 6:
[1209] Continuing the dialogue
[1210] The server sends the next question or comment to the terminal, which then displays it to the user. This process repeats until the interaction is complete.
[1211] Input: Next question or comment
[1212] Output: User response
[1213] Specific operation: The terminal displays the next question received from the server to the user, and the user responds and sends the response back to the server.
[1214] Step 7:
[1215] End of dialogue and evaluation
[1216] When the conversation ends, the server analyzes the entire conversation and evaluates the user's performance. This evaluation includes aspects such as politeness, fluency, and accuracy. Based on the evaluation results, the server generates feedback and sends it to the terminal.
[1217] Input: Entire dialogue
[1218] Output: Evaluation results, feedback
[1219] Specific operation: The server analyzes the entire conversation using an AI algorithm, calculates an evaluation score based on that analysis, generates a feedback message, and sends it to the terminal.
[1220] In this way, users can receive customer service training in a realistic setting and efficiently improve their skills.
[1221] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1222] This invention is a system that supports customer service training, providing a realistic customer service experience through user role selection and dialogue simulation. Furthermore, by combining it with an emotion engine that recognizes user emotions, it achieves a more responsive and realistic interaction.
[1223] System Configuration
[1224] This system consists of three main elements: a server, a terminal, and a user. The server integrates an artificial intelligence model and an emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[1225] Program processing flow
[1226] User role settings
[1227] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[1228] AI Model and Emotion Engine Selection
[1229] The server selects the appropriate artificial intelligence model and emotion engine based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer model" and emotion engine. Conversely, if the user selects "Customer," the server will select the "Crew model" and emotion engine.
[1230] Start of dialogue
[1231] The server causes the selected artificial intelligence model to generate initial questions or topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[1232] Continuing the dialogue
[1233] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion engine to recognize the user's emotions. For example, if the user has an anxious expression, the emotion engine identifies this as "anxiety." The AI model then generates the next question or comment, and the generated next question or comment is sent to the device. The device displays this to the user again. This process is repeated until the dialogue reaches its goal.
[1234] Evaluation and Feedback
[1235] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes the user's courtesy, fluency, accuracy, and sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1236] Specific example
[1237] User acting as crew member and AI acting as customer.
[1238] Let's say a user uses the system to improve their customer service skills.
[1239] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[1240] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1241] 3. The server selects a "customer role model" and an emotion engine based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[1242] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1243] 5. The server receives the response and uses its emotion engine to recognize the user's emotion (for example, detecting a smile and identifying it as "joy"). It then generates the next question, "Are there any seats available?", and the conversation continues.
[1244] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[1245] Thus, by using this invention, users can conduct customer service training in a realistic manner and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for practical training.
[1246] The following describes the processing flow.
[1247] Step 1:
[1248] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[1249] Step 2:
[1250] The terminal displays the generated role selection screen to the user.
[1251] Step 3:
[1252] Users can choose to play either the role of a "crew member" or a "customer."
[1253] Step 4:
[1254] The device sends the user's selection results to the server.
[1255] Step 5:
[1256] The server reviews the received selection results and selects an artificial intelligence model and emotion engine appropriate to the chosen role.
[1257] Step 6:
[1258] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[1259] Step 7:
[1260] The server sends the initial generated questions and topics to the terminal.
[1261] Step 8:
[1262] The device displays received questions and topics to the user.
[1263] Step 9:
[1264] The user enters their response to the displayed question.
[1265] Step 10:
[1266] The device records user response data and emotional expression data and sends it to the server.
[1267] Step 11:
[1268] The server analyzes the received user response data and uses an emotion engine to recognize the user's emotions.
[1269] Step 12:
[1270] Based on the recognized sentiment information, the server uses an artificial intelligence model to generate the next questions and comments.
[1271] Step 13:
[1272] The server sends the next generated question or comment to the terminal.
[1273] Step 14:
[1274] The device displays the next question or comment received by the user.
[1275] Step 15:
[1276] The device supplies the user's voice data and facial expression data to the emotion engine in real time and records the emotion recognition results.
[1277] Step 16:
[1278] Repeat the process from Step 9 to Step 15 until the dialogue reaches its goal.
[1279] Step 17:
[1280] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[1281] Step 18:
[1282] The server generates detailed feedback based on the evaluation results.
[1283] Step 19:
[1284] The server sends the generated feedback to the terminal.
[1285] Step 20:
[1286] The device displays feedback to the user.
[1287] (Example 2)
[1288] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1289] The present invention aims to provide a system for efficiently conducting customer service training. Specifically, it solves the problem of supplementing aspects that conventional training methods could not fully cover by providing a system that can perform realistic dialogue simulations according to the role selected by the user, recognize the user's emotions, and generate appropriate feedback.
[1290] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a display screen for the user to select a role; means for selecting an artificial intelligence engine that generates initial questions and topics to initiate a dialogue with the client device according to the selected role; means for displaying the generated questions and topics on the client device; means for receiving responses from the user and using an emotion analysis engine to recognize the user's emotions; means for using an artificial intelligence engine that generates the next questions and comments based on the emotion analysis results; and means for analyzing and evaluating the user's responses and generating feedback when the dialogue ends. This makes it possible for the user to receive practical training through realistic dialogue simulations.
[1291] A "display screen" is a screen that provides an interface for users to select roles.
[1292] An "artificial intelligence engine" is an engine that has algorithms for generating initial questions and topics based on the user's choices.
[1293] A "client device" is a terminal operated by a user, which communicates with a server to display information and receive input.
[1294] "Initial questions or topics to start a conversation" refer to questions or topics used in the introductory part of a conversation, designed to allow the user to smoothly begin the customer service simulation.
[1295] An "emotion analysis engine" is an engine that analyzes and identifies emotions from user responses and facial expression data.
[1296] "Means for evaluation and generating feedback" refers to a means of comprehensively evaluating the user's performance at the end of an interaction and informing the user of areas for improvement and weaknesses based on the results.
[1297] "Means for generating initial questions and topics" refer to algorithms or programs used to automatically generate appropriate questions and topics at the start of a dialogue.
[1298] This invention relates to a customer service training system and aims to provide realistic dialogue simulations based on roles selected by the user. The system consists of three main elements: a server, a terminal, and the user.
[1299] The server primarily integrates an artificial intelligence engine (generative AI model) and an emotion analysis engine, and is responsible for generating, analyzing, and evaluating dialogue. Specific software used includes large-scale language models such as GPT-3 and emotion recognition tools such as the Affectiva SDK. The terminal is a device operated by the user, receiving user selections and inputs, and displaying information received from the server. Users interact with the system as either a "crew member" or a "customer," and receive training based on the content of those interactions.
[1300] 1. User role settings
[1301] When a user accesses the system from their device, the server generates a role selection screen and sends it to the device. The device displays the role selection screen using HTML and JavaScript. The user selects either "Crew" or "Customer" and sends the selection result from the device to the server. The server analyzes this selection result and selects the appropriate dialogue model and sentiment analysis engine.
[1302] 2. Selection of AI Model and Emotion Engine
[1303] The server selects and initializes the appropriate artificial intelligence engine (generative AI model) and sentiment analysis engine based on the user's selection. For example, if the user selects "Crew Role," the server initializes the "Customer Role Model" and sentiment analysis engine.
[1304] 3. Starting the dialogue
[1305] The server inputs prompt text into the selected artificial intelligence engine, generating an initial question or topic. For example, it might generate the question, "Hello, what coffee do you recommend here?" The generated question is sent from the server to the terminal and displayed to the user on the terminal.
[1306] 4. Continue the dialogue
[1307] The user enters a response to a displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion analysis engine to recognize the user's emotion. For example, if the user displays "smile," the emotion analysis engine identifies this as "joy." The server then generates the next question or comment and sends it back to the device. This process is repeated until the dialogue reaches its goal.
[1308] 5. Evaluation and Feedback
[1309] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, accuracy, and the outcome of sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1310] Specific example
[1311] When a user uses the system to improve their customer service skills, the system operates as follows:
[1312] 1. The user accesses the system by operating a terminal. The server generates a role selection screen and sends it to the terminal.
[1313] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1314] 3. The server selects a "customer role model" and an emotion analysis engine, and generates the initial question, "Hi, what coffee do you recommend here?"
[1315] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1316] 5. The server receives the response, performs emotion recognition, and for example, detects a smile and identifies it as "joy." Based on this, it generates the next question, "Are there any seats available?", and the dialogue continues.
[1317] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. The feedback is sent to the terminal and displayed to the user.
[1318] Example prompt: "Generate an initial question for when the user chooses to play the role of a crew member. For example, 'Hi, what coffee do you recommend here?'"
[1319] Thus, by utilizing this invention, users can conduct customer service training in a realistic setting and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for more practical training.
[1320] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1321] Step 1: Generate the user role selection screen.
[1322] Input: The user accesses the system via the terminal's web browser.
[1323] Processing: The server uses HTML and JavaScript to generate a role selection screen, which includes the options "Crew Role" and "Customer Role".
[1324] Output: The server sends the generated role selection screen to the terminal as an HTTP response.
[1325] Specific operation: When a user accesses a URL, the server sends a response to the client containing the appropriate HTML and JavaScript code, which the terminal then renders and displays to the user.
[1326] Step 2: Submit the role selection results
[1327] Input: The user selects either "Crew Role" or "Customer Role" on the role selection screen.
[1328] Processing: The terminal retrieves the user's selection results using JavaScript and sends the selection data to the server as a JSON payload.
[1329] Output: The server receives the selected data.
[1330] Specific operation: When the user clicks a button, a JavaScript event handler retrieves the selection result and sends an AJAX request to the server.
[1331] Step 3: Select and initialize the AI model and emotion engine.
[1332] Input: The server receives the user's role selection results.
[1333] Processing: Based on the selected role, the server selects an appropriate AI generative model (e.g., GPT-3) and sentiment analysis engine (e.g., Affectiva SDK), and initializes them.
[1334] Output: The server generates an instance of the pre-configured AI model and emotion engine.
[1335] Specific operation: The server analyzes the received selection data, and if "Crew Role" is selected, it initializes the "Customer Role Model" and emotion engine, and loads the respective libraries and APIs.
[1336] Step 4: Generate and submit initial questions
[1337] Input: An AI model initialized on the server.
[1338] Processing: The server inputs a prompt sentence into the AI model for initial question generation. For example, it generates the question, "Hello, what coffee do you recommend here?"
[1339] Output: The server retrieves the generated question in text format and sends it to the terminal.
[1340] Specific operation: The server provides prompt text to the AI model, retrieves the generated question, and sends it to the client as an HTTP response.
[1341] Step 5: Display the initial question
[1342] Input: The initial question received by the user's device from the server.
[1343] Processing: The terminal displays the initial question it received on the screen.
[1344] Output: The screen displaying the initial question.
[1345] Specific operation: The device's browser receives the server response and displays the question text to the user as an HTML element.
[1346] Step 6: Sending the User Response
[1347] Input: The user enters their response to the initial question displayed on the screen.
[1348] Processing: The terminal receives the input response and sends it to the server as a JSON payload.
[1349] Output: The server receives the user's response data.
[1350] Specific operation: When the user enters a response in the text box and presses the submit button, the device sends an AJAX request to the server.
[1351] Step 7: Emotion recognition and next question generation
[1352] Input: Response data from users who reached the server.
[1353] Processing: The server uses an emotion analysis engine to recognize the user's emotions and, based on the results, has the AI model generate the next question or comment.
[1354] Output: Sentiment recognition results and generated questions or comments.
[1355] Specific operation: The server analyzes the response text and identifies emotions such as "joy" using emotion data, for example, "smile." Then, it generates the next question and sends it to the terminal.
[1356] Step 8: Continuing the Dialogue
[1357] Input: The following questions and comments generated from the server.
[1358] Processing: The terminal receives the next question or comment and displays it on the screen.
[1359] Output: The screen showing the following questions and comments.
[1360] Specific operation: The device's browser displays the new question text to the user again as an HTML element. This process is repeated until the interaction reaches its goal.
[1361] Step 9: Evaluation and feedback on the entire dialogue
[1362] Input: All dialogue data stored on the server when the dialogue ends.
[1363] Processing: The server runs an algorithm that evaluates the entire interaction, comprehensively assessing the user's politeness, fluency, accuracy, and sentiment recognition.
[1364] Output: Evaluation results and feedback messages.
[1365] Specific operation: The server analyzes the dialogue data, evaluates the user's performance using an evaluation algorithm, generates feedback including what to do next and areas for improvement, and sends it to the terminal.
[1366] Step 10: Displaying Feedback
[1367] Input: Feedback message received from the server.
[1368] Processing: The device displays the received feedback message to the user.
[1369] Output: Screen showing feedback.
[1370] Specific operation: The device's browser displays feedback text to the user as an HTML element, allowing the user to check their performance.
[1371] This detailed, step-by-step flow allows users to effectively conduct realistic and responsive customer service training. The introduction of an emotion analysis engine enables appropriate responses based on the user's emotions, resulting in more practical training.
[1372] (Application Example 2)
[1373] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1374] In real-world customer service training, it is crucial for trainees to practice in truly realistic scenarios and receive feedback based on their own emotional responses. However, conventional systems have limitations in training effectiveness due to insufficient emotional recognition and interactions that are not always real-time and responsive. Furthermore, for users to improve their conversational skills as store staff, training in an environment closer to actual customer service situations is required.
[1375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1376] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using an artificial intelligence model to generate subsequent questions and comments using an emotion engine that recognizes emotions based on the generated responses; and means for evaluating the user's responses and generating feedback when the interaction ends. This enables the user to receive realistic customer service training, is provided with emotion-based interaction and feedback, and enables practical skill improvement.
[1377] A "role" refers to the specific actions or functions that a user is responsible for within a system.
[1378] "Means for generating an interface" refers to means that have the function of creating screens and operating mechanisms for users to access and operate a system.
[1379] An "artificial intelligence model for generating initial questions and topics" is an artificial intelligence model that generates questions and topics presented at the start of a conversation, based on the role selected by the user.
[1380] A "client terminal" refers to a computer or digital device that a user directly operates.
[1381] "Means for displaying questions and topics" refers to means that have the functionality to display generated questions and topics on a client terminal.
[1382] An "emotion-recognizing emotion engine" is an engine that analyzes a user's emotions from their facial expressions and voice, and generates an appropriate response based on the analysis results.
[1383] An "artificial intelligence model for generating the next question or comment" is an artificial intelligence-powered model that automatically generates the next question or comment based on the user's response and the results of sentiment analysis.
[1384] "Means for evaluating user responses and generating feedback" refers to methods for analyzing user interactions, evaluating their performance, and providing feedback on areas for improvement and positive aspects.
[1385] This system was built to help users improve their customer service skills in physical stores. The embodiments for carrying out the invention are shown below.
[1386] System Configuration
[1387] This system consists of three main elements: a server, a terminal (such as a smartphone or tablet), and the user. The server integrates an artificial intelligence model and emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[1388] Hardware and software to use
[1389] Server: A high-performance computer system used to run AI models and emotion engines.
[1390] Terminal: A user device such as a smartphone or tablet, which is operated by the user through an interface.
[1391] Software: AI models written in Python, an emotion engine for emotion recognition, and an application for the user interface.
[1392] Data processing and data calculation
[1393] Role Selection: When a user operates a terminal to select either the role of crew member or customer, that information is sent to the server.
[1394] Initial Question Generation: Based on the user's selection, the server selects an appropriate AI model (generative AI model) and generates initial questions and topics.
[1395] Display and Response: Generated questions and topics are displayed on the terminal, and responses from the user are retrieved. These responses are then sent to the server.
[1396] Emotion Recognition: The server processes user responses using an emotion engine and analyzes the emotion data. Based on this analysis, the next questions and comments are generated.
[1397] Evaluation and Feedback: After the interaction ends, the server comprehensively evaluates the user's responses and generates performance feedback. This feedback is provided to the user via the terminal, suggesting areas for improvement in training.
[1398] Specific Scenario Examples
[1399] When new staff members at a cafe undergo customer service training, the following steps are taken:
[1400] 1. The user selects a crew role and accesses the system.
[1401] 2. The server selects the "customer role model" and generates the initial question, "Hello, what coffee do you recommend here?"
[1402] 3. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1403] 4. The server analyzes this response using an emotion engine and recognizes the user's emotion (e.g., "joy").
[1404] 5. Generate the following question or comment. For example, "Are there any seats available?"
[1405] 6. When the interaction ends, the server evaluates the entire interaction and generates feedback. This feedback is displayed to the user through the terminal.
[1406] Example of a prompt
[1407] The following are examples of prompt statements that can actually be used.
[1408] "Please generate questions that can be used for customer service training at a cafe. Please use native-level language, and keep them simple and polite."
[1409] Thus, by using the system of the present invention, users can efficiently improve their skills by conducting customer service training in a realistic manner and receiving feedback that takes emotions into account.
[1410] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1411] Step 1:
[1412] The user operates the terminal to access the system, and the role selection screen is displayed.
[1413] Specifically, the terminal sends a request to the server, asking for a role selection screen. Upon receiving this request, the server generates an interface and sends it to the terminal. The user selects either the "Crew Role" or the "Customer Role," and this selection is sent to the server via the terminal.
[1414] Input: User role selection request
[1415] Output: Role selection screen
[1416] Step 2:
[1417] The server selects an appropriate artificial intelligence model based on the user's choices and generates initial questions and topics.
[1418] Specifically, the server selects either a "crew role model" or a "customer role model" and an emotion engine based on the role selected by the user. Next, it uses a generative AI model to generate initial questions and topics based on the prompt text.
[1419] Input: Select user role
[1420] Output: Initial questions and topics
[1421] Step 3:
[1422] The terminal displays initial questions and topics received from the server to the user.
[1423] In terms of specific operations, the server sends the initial generated questions and topics to the terminal, which then displays them to the user. The user then responds to the displayed questions.
[1424] Input: Initial questions or topics
[1425] Output: User response
[1426] Step 4:
[1427] The server receives responses from users and uses an emotion engine to recognize those emotions.
[1428] In terms of specific operations, the terminal sends the user's response to the server, where an emotion engine analyzes the response and recognizes the emotion. For example, it performs facial expression analysis and voice analysis to determine the user's emotional state.
[1429] Input: User response
[1430] Output: Emotion recognition result
[1431] Step 5:
[1432] Based on the results of the sentiment engine, the server generates the following questions and comments.
[1433] In terms of specific operation, the server uses emotion recognition data as input and a generative AI model to generate the next question or comment. This allows the dialogue to continue.
[1434] Input: Sentiment recognition result
[1435] Output: The following questions and comments
[1436] Step 6:
[1437] The terminal displays the next question or comment received from the server to the user.
[1438] In terms of specific actions, this process is carried out similarly to step 3, with the terminal receiving data from the server and displaying it to the user. The user then responds again.
[1439] Input: Next question or comment
[1440] Output: User response
[1441] Step 7:
[1442] Once the interaction is complete, the server evaluates the entire interaction and generates feedback.
[1443] Specifically, the server analyzes the entire dialogue script and evaluates it based on the user's politeness, fluency, accuracy, and emotion recognition results. Based on this evaluation, it generates feedback and sends it to the terminal.
[1444] Input: Content of the entire dialogue
[1445] Output: User feedback
[1446] Step 8:
[1447] The terminal displays feedback from the server to the user.
[1448] In terms of specific actions, the device displays feedback received from the server to the user, allowing the user to check their performance and understand areas for improvement for the next training session.
[1449] Input: Feedback
[1450] Output: Feedback display
[1451] Through this series of processing steps, users can conduct customer service training in a realistic setting and receive emotion-based interaction and feedback.
[1452] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1453] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1454] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1455] [Fourth Embodiment]
[1456] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1457] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1458] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1459] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1460] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1462] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1463] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1464] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1465] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1466] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1467] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1468] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1469] This invention provides a system for using AI to perform realistic dialogue simulations when users conduct customer service training. Specific embodiments and their processes are described below.
[1470] System Configuration
[1471] This system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting AI models, generating dialogues, and evaluation, while the terminal accepts user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[1472] Program processing flow
[1473] User role settings
[1474] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[1475] AI Model Selection
[1476] The server selects the appropriate AI model based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer" model. Conversely, if the user selects "Customer," the server will select the "Crew" model.
[1477] Start of dialogue
[1478] The server prompts the selected AI model to generate initial questions and topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[1479] Continuing the dialogue
[1480] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an AI model to generate the next question or comment. For example, if the user responds, "Hello! Our recommendation is the cappuccino. Let me show you around," the AI (the customer model) will generate the next question, "Do you have any seats available?" This process is repeated until the dialogue reaches its goal.
[1481] Evaluation and Feedback
[1482] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, and accuracy. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1483] Specific example
[1484] User acting as crew member and AI acting as customer.
[1485] Let's say a user uses the system to improve their customer service skills.
[1486] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[1487] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1488] 3. The server selects a "customer role model" based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[1489] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1490] 5. The server receives the response and generates the next question, "Are there any seats available?", and the conversation continues.
[1491] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[1492] Thus, by using the present invention, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[1493] The following describes the processing flow.
[1494] Step 1:
[1495] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[1496] Step 2:
[1497] The terminal displays the generated role selection screen to the user.
[1498] Step 3:
[1499] Users can choose to play either the role of a "crew member" or a "customer."
[1500] Step 4:
[1501] The device sends the user's selection results to the server.
[1502] Step 5:
[1503] The server checks the received role selection results and selects an appropriate artificial intelligence model according to the selected role.
[1504] Step 6:
[1505] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[1506] Step 7:
[1507] The server sends the initial generated questions and topics to the terminal.
[1508] Step 8:
[1509] The device displays received questions and topics to the user.
[1510] Step 9:
[1511] The user enters their response to the displayed question.
[1512] Step 10:
[1513] The terminal sends the user's response to the server.
[1514] Step 11:
[1515] The server analyzes the user's response and uses an artificial intelligence model to generate the next question or comment.
[1516] Step 12:
[1517] The server sends the next generated question or comment to the terminal.
[1518] Step 13:
[1519] The device displays the next question or comment received by the user.
[1520] Step 14:
[1521] Repeat the process from Step 9 to Step 13 until the dialogue reaches its goal.
[1522] Step 15:
[1523] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[1524] Step 16:
[1525] The server generates feedback based on the evaluation results.
[1526] Step 17:
[1527] The server sends the generated feedback to the terminal.
[1528] Step 18:
[1529] The device displays feedback to the user.
[1530] (Example 1)
[1531] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1532] In today's service industry, customer service skills are a crucial element. However, traditional customer service training is time-consuming and expensive, limiting opportunities for implementation. Furthermore, it is difficult to accurately replicate real-world situations, resulting in insufficient skill improvement among employees. Therefore, there is a need for a realistic dialogue simulation system that enables employees to efficiently and effectively improve their customer service skills.
[1533] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1534] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the terminal according to the selected role; means for displaying the generated questions and topics on the terminal; means for receiving responses from the user and using a natural language processing algorithm to generate subsequent questions and comments; and means for evaluating the user's responses when the dialogue ends and generating feedback based on the evaluation. This allows the user to have a realistic dialogue simulation while saving time and money.
[1535] A "user" refers to a person who operates the system, selects a role, and participates in dialogue simulations.
[1536] A "role" refers to the role that a user chooses in a dialogue simulation, and includes roles such as staff member or visitor.
[1537] "Interface" refers to the screens and functions that users use to operate a system, and specifically includes screens for selecting roles.
[1538] "Terminal" refers to a device operated by a user, and includes personal computers, smartphones, tablets, and other similar devices.
[1539] "Initial questions or topics" refer to the inquiries or topics that the system initially generates when starting a dialogue simulation.
[1540] An "artificial intelligence model" is a model that operates based on machine learning algorithms and is used to generate dialogues.
[1541] A "natural language processing algorithm" refers to a technology that analyzes user input and generates an appropriate response based on that analysis.
[1542] "Evaluation" refers to the process of analyzing the user's responses after the dialogue simulation is completed and measuring performance based on items such as politeness, fluency, and accuracy.
[1543] "Feedback" refers to the improvement suggestions and performance evaluations provided to users based on the evaluation results.
[1544] Modes for carrying out the invention
[1545] This invention is a system that uses artificial intelligence to conduct realistic dialogue simulations when users undergo customer service training. The system consists of three main elements: a server, a terminal, and a user. The server is responsible for major processes such as selecting an AI model, generating dialogues, and evaluation, while the terminal receives user input and displays questions and feedback. The user interacts with the system as either a crew member or a customer.
[1546] When a user accesses the system, the server first generates a role selection screen. This uses web technologies such as HTML, CSS, and JavaScript. The generated role selection screen is sent to the terminal, which displays it on the user's screen. The user selects either "Crew" or "Customer" on the role selection screen. The selection result is sent from the terminal to the server.
[1547] Next, the server selects an appropriate AI model based on the user's chosen role. Specifically, for example, if the user selects "Crew," it will select the "Customer model," and conversely, if the user selects "Customer," it will select the "Crew model." This utilizes AI models that have been pre-trained using machine learning frameworks such as TensorFlow or PyTorch.
[1548] The server initializes the selected AI model and generates the first question or topic. For example, the "customer model" would generate the question, "Hello, what coffee do you recommend here?" This information is sent to the device in JSON format, and the device displays it to the user. The user enters a response to the question, and that response is sent from the device to the server.
[1549] The server analyzes the user's responses and uses natural language processing algorithms to generate subsequent questions and comments. This process is repeated until the dialogue ends. After the dialogue ends, the server analyzes the user's entire dialogue log and generates feedback based on evaluation criteria such as politeness, fluency, and accuracy. The feedback is sent to the terminal in JSON format and displayed to the user.
[1550] Specific example
[1551] Specifically, the system operates in the following steps:
[1552] 1. When a user accesses the system, the server generates a role selection screen and sends it to the terminal.
[1553] 2. The terminal displays a role selection screen to the user.
[1554] 3. When a user selects a "crew member" role, the selection is sent to the server.
[1555] 4. The server selects the "customer role model" and generates the first question, "Hello, what coffee do you recommend here?" and sends it to the terminal.
[1556] 5. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1557] 6. The server receives the response, generates the next question "Are there any seats available?", and sends it back to the terminal.
[1558] In this way, users can efficiently and effectively improve their customer service skills through realistic dialogue simulations.
[1559] Example of a prompt
[1560] Examples of specific prompt messages include the following:
[1561] Prompt for users who chose the crew role: "As a crew member, please recommend a coffee to a customer who has come into the store."
[1562] Prompt for users who chose to play the customer role: "As a customer, act out a scenario where you visit a cafe and ask the staff for their coffee recommendation."
[1563] This system allows users to conduct realistic customer service training while saving time and money.
[1564] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1565] System processing flow
[1566] Step 1: Generate and display the role selection screen.
[1567] The server generates a role selection screen using HTML, CSS, and JavaScript when a user accesses the system.
[1568] Output: Data from the role selection screen (HTML, CSS, JavaScript) sent from the server to the terminal.
[1569] The terminal displays the received role selection screen data on the user's display.
[1570] • Users can choose to play either the role of a "crew member" or a "customer."
[1571] • Input: Select user role
[1572] Output: POST request from terminal to server containing the user's role selection results.
[1573] Step 2: Select and initialize the AI model
[1574] The server receives the user's selection results, parses them in JSON format, and selects the appropriate artificial intelligence model based on the selected role.
[1575] • Input: User role selection results (JSON format)
[1576] Output: Selected AI model
[1577] For example, if the user selects the "Crew Role," select the "Customer Role Model."
[1578] The server initializes the selected AI model and prepares it for the start of the interaction.
[1579] • Specific example: Loading a pre-trained model from TensorFlow or PyTorch.
[1580] Step 3: Generate and display the first question.
[1581] The server generates initial questions and topics. For example, for a selected "customer role model," it generates the initial question, "Hi, what coffee do you recommend here?"
[1582] • Input: Selected AI model
[1583] Output: Initial Question
[1584] The server sends the generated initial questions to the terminal in JSON format.
[1585] The terminal displays the initial question received to the user.
[1586] Step 4: Receiving and analyzing user responses
[1587] • Users enter appropriate responses to the displayed questions. For example, they might answer, "Our recommendation is the cappuccino."
[1588] • Input: User response
[1589] Output: User response data (JSON format) sent from the terminal to the server.
[1590] The server analyzes the user's response and uses natural language processing algorithms to generate the next question or comment.
[1591] • Input: User response data
[1592] Output: Next question or comment
[1593] Step 5: Continue the dialogue
[1594] The server sends the generated next question or comment to the terminal in JSON format.
[1595] The device displays received questions or comments to the user.
[1596] This process is repeated until the dialogue is finished.
[1597] • Input: User responses and generated questions or comments
[1598] • Output: Generate and display the next question or comment.
[1599] Step 6: Evaluate the dialogue and generate feedback
[1600] • Once the conversation ends, the server analyzes the entire conversation log and evaluates the user's performance based on evaluation criteria such as politeness, fluency, and accuracy.
[1601] • Input: Overall dialogue log
[1602] • Output: Evaluation results
[1603] The server generates feedback based on the evaluation results and sends it to the terminal in JSON format.
[1604] The device displays the generated feedback to the user, allowing them to self-evaluate and understand areas for improvement next time.
[1605] • Specific example: The evaluation is displayed in a score format, such as "Courtesy: 80 points, Fluency: 70 points, Accuracy: 90 points."
[1606] This allows users to improve their customer service skills through realistic dialogue simulations.
[1607] (Application Example 1)
[1608] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1609] Traditional customer service training methods have been difficult to implement effectively due to limited opportunities for direct interaction with real customers. Furthermore, the consistency and timeliness of feedback were also problematic, hindering efficient training. In addition, in real-world store environments, securing training time during busy periods was difficult, leading to delays in skill development.
[1610] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1611] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on a client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using the artificial intelligence model to generate subsequent questions and comments; means for evaluating the user's responses and generating feedback when the interaction ends; and means for the user to select a role as a crew member or a customer, obtain an artificial intelligence model corresponding to the selected role on the server, and run it on a smart device. This enables the user to conduct customer service training in a realistic manner and efficiently improve their skills regardless of time or location.
[1612] A "role" refers to the role a user chooses within a system.
[1613] "Interface" refers to the screens and input methods that users use to interact with a system.
[1614] A "client terminal" refers to a smart device operated by a user.
[1615] An "artificial intelligence model" refers to a machine learning algorithm used to generate and analyze dialogue through natural language processing.
[1616] A "server" refers to a remote computer that is responsible for the main processing of a system.
[1617] A "question or topic" refers to the starting point of a dialogue that is generated by the system and presented to the user.
[1618] "Response" refers to the reply that a user enters into the system.
[1619] "Feedback" refers to the improvement suggestions and evaluation information provided after evaluating the user's interaction results.
[1620] "Crew role" refers to a user who takes on the role of providing customer service during customer service training.
[1621] "Customer role" refers to the user who takes on the role of the person being served in customer service training.
[1622] "Smart devices" refer to mobile information terminals connected to the internet, such as smartphones and tablet devices.
[1623] "Retrieving on a server" refers to downloading specific information or models from a remote computer and using them.
[1624] "Execution" refers to a system carrying out a specific process.
[1625] The system in this invention provides the necessary functions for users to conduct customer service training. The system mainly consists of a server, a client terminal (smart device), and the user. The server is responsible for major data processing and execution of AI models, while the client terminal functions as the user interface.
[1626] Hardware and software to be used
[1627] hardware
[1628] Smart devices (smartphones and tablet devices)
[1629] Server (using a cloud server, e.g., AWS)
[1630] software
[1631] Mobile applications (developed in Swift for iOS and Kotlin for Android)
[1632] Server-side platform (combination of Node.js and Python)
[1633] Artificial intelligence model (using GPT-4)
[1634] System program processing flow
[1635] User Interface
[1636] The terminal generates and displays an interface for the user to select a role (crew member or customer). This allows the user to choose which role to simulate.
[1637] Server-side processing
[1638] When a user selects a role, that information is sent from the terminal to the server. The server then selects the corresponding artificial intelligence model based on the selected role. For example, if the user chooses the "crew" role, the server will select the "customer" model.
[1639] Initial question generation and display
[1640] The server uses the selected AI model to generate initial questions and topics. These questions are sent to the client terminal, which then displays them to the user. For example, the question "Hi, what coffee do you recommend here?" might be displayed.
[1641] Continuing the dialogue
[1642] When the user responds to a question, the device sends that response to the server. The server uses an AI model to generate the next question or comment and sends it back to the device. The system repeats this process, continuing the dialogue.
[1643] Evaluation and Feedback
[1644] Once the conversation ends, the server evaluates the entire interaction and generates feedback. This feedback evaluates the user's courtesy, fluency, and accuracy, among other things. This feedback is sent to the terminal and displayed to the user.
[1645] Adding specific examples
[1646] When a user begins customer service training using their smartphone, the following steps are taken:
[1647] 1. The user launches the app and logs in.
[1648] 2. On the role selection screen, select either "Crew Member" or "Customer."
[1649] 3. The server selects either a "customer model" or a "crew model" and generates an initial question (e.g., "Hello, what coffee do you recommend here?").
[1650] 4. When the user responds, the server generates the next question (e.g., User: "Our recommendation is the cappuccino" → Server: "Do you have any seats available?").
[1651] 5. After the interaction ends, the server evaluates the user's response and generates and returns feedback.
[1652] Example of a prompt
[1653] You will act as an AI for customer service training. When a user responds as a crew member, generate the following questions and comments as needed. For example, if the user answers "Our recommendation is the cappuccino," then ask "Do you have any seats available?"
[1654] In this way, users can conduct customer service training in a realistic setting and efficiently improve their skills.
[1655] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1656] Step 1:
[1657] User login
[1658] The user launches the app on their smartphone and enters their email address and password on the login screen. The device sends the entered authentication information to the server, which then verifies it against its database. If authentication is successful, the server returns an authentication success message to the device.
[1659] Enter: Email address, password
[1660] Output: Authentication success message
[1661] Specific operation: The terminal sends the entered email address and password to the server as an HTTP request, and the server checks the database and returns the authentication result.
[1662] Step 2:
[1663] Role Selection
[1664] Once the user successfully logs in, the terminal displays a role selection screen. The user chooses either "Crew Member" or "Customer." The terminal then sends the selected role information to the server.
[1665] Input: Role selection (Crew member or customer)
[1666] Output: Role information
[1667] Specific operation: The terminal sends the user's role selection information to the server as an HTTP request. The server receives the role information and starts the process of selecting the appropriate AI model.
[1668] Step 3:
[1669] Selection of an artificial intelligence model
[1670] The server selects the corresponding artificial intelligence model based on the role information chosen by the user. If the user chooses the "crew" role, the server selects the "customer" model; otherwise, it selects the "crew" model.
[1671] Input: Role information
[1672] Output: Corresponding AI model
[1673] Specific operation: The server selects the appropriate AI model (e.g., a GPT-4 based model) based on its internal logic and loads the model.
[1674] Step 4:
[1675] Generating initial questions
[1676] The server uses the selected AI model to generate initial questions and topics. The generated questions are then sent to the terminal.
[1677] Input: AI model, initial prompt
[1678] Output: Initial Question
[1679] Specific operation: The server inputs a prompt message into the AI model and sends the initial question generated by the AI to the terminal (e.g., "Hello, what coffee do you recommend here?").
[1680] Step 5:
[1681] Receiving and analyzing user responses
[1682] The terminal receives the user's response and sends it to the server. The server analyzes the user's response and generates the next question or comment.
[1683] Input: User response
[1684] Output: The following questions and comments
[1685] Specific operation: The terminal sends the user response as text data to the server, and the server uses an AI model to generate the next question (e.g., User: "Our recommendation is the cappuccino" → AI: "Do you have any seats available?").
[1686] Step 6:
[1687] Continuing the dialogue
[1688] The server sends the next question or comment to the terminal, which then displays it to the user. This process repeats until the interaction is complete.
[1689] Input: Next question or comment
[1690] Output: User response
[1691] Specific operation: The terminal displays the next question received from the server to the user, and the user responds and sends the response back to the server.
[1692] Step 7:
[1693] End of dialogue and evaluation
[1694] When the conversation ends, the server analyzes the entire conversation and evaluates the user's performance. This evaluation includes aspects such as politeness, fluency, and accuracy. Based on the evaluation results, the server generates feedback and sends it to the terminal.
[1695] Input: Entire dialogue
[1696] Output: Evaluation results, feedback
[1697] Specific operation: The server analyzes the entire conversation using an AI algorithm, calculates an evaluation score based on that analysis, generates a feedback message, and sends it to the terminal.
[1698] In this way, users can receive customer service training in a realistic setting and efficiently improve their skills.
[1699] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1700] This invention is a system that supports customer service training, providing a realistic customer service experience through user role selection and dialogue simulation. Furthermore, by combining it with an emotion engine that recognizes user emotions, it achieves a more responsive and realistic interaction.
[1701] System Configuration
[1702] This system consists of three main elements: a server, a terminal, and a user. The server integrates an artificial intelligence model and an emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[1703] Program processing flow
[1704] User role settings
[1705] When a user accesses the server, it generates a role selection screen and sends it to the terminal. The terminal displays the generated role selection screen to the user, who then selects either the "Crew Role" or the "Customer Role." The terminal then sends this selection result to the server.
[1706] AI Model and Emotion Engine Selection
[1707] The server selects the appropriate artificial intelligence model and emotion engine based on the user's chosen role. For example, if the user selects "Crew," the server will select the "Customer model" and emotion engine. Conversely, if the user selects "Customer," the server will select the "Crew model" and emotion engine.
[1708] Start of dialogue
[1709] The server causes the selected artificial intelligence model to generate initial questions or topics. For example, the "customer role model" would generate the question, "Hello, what coffee do you recommend here?" The generated question is sent to the terminal via the server, and the terminal displays it to the user.
[1710] Continuing the dialogue
[1711] The user enters an appropriate response to the displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion engine to recognize the user's emotions. For example, if the user has an anxious expression, the emotion engine identifies this as "anxiety." The AI model then generates the next question or comment, and the generated next question or comment is sent to the device. The device displays this to the user again. This process is repeated until the dialogue reaches its goal.
[1712] Evaluation and Feedback
[1713] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes the user's courtesy, fluency, accuracy, and sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1714] Specific example
[1715] User acting as crew member and AI acting as customer.
[1716] Let's say a user uses the system to improve their customer service skills.
[1717] 1. When a user operates a terminal and accesses the system, the server generates a role selection screen and sends it to the terminal.
[1718] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1719] 3. The server selects a "customer role model" and an emotion engine based on the user's role and generates the initial question, "Hi, what coffee do you recommend here?"
[1720] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1721] 5. The server receives the response and uses its emotion engine to recognize the user's emotion (for example, detecting a smile and identifying it as "joy"). It then generates the next question, "Are there any seats available?", and the conversation continues.
[1722] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. This feedback is displayed to the user via the terminal.
[1723] Thus, by using this invention, users can conduct customer service training in a realistic manner and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for practical training.
[1724] The following describes the processing flow.
[1725] Step 1:
[1726] When a user accesses the server, it generates a role selection screen and sends it to the terminal.
[1727] Step 2:
[1728] The terminal displays the generated role selection screen to the user.
[1729] Step 3:
[1730] Users can choose to play either the role of a "crew member" or a "customer."
[1731] Step 4:
[1732] The device sends the user's selection results to the server.
[1733] Step 5:
[1734] The server reviews the received selection results and selects an artificial intelligence model and emotion engine appropriate to the chosen role.
[1735] Step 6:
[1736] The server instructs the selected artificial intelligence model to generate initial questions and topics.
[1737] Step 7:
[1738] The server sends the initial generated questions and topics to the terminal.
[1739] Step 8:
[1740] The device displays received questions and topics to the user.
[1741] Step 9:
[1742] The user enters their response to the displayed question.
[1743] Step 10:
[1744] The device records user response data and emotional expression data and sends it to the server.
[1745] Step 11:
[1746] The server analyzes the received user response data and uses an emotion engine to recognize the user's emotions.
[1747] Step 12:
[1748] Based on the recognized sentiment information, the server uses an artificial intelligence model to generate the next questions and comments.
[1749] Step 13:
[1750] The server sends the next generated question or comment to the terminal.
[1751] Step 14:
[1752] The device displays the next question or comment received by the user.
[1753] Step 15:
[1754] The device supplies the user's voice data and facial expression data to the emotion engine in real time and records the emotion recognition results.
[1755] Step 16:
[1756] Repeat the process from Step 9 to Step 15 until the dialogue reaches its goal.
[1757] Step 17:
[1758] When the interaction ends, the server runs an algorithm that evaluates the entire interaction.
[1759] Step 18:
[1760] The server generates detailed feedback based on the evaluation results.
[1761] Step 19:
[1762] The server sends the generated feedback to the terminal.
[1763] Step 20:
[1764] The device displays feedback to the user.
[1765] (Example 2)
[1766] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1767] The present invention aims to provide a system for efficiently conducting customer service training. Specifically, it solves the problem of supplementing aspects that conventional training methods could not fully cover by providing a system that can perform realistic dialogue simulations according to the role selected by the user, recognize the user's emotions, and generate appropriate feedback.
[1768] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for generating a display screen for the user to select a role; means for selecting an artificial intelligence engine that generates initial questions and topics to initiate a dialogue with the client device according to the selected role; means for displaying the generated questions and topics on the client device; means for receiving responses from the user and using an emotion analysis engine to recognize the user's emotions; means for using an artificial intelligence engine that generates the next questions and comments based on the emotion analysis results; and means for analyzing and evaluating the user's responses and generating feedback when the dialogue ends. This makes it possible for the user to receive practical training through realistic dialogue simulations.
[1769] A "display screen" is a screen that provides an interface for users to select roles.
[1770] An "artificial intelligence engine" is an engine that has algorithms for generating initial questions and topics based on the user's choices.
[1771] A "client device" is a terminal operated by a user, which communicates with a server to display information and receive input.
[1772] "Initial questions or topics to start a conversation" refer to questions or topics used in the introductory part of a conversation, designed to allow the user to smoothly begin the customer service simulation.
[1773] An "emotion analysis engine" is an engine that analyzes and identifies emotions from user responses and facial expression data.
[1774] "Means for evaluation and generating feedback" refers to a means of comprehensively evaluating the user's performance at the end of an interaction and informing the user of areas for improvement and weaknesses based on the results.
[1775] "Means for generating initial questions and topics" refer to algorithms or programs used to automatically generate appropriate questions and topics at the start of a dialogue.
[1776] This invention relates to a customer service training system and aims to provide realistic dialogue simulations based on roles selected by the user. The system consists of three main elements: a server, a terminal, and the user.
[1777] The server primarily integrates an artificial intelligence engine (generative AI model) and an emotion analysis engine, and is responsible for generating, analyzing, and evaluating dialogue. Specific software used includes large-scale language models such as GPT-3 and emotion recognition tools such as the Affectiva SDK. The terminal is a device operated by the user, receiving user selections and inputs, and displaying information received from the server. Users interact with the system as either a "crew member" or a "customer," and receive training based on the content of those interactions.
[1778] 1. User role settings
[1779] When a user accesses the system from their device, the server generates a role selection screen and sends it to the device. The device displays the role selection screen using HTML and JavaScript. The user selects either "Crew" or "Customer" and sends the selection result from the device to the server. The server analyzes this selection result and selects the appropriate dialogue model and sentiment analysis engine.
[1780] 2. Selection of AI Model and Emotion Engine
[1781] The server selects and initializes the appropriate artificial intelligence engine (generative AI model) and sentiment analysis engine based on the user's selection. For example, if the user selects "Crew Role," the server initializes the "Customer Role Model" and sentiment analysis engine.
[1782] 3. Starting the dialogue
[1783] The server inputs prompt text into the selected artificial intelligence engine, generating an initial question or topic. For example, it might generate the question, "Hello, what coffee do you recommend here?" The generated question is sent from the server to the terminal and displayed to the user on the terminal.
[1784] 4. Continue the dialogue
[1785] The user enters a response to a displayed question, and the device sends that response to the server. The server analyzes the user's response and uses an emotion analysis engine to recognize the user's emotion. For example, if the user displays "smile," the emotion analysis engine identifies this as "joy." The server then generates the next question or comment and sends it back to the device. This process is repeated until the dialogue reaches its goal.
[1786] 5. Evaluation and Feedback
[1787] When a conversation ends, the server runs an algorithm that evaluates the entire conversation. This evaluation includes aspects such as the user's courtesy, fluency, accuracy, and the outcome of sentiment recognition. Based on the evaluation, the server generates feedback and sends it to the terminal. The terminal displays the feedback to the user, allowing them to review their performance and understand areas for improvement next time.
[1788] Specific example
[1789] When a user uses the system to improve their customer service skills, the system operates as follows:
[1790] 1. The user accesses the system by operating a terminal. The server generates a role selection screen and sends it to the terminal.
[1791] 2. When the user selects "Crew Role," the terminal sends that selection to the server.
[1792] 3. The server selects a "customer role model" and an emotion analysis engine, and generates the initial question, "Hi, what coffee do you recommend here?"
[1793] 4. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1794] 5. The server receives the response, performs emotion recognition, and for example, detects a smile and identifies it as "joy." Based on this, it generates the next question, "Are there any seats available?", and the dialogue continues.
[1795] 6. After the interaction ends, the server evaluates the user's performance and generates feedback. The feedback is sent to the terminal and displayed to the user.
[1796] Example prompt: "Generate an initial question for when the user chooses to play the role of a crew member. For example, 'Hi, what coffee do you recommend here?'"
[1797] Thus, by utilizing this invention, users can conduct customer service training in a realistic setting and efficiently improve their skills. The introduction of an emotion engine enables responses that respond to the user's emotions, allowing for more practical training.
[1798] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1799] Step 1: Generate the user role selection screen.
[1800] Input: The user accesses the system via the terminal's web browser.
[1801] Processing: The server uses HTML and JavaScript to generate a role selection screen, which includes the options "Crew Role" and "Customer Role".
[1802] Output: The server sends the generated role selection screen to the terminal as an HTTP response.
[1803] Specific operation: When a user accesses a URL, the server sends a response to the client containing the appropriate HTML and JavaScript code, which the terminal then renders and displays to the user.
[1804] Step 2: Submit the role selection results
[1805] Input: The user selects either "Crew Role" or "Customer Role" on the role selection screen.
[1806] Processing: The terminal retrieves the user's selection results using JavaScript and sends the selection data to the server as a JSON payload.
[1807] Output: The server receives the selected data.
[1808] Specific operation: When the user clicks a button, a JavaScript event handler retrieves the selection result and sends an AJAX request to the server.
[1809] Step 3: Select and initialize the AI model and emotion engine.
[1810] Input: The server receives the user's role selection results.
[1811] Processing: Based on the selected role, the server selects an appropriate AI generative model (e.g., GPT-3) and sentiment analysis engine (e.g., Affectiva SDK), and initializes them.
[1812] Output: The server generates an instance of the pre-configured AI model and emotion engine.
[1813] Specific operation: The server analyzes the received selection data, and if "Crew Role" is selected, it initializes the "Customer Role Model" and emotion engine, and loads the respective libraries and APIs.
[1814] Step 4: Generate and submit initial questions
[1815] Input: An AI model initialized on the server.
[1816] Processing: The server inputs a prompt sentence into the AI model for initial question generation. For example, it generates the question, "Hello, what coffee do you recommend here?"
[1817] Output: The server retrieves the generated question in text format and sends it to the terminal.
[1818] Specific operation: The server provides prompt text to the AI model, retrieves the generated question, and sends it to the client as an HTTP response.
[1819] Step 5: Display the initial question
[1820] Input: The initial question received by the user's device from the server.
[1821] Processing: The terminal displays the initial question it received on the screen.
[1822] Output: The screen displaying the initial question.
[1823] Specific operation: The device's browser receives the server response and displays the question text to the user as an HTML element.
[1824] Step 6: Sending the User Response
[1825] Input: The user enters their response to the initial question displayed on the screen.
[1826] Processing: The terminal receives the input response and sends it to the server as a JSON payload.
[1827] Output: The server receives the user's response data.
[1828] Specific operation: When the user enters a response in the text box and presses the submit button, the device sends an AJAX request to the server.
[1829] Step 7: Emotion recognition and next question generation
[1830] Input: Response data from users who reached the server.
[1831] Processing: The server uses an emotion analysis engine to recognize the user's emotions and, based on the results, has the AI model generate the next question or comment.
[1832] Output: Sentiment recognition results and generated questions or comments.
[1833] Specific operation: The server analyzes the response text and identifies emotions such as "joy" using emotion data, for example, "smile." Then, it generates the next question and sends it to the terminal.
[1834] Step 8: Continuing the Dialogue
[1835] Input: The following questions and comments generated from the server.
[1836] Processing: The terminal receives the next question or comment and displays it on the screen.
[1837] Output: The screen showing the following questions and comments.
[1838] Specific operation: The device's browser displays the new question text to the user again as an HTML element. This process is repeated until the interaction reaches its goal.
[1839] Step 9: Evaluation and feedback on the entire dialogue
[1840] Input: All dialogue data stored on the server when the dialogue ends.
[1841] Processing: The server runs an algorithm that evaluates the entire interaction, comprehensively assessing the user's politeness, fluency, accuracy, and sentiment recognition.
[1842] Output: Evaluation results and feedback messages.
[1843] Specific operation: The server analyzes the dialogue data, evaluates the user's performance using an evaluation algorithm, generates feedback including what to do next and areas for improvement, and sends it to the terminal.
[1844] Step 10: Displaying Feedback
[1845] Input: Feedback message received from the server.
[1846] Processing: The device displays the received feedback message to the user.
[1847] Output: Screen showing feedback.
[1848] Specific operation: The device's browser displays feedback text to the user as an HTML element, allowing the user to check their performance.
[1849] This detailed, step-by-step flow allows users to effectively conduct realistic and responsive customer service training. The introduction of an emotion analysis engine enables appropriate responses based on the user's emotions, resulting in more practical training.
[1850] (Application Example 2)
[1851] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1852] In real-world customer service training, it is crucial for trainees to practice in truly realistic scenarios and receive feedback based on their own emotional responses. However, conventional systems have limitations in training effectiveness due to insufficient emotional recognition and interactions that are not always real-time and responsive. Furthermore, for users to improve their conversational skills as store staff, training in an environment closer to actual customer service situations is required.
[1853] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1854] In this invention, the server includes means for generating an interface for the user to select a role; means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal according to the selected role; means for displaying the generated questions and topics on the client terminal; means for receiving responses from the user and using an artificial intelligence model to generate subsequent questions and comments using an emotion engine that recognizes emotions based on the generated responses; and means for evaluating the user's responses and generating feedback when the interaction ends. This enables the user to receive realistic customer service training, is provided with emotion-based interaction and feedback, and enables practical skill improvement.
[1855] A "role" refers to the specific actions or functions that a user is responsible for within a system.
[1856] "Means for generating an interface" refers to means that have the function of creating screens and operating mechanisms for users to access and operate a system.
[1857] An "artificial intelligence model for generating initial questions and topics" is an artificial intelligence model that generates questions and topics presented at the start of a conversation, based on the role selected by the user.
[1858] A "client terminal" refers to a computer or digital device that a user directly operates.
[1859] "Means for displaying questions and topics" refers to means that have the functionality to display generated questions and topics on a client terminal.
[1860] An "emotion-recognizing emotion engine" is an engine that analyzes a user's emotions from their facial expressions and voice, and generates an appropriate response based on the analysis results.
[1861] An "artificial intelligence model for generating the next question or comment" is an artificial intelligence-powered model that automatically generates the next question or comment based on the user's response and the results of sentiment analysis.
[1862] "Means for evaluating user responses and generating feedback" refers to methods for analyzing user interactions, evaluating their performance, and providing feedback on areas for improvement and positive aspects.
[1863] This system was built to help users improve their customer service skills in physical stores. The embodiments for carrying out the invention are shown below.
[1864] System Configuration
[1865] This system consists of three main elements: a server, a terminal (such as a smartphone or tablet), and the user. The server integrates an artificial intelligence model and emotion engine, and is responsible for dialogue generation, analysis, and evaluation. The terminal handles user interaction and collects emotion data, displaying various information through its interface. The user interacts with the system as either a crew member or a customer.
[1866] Hardware and software to use
[1867] Server: A high-performance computer system used to run AI models and emotion engines.
[1868] Terminal: A user device such as a smartphone or tablet, which is operated by the user through an interface.
[1869] Software: AI models written in Python, an emotion engine for emotion recognition, and an application for the user interface.
[1870] Data processing and data calculation
[1871] Role Selection: When a user operates a terminal to select either the role of crew member or customer, that information is sent to the server.
[1872] Initial Question Generation: Based on the user's selection, the server selects an appropriate AI model (generative AI model) and generates initial questions and topics.
[1873] Display and Response: Generated questions and topics are displayed on the terminal, and responses from the user are retrieved. These responses are then sent to the server.
[1874] Emotion Recognition: The server processes user responses using an emotion engine and analyzes the emotion data. Based on this analysis, the next questions and comments are generated.
[1875] Evaluation and Feedback: After the interaction ends, the server comprehensively evaluates the user's responses and generates performance feedback. This feedback is provided to the user via the terminal, suggesting areas for improvement in training.
[1876] Specific Scenario Examples
[1877] When new staff members at a cafe undergo customer service training, the following steps are taken:
[1878] 1. The user selects a crew role and accesses the system.
[1879] 2. The server selects the "customer role model" and generates the initial question, "Hello, what coffee do you recommend here?"
[1880] 3. The terminal displays this question to the user, who responds, "Our recommendation is the cappuccino."
[1881] 4. The server analyzes this response using an emotion engine and recognizes the user's emotion (e.g., "joy").
[1882] 5. Generate the following question or comment. For example, "Are there any seats available?"
[1883] 6. When the interaction ends, the server evaluates the entire interaction and generates feedback. This feedback is displayed to the user through the terminal.
[1884] Example of a prompt
[1885] The following are examples of prompt statements that can actually be used.
[1886] "Please generate questions that can be used for customer service training at a cafe. Please use native-level language, and keep them simple and polite."
[1887] Thus, by using the system of the present invention, users can efficiently improve their skills by conducting customer service training in a realistic manner and receiving feedback that takes emotions into account.
[1888] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1889] Step 1:
[1890] The user operates the terminal to access the system, and the role selection screen is displayed.
[1891] Specifically, the terminal sends a request to the server, asking for a role selection screen. Upon receiving this request, the server generates an interface and sends it to the terminal. The user selects either the "Crew Role" or the "Customer Role," and this selection is sent to the server via the terminal.
[1892] Input: User role selection request
[1893] Output: Role selection screen
[1894] Step 2:
[1895] The server selects an appropriate artificial intelligence model based on the user's choices and generates initial questions and topics.
[1896] Specifically, the server selects either a "crew role model" or a "customer role model" and an emotion engine based on the role selected by the user. Next, it uses a generative AI model to generate initial questions and topics based on the prompt text.
[1897] Input: Select user role
[1898] Output: Initial questions and topics
[1899] Step 3:
[1900] The terminal displays initial questions and topics received from the server to the user.
[1901] In terms of specific operations, the server sends the initial generated questions and topics to the terminal, which then displays them to the user. The user then responds to the displayed questions.
[1902] Input: Initial questions or topics
[1903] Output: User response
[1904] Step 4:
[1905] The server receives responses from users and uses an emotion engine to recognize those emotions.
[1906] In terms of specific operations, the terminal sends the user's response to the server, where an emotion engine analyzes the response and recognizes the emotion. For example, it performs facial expression analysis and voice analysis to determine the user's emotional state.
[1907] Input: User response
[1908] Output: Emotion recognition result
[1909] Step 5:
[1910] Based on the results of the sentiment engine, the server generates the following questions and comments.
[1911] In terms of specific operation, the server uses emotion recognition data as input and a generative AI model to generate the next question or comment. This allows the dialogue to continue.
[1912] Input: Sentiment recognition result
[1913] Output: The following questions and comments
[1914] Step 6:
[1915] The terminal displays the next question or comment received from the server to the user.
[1916] In terms of specific actions, this process is carried out similarly to step 3, with the terminal receiving data from the server and displaying it to the user. The user then responds again.
[1917] Input: Next question or comment
[1918] Output: User response
[1919] Step 7:
[1920] Once the interaction is complete, the server evaluates the entire interaction and generates feedback.
[1921] Specifically, the server analyzes the entire dialogue script and evaluates it based on the user's politeness, fluency, accuracy, and emotion recognition results. Based on this evaluation, it generates feedback and sends it to the terminal.
[1922] Input: Content of the entire dialogue
[1923] Output: User feedback
[1924] Step 8:
[1925] The terminal displays feedback from the server to the user.
[1926] In terms of specific actions, the device displays feedback received from the server to the user, allowing the user to check their performance and understand areas for improvement for the next training session.
[1927] Input: Feedback
[1928] Output: Feedback display
[1929] Through this series of processing steps, users can conduct customer service training in a realistic setting and receive emotion-based interaction and feedback.
[1930] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1931] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1932] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1933] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1934] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1935] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1936] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1937] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1938] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1939] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1940] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1941] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1942] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1943] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1944] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1945] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1946] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1947] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1948] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1949] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1950] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1951] The following is further disclosed regarding the embodiments described above.
[1952] (Claim 1)
[1953] A means of generating an interface for users to select roles,
[1954] A means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal, depending on the selected role,
[1955] A means of displaying questions and topics generated on the client terminal,
[1956] A means of receiving user responses and using an artificial intelligence model to generate subsequent questions or comments,
[1957] A means of evaluating the user's response and generating feedback when the dialogue ends,
[1958] A system that includes this.
[1959] (Claim 2)
[1960] The system according to claim 1, comprising means for a user to select a role as a crew member or a customer.
[1961] (Claim 3)
[1962] The system according to claim 1, comprising means for operating an artificial intelligence model based on a machine learning algorithm to generate subsequent questions or comments based on user responses.
[1963] "Example 1"
[1964] (Claim 1)
[1965] A means of generating an interface for users to select roles,
[1966] A means of selecting an artificial intelligence model to generate initial questions and topics on the terminal, depending on the selected role,
[1967] A means of displaying questions and topics generated on the terminal,
[1968] A means of receiving user responses and using a natural language processing algorithm to generate subsequent questions or comments,
[1969] A means of evaluating the user's response when the dialogue ends and generating feedback based on the evaluation,
[1970] A system that includes this.
[1971] (Claim 2)
[1972] The system according to claim 1, comprising means for a user to select a role as staff or a guest.
[1973] (Claim 3)
[1974] The system according to claim 1, comprising means for operating an artificial intelligence model based on a machine learning algorithm to generate subsequent questions or comments based on user responses.
[1975] "Application Example 1"
[1976] (Claim 1)
[1977] A means of generating an interface for users to select roles,
[1978] A means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal, depending on the selected role,
[1979] A means of displaying questions and topics generated on the client terminal,
[1980] A means of receiving user responses and using an artificial intelligence model to generate subsequent questions or comments,
[1981] A means of evaluating the user's response and generating feedback when the dialogue ends,
[1982] A means by which a user selects a role, either as a crew member or a customer, retrieves an artificial intelligence model corresponding to the selected role on the server, and executes it on a smart device,
[1983] A system that includes this.
[1984] (Claim 2)
[1985] The system according to claim 1, comprising means for a user to select a role as a crew member or a customer.
[1986] (Claim 3)
[1987] The system according to claim 1, comprising means of operating an artificial intelligence model on a smart device, based on a machine learning algorithm, for generating subsequent questions or comments based on user responses.
[1988] "Example 2 of combining an emotion engine"
[1989] (Claim 1)
[1990] A means for generating a display screen for the user to select a role,
[1991] A means for selecting an artificial intelligence engine that generates initial questions and topics to initiate interaction with the client device, depending on the selected role,
[1992] A means of displaying questions and topics generated on the client device,
[1993] A means of receiving user responses and using an emotion analysis engine to recognize the user's emotions,
[1994] A means of using an artificial intelligence engine that generates the next questions or comments based on the sentiment analysis results,
[1995] A means to analyze and evaluate the user's response when the conversation ends and generate feedback,
[1996] A system that includes this.
[1997] (Claim 2)
[1998] The system according to claim 1, comprising means for a user to select a role as a crew member or a customer.
[1999] (Claim 3)
[2000] The system according to claim 1, comprising means for operating an artificial intelligence engine based on a machine learning algorithm to generate the next question or comment based on the user's response.
[2001] "Application example 2 of combining emotional engines"
[2002] (Claim 1)
[2003] A means of generating an interface for users to select roles,
[2004] A means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal, depending on the selected role,
[2005] A means of displaying questions and topics generated on the client terminal,
[2006] A means of using an artificial intelligence model to generate subsequent questions or comments using an emotion engine that receives user responses and recognizes emotions based on the generated responses,
[2007] A means of evaluating the user's response and generating feedback when the dialogue ends,
[2008] A system that includes this.
[2009] (Claim 2)
[2010] The system according to claim 1, comprising means for a user to select a role as a crew member or a customer.
[2011] (Claim 3)
[2012] The system according to claim 1, further comprising means for operating an artificial intelligence model based on a machine learning algorithm to generate the next question or comment based on the user's response, and for recognizing the user's emotions using an emotion engine. [Explanation of Symbols]
[2013] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for generating an interface for users to select a role, A means for selecting an artificial intelligence model to generate initial questions and topics on the client terminal, depending on the selected role, A means of displaying questions and topics generated on the client terminal, A means of receiving user responses and using an artificial intelligence model to generate subsequent questions or comments, A means of evaluating the user's response and generating feedback when the dialogue ends, A system that includes this.
2. The system according to claim 1, comprising means for a user to select a role as a crew member or a customer.
3. The system according to claim 1, comprising means for operating an artificial intelligence model based on a machine learning algorithm to generate subsequent questions or comments based on user responses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A