system
The system addresses the challenge of training inexperienced engineers by generating scenarios for AI-driven interactions, analyzing user responses, and offering personalized feedback, resulting in efficient skill development.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Inexperienced engineers face challenges in acquiring practical knowledge and skills due to the lack of diverse training materials and effective feedback, leading to inefficient skill development.
A system that generates scenarios for AI-driven customer interactions, records and analyzes user responses, and provides personalized feedback to improve practical skills, while ensuring secure information management.
Enables efficient and safe training by simulating realistic customer interactions, providing specific feedback, and reducing the burden on supervisors, thus enhancing skill acquisition.
Smart Images

Figure 2026074902000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is an object to solve problems such as the diversity of commercial materials faced by inexperienced engineers, the difficulty in acquiring the associated knowledge, and the lack of time for practical training. Further, it is an object to improve the education process by providing a system that can reduce the burden on supervisors who guide engineers and enable efficient and safe training.
Means for Solving the Problems
[0005] This invention provides a system that automatically generates scenarios based on user input and conducts dialogues in which an AI acts as a customer according to those scenarios. The system records the dialogue as audio data, analyzes this data for evaluation, and generates specific feedback for the user. Furthermore, it incorporates means for securely managing all information, reducing the risk of information leakage while enabling practical training.
[0006] A "user" refers to an inexperienced technician who uses the system to learn about products and receive practical training.
[0007] A "scenario" refers to a setting that simulates a customer interaction scenario or situation, generated based on product information entered by the user.
[0008] "Dialogue" refers to the process of communication that the AI engages with the user as a customer, based on a generated scenario.
[0009] "Audio data" refers to digital data in audio format used to record the content of a conversation.
[0010] "Analysis" refers to the process of evaluating user responses using recorded audio data.
[0011] "Feedback" refers to advice, including evaluations and suggestions for improvement of user responses, generated based on analysis.
[0012] "Information data" refers to product information entered by users and all digital information generated during that process.
[0013] "Means of secure management" refers to methods for protecting information data from unauthorized access and leakage, and for properly storing and managing it within a system. [Brief explanation of the drawing]
[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is a system for improving the practical skills of inexperienced engineers regarding commercial products, and it uses AI to simulate customer interactions and provide feedback. Embodiments of this invention are described below.
[0036] System Configuration
[0037] The system consists of a terminal used by the user, a server that generates scenarios and processes information, and an AI that simulates dialogue.
[0038] Program Processing Overview
[0039] User: Log in to the system via your device, select the product you want to train on, and enter its details.
[0040] Server: Based on information received from the user, the system generates scenarios for use by the AI. In this process, the system prepares common questions and problem situations related to the product.
[0041] AI: Based on the generated scenario, it begins a conversation with the user as a customer. The AI asks questions related to the product and waits for the user's response.
[0042] Terminal: User responses are recorded in audio format, and the data is sent to the server.
[0043] Server: Analyzes voice data and evaluates the user's response. This evaluation is based on the accuracy and specificity of the response.
[0044] Feedback Provision: The server uses the analyzed data to generate feedback, including specific areas for improvement, and presents it to the user via the terminal.
[0045] Specific example
[0046] For example, if a user requests training on "cloud storage services," the AI, based on a scenario, will ask questions such as, "Please explain the backup function of cloud storage in detail." The user will answer, and their answer will be recorded. The server will analyze the recording and generate feedback such as, "The explanation is abstract; it would be better to mention specific steps," which will then be provided to the user via their device. Through this process, users can develop more specific and practical response skills.
[0047] The following describes the processing flow.
[0048] Step 1:
[0049] The user logs into the device and enters information about the product they wish to train on. The device receives this information and sends it to the server.
[0050] Step 2:
[0051] The server generates a scenario for use by the AI based on the product information it receives. The scenario includes product features, common problems, and anticipated customer questions. The generated scenario is then sent to the terminal.
[0052] Step 3:
[0053] The user initiates a conversational session with the AI through their device. Based on a scenario received from the server, the AI, acting as a customer, asks the user questions.
[0054] Step 4:
[0055] The user responds to questions from the AI. This interaction is recorded as audio, and the recorded data is sent from the device to the server.
[0056] Step 5:
[0057] The server analyzes the voice data. Specifically, it uses natural language processing technology to evaluate the accuracy, speed, and specificity of the user's responses.
[0058] Step 6:
[0059] The server generates specific feedback based on the analysis results. This feedback includes guidelines on what areas the user should improve.
[0060] Step 7:
[0061] The device receives feedback from the server and displays it to the user. The user can use this as a reference to help with their next learning.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] It is difficult for inexperienced engineers to effectively improve their practical skills related to products in a short period of time. Traditional training methods have difficulty providing realistic customer interaction simulations, and opportunities to receive specific feedback are limited, thus hindering efficient skill improvement.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for generating scenarios based on information input from the user, means for engaging in dialogue based on the generated scenarios, and means for generating general questions and problem situations related to the product using a generating AI model. This provides a realistic customer interaction simulation and enables efficient skill improvement through specific and useful feedback to the user.
[0067] A "user" refers to an individual or group that uses the system to receive training related to the product or service.
[0068] A "scenario" refers to a series of questions and situations prepared by a generative AI model to structure the flow of interaction with the user.
[0069] "Dialogue" refers to a series of communication processes in which the AI asks the user questions about the product, and the user responds to those questions.
[0070] "Audio data" refers to digital data in audio format used to record user responses.
[0071] "Evaluation" refers to the process of analyzing user responses based on recorded audio data to measure their accuracy and specificity.
[0072] "Feedback" refers to information generated based on evaluation results, including specific areas for improvement and advice regarding the response.
[0073] A "generative AI model" refers to an artificial intelligence model used to generate questions and problem situations related to a product or service.
[0074] A "terminal" is a device used by a user to access a system, and it has functions for information input and voice recording.
[0075] A "server" refers to a central computing device that handles tasks such as scenario generation, dialogue management, audio data analysis, and feedback generation.
[0076] This invention is designed to help inexperienced engineers improve their practical skills with specific products. The system consists of a terminal accessed by the user, a server for information processing, and an AI for managing the interaction. The user logs into the system through the terminal and selects the product to be trained on. The terminal transmits detailed information about the product entered by the user to the server. Based on this information, the server generates an interaction scenario using a generative AI model.
[0077] Based on the generated scenario, the AI simulates a realistic conversation with the user. The AI creates questions about the product and asks them sequentially to the user. Specific questions might include things like, "Please explain the cloud storage backup function in detail."
[0078] The device has the function of recording the user's voice responses in real time and sending that data to the server. The server uses speech recognition software to analyze this voice data and evaluate the accuracy and specificity of the user's responses. Based on this evaluation, the server generates feedback that includes specific areas for improvement and provides this information to the user through the device.
[0079] As a concrete example in this system, if a user requests training on cloud storage services, the following prompt message would be used: "We want to simulate explaining cloud storage services to a customer. Please create a scenario where the user requests a specific explanation about the backup function."
[0080] This procedure allows users to efficiently acquire practical and specific skills. Furthermore, by using a generative AI model, this system can flexibly design product-specific questions and simulations, making it suitable for a wide variety of products.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user logs into the system. The user uses their device to enter their login information and access the system. The server receives the entered user information and performs authentication. If successful, the user is redirected to the home screen.
[0084] Step 2:
[0085] The user selects a product. The user chooses the product they want to train on and enters its details into the terminal. The server then queries the database for product information and receives information about the product specified by the user.
[0086] Step 3:
[0087] The server generates the scenario. Based on the received product information, it uses a generating AI model to generate a scenario that includes common questions and problem situations related to the product. Specifically, prompt sentences are input to the AI model, and a scenario is obtained as output. This scenario includes a list of questions related to the product.
[0088] Step 4:
[0089] The AI initiates a conversation. Based on the generated scenario, the AI sequentially presents the user with questions about the product. The AI selects a question from the scenario, sends it to the device in voice or text format, and waits for the user's response.
[0090] Step 5:
[0091] The device records the user's responses. When the user responds to the AI's questions verbally, the device records the audio in real time and sends it to the server as digital audio data. The recorded audio data is obtained as output.
[0092] Step 6:
[0093] The server analyzes the audio data. The server uses speech recognition software to analyze the received audio data and convert it into text. Based on this text data, the server evaluates the accuracy and specificity of the response. At this point, an evaluation score and areas for improvement are output as part of the analysis results.
[0094] Step 7:
[0095] The server generates and provides feedback. Based on the evaluation results, the server generates feedback that includes specific areas for improvement. This feedback is sent to the terminal and provided to the user. As output, the user is presented with areas for improvement and advice.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] Sales staff in physical stores need to quickly improve their specialized product knowledge and customer service skills. However, traditional training methods lacked opportunities for practical, scenario-based training, making it difficult to acquire these skills efficiently.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for generating a scenario based on information input by the user, means for conducting a dialogue based on the generated scenario, and means for recording the content of the dialogue as audio data. This enables sales staff to receive practical training tailored to the conditions of a physical store.
[0101] A "user" is an individual who logs into the system, selects products and information, and participates in training.
[0102] A "scenario" is a virtual flow of dialogue, including conversations and questions related to the product or service, that is generated to support user training.
[0103] "Dialogue" refers to two-way communication between a user and AI, conducted via voice or text.
[0104] "Audio data" refers to digital audio information that records a user's speech.
[0105] "Evaluation" is the process of analyzing user responses and making judgments based on their accuracy and specificity.
[0106] "Feedback" refers to information based on evaluation results, including suggestions for improvement and advice regarding user responses.
[0107] "Speech synthesis" is a technology that generates speech based on text information and is used when generating questions from a virtual customer.
[0108] "Customer service skills" refer to the ability to provide appropriate information in response to customer requests and to build relationships.
[0109] "Artificial intelligence" refers to the intelligent behavior and learning abilities that computer programs simulate, and in dialogue scenarios, it takes on the role of the customer.
[0110] The system for realizing this application consists of a user terminal, a server, and artificial intelligence functions. The user uses a device such as a smartphone or tablet to log in to the system and select the product category to train. Based on the information received, the server generates a conversational scenario suitable for the user using a generative AI model.
[0111] Based on the generated scenario, the server utilizes speech synthesis to create questions from a virtual customer's perspective and outputs them as audio to the user's terminal. The user responds to these questions, and the responses are recorded as audio data on the terminal and sent to the server. The server uses speech recognition technology to convert the audio data into text and analyzes the content of the responses.
[0112] Based on the analysis results, the server evaluates the user's response and generates specific feedback, including areas for improvement. This feedback is then sent back to the user's terminal and presented as part of their training. This allows the user to acquire more practical customer service skills.
[0113] As a concrete example, consider a scenario where a user receives the prompt, "Explain to the customer how to choose colors and materials for the new smartphone case." Through responding to this question, the user can gain practical experience in explaining product features in detail. Through this process, salespeople can naturally improve their advanced customer service skills.
[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0115] Step 1:
[0116] The user logs into the system using a terminal and selects the product category they wish to train in. Inputs include the user ID, password, and the selected product category. The server outputs this selection information for the next processing step.
[0117] Step 2:
[0118] The server uses a generative AI model to generate dialogue scenarios based on the received product category information. This involves setting up scenarios based on typical customer questions and scenario-based situations. The input is product category information, and the output is the generated scenario.
[0119] Step 3:
[0120] Based on the generated scenario, the server uses speech synthesis to create voice-based questions from a virtual customer and sends them to the terminal. The input is scenario information, and the output is voice data.
[0121] Step 4:
[0122] The user responds to questions from a virtual customer presented as audio from the terminal. The terminal records the user's responses as audio data. The input is the audio questions, and the output is the user's response audio data.
[0123] Step 5:
[0124] The terminal sends the recorded audio data to the server. The input is the user's response audio data, and the output is the data transferred to the server.
[0125] Step 6:
[0126] The server uses speech recognition technology to convert the audio data into text data and then performs analysis. This analysis evaluates the accuracy and specificity of the user's responses. The input is the user's audio response data, and the output is text data and the evaluation results.
[0127] Step 7:
[0128] The server generates feedback based on the evaluation results. This feedback includes areas for improvement and specific advice. The input is the evaluation result, and the output is the feedback text.
[0129] Step 8:
[0130] The server sends the generated feedback to the terminal and presents it to the user. The input is the text of the feedback, and the output is its display on the user's terminal.
[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0132] The present invention is an interactive training system for inexperienced engineers to acquire practical skills, and in particular includes technology that recognizes user emotions and adjusts responses accordingly. Embodiments of the present invention are described in detail below.
[0133] System Configuration
[0134] This system consists of a user terminal, a server that generates and manages scenarios, an emotion engine that performs emotion recognition, and an AI that acts as the customer. The emotion engine evaluates the user's emotional state in real time and dynamically adjusts the feedback and scenarios.
[0135] Program Processing Overview
[0136] User: Log in to the terminal, select the product to be trained on, and enter the information. This information is immediately sent to the server.
[0137] Server: Based on the received product information, it generates training scenarios for the AI to use. The emotion engine is also initialized and prepared to analyze the user's voice data and speech characteristics.
[0138] AI: Following a scenario generated on the server, the AI begins interacting with the user as a customer. The AI tests the user's understanding by asking questions about the product.
[0139] Emotion Engine: Analyzes user speech data in real time to estimate the user's emotional state based on voice tone, speech rate, pronunciation characteristics, etc.
[0140] Terminal: During this dialogue process, all user responses are recorded as audio data, and the results of the emotion engine's analysis are stored on the server.
[0141] Server: Analyzes the results of the emotion engine and the dialogue content, and generates feedback based on the results. The feedback content is adjusted to match the user's emotional state.
[0142] Feedback Provision: Adjusted feedback is presented to the user via the device, and the user uses it to improve their next training session.
[0143] Specific example
[0144] For example, if a user is undergoing training on a "network security platform," the AI might ask, "Could you explain the security protocols of this platform?" If the user is nervous and their voice is trembling, the emotion engine recognizes this emotion as "nervousness." In response, the server can adjust the feedback, providing softer language such as, "It's okay to answer more relaxed; the goal is to deepen your understanding," thereby supporting the user's confidence. In this way, the system can leverage real-time emotion recognition to provide a more personalized training experience.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] The user logs into the terminal, selects and enters information about the product they wish to train on. The terminal then sends the entered information to the server.
[0148] Step 2:
[0149] Based on the product information received by the server, the AI generates a dialogue scenario. This scenario includes typical customer questions and situations related to the product.
[0150] Step 3:
[0151] The user initiates a dialogue session with the AI based on instructions from their device. The AI, acting as a customer, asks the user questions about the product, following a scenario generated on the server.
[0152] Step 4:
[0153] The emotion engine analyzes the user's voice in real time. It analyzes tone, tempo, and voice intensity from the audio data to infer the user's emotional state.
[0154] Step 5:
[0155] The user answers the AI's questions. The device records these responses as audio data and sends it to the server. The server also stores the analysis data from the emotion engine.
[0156] Step 6:
[0157] The server analyzes recorded audio data and emotional states, and provides an evaluation based on the user's responses and emotions. Natural language processing techniques are used for detailed analysis.
[0158] Step 7:
[0159] The server generates feedback based on the analysis results, tailored to the user's emotional state. This feedback includes areas for improvement and positive reinforcement.
[0160] Step 8:
[0161] The device displays the generated feedback to the user. The user can then use this feedback to further improve their training.
[0162] (Example 2)
[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0164] Traditional training systems have a problem in that they provide uniform feedback without considering the user's emotional state, resulting in insufficient individual skill improvement and deepened understanding. Furthermore, the automatic generation of dialogue scenarios based on user-selected items is often inefficient, limiting the practical training effectiveness.
[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0166] In this invention, the server includes emotion recognition means for analyzing the user's emotional state in real time, means for adjusting feedback according to the emotional state, and means for generating dialogue scenarios based on information input from the user. This enables personalized feedback tailored to each user's emotional state and specific dialogue training suitable for the selected item.
[0167] A "user" refers to an individual who uses the system and is the recipient of individually provided services and feedback.
[0168] "Information" refers to data related to products and goods that users input into the system. This information is used to generate dialogue scenarios.
[0169] A "scenario" refers to the plan and flow of the dialogue between the user and artificial intelligence, and its pre-generation enables smoother interactions.
[0170] "Dialogue" refers to two-way communication between the user and artificial intelligence, and is a means of evaluating the user's level of understanding and skills.
[0171] "Voice data" refers to digital recordings of user speech information, which are used for analysis and evaluation.
[0172] "Emotion recognition" refers to the process of estimating a user's emotional state by analyzing their voice data.
[0173] "Feedback" refers to evaluations and advice provided to users based on the results of dialogue and sentiment analysis.
[0174] "Adjustment" refers to optimizing the content of feedback according to the user's emotional state, and is a means of providing a personalized training experience.
[0175] "Artificial intelligence" refers to an automated system that acts as a customer in conversations with users, assisting in user training through questions and feedback.
[0176] The system according to the present invention provides a training environment in which inexperienced engineers can acquire practical skills by dynamically adjusting their responses while recognizing the user's emotional state. Specific embodiments of the present invention are described below.
[0177] The user first logs into the system using a terminal and enters information about the item they wish to train. This information is immediately sent to the server. The server uses a generative AI model to generate dialogue scenarios based on the received information. The server also initializes an emotion engine and prepares to analyze the user's voice data in real time. This emotion engine has the ability to analyze the user's voice tone, speaking speed, and pronunciation patterns to estimate their emotional state.
[0178] Following the generated scenario, the artificial intelligence takes on the role of a customer and initiates a conversation with the user. During the conversation, the user answers questions about the items, and speech data is collected as knowledge is verified. The device records this data and sends the results of the emotion engine's analysis to the server to understand the user's emotional state.
[0179] The server generates feedback based on collected data and sentiment assessment. This feedback is tailored to the user's emotions and delivered to the user via their device. This allows the user to understand areas for improvement needed in their next training session and use that information to enhance their skills.
[0180] As a concrete example, let's assume a user is undergoing training on a "network security platform." The artificial intelligence asks, "Could you explain the security protocols of this platform?" If the emotion engine detects that the user is nervous, the server provides feedback such as, "You can answer more relaxed; the goal is to deepen your understanding."
[0181] An example of a prompt would be, "As a customer, ask the user questions about the network security platform. Use the emotion engine to determine if the user is nervous and adjust your feedback accordingly." This allows the system to create a personalized and effective training experience for the user.
[0182] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0183] Step 1:
[0184] The user logs into the terminal. This prompts the terminal to input the user's authentication information. The terminal uses this information to authenticate the user and displays a dashboard that the user can access. The output is a training dashboard.
[0185] Step 2:
[0186] The user selects information about the item they wish to train on and enters it via a terminal. This information constitutes the input. The terminal sends this information to the server, which uses it as data necessary to generate the dialogue scenario. The output is the transmission of product information to the server.
[0187] Step 3:
[0188] The server uses an AI model based on the received item information to generate a dialogue scenario. The input is the product information received from the user, and the output is the generated dialogue scenario. The server provides this scenario to the artificial intelligence, and the dialogue is ready.
[0189] Step 4:
[0190] The server initializes the emotion engine. The input is the speech analysis model and algorithms necessary for the dialogue. The emotion engine analyzes the user's voice data in real time and prepares to estimate the emotional state. The output is the ready emotion engine.
[0191] Step 5:
[0192] The artificial intelligence, following a generated dialogue scenario, begins a conversation with the user as a customer. The input consists of the generated scenario and the user's utterances, while the output consists of questions posed to the user and the progress of the conversation. The device displays this dialogue on the screen, and the user responds.
[0193] Step 6:
[0194] The emotion engine analyzes user speech data in real time. The input is user voice data, and data processing includes analysis of voice tone, speech rate, and pronunciation characteristics. The output is an estimated result of the user's emotional state.
[0195] Step 7:
[0196] The terminal records the user's responses as audio data and sends it to the server along with the analysis results from the emotion engine. The input is the speech data and analysis results, and the output is the recorded dialogue data and emotional state.
[0197] Step 8:
[0198] The server generates feedback using recorded dialogue content and sentiment analysis results. The input consists of dialogue data and emotional state, and the feedback generation includes adjustments based on the user's emotional state. The output is the generated feedback.
[0199] Step 9:
[0200] The terminal displays feedback sent from the server to the user. The input is adjusted feedback, and the output provides information that the user can use to improve their next training session.
[0201] (Application Example 2)
[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0203] When inexperienced engineers efficiently acquire practical skills, there are challenges in providing emotionally responsive feedback and personalized training tailored to individual characteristics. Furthermore, to enhance the effectiveness of training, there is a need for appropriate real-time analysis of emotional states and the provision of feedback based on that analysis.
[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0205] In this invention, the server includes means for generating a scenario based on information input by the user, means for engaging in dialogue based on the generated scenario, and means for estimating the user's emotional state in real time from the user's voice. This makes it possible for inexperienced technicians to effectively acquire practical skills while receiving feedback tailored to their own emotional state.
[0206] A "user" refers to an individual who inputs information and participates in dialogue in order to acquire practical skills using a training system.
[0207] "Means for generating scenarios" refers to a function that automatically creates the flow and content of conversations used for training based on information entered by the user.
[0208] "Means of dialogue" refers to functions for performing communication with users based on generated scenarios.
[0209] "Means of recording as audio data" refers to a system that saves the content of conversations with the user in audio format.
[0210] "Means of analysis and evaluation" refers to a system function that analyzes recorded audio data and evaluates the user's responses and emotions.
[0211] "Means of generating feedback" refers to a function that creates information to provide users with suggestions for improvement and advice based on evaluation results.
[0212] "Means of adjusting feedback" refers to functions that optimize feedback content according to the user's emotional state and provide it in an appropriate format.
[0213] "Means for securely managing information data" refers to functions that protect user and system data and safeguard it from unauthorized access and leakage.
[0214] "Methods for estimating emotional states in real time" refers to a function that analyzes the user's voice and instantly recognizes and judges their emotions during a conversation.
[0215] The system for realizing this invention mainly consists of a server, a user terminal, an emotion recognition engine, and a conversational AI. Each of these elements is described in detail below.
[0216] The user first accesses the application on their device, logs in, and begins training. The device used is a smartphone, and the microphone and camera capture the user's voice and video. In the initial stages of training, the user selects product candidates and enters initial information about them. All of this information is sent to the server.
[0217] The servers are hosted in a cloud environment and use AI technology to generate scenarios. Specifically, an AI training model is used to build individual conversation scenarios based on each user's choices. In these scenarios, the AI takes on the role of a virtual customer and initiates communication with the user.
[0218] The emotion recognition engine first analyzes the user's voice data in real time. For example, it uses the Google® Speech-to-Text API to convert the speech to text and estimates emotions by analyzing speech speed and tone. The estimated emotional state is instantly sent to the server and used as basic data for generating feedback.
[0219] The feedback generation system creates personalized feedback based on analysis results. This feedback is sent to the user's device at the appropriate time and presented to the user visually and audibly. This allows users to flexibly adjust their responses and acquire effective skill sets.
[0220] For example, if a user chooses a setting to train on "fine wines," the AI might ask, "Can you describe the region where this wine is produced and its aromatic characteristics?" If the emotion recognition engine determines the user is "anxious" due to an unconvincing answer, the server will generate feedback such as, "Let's try to calm down and speak slowly. It would be good to pick out a few key points and explain them in order."
[0221] The following are examples of prompts for the generated AI model:
[0222] "Analyze user voice data in real time and estimate their emotional state. When a user says, 'Thank you for visiting. What products are you looking for today?', evaluate their emotion based on their tone and speed of voice, and generate appropriate feedback."
[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0224] Step 1:
[0225] The user logs into their device, selects the product they want to train on, enters the information, and sends it to the server. The entered product information is used as basic data to generate the user's training scenario.
[0226] Step 2:
[0227] The server generates dialogue scenarios using a generative AI model based on the received product information. This generation process constructs dialogue patterns for product-related questions and virtual customers, and provides the generated scenarios to the AI.
[0228] Step 3:
[0229] Based on the dialogue scenario received from the server, the AI initiates questions about the product to the user, acting as a virtual customer. The user's responses are sent to the server in real time and recorded as audio data.
[0230] Step 4:
[0231] The emotion recognition engine analyzes the user's voice data in real time. The input voice is converted to text via the Google Speech-to-Text API, and the user's emotional state is estimated by analyzing the text and acoustic features.
[0232] Step 5:
[0233] Based on the output of the emotion recognition engine, the server generates feedback adapted to the user's emotional state. This feedback generation process concretizes improvement suggestions and advice tailored to the user's emotional state.
[0234] Step 6:
[0235] The server sends the generated feedback to the terminal, which then presents the feedback to the user visually or audibly. The user can then use the received feedback to review their responses in subsequent interactions.
[0236] Step 7:
[0237] Users utilize the feedback they receive to improve their skills and initiate new dialogue sessions. Training progresses by repeating this cycle, promoting effective skill acquisition.
[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0241] [Second Embodiment]
[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0254] This invention is a system for improving the practical skills of inexperienced engineers regarding commercial products, and it uses AI to simulate customer interactions and provide feedback. Embodiments of this invention are described below.
[0255] System Configuration
[0256] The system consists of a terminal used by the user, a server that generates scenarios and processes information, and an AI that simulates dialogue.
[0257] Program Processing Overview
[0258] User: Log in to the system via your device, select the product you want to train on, and enter its details.
[0259] Server: Based on information received from the user, the system generates scenarios for use by the AI. In this process, the system prepares common questions and problem situations related to the product.
[0260] AI: Based on the generated scenario, it begins a conversation with the user as a customer. The AI asks questions related to the product and waits for the user's response.
[0261] Terminal: User responses are recorded in audio format, and the data is sent to the server.
[0262] Server: Analyzes voice data and evaluates the user's response. This evaluation is based on the accuracy and specificity of the response.
[0263] Feedback Provision: The server uses the analyzed data to generate feedback, including specific areas for improvement, and presents it to the user via the terminal.
[0264] Specific example
[0265] For example, if a user requests training on "cloud storage services," the AI, based on a scenario, will ask questions such as, "Please explain the backup function of cloud storage in detail." The user will answer, and their answer will be recorded. The server will analyze the recording and generate feedback such as, "The explanation is abstract; it would be better to mention specific steps," which will then be provided to the user via their device. Through this process, users can develop more specific and practical response skills.
[0266] The following describes the processing flow.
[0267] Step 1:
[0268] The user logs into the device and enters information about the product they wish to train on. The device receives this information and sends it to the server.
[0269] Step 2:
[0270] The server generates a scenario for use by the AI based on the product information it receives. The scenario includes product features, common problems, and anticipated customer questions. The generated scenario is then sent to the terminal.
[0271] Step 3:
[0272] The user initiates a conversational session with the AI through their device. Based on a scenario received from the server, the AI, acting as a customer, asks the user questions.
[0273] Step 4:
[0274] The user responds to questions from the AI. This interaction is recorded as audio, and the recorded data is sent from the device to the server.
[0275] Step 5:
[0276] The server analyzes the voice data. Specifically, it uses natural language processing technology to evaluate the accuracy, speed, and specificity of the user's responses.
[0277] Step 6:
[0278] The server generates specific feedback based on the analysis results. This feedback includes guidelines on what areas the user should improve.
[0279] Step 7:
[0280] The device receives feedback from the server and displays it to the user. The user can use this as a reference to help with their next learning.
[0281] (Example 1)
[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0283] It is difficult for an inexperienced engineer to effectively improve practical skills related to a commercial product in a short period of time. In conventional training methods, it is difficult to provide a realistic customer response simulation, and the opportunity to obtain specific feedback is limited, so efficient skill improvement has been hindered.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0285] In this invention, the server includes means for generating a scenario based on information input from a user, means for conducting a dialogue based on the generated scenario, and means for generating general questions and problem situations related to a commercial product using a generation AI model. Thereby, a realistic customer response simulation is provided, and efficient skill improvement is possible through specific and useful feedback to the user.
[0286] The "user" refers to an individual or group that receives training related to a commercial product using the system.
[0287] The "scenario" represents a series of questions and situation settings prepared by the generation AI model to constitute the flow of dialogue with the user.
[0288] The "dialogue" refers to a series of communication processes in which the AI asks the user questions related to the commercial product and the user responds to the questions.
[0289] The "voice data" refers to digital data in voice format used to record the response content of the user.
[0290] "Evaluation" refers to the process of analyzing user responses based on recorded audio data to measure their accuracy and specificity.
[0291] "Feedback" refers to information generated based on evaluation results, including specific areas for improvement and advice regarding the response.
[0292] A "generative AI model" refers to an artificial intelligence model used to generate questions and problem situations related to a product or service.
[0293] A "terminal" is a device used by a user to access a system, and it has functions for information input and voice recording.
[0294] A "server" refers to a central computing device that handles tasks such as scenario generation, dialogue management, audio data analysis, and feedback generation.
[0295] This invention is designed to help inexperienced engineers improve their practical skills with specific products. The system consists of a terminal accessed by the user, a server for information processing, and an AI for managing the interaction. The user logs into the system through the terminal and selects the product to be trained on. The terminal transmits detailed information about the product entered by the user to the server. Based on this information, the server generates an interaction scenario using a generative AI model.
[0296] Based on the generated scenario, the AI simulates a realistic conversation with the user. The AI creates questions about the product and asks them sequentially to the user. Specific questions might include things like, "Please explain the cloud storage backup function in detail."
[0297] The device has the function of recording the user's voice responses in real time and sending that data to the server. The server uses speech recognition software to analyze this voice data and evaluate the accuracy and specificity of the user's responses. Based on this evaluation, the server generates feedback that includes specific areas for improvement and provides this information to the user through the device.
[0298] As a concrete example in this system, if a user requests training on cloud storage services, the following prompt message would be used: "We want to simulate explaining cloud storage services to a customer. Please create a scenario where the user requests a specific explanation about the backup function."
[0299] This procedure allows users to efficiently acquire practical and specific skills. Furthermore, by using a generative AI model, this system can flexibly design product-specific questions and simulations, making it suitable for a wide variety of products.
[0300] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0301] Step 1:
[0302] The user logs into the system. The user uses their device to enter their login information and access the system. The server receives the entered user information and performs authentication. If successful, the user is redirected to the home screen.
[0303] Step 2:
[0304] The user selects a product. The user chooses the product they want to train on and enters its details into the terminal. The server then queries the database for product information and receives information about the product specified by the user.
[0305] Step 3:
[0306] The server generates a scenario. Based on the received product information, it uses a generation AI model to generate a scenario that includes common questions and problem situations related to the product. Specifically, it inputs a prompt sentence into the AI model and obtains the scenario as the output. This scenario includes a list of questions related to the product.
[0307] Step 4:
[0308] The AI starts the conversation. Based on the generated scenario, the AI sequentially presents questions about the product to the user. The AI selects a question from the scenario, sends it to the terminal in voice or text form, and waits for the user's response.
[0309] Step 5:
[0310] The terminal records the user's response. When the user responds to the AI's question by voice, the terminal records the voice in real time and sends it to the server as digital voice data. The recorded voice data is obtained as the output.
[0311] Step 6:
[0312] The server analyzes the voice data. The server uses voice recognition software to analyze the received voice data and convert it into text. Based on this text data, it evaluates the accuracy and specificity of the response. Here, an evaluation score and improvement points are output as the analysis result.
[0313] Step 7:
[0314] The server generates and provides feedback. Based on the evaluation result, the server generates feedback that includes specific improvement points. It sends the feedback to the terminal and provides it to the user. As the output, improvement points and advice are presented to the user. <000Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0317] Sales staff in physical stores need to quickly improve their specialized product knowledge and customer service skills. However, traditional training methods lacked opportunities for practical, scenario-based training, making it difficult to acquire these skills efficiently.
[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0319] In this invention, the server includes means for generating a scenario based on information input by the user, means for conducting a dialogue based on the generated scenario, and means for recording the content of the dialogue as audio data. This enables sales staff to receive practical training tailored to the conditions of a physical store.
[0320] A "user" is an individual who logs into the system, selects products and information, and participates in training.
[0321] A "scenario" is a virtual flow of dialogue, including conversations and questions related to the product or service, that is generated to support user training.
[0322] "Dialogue" refers to two-way communication between a user and AI, conducted via voice or text.
[0323] "Audio data" refers to digital audio information that records a user's speech.
[0324] "Evaluation" is the process of analyzing user responses and making judgments based on their accuracy and specificity.
[0325] "Feedback" refers to information based on evaluation results, including suggestions for improvement and advice regarding user responses.
[0326] "Speech synthesis" is a technology that generates speech based on text information and is used when generating questions from a virtual customer.
[0327] "Customer service skills" refer to the ability to provide appropriate information in response to customer requests and to build relationships.
[0328] "Artificial intelligence" refers to the intelligent behavior and learning abilities that computer programs simulate, and in dialogue scenarios, it takes on the role of the customer.
[0329] The system for realizing this application consists of a user terminal, a server, and artificial intelligence functions. The user uses a device such as a smartphone or tablet to log in to the system and select the product category to train. Based on the information received, the server generates a conversational scenario suitable for the user using a generative AI model.
[0330] Based on the generated scenario, the server utilizes speech synthesis to create questions from a virtual customer's perspective and outputs them as audio to the user's terminal. The user responds to these questions, and the responses are recorded as audio data on the terminal and sent to the server. The server uses speech recognition technology to convert the audio data into text and analyzes the content of the responses.
[0331] Based on the analysis results, the server evaluates the user's response and generates specific feedback, including areas for improvement. This feedback is then sent back to the user's terminal and presented as part of their training. This allows the user to acquire more practical customer service skills.
[0332] As a concrete example, consider a scenario where a user receives the prompt, "Explain to the customer how to choose colors and materials for the new smartphone case." Through responding to this question, the user can gain practical experience in explaining product features in detail. Through this process, salespeople can naturally improve their advanced customer service skills.
[0333] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0334] Step 1:
[0335] The user logs into the system using a terminal and selects the product category they wish to train in. Inputs include the user ID, password, and the selected product category. The server outputs this selection information for the next processing step.
[0336] Step 2:
[0337] The server uses a generative AI model to generate dialogue scenarios based on the received product category information. This involves setting up scenarios based on typical customer questions and scenario-based situations. The input is product category information, and the output is the generated scenario.
[0338] Step 3:
[0339] Based on the generated scenario, the server uses speech synthesis to create questions in voice format as a virtual customer and sends them to the terminal as output. The input is scenario information, and the output is voice data.
[0340] Step 4:
[0341] The user responds to questions from a virtual customer presented as audio from the terminal. The terminal records the user's responses as audio data. The input is the audio questions, and the output is the user's response audio data.
[0342] Step 5:
[0343] The terminal sends the recorded audio data to the server. The input is the user's response audio data, and the output is the data transferred to the server.
[0344] Step 6:
[0345] The server uses speech recognition technology to convert audio data into text data and performs analysis. This analysis evaluates the accuracy and specificity of the user's responses. The input is the user's audio response data, and the output is text data and the evaluation results.
[0346] Step 7:
[0347] The server generates feedback based on the evaluation results. This feedback includes areas for improvement and specific advice. The input is the evaluation result, and the output is the feedback text.
[0348] Step 8:
[0349] The server sends the generated feedback to the terminal and presents it to the user. The input is the text of the feedback, and the output is its display on the user's terminal.
[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0351] The present invention is an interactive training system for inexperienced engineers to acquire practical skills, and in particular includes technology that recognizes user emotions and adjusts responses accordingly. Embodiments of the present invention are described in detail below.
[0352] System Configuration
[0353] This system consists of a user terminal, a server that generates and manages scenarios, an emotion engine that performs emotion recognition, and an AI that acts as the customer. The emotion engine evaluates the user's emotional state in real time and dynamically adjusts the feedback and scenarios.
[0354] Program Processing Overview
[0355] User: Log in to the terminal, select the product to be trained on, and enter the information. This information is immediately sent to the server.
[0356] Server: Based on the received product information, it generates training scenarios for the AI to use. The emotion engine is also initialized and prepared to analyze the user's voice data and speech characteristics.
[0357] AI: Following a scenario generated on the server, the AI begins interacting with the user as a customer. The AI tests the user's understanding by asking questions about the product.
[0358] Emotion Engine: Analyzes user speech data in real time to estimate the user's emotional state based on voice tone, speech rate, pronunciation characteristics, etc.
[0359] Terminal: During this dialogue process, all user responses are recorded as audio data, and the results of the emotion engine's analysis are stored on the server.
[0360] Server: Analyzes the results of the emotion engine and the dialogue content, and generates feedback based on the results. The feedback content is adjusted to match the user's emotional state.
[0361] Feedback Provision: Adjusted feedback is presented to the user via the device, and the user uses it to improve their next training session.
[0362] Specific example
[0363] For example, if a user is undergoing training on a "network security platform," the AI might ask, "Could you explain the security protocols of this platform?" If the user is nervous and their voice is trembling, the emotion engine recognizes this emotion as "nervousness." In response, the server can adjust the feedback, providing softer language such as, "It's okay to answer more relaxed; the goal is to deepen your understanding," thereby supporting the user's confidence. In this way, the system can leverage real-time emotion recognition to provide a more personalized training experience.
[0364] The following describes the processing flow.
[0365] Step 1:
[0366] The user logs into the terminal, selects and enters information about the product they wish to train on. The terminal then sends the entered information to the server.
[0367] Step 2:
[0368] Based on the product information received by the server, the AI generates a dialogue scenario. This scenario includes typical customer questions and situations related to the product.
[0369] Step 3:
[0370] The user initiates a dialogue session with the AI based on instructions from their device. The AI, acting as a customer, asks the user questions about the product, following a scenario generated on the server.
[0371] Step 4:
[0372] The emotion engine analyzes the user's voice in real time. It analyzes tone, tempo, and voice intensity from the audio data to infer the user's emotional state.
[0373] Step 5:
[0374] The user answers the AI's questions. The device records these responses as audio data and sends it to the server. The server also stores the analysis data from the emotion engine.
[0375] Step 6:
[0376] The server analyzes recorded audio data and emotional states, and provides an evaluation based on the user's responses and emotions. Natural language processing techniques are used for detailed analysis.
[0377] Step 7:
[0378] The server generates feedback based on the analysis results, tailored to the user's emotional state. This feedback includes areas for improvement and positive reinforcement.
[0379] Step 8:
[0380] The device displays the generated feedback to the user. The user can then use this feedback to further improve their training.
[0381] (Example 2)
[0382] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0383] Traditional training systems have a problem in that they provide uniform feedback without considering the user's emotional state, resulting in insufficient individual skill improvement and deepened understanding. Furthermore, the automatic generation of dialogue scenarios based on user-selected items is often inefficient, limiting the practical training effectiveness.
[0384] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0385] In this invention, the server includes emotion recognition means for analyzing the user's emotional state in real time, means for adjusting feedback according to the emotional state, and means for generating dialogue scenarios based on information input from the user. This enables personalized feedback tailored to each user's emotional state and specific dialogue training suitable for the selected item.
[0386] A "user" refers to an individual who uses the system and is the recipient of individually provided services and feedback.
[0387] "Information" refers to data related to products and goods that users input into the system. This information is used to generate dialogue scenarios.
[0388] A "scenario" refers to the plan and flow of the dialogue between the user and artificial intelligence, and its pre-generation enables smoother interactions.
[0389] "Dialogue" refers to two-way communication between the user and artificial intelligence, and is a means of evaluating the user's level of understanding and skills.
[0390] "Voice data" refers to digital recordings of user speech information, which are used for analysis and evaluation.
[0391] "Emotion recognition" refers to the process of estimating a user's emotional state by analyzing their voice data.
[0392] "Feedback" refers to evaluations and advice provided to users based on the results of dialogue and sentiment analysis.
[0393] "Adjustment" refers to optimizing the content of feedback according to the user's emotional state, and is a means of providing a personalized training experience.
[0394] "Artificial intelligence" refers to an automated system that acts as a customer in conversations with users, assisting in user training through questions and feedback.
[0395] The system according to the present invention provides a training environment in which inexperienced engineers can acquire practical skills by dynamically adjusting their responses while recognizing the user's emotional state. Specific embodiments of the present invention are described below.
[0396] The user first logs into the system using a terminal and enters information about the item they wish to train. This information is immediately sent to the server. The server uses a generative AI model to generate dialogue scenarios based on the received information. The server also initializes an emotion engine and prepares to analyze the user's voice data in real time. This emotion engine has the ability to analyze the user's voice tone, speaking speed, and pronunciation patterns to estimate their emotional state.
[0397] Following the generated scenario, the artificial intelligence takes on the role of a customer and initiates a conversation with the user. During the conversation, the user answers questions about the items, and speech data is collected as knowledge is verified. The device records this data and sends the results of the emotion engine's analysis to the server to understand the user's emotional state.
[0398] The server generates feedback based on collected data and sentiment assessment. This feedback is tailored to the user's emotions and delivered to the user via their device. This allows the user to understand areas for improvement needed in their next training session and use that information to enhance their skills.
[0399] As a concrete example, let's assume a user is undergoing training on a "network security platform." The artificial intelligence asks, "Could you explain the security protocols of this platform?" If the emotion engine detects that the user is nervous, the server provides feedback such as, "You can answer more relaxed; the goal is to deepen your understanding."
[0400] An example of a prompt would be, "As a customer, ask the user questions about the network security platform. Use the emotion engine to determine if the user is nervous and adjust your feedback accordingly." This allows the system to create a personalized and effective training experience for the user.
[0401] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0402] Step 1:
[0403] The user logs into the terminal. This prompts the terminal to input the user's authentication information. The terminal uses this information to authenticate the user and displays a dashboard that the user can access. The output is a training dashboard.
[0404] Step 2:
[0405] The user selects information about the item they wish to train on and enters it via a terminal. This information constitutes the input. The terminal sends this information to the server, which uses it as data necessary to generate the dialogue scenario. The output is the transmission of product information to the server.
[0406] Step 3:
[0407] The server uses an AI model based on the received item information to generate a dialogue scenario. The input is the product information received from the user, and the output is the generated dialogue scenario. The server provides this scenario to the artificial intelligence, and the dialogue is ready.
[0408] Step 4:
[0409] The server initializes the emotion engine. The input is the speech analysis model and algorithms necessary for the dialogue. The emotion engine analyzes the user's voice data in real time and prepares to estimate the emotional state. The output is the ready emotion engine.
[0410] Step 5:
[0411] The artificial intelligence, following a generated dialogue scenario, begins a conversation with the user as a customer. The input consists of the generated scenario and the user's utterances, while the output consists of questions posed to the user and the progress of the conversation. The device displays this dialogue on the screen, and the user responds.
[0412] Step 6:
[0413] The emotion engine analyzes user speech data in real time. The input is user voice data, and data processing includes analysis of voice tone, speech rate, and pronunciation features. The output is an estimated result of the user's emotional state.
[0414] Step 7:
[0415] The terminal records the user's responses as audio data and sends it to the server along with the analysis results from the emotion engine. The input is the speech data and analysis results, and the output is the recorded dialogue data and emotional state.
[0416] Step 8:
[0417] The server generates feedback using recorded dialogue content and sentiment analysis results. The input consists of dialogue data and emotional state, and the feedback generation includes adjustments based on the user's emotional state. The output is the generated feedback.
[0418] Step 9:
[0419] The terminal displays feedback sent from the server to the user. The input is adjusted feedback, and the output provides information that the user can use to improve their next training session.
[0420] (Application Example 2)
[0421] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0422] When inexperienced engineers efficiently acquire practical skills, there are challenges in providing emotionally responsive feedback and personalized training tailored to individual characteristics. Furthermore, to enhance the effectiveness of training, there is a need for appropriate real-time analysis of emotional states and the provision of feedback based on that analysis.
[0423] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0424] In this invention, the server includes means for generating a scenario based on information input by the user, means for engaging in dialogue based on the generated scenario, and means for estimating the user's emotional state in real time from the user's voice. This makes it possible for inexperienced technicians to effectively acquire practical skills while receiving feedback tailored to their own emotional state.
[0425] A "user" refers to an individual who inputs information and participates in dialogue in order to acquire practical skills using a training system.
[0426] "Means for generating scenarios" refers to a function that automatically creates the flow and content of conversations used for training based on information entered by the user.
[0427] "Means of dialogue" refers to functions for performing communication with users based on generated scenarios.
[0428] "Means of recording as audio data" refers to a system that saves the content of conversations with the user in audio format.
[0429] "Means of analysis and evaluation" refers to a system function that analyzes recorded audio data and evaluates the user's responses and emotions.
[0430] "Means of generating feedback" refers to a function that creates information to provide users with suggestions for improvement and advice based on evaluation results.
[0431] "Means of adjusting feedback" refers to functions that optimize feedback content according to the user's emotional state and provide it in an appropriate format.
[0432] "Means for securely managing information data" refers to functions that protect user and system data and safeguard it from unauthorized access and leakage.
[0433] "Methods for estimating emotional states in real time" refers to a function that analyzes the user's voice and instantly recognizes and judges their emotions during a conversation.
[0434] The system for realizing this invention mainly consists of a server, a user terminal, an emotion recognition engine, and a conversational AI. Each of these elements is described in detail below.
[0435] The user first accesses the application on their device, logs in, and begins training. The device used is a smartphone, and the microphone and camera capture the user's voice and video. In the initial stages of training, the user selects product candidates and enters initial information about them. All of this information is sent to the server.
[0436] The servers are hosted in a cloud environment and use AI technology to generate scenarios. Specifically, an AI training model is used to build individual conversation scenarios based on each user's choices. In these scenarios, the AI takes on the role of a virtual customer and initiates communication with the user.
[0437] The emotion recognition engine first analyzes the user's voice data in real time. For example, it uses the Google Speech-to-Text API to convert the speech to text and estimates emotions by analyzing speech speed and tone. The estimated emotional state is instantly sent to the server and used as basic data for generating feedback.
[0438] The feedback generation system creates personalized feedback based on analysis results. This feedback is sent to the user's device at the appropriate time and presented to the user visually and audibly. This allows users to flexibly adjust their responses and acquire effective skill sets.
[0439] For example, if a user chooses a setting to train on "fine wines," the AI might ask, "Can you describe the region where this wine is produced and its aromatic characteristics?" If the emotion recognition engine determines the user is "anxious" due to an unconvincing answer, the server will generate feedback such as, "Let's try to calm down and speak slowly. It would be good to pick out a few key points and explain them in order."
[0440] The following are examples of prompts for the generated AI model:
[0441] "Analyze user voice data in real time and estimate their emotional state. When a user says, 'Thank you for visiting. What products are you looking for today?', evaluate their emotion based on their tone and speed of voice, and generate appropriate feedback."
[0442] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0443] Step 1:
[0444] The user logs into their device, selects the product they want to train on, enters the information, and sends it to the server. The entered product information is used as basic data to generate the user's training scenario.
[0445] Step 2:
[0446] The server generates dialogue scenarios using a generative AI model based on the received product information. This generation process constructs dialogue patterns for product-related questions and virtual customers, and provides the generated scenarios to the AI.
[0447] Step 3:
[0448] Based on the dialogue scenario received from the server, the AI initiates questions about the product to the user, acting as a virtual customer. The user's responses are sent to the server in real time and recorded as audio data.
[0449] Step 4:
[0450] The emotion recognition engine analyzes the user's voice data in real time. The input voice is converted to text via the Google Speech-to-Text API, and the user's emotional state is estimated by analyzing the text and acoustic features.
[0451] Step 5:
[0452] Based on the output of the emotion recognition engine, the server generates feedback adapted to the user's emotional state. This feedback generation process concretizes improvement suggestions and advice tailored to the user's emotional state.
[0453] Step 6:
[0454] The server sends the generated feedback to the terminal, which then presents the feedback to the user visually or audibly. The user can then use the received feedback to review their responses in subsequent interactions.
[0455] Step 7:
[0456] Users utilize the feedback they receive to improve their skills and initiate new dialogue sessions. Training progresses by repeating this cycle, promoting effective skill acquisition.
[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0460] [Third Embodiment]
[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0473] This invention is a system for improving the practical skills of inexperienced engineers regarding commercial products, and it uses AI to simulate customer interactions and provide feedback. Embodiments of this invention are described below.
[0474] System Configuration
[0475] The system consists of a terminal used by the user, a server that generates scenarios and processes information, and an AI that simulates dialogue.
[0476] Program Processing Overview
[0477] User: Log in to the system via your device, select the product you want to train on, and enter its details.
[0478] Server: Based on information received from the user, the system generates scenarios for use by the AI. In this process, the system prepares common questions and problem situations related to the product.
[0479] AI: Based on the generated scenario, it begins a conversation with the user as a customer. The AI asks questions related to the product and waits for the user's response.
[0480] Terminal: User responses are recorded in audio format, and the data is sent to the server.
[0481] Server: Analyzes voice data and evaluates the user's response. This evaluation is based on the accuracy and specificity of the response.
[0482] Feedback Provision: The server uses the analyzed data to generate feedback, including specific areas for improvement, and presents it to the user via the terminal.
[0483] Specific example
[0484] For example, if a user requests training on "cloud storage services," the AI, based on a scenario, will ask questions such as, "Please explain the backup function of cloud storage in detail." The user will answer, and their answer will be recorded. The server will analyze the recording and generate feedback such as, "The explanation is abstract; it would be better to mention specific steps," which will then be provided to the user via their device. Through this process, users can develop more specific and practical response skills.
[0485] The following describes the processing flow.
[0486] Step 1:
[0487] The user logs into the device and enters information about the product they wish to train on. The device receives this information and sends it to the server.
[0488] Step 2:
[0489] The server generates a scenario for use by the AI based on the product information it receives. The scenario includes product features, common problems, and anticipated customer questions. The generated scenario is then sent to the terminal.
[0490] Step 3:
[0491] The user initiates a conversational session with the AI through their device. Based on a scenario received from the server, the AI, acting as a customer, asks the user questions.
[0492] Step 4:
[0493] The user responds to questions from the AI. This interaction is recorded as audio, and the recorded data is sent from the device to the server.
[0494] Step 5:
[0495] The server analyzes the voice data. Specifically, it uses natural language processing technology to evaluate the accuracy, speed, and specificity of the user's responses.
[0496] Step 6:
[0497] The server generates specific feedback based on the analysis results. This feedback includes guidelines on what areas the user should improve.
[0498] Step 7:
[0499] The device receives feedback from the server and displays it to the user. The user can use this as a reference to help with their next learning.
[0500] (Example 1)
[0501] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] It is difficult for inexperienced engineers to effectively improve their practical skills related to products in a short period of time. Traditional training methods have difficulty providing realistic customer interaction simulations, and opportunities to receive specific feedback are limited, thus hindering efficient skill improvement.
[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0504] In this invention, the server includes means for generating scenarios based on information input from the user, means for engaging in dialogue based on the generated scenarios, and means for generating general questions and problem situations related to the product using a generating AI model. This provides a realistic customer interaction simulation and enables efficient skill improvement through specific and useful feedback to the user.
[0505] A "user" refers to an individual or group that uses the system to receive training related to the product or service.
[0506] A "scenario" refers to a series of questions and situations prepared by a generative AI model to structure the flow of interaction with the user.
[0507] "Dialogue" refers to a series of communication processes in which the AI asks the user questions about the product, and the user responds to those questions.
[0508] "Audio data" refers to digital data in audio format used to record user responses.
[0509] "Evaluation" refers to the process of analyzing user responses based on recorded audio data to measure their accuracy and specificity.
[0510] "Feedback" refers to information generated based on evaluation results, including specific areas for improvement and advice regarding the response.
[0511] A "generative AI model" refers to an artificial intelligence model used to generate questions and problem situations related to a product or service.
[0512] A "terminal" is a device used by a user to access a system, and it has functions for information input and voice recording.
[0513] A "server" refers to a central computing device that handles tasks such as scenario generation, dialogue management, audio data analysis, and feedback generation.
[0514] This invention is designed to help inexperienced engineers improve their practical skills with specific products. The system consists of a terminal accessed by the user, a server for information processing, and an AI for managing the interaction. The user logs into the system through the terminal and selects the product to be trained on. The terminal transmits detailed information about the product entered by the user to the server. Based on this information, the server generates an interaction scenario using a generative AI model.
[0515] Based on the generated scenario, the AI simulates a realistic conversation with the user. The AI creates questions about the product and asks them sequentially to the user. Specific questions might include things like, "Please explain the cloud storage backup function in detail."
[0516] The device has the function of recording the user's voice responses in real time and sending that data to the server. The server uses speech recognition software to analyze this voice data and evaluate the accuracy and specificity of the user's responses. Based on this evaluation, the server generates feedback that includes specific areas for improvement and provides this information to the user through the device.
[0517] As a concrete example in this system, if a user requests training on cloud storage services, the following prompt message would be used: "We want to simulate explaining cloud storage services to a customer. Please create a scenario where the user requests a specific explanation about the backup function."
[0518] This procedure allows users to efficiently acquire practical and specific skills. Furthermore, by using a generative AI model, this system can flexibly design product-specific questions and simulations, making it suitable for a wide variety of products.
[0519] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0520] Step 1:
[0521] The user logs into the system. The user uses their device to enter their login information and access the system. The server receives the entered user information and performs authentication. If successful, the user is redirected to the home screen.
[0522] Step 2:
[0523] The user selects a product. The user chooses the product they want to train on and enters its details into the terminal. The server then queries the database for product information and receives information about the product specified by the user.
[0524] Step 3:
[0525] The server generates the scenario. Based on the received product information, it uses a generating AI model to generate a scenario that includes common questions and problem situations related to the product. Specifically, prompt sentences are input to the AI model, and a scenario is obtained as output. This scenario includes a list of questions related to the product.
[0526] Step 4:
[0527] The AI initiates the conversation. Based on the generated scenario, the AI sequentially presents the user with questions about the product. The AI selects a question from the scenario, sends it to the device in voice or text format, and waits for the user's response.
[0528] Step 5:
[0529] The device records the user's responses. When the user responds to the AI's questions verbally, the device records the audio in real time and sends it to the server as digital audio data. The recorded audio data is obtained as output.
[0530] Step 6:
[0531] The server analyzes the audio data. The server uses speech recognition software to analyze the received audio data and convert it into text. Based on this text data, the server evaluates the accuracy and specificity of the response. At this point, an evaluation score and areas for improvement are output as part of the analysis results.
[0532] Step 7:
[0533] The server generates and provides feedback. Based on the evaluation results, the server generates feedback that includes specific areas for improvement. This feedback is sent to the terminal and provided to the user. As output, the user is presented with areas for improvement and advice.
[0534] (Application Example 1)
[0535] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] Sales staff in physical stores need to quickly improve their specialized product knowledge and customer service skills. However, traditional training methods lacked opportunities for practical, scenario-based training, making it difficult to acquire these skills efficiently.
[0537] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0538] In this invention, the server includes means for generating a scenario based on information input by the user, means for conducting a dialogue based on the generated scenario, and means for recording the content of the dialogue as audio data. This enables sales staff to receive practical training tailored to the conditions of a physical store.
[0539] A "user" is an individual who logs into the system, selects products and information, and participates in training.
[0540] A "scenario" is a virtual flow of dialogue, including conversations and questions related to the product or service, that is generated to support user training.
[0541] "Dialogue" refers to two-way communication between a user and AI, conducted via voice or text.
[0542] "Audio data" refers to digital audio information that records a user's speech.
[0543] "Evaluation" is the process of analyzing user responses and making judgments based on their accuracy and specificity.
[0544] "Feedback" refers to information based on evaluation results, including suggestions for improvement and advice regarding user responses.
[0545] "Speech synthesis" is a technology that generates speech based on text information and is used when generating questions from a virtual customer.
[0546] "Customer service skills" refer to the ability to provide appropriate information in response to customer requests and to build relationships.
[0547] "Artificial intelligence" refers to the intelligent behavior and learning abilities that computer programs simulate, and in dialogue scenarios, it takes on the role of the customer.
[0548] The system for realizing this application consists of a user terminal, a server, and artificial intelligence functions. The user uses a device such as a smartphone or tablet to log in to the system and select the product category to train. Based on the information received, the server generates a conversational scenario suitable for the user using a generative AI model.
[0549] Based on the generated scenario, the server utilizes speech synthesis to create questions from a virtual customer's perspective and outputs them as audio to the user's terminal. The user responds to these questions, and the responses are recorded as audio data on the terminal and sent to the server. The server uses speech recognition technology to convert the audio data into text and analyzes the content of the responses.
[0550] Based on the analysis results, the server evaluates the user's response and generates specific feedback, including areas for improvement. This feedback is then sent back to the user's terminal and presented as part of their training. This allows the user to acquire more practical customer service skills.
[0551] As a concrete example, consider a scenario where a user receives the prompt, "Explain to the customer how to choose colors and materials for the new smartphone case." Through responding to this question, the user can gain practical experience in explaining product features in detail. Through this process, salespeople can naturally improve their advanced customer service skills.
[0552] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0553] Step 1:
[0554] The user logs into the system using a terminal and selects the product category they wish to train in. Inputs include the user ID, password, and the selected product category. The server outputs this selection information for the next processing step.
[0555] Step 2:
[0556] The server uses a generative AI model to generate dialogue scenarios based on the received product category information. This involves setting up scenarios based on typical customer questions and scenario-based situations. The input is product category information, and the output is the generated scenario.
[0557] Step 3:
[0558] Based on the generated scenario, the server uses speech synthesis to create questions in voice format as a virtual customer and sends them to the terminal as output. The input is scenario information, and the output is voice data.
[0559] Step 4:
[0560] The user responds to questions from a virtual customer presented as audio from the terminal. The terminal records the user's responses as audio data. The input is the audio questions, and the output is the user's response audio data.
[0561] Step 5:
[0562] The terminal sends the recorded audio data to the server. The input is the user's response audio data, and the output is the data transferred to the server.
[0563] Step 6:
[0564] The server uses speech recognition technology to convert audio data into text data and performs analysis. This analysis evaluates the accuracy and specificity of the user's responses. The input is the user's audio response data, and the output is text data and the evaluation results.
[0565] Step 7:
[0566] The server generates feedback based on the evaluation results. This feedback includes areas for improvement and specific advice. The input is the evaluation result, and the output is the feedback text.
[0567] Step 8:
[0568] The server sends the generated feedback to the terminal and presents it to the user. The input is the text of the feedback, and the output is its display on the user's terminal.
[0569] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0570] The present invention is an interactive training system for inexperienced engineers to acquire practical skills, and in particular includes technology that recognizes user emotions and adjusts responses accordingly. Embodiments of the present invention are described in detail below.
[0571] System Configuration
[0572] This system consists of a user terminal, a server that generates and manages scenarios, an emotion engine that performs emotion recognition, and an AI that acts as the customer. The emotion engine evaluates the user's emotional state in real time and dynamically adjusts the feedback and scenarios.
[0573] Program Processing Overview
[0574] User: Log in to the terminal, select the product to be trained on, and enter the information. This information is immediately sent to the server.
[0575] Server: Based on the received product information, it generates training scenarios for the AI to use. The emotion engine is also initialized and prepared to analyze the user's voice data and speech characteristics.
[0576] AI: Following a scenario generated on the server, the AI begins interacting with the user as a customer. The AI tests the user's understanding by asking questions about the product.
[0577] Emotion Engine: Analyzes user speech data in real time to estimate the user's emotional state based on voice tone, speech rate, pronunciation characteristics, etc.
[0578] Terminal: During this dialogue process, all user responses are recorded as audio data, and the results of the emotion engine's analysis are stored on the server.
[0579] Server: Analyzes the results of the emotion engine and the dialogue content, and generates feedback based on the results. The feedback content is adjusted to match the user's emotional state.
[0580] Feedback Provision: Adjusted feedback is presented to the user via the device, and the user uses it to improve their next training session.
[0581] Specific example
[0582] For example, if a user is undergoing training on a "network security platform," the AI might ask, "Could you explain the security protocols of this platform?" If the user is nervous and their voice is trembling, the emotion engine recognizes this emotion as "nervousness." In response, the server can adjust the feedback, providing softer language such as, "It's okay to answer more relaxed; the goal is to deepen your understanding," thereby supporting the user's confidence. In this way, the system can leverage real-time emotion recognition to provide a more personalized training experience.
[0583] The following describes the processing flow.
[0584] Step 1:
[0585] The user logs into the terminal, selects and enters information about the product they wish to train on. The terminal then sends the entered information to the server.
[0586] Step 2:
[0587] Based on the product information received by the server, the AI generates a dialogue scenario. This scenario includes typical customer questions and situations related to the product.
[0588] Step 3:
[0589] The user initiates a dialogue session with the AI based on instructions from their device. The AI, acting as a customer, asks the user questions about the product, following a scenario generated on the server.
[0590] Step 4:
[0591] The emotion engine analyzes the user's voice in real time. It analyzes tone, tempo, and voice intensity from the audio data to infer the user's emotional state.
[0592] Step 5:
[0593] The user answers the AI's questions. The device records these responses as audio data and sends it to the server. The server also stores the analysis data from the emotion engine.
[0594] Step 6:
[0595] The server analyzes recorded audio data and emotional states, and provides an evaluation based on the user's responses and emotions. Natural language processing techniques are used for detailed analysis.
[0596] Step 7:
[0597] The server generates feedback based on the analysis results, tailored to the user's emotional state. This feedback includes areas for improvement and positive reinforcement.
[0598] Step 8:
[0599] The device displays the generated feedback to the user. The user can then use this feedback to further improve their training.
[0600] (Example 2)
[0601] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0602] Traditional training systems have a problem in that they provide uniform feedback without considering the user's emotional state, resulting in insufficient individual skill improvement and deepened understanding. Furthermore, the automatic generation of dialogue scenarios based on user-selected items is often inefficient, limiting the practical training effectiveness.
[0603] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0604] In this invention, the server includes emotion recognition means for analyzing the user's emotional state in real time, means for adjusting feedback according to the emotional state, and means for generating dialogue scenarios based on information input from the user. This enables personalized feedback tailored to each user's emotional state and specific dialogue training suitable for the selected item.
[0605] A "user" refers to an individual who uses the system and is the recipient of individually provided services and feedback.
[0606] "Information" refers to data related to products and goods that users input into the system. This information is used to generate dialogue scenarios.
[0607] A "scenario" refers to the plan and flow of the dialogue between the user and artificial intelligence, and its pre-generation enables smoother interactions.
[0608] "Dialogue" refers to two-way communication between the user and artificial intelligence, and is a means of evaluating the user's level of understanding and skills.
[0609] "Voice data" refers to digital recordings of user speech information, which are used for analysis and evaluation.
[0610] "Emotion recognition" refers to the process of estimating a user's emotional state by analyzing their voice data.
[0611] "Feedback" refers to evaluations and advice provided to users based on the results of dialogue and sentiment analysis.
[0612] "Adjustment" refers to optimizing the content of feedback according to the user's emotional state, and is a means of providing a personalized training experience.
[0613] "Artificial intelligence" refers to an automated system that acts as a customer in conversations with users, assisting in user training through questions and feedback.
[0614] The system according to the present invention provides a training environment in which inexperienced engineers can acquire practical skills by dynamically adjusting their responses while recognizing the user's emotional state. Specific embodiments of the present invention are described below.
[0615] The user first logs into the system using a terminal and enters information about the item they wish to train. This information is immediately sent to the server. The server uses a generative AI model to generate dialogue scenarios based on the received information. The server also initializes an emotion engine and prepares to analyze the user's voice data in real time. This emotion engine has the ability to analyze the user's voice tone, speaking speed, and pronunciation patterns to estimate their emotional state.
[0616] Following the generated scenario, the artificial intelligence takes on the role of a customer and initiates a conversation with the user. During the conversation, the user answers questions about the items, and speech data is collected as knowledge is verified. The device records this data and sends the results of the emotion engine's analysis to the server to understand the user's emotional state.
[0617] The server generates feedback based on collected data and sentiment assessment. This feedback is tailored to the user's emotions and delivered to the user via their device. This allows the user to understand areas for improvement needed in their next training session and use that information to enhance their skills.
[0618] As a concrete example, let's assume a user is undergoing training on a "network security platform." The artificial intelligence asks, "Could you explain the security protocols of this platform?" If the emotion engine detects that the user is nervous, the server provides feedback such as, "You can answer more relaxed; the goal is to deepen your understanding."
[0619] An example of a prompt would be, "As a customer, ask the user questions about the network security platform. Use the emotion engine to determine if the user is nervous and adjust your feedback accordingly." This allows the system to create a personalized and effective training experience for the user.
[0620] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0621] Step 1:
[0622] The user logs into the terminal. This prompts the terminal to input the user's authentication information. The terminal uses this information to authenticate the user and displays a dashboard that the user can access. The output is a training dashboard.
[0623] Step 2:
[0624] The user selects information about the item they wish to train on and enters it via a terminal. This information constitutes the input. The terminal sends this information to the server, which uses it as data necessary to generate the dialogue scenario. The output is the transmission of product information to the server.
[0625] Step 3:
[0626] The server uses an AI model based on the received item information to generate a dialogue scenario. The input is the product information received from the user, and the output is the generated dialogue scenario. The server provides this scenario to the artificial intelligence, and the dialogue is ready.
[0627] Step 4:
[0628] The server initializes the emotion engine. The input is the speech analysis model and algorithms necessary for the dialogue. The emotion engine analyzes the user's voice data in real time and prepares to estimate the emotional state. The output is the ready emotion engine.
[0629] Step 5:
[0630] The artificial intelligence, following a generated dialogue scenario, begins a conversation with the user as a customer. The input consists of the generated scenario and the user's utterances, while the output consists of questions posed to the user and the progress of the conversation. The device displays this dialogue on the screen, and the user responds.
[0631] Step 6:
[0632] The emotion engine analyzes user speech data in real time. The input is user voice data, and data processing includes analysis of voice tone, speech rate, and pronunciation features. The output is an estimated result of the user's emotional state.
[0633] Step 7:
[0634] The terminal records the user's responses as audio data and sends it to the server along with the analysis results from the emotion engine. The input is the speech data and analysis results, and the output is the recorded dialogue data and emotional state.
[0635] Step 8:
[0636] The server generates feedback using recorded dialogue content and sentiment analysis results. The input consists of dialogue data and emotional state, and the feedback generation includes adjustments based on the user's emotional state. The output is the generated feedback.
[0637] Step 9:
[0638] The terminal displays feedback sent from the server to the user. The input is adjusted feedback, and the output provides information that the user can use to improve their next training session.
[0639] (Application Example 2)
[0640] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0641] When inexperienced engineers efficiently acquire practical skills, there are challenges in providing emotionally responsive feedback and personalized training tailored to individual characteristics. Furthermore, to enhance the effectiveness of training, there is a need for appropriate real-time analysis of emotional states and the provision of feedback based on that analysis.
[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0643] In this invention, the server includes means for generating a scenario based on information input by the user, means for engaging in dialogue based on the generated scenario, and means for estimating the user's emotional state in real time from the user's voice. This makes it possible for inexperienced technicians to effectively acquire practical skills while receiving feedback tailored to their own emotional state.
[0644] A "user" refers to an individual who inputs information and participates in dialogue in order to acquire practical skills using a training system.
[0645] "Means for generating scenarios" refers to a function that automatically creates the flow and content of conversations used for training based on information entered by the user.
[0646] "Means of dialogue" refers to functions for performing communication with users based on generated scenarios.
[0647] "Means of recording as audio data" refers to a system that saves the content of conversations with the user in audio format.
[0648] "Means of analysis and evaluation" refers to a system function that analyzes recorded audio data and evaluates the user's responses and emotions.
[0649] "Means of generating feedback" refers to a function that creates information to provide users with suggestions for improvement and advice based on evaluation results.
[0650] "Means of adjusting feedback" refers to functions that optimize feedback content according to the user's emotional state and provide it in an appropriate format.
[0651] "Means for securely managing information data" refers to functions that protect user and system data and safeguard it from unauthorized access and leakage.
[0652] "Methods for estimating emotional states in real time" refers to a function that analyzes the user's voice and instantly recognizes and judges their emotions during a conversation.
[0653] The system for realizing this invention mainly consists of a server, a user terminal, an emotion recognition engine, and a conversational AI. Each of these elements is described in detail below.
[0654] The user first accesses the application on their device, logs in, and begins training. The device used is a smartphone, and the microphone and camera capture the user's voice and video. In the initial stages of training, the user selects product candidates and enters initial information about them. All of this information is sent to the server.
[0655] The servers are hosted in a cloud environment and use AI technology to generate scenarios. Specifically, an AI training model is used to build individual conversation scenarios based on each user's choices. In these scenarios, the AI takes on the role of a virtual customer and initiates communication with the user.
[0656] The emotion recognition engine first analyzes the user's voice data in real time. For example, it uses the Google Speech-to-Text API to convert the speech to text and estimates emotions by analyzing speech speed and tone. The estimated emotional state is instantly sent to the server and used as basic data for generating feedback.
[0657] The feedback generation system creates personalized feedback based on analysis results. This feedback is sent to the user's device at the appropriate time and presented to the user visually and audibly. This allows users to flexibly adjust their responses and acquire effective skill sets.
[0658] For example, if a user chooses a setting to train on "fine wines," the AI might ask, "Can you describe the region where this wine is produced and its aromatic characteristics?" If the emotion recognition engine determines the user is "anxious" due to an unconvincing answer, the server will generate feedback such as, "Let's try to calm down and speak slowly. It would be good to pick out a few key points and explain them in order."
[0659] The following are examples of prompts for the generated AI model:
[0660] "Analyze user voice data in real time and estimate their emotional state. When a user says, 'Thank you for visiting. What products are you looking for today?', evaluate their emotion based on their tone and speed of voice, and generate appropriate feedback."
[0661] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0662] Step 1:
[0663] The user logs into their device, selects the product they want to train on, enters the information, and sends it to the server. The entered product information is used as basic data to generate the user's training scenario.
[0664] Step 2:
[0665] The server generates dialogue scenarios using a generative AI model based on the received product information. This generation process constructs dialogue patterns for product-related questions and virtual customers, and provides the generated scenarios to the AI.
[0666] Step 3:
[0667] Based on the dialogue scenario received from the server, the AI initiates questions about the product to the user, acting as a virtual customer. The user's responses are sent to the server in real time and recorded as audio data.
[0668] Step 4:
[0669] The emotion recognition engine analyzes the user's voice data in real time. The input voice is converted to text via the Google Speech-to-Text API, and the user's emotional state is estimated by analyzing the text and acoustic features.
[0670] Step 5:
[0671] Based on the output of the emotion recognition engine, the server generates feedback adapted to the user's emotional state. This feedback generation process concretizes improvement suggestions and advice tailored to the user's emotional state.
[0672] Step 6:
[0673] The server sends the generated feedback to the terminal, which then presents the feedback to the user visually or audibly. The user can then use the received feedback to review their responses in subsequent interactions.
[0674] Step 7:
[0675] Users utilize the feedback they receive to improve their skills and initiate new dialogue sessions. Training progresses by repeating this cycle, promoting effective skill acquisition.
[0676] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0678] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0679] [Fourth Embodiment]
[0680] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0681] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0682] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0683] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0684] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0685] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0686] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0687] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0688] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0691] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] This invention is a system for improving the practical skills of inexperienced engineers regarding commercial products, and it uses AI to simulate customer interactions and provide feedback. Embodiments of this invention are described below.
[0694] System Configuration
[0695] The system consists of a terminal used by the user, a server that generates scenarios and processes information, and an AI that simulates dialogue.
[0696] Program Processing Overview
[0697] User: Log in to the system via your device, select the product you want to train on, and enter its details.
[0698] Server: Based on information received from the user, the system generates scenarios for use by the AI. In this process, the system prepares common questions and problem situations related to the product.
[0699] AI: Based on the generated scenario, it begins a conversation with the user as a customer. The AI asks questions related to the product and waits for the user's response.
[0700] Terminal: User responses are recorded in audio format, and the data is sent to the server.
[0701] Server: Analyzes voice data and evaluates the user's response. This evaluation is based on the accuracy and specificity of the response.
[0702] Feedback Provision: The server uses the analyzed data to generate feedback, including specific areas for improvement, and presents it to the user via the terminal.
[0703] Specific example
[0704] For example, if a user requests training on "cloud storage services," the AI, based on a scenario, will ask questions such as, "Please explain the backup function of cloud storage in detail." The user will answer, and their answer will be recorded. The server will analyze the recording and generate feedback such as, "The explanation is abstract; it would be better to mention specific steps," which will then be provided to the user via their device. Through this process, users can develop more specific and practical response skills.
[0705] The following describes the processing flow.
[0706] Step 1:
[0707] The user logs into the device and enters information about the product they wish to train on. The device receives this information and sends it to the server.
[0708] Step 2:
[0709] The server generates a scenario for use by the AI based on the product information it receives. The scenario includes product features, common problems, and anticipated customer questions. The generated scenario is then sent to the terminal.
[0710] Step 3:
[0711] The user initiates a conversational session with the AI through their device. Based on a scenario received from the server, the AI, acting as a customer, asks the user questions.
[0712] Step 4:
[0713] The user responds to questions from the AI. This interaction is recorded as audio, and the recorded data is sent from the device to the server.
[0714] Step 5:
[0715] The server analyzes the voice data. Specifically, it uses natural language processing technology to evaluate the accuracy, speed, and specificity of the user's responses.
[0716] Step 6:
[0717] The server generates specific feedback based on the analysis results. This feedback includes guidelines on what areas the user should improve.
[0718] Step 7:
[0719] The device receives feedback from the server and displays it to the user. The user can use this as a reference to help with their next learning.
[0720] (Example 1)
[0721] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0722] It is difficult for inexperienced engineers to effectively improve their practical skills related to products in a short period of time. Traditional training methods have difficulty providing realistic customer interaction simulations, and opportunities to receive specific feedback are limited, thus hindering efficient skill improvement.
[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0724] In this invention, the server includes means for generating scenarios based on information input from the user, means for engaging in dialogue based on the generated scenarios, and means for generating general questions and problem situations related to the product using a generating AI model. This provides a realistic customer interaction simulation and enables efficient skill improvement through specific and useful feedback to the user.
[0725] A "user" refers to an individual or group that uses the system to receive training related to the product or service.
[0726] A "scenario" refers to a series of questions and situations prepared by a generative AI model to structure the flow of interaction with the user.
[0727] "Dialogue" refers to a series of communication processes in which the AI asks the user questions about the product, and the user responds to those questions.
[0728] "Audio data" refers to digital data in audio format used to record user responses.
[0729] "Evaluation" refers to the process of analyzing user responses based on recorded audio data to measure their accuracy and specificity.
[0730] "Feedback" refers to information generated based on evaluation results, including specific areas for improvement and advice regarding the response.
[0731] A "generative AI model" refers to an artificial intelligence model used to generate questions and problem situations related to a product or service.
[0732] A "terminal" is a device used by a user to access a system, and it has functions for information input and voice recording.
[0733] A "server" refers to a central computing device that handles tasks such as scenario generation, dialogue management, audio data analysis, and feedback generation.
[0734] This invention is designed to help inexperienced engineers improve their practical skills with specific products. The system consists of a terminal accessed by the user, a server for information processing, and an AI for managing the interaction. The user logs into the system through the terminal and selects the product to be trained on. The terminal transmits detailed information about the product entered by the user to the server. Based on this information, the server generates an interaction scenario using a generative AI model.
[0735] Based on the generated scenario, the AI simulates a realistic conversation with the user. The AI creates questions about the product and asks them sequentially to the user. Specific questions might include things like, "Please explain the cloud storage backup function in detail."
[0736] The device has the function of recording the user's voice responses in real time and sending that data to the server. The server uses speech recognition software to analyze this voice data and evaluate the accuracy and specificity of the user's responses. Based on this evaluation, the server generates feedback that includes specific areas for improvement and provides this information to the user through the device.
[0737] As a concrete example in this system, if a user requests training on cloud storage services, the following prompt message would be used: "We want to simulate explaining cloud storage services to a customer. Please create a scenario where the user requests a specific explanation about the backup function."
[0738] This procedure allows users to efficiently acquire practical and specific skills. Furthermore, by using a generative AI model, this system can flexibly design product-specific questions and simulations, making it suitable for a wide variety of products.
[0739] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0740] Step 1:
[0741] The user logs into the system. The user uses their device to enter their login information and access the system. The server receives the entered user information and performs authentication. If successful, the user is redirected to the home screen.
[0742] Step 2:
[0743] The user selects a product. The user chooses the product they want to train on and enters its details into the terminal. The server then queries the database for product information and receives information about the product specified by the user.
[0744] Step 3:
[0745] The server generates the scenario. Based on the received product information, it uses a generating AI model to generate a scenario that includes common questions and problem situations related to the product. Specifically, prompt sentences are input to the AI model, and a scenario is obtained as output. This scenario includes a list of questions related to the product.
[0746] Step 4:
[0747] The AI initiates the conversation. Based on the generated scenario, the AI sequentially presents the user with questions about the product. The AI selects a question from the scenario, sends it to the device in voice or text format, and waits for the user's response.
[0748] Step 5:
[0749] The device records the user's responses. When the user responds to the AI's questions verbally, the device records the audio in real time and sends it to the server as digital audio data. The recorded audio data is obtained as output.
[0750] Step 6:
[0751] The server analyzes the audio data. The server uses speech recognition software to analyze the received audio data and convert it into text. Based on this text data, the server evaluates the accuracy and specificity of the response. At this point, an evaluation score and areas for improvement are output as part of the analysis results.
[0752] Step 7:
[0753] The server generates and provides feedback. Based on the evaluation results, the server generates feedback that includes specific areas for improvement. This feedback is sent to the terminal and provided to the user. As output, the user is presented with areas for improvement and advice.
[0754] (Application Example 1)
[0755] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0756] Sales staff in physical stores need to quickly improve their specialized product knowledge and customer service skills. However, traditional training methods lacked opportunities for practical, scenario-based training, making it difficult to acquire these skills efficiently.
[0757] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0758] In this invention, the server includes means for generating a scenario based on information input by the user, means for conducting a dialogue based on the generated scenario, and means for recording the content of the dialogue as audio data. This enables sales staff to receive practical training tailored to the conditions of a physical store.
[0759] A "user" is an individual who logs into the system, selects products and information, and participates in training.
[0760] A "scenario" is a virtual flow of dialogue, including conversations and questions related to the product or service, that is generated to support user training.
[0761] "Dialogue" refers to two-way communication between a user and AI, conducted via voice or text.
[0762] "Audio data" refers to digital audio information that records a user's speech.
[0763] "Evaluation" is the process of analyzing user responses and making judgments based on their accuracy and specificity.
[0764] "Feedback" refers to information based on evaluation results, including suggestions for improvement and advice regarding user responses.
[0765] "Speech synthesis" is a technology that generates speech based on text information and is used when generating questions from a virtual customer.
[0766] "Customer service skills" refer to the ability to provide appropriate information in response to customer requests and to build relationships.
[0767] "Artificial intelligence" refers to the intelligent behavior and learning abilities that computer programs simulate, and in dialogue scenarios, it takes on the role of the customer.
[0768] The system for realizing this application consists of a user terminal, a server, and artificial intelligence functions. The user uses a device such as a smartphone or tablet to log in to the system and select the product category to train. Based on the information received, the server generates a conversational scenario suitable for the user using a generative AI model.
[0769] Based on the generated scenario, the server utilizes speech synthesis to create questions from a virtual customer's perspective and outputs them as audio to the user's terminal. The user responds to these questions, and the responses are recorded as audio data on the terminal and sent to the server. The server uses speech recognition technology to convert the audio data into text and analyzes the content of the responses.
[0770] Based on the analysis results, the server evaluates the user's response and generates specific feedback, including areas for improvement. This feedback is then sent back to the user's terminal and presented as part of their training. This allows the user to acquire more practical customer service skills.
[0771] As a concrete example, consider a scenario where a user receives the prompt, "Explain to the customer how to choose colors and materials for the new smartphone case." Through responding to this question, the user can gain practical experience in explaining product features in detail. Through this process, salespeople can naturally improve their advanced customer service skills.
[0772] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0773] Step 1:
[0774] The user logs into the system using a terminal and selects the product category they wish to train in. Inputs include the user ID, password, and the selected product category. The server outputs this selection information for the next processing step.
[0775] Step 2:
[0776] The server uses a generative AI model to generate dialogue scenarios based on the received product category information. This involves setting up scenarios based on typical customer questions and scenario-based situations. The input is product category information, and the output is the generated scenario.
[0777] Step 3:
[0778] Based on the generated scenario, the server uses speech synthesis to create questions in voice format as a virtual customer and sends them to the terminal as output. The input is scenario information, and the output is voice data.
[0779] Step 4:
[0780] The user responds to questions from a virtual customer presented as audio from the terminal. The terminal records the user's responses as audio data. The input is the audio questions, and the output is the user's response audio data.
[0781] Step 5:
[0782] The terminal sends the recorded audio data to the server. The input is the user's response audio data, and the output is the data transferred to the server.
[0783] Step 6:
[0784] The server uses speech recognition technology to convert audio data into text data and performs analysis. This analysis evaluates the accuracy and specificity of the user's responses. The input is the user's audio response data, and the output is text data and the evaluation results.
[0785] Step 7:
[0786] The server generates feedback based on the evaluation results. This feedback includes areas for improvement and specific advice. The input is the evaluation result, and the output is the feedback text.
[0787] Step 8:
[0788] The server sends the generated feedback to the terminal and presents it to the user. The input is the text of the feedback, and the output is its display on the user's terminal.
[0789] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0790] The present invention is an interactive training system for inexperienced engineers to acquire practical skills, and in particular includes technology that recognizes user emotions and adjusts responses accordingly. Embodiments of the present invention are described in detail below.
[0791] System Configuration
[0792] This system consists of a user terminal, a server that generates and manages scenarios, an emotion engine that performs emotion recognition, and an AI that acts as the customer. The emotion engine evaluates the user's emotional state in real time and dynamically adjusts the feedback and scenarios.
[0793] Program Processing Overview
[0794] User: Log in to the terminal, select the product to be trained on, and enter the information. This information is immediately sent to the server.
[0795] Server: Based on the received product information, it generates training scenarios for the AI to use. The emotion engine is also initialized and prepared to analyze the user's voice data and speech characteristics.
[0796] AI: Following a scenario generated on the server, the AI begins interacting with the user as a customer. The AI tests the user's understanding by asking questions about the product.
[0797] Emotion Engine: Analyzes user speech data in real time to estimate the user's emotional state based on voice tone, speech rate, pronunciation characteristics, etc.
[0798] Terminal: During this dialogue process, all user responses are recorded as audio data, and the results of the emotion engine's analysis are stored on the server.
[0799] Server: Analyzes the results of the emotion engine and the dialogue content, and generates feedback based on the results. The feedback content is adjusted to match the user's emotional state.
[0800] Feedback Provision: Adjusted feedback is presented to the user via the device, and the user uses it to improve their next training session.
[0801] Specific example
[0802] For example, if a user is undergoing training on a "network security platform," the AI might ask, "Could you explain the security protocols of this platform?" If the user is nervous and their voice is trembling, the emotion engine recognizes this emotion as "nervousness." In response, the server can adjust the feedback, providing softer language such as, "It's okay to answer more relaxed; the goal is to deepen your understanding," thereby supporting the user's confidence. In this way, the system can leverage real-time emotion recognition to provide a more personalized training experience.
[0803] The following describes the processing flow.
[0804] Step 1:
[0805] The user logs into the terminal, selects and enters information about the product they wish to train on. The terminal then sends the entered information to the server.
[0806] Step 2:
[0807] Based on the product information received by the server, the AI generates a dialogue scenario. This scenario includes typical customer questions and situations related to the product.
[0808] Step 3:
[0809] The user initiates a dialogue session with the AI based on instructions from their device. The AI, acting as a customer, asks the user questions about the product, following a scenario generated on the server.
[0810] Step 4:
[0811] The emotion engine analyzes the user's voice in real time. It analyzes tone, tempo, and voice intensity from the audio data to infer the user's emotional state.
[0812] Step 5:
[0813] The user answers the AI's questions. The device records these responses as audio data and sends it to the server. The server also stores the analysis data from the emotion engine.
[0814] Step 6:
[0815] The server analyzes recorded audio data and emotional states, and provides an evaluation based on the user's responses and emotions. Natural language processing techniques are used for detailed analysis.
[0816] Step 7:
[0817] The server generates feedback based on the analysis results, tailored to the user's emotional state. This feedback includes areas for improvement and positive reinforcement.
[0818] Step 8:
[0819] The device displays the generated feedback to the user. The user can then use this feedback to further improve their training.
[0820] (Example 2)
[0821] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0822] Traditional training systems have a problem in that they provide uniform feedback without considering the user's emotional state, resulting in insufficient individual skill improvement and deepened understanding. Furthermore, the automatic generation of dialogue scenarios based on user-selected items is often inefficient, limiting the practical training effectiveness.
[0823] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0824] In this invention, the server includes emotion recognition means for analyzing the user's emotional state in real time, means for adjusting feedback according to the emotional state, and means for generating dialogue scenarios based on information input from the user. This enables personalized feedback tailored to each user's emotional state and specific dialogue training suitable for the selected item.
[0825] A "user" refers to an individual who uses the system and is the recipient of individually provided services and feedback.
[0826] "Information" refers to data related to products and goods that users input into the system. This information is used to generate dialogue scenarios.
[0827] A "scenario" refers to the plan and flow of the dialogue between the user and artificial intelligence, and its pre-generation enables smoother interactions.
[0828] "Dialogue" refers to two-way communication between the user and artificial intelligence, and is a means of evaluating the user's level of understanding and skills.
[0829] "Voice data" refers to digital recordings of user speech information, which are used for analysis and evaluation.
[0830] "Emotion recognition" refers to the process of estimating a user's emotional state by analyzing their voice data.
[0831] "Feedback" refers to evaluations and advice provided to users based on the results of dialogue and sentiment analysis.
[0832] "Adjustment" refers to optimizing the content of feedback according to the user's emotional state, and is a means of providing a personalized training experience.
[0833] "Artificial intelligence" refers to an automated system that acts as a customer in conversations with users, assisting in user training through questions and feedback.
[0834] The system according to the present invention provides a training environment in which inexperienced engineers can acquire practical skills by dynamically adjusting their responses while recognizing the user's emotional state. Specific embodiments of the present invention are described below.
[0835] The user first logs into the system using a terminal and enters information about the item they wish to train. This information is immediately sent to the server. The server uses a generative AI model to generate dialogue scenarios based on the received information. The server also initializes an emotion engine and prepares to analyze the user's voice data in real time. This emotion engine has the ability to analyze the user's voice tone, speaking speed, and pronunciation patterns to estimate their emotional state.
[0836] Following the generated scenario, the artificial intelligence takes on the role of a customer and initiates a conversation with the user. During the conversation, the user answers questions about the items, and speech data is collected as knowledge is verified. The device records this data and sends the results of the emotion engine's analysis to the server to understand the user's emotional state.
[0837] The server generates feedback based on collected data and sentiment assessment. This feedback is tailored to the user's emotions and delivered to the user via their device. This allows the user to understand areas for improvement needed in their next training session and use that information to enhance their skills.
[0838] As a concrete example, let's assume a user is undergoing training on a "network security platform." The artificial intelligence asks, "Could you explain the security protocols of this platform?" If the emotion engine detects that the user is nervous, the server provides feedback such as, "You can answer more relaxed; the goal is to deepen your understanding."
[0839] An example of a prompt would be, "As a customer, ask the user questions about the network security platform. Use the emotion engine to determine if the user is nervous and adjust your feedback accordingly." This allows the system to create a personalized and effective training experience for the user.
[0840] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0841] Step 1:
[0842] The user logs into the terminal. This prompts the terminal to input the user's authentication information. The terminal uses this information to authenticate the user and displays a dashboard that the user can access. The output is a training dashboard.
[0843] Step 2:
[0844] The user selects information about the item they wish to train on and enters it via a terminal. This information constitutes the input. The terminal sends this information to the server, which uses it as data necessary to generate the dialogue scenario. The output is the transmission of product information to the server.
[0845] Step 3:
[0846] The server uses an AI model based on the received item information to generate a dialogue scenario. The input is the product information received from the user, and the output is the generated dialogue scenario. The server provides this scenario to the artificial intelligence, and the dialogue is ready.
[0847] Step 4:
[0848] The server initializes the emotion engine. The input is the speech analysis model and algorithms necessary for the dialogue. The emotion engine analyzes the user's voice data in real time and prepares to estimate the emotional state. The output is the ready emotion engine.
[0849] Step 5:
[0850] The artificial intelligence, following a generated dialogue scenario, begins a conversation with the user as a customer. The input consists of the generated scenario and the user's utterances, while the output consists of questions posed to the user and the progress of the conversation. The device displays this dialogue on the screen, and the user responds.
[0851] Step 6:
[0852] The emotion engine analyzes user speech data in real time. The input is user voice data, and data processing includes analysis of voice tone, speech rate, and pronunciation features. The output is an estimated result of the user's emotional state.
[0853] Step 7:
[0854] The terminal records the user's responses as audio data and sends it to the server along with the analysis results from the emotion engine. The input is the speech data and analysis results, and the output is the recorded dialogue data and emotional state.
[0855] Step 8:
[0856] The server generates feedback using recorded dialogue content and sentiment analysis results. The input consists of dialogue data and emotional state, and the feedback generation includes adjustments based on the user's emotional state. The output is the generated feedback.
[0857] Step 9:
[0858] The terminal displays feedback sent from the server to the user. The input is adjusted feedback, and the output provides information that the user can use to improve their next training session.
[0859] (Application Example 2)
[0860] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0861] When inexperienced engineers efficiently acquire practical skills, there are challenges in providing emotionally responsive feedback and personalized training tailored to individual characteristics. Furthermore, to enhance the effectiveness of training, there is a need for appropriate real-time analysis of emotional states and the provision of feedback based on that analysis.
[0862] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0863] In this invention, the server includes means for generating a scenario based on information input by the user, means for engaging in dialogue based on the generated scenario, and means for estimating the user's emotional state in real time from the user's voice. This makes it possible for inexperienced technicians to effectively acquire practical skills while receiving feedback tailored to their own emotional state.
[0864] A "user" refers to an individual who inputs information and participates in dialogue in order to acquire practical skills using a training system.
[0865] "Means for generating scenarios" refers to a function that automatically creates the flow and content of conversations used for training based on information entered by the user.
[0866] "Means of dialogue" refers to functions for performing communication with users based on generated scenarios.
[0867] "Means of recording as audio data" refers to a system that saves the content of conversations with the user in audio format.
[0868] "Means of analysis and evaluation" refers to a system function that analyzes recorded audio data and evaluates the user's responses and emotions.
[0869] "Means of generating feedback" refers to a function that creates information to provide users with suggestions for improvement and advice based on evaluation results.
[0870] "Means of adjusting feedback" refers to functions that optimize feedback content according to the user's emotional state and provide it in an appropriate format.
[0871] "Means for securely managing information data" refers to functions that protect user and system data and safeguard it from unauthorized access and leakage.
[0872] "Methods for estimating emotional states in real time" refers to a function that analyzes the user's voice and instantly recognizes and judges their emotions during a conversation.
[0873] The system for realizing this invention mainly consists of a server, a user terminal, an emotion recognition engine, and a conversational AI. Each of these elements is described in detail below.
[0874] The user first accesses the application on their device, logs in, and begins training. The device used is a smartphone, and the microphone and camera capture the user's voice and video. In the initial stages of training, the user selects product candidates and enters initial information about them. All of this information is sent to the server.
[0875] The servers are hosted in a cloud environment and use AI technology to generate scenarios. Specifically, an AI training model is used to build individual conversation scenarios based on each user's choices. In these scenarios, the AI takes on the role of a virtual customer and initiates communication with the user.
[0876] The emotion recognition engine first analyzes the user's voice data in real time. For example, it uses the Google Speech-to-Text API to convert the speech to text and estimates emotions by analyzing speech speed and tone. The estimated emotional state is instantly sent to the server and used as basic data for generating feedback.
[0877] The feedback generation system creates personalized feedback based on analysis results. This feedback is sent to the user's device at the appropriate time and presented to the user visually and audibly. This allows users to flexibly adjust their responses and acquire effective skill sets.
[0878] For example, if a user chooses a setting to train on "fine wines," the AI might ask, "Can you describe the region where this wine is produced and its aromatic characteristics?" If the emotion recognition engine determines the user is "anxious" due to an unconvincing answer, the server will generate feedback such as, "Let's try to calm down and speak slowly. It would be good to pick out a few key points and explain them in order."
[0879] The following are examples of prompts for the generated AI model:
[0880] "Analyze user voice data in real time and estimate their emotional state. When a user says, 'Thank you for visiting. What products are you looking for today?', evaluate their emotion based on their tone and speed of voice, and generate appropriate feedback."
[0881] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0882] Step 1:
[0883] The user logs into their device, selects the product they want to train on, enters the information, and sends it to the server. The entered product information is used as basic data to generate the user's training scenario.
[0884] Step 2:
[0885] The server generates dialogue scenarios using a generative AI model based on the received product information. This generation process constructs dialogue patterns for product-related questions and virtual customers, and provides the generated scenarios to the AI.
[0886] Step 3:
[0887] Based on the dialogue scenario received from the server, the AI initiates questions about the product to the user, acting as a virtual customer. The user's responses are sent to the server in real time and recorded as audio data.
[0888] Step 4:
[0889] The emotion recognition engine analyzes the user's voice data in real time. The input voice is converted to text via the Google Speech-to-Text API, and the user's emotional state is estimated by analyzing the text and acoustic features.
[0890] Step 5:
[0891] Based on the output of the emotion recognition engine, the server generates feedback adapted to the user's emotional state. This feedback generation process concretizes improvement suggestions and advice tailored to the user's emotional state.
[0892] Step 6:
[0893] The server sends the generated feedback to the terminal, which then presents the feedback to the user visually or audibly. The user can then use the received feedback to review their responses in subsequent interactions.
[0894] Step 7:
[0895] Users utilize the feedback they receive to improve their skills and initiate new dialogue sessions. Training progresses by repeating this cycle, promoting effective skill acquisition.
[0896] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0897] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0898] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0899] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0900] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0901] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0902] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0903] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0904] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0905] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0906] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0907] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0908] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0909] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0910] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0911] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0912] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0913] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0914] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0915] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0916] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0917] The following is further disclosed regarding the embodiments described above.
[0918] (Claim 1)
[0919] A means of generating a scenario based on information entered by the user,
[0920] A means of conducting dialogue based on the generated scenario,
[0921] A means of recording the content of the conversation as audio data,
[0922] A means of analyzing and evaluating recorded audio data,
[0923] A means for generating feedback on the response based on the evaluation results,
[0924] Means of providing feedback to users,
[0925] Means for securely managing information data,
[0926] A system that includes this.
[0927] (Claim 2)
[0928] The system according to claim 1, which includes means for receiving information about a product selected by the user and automatically generating a dialogue scenario based on that information.
[0929] (Claim 3)
[0930] The system according to claim 1, comprising means by which the AI takes on the role of a customer in a generated dialogue scenario, and the AI asks the user questions about the product.
[0931] "Example 1"
[0932] (Claim 1)
[0933] A means of generating a scenario based on information entered by the user,
[0934] A means of conducting dialogue based on the generated scenario,
[0935] A means of recording the content of the conversation as audio data,
[0936] A means of analyzing and evaluating recorded audio data,
[0937] A means for generating feedback on the response based on the evaluation results,
[0938] Means of providing feedback to users,
[0939] Means for securely managing information data,
[0940] A means of generating common questions and problem situations related to products using a generative AI model,
[0941] A means of recording the user's response in audio format on the terminal,
[0942] A system that includes this.
[0943] (Claim 2)
[0944] The system according to claim 1, which includes means for receiving information about a product selected by the user and automatically generating a dialogue scenario based on that information.
[0945] (Claim 3)
[0946] The system according to claim 1, comprising means by which the AI takes on the role of a customer in a generated dialogue scenario, and the AI asks the user questions about the product.
[0947] "Application Example 1"
[0948] (Claim 1)
[0949] A means of generating a scenario based on information entered by the user,
[0950] A means of conducting dialogue based on the generated scenario,
[0951] A means of recording the content of the conversation as audio data,
[0952] A means of analyzing and evaluating recorded audio data,
[0953] A means for generating feedback on the response based on the evaluation results,
[0954] Means for securely managing information data,
[0955] A means of generating questions as a virtual customer using speech synthesis functionality,
[0956] A means of providing scenario-based training to improve customer service skills,
[0957] A system that includes this.
[0958] (Claim 2)
[0959] The system according to claim 1, comprising means for receiving information about a product category selected by the user and for automatically generating a dialogue scenario based on that information.
[0960] (Claim 3)
[0961] The system according to claim 1, comprising means by which artificial intelligence plays the role of a customer in a generated dialogue scenario, and the artificial intelligence asks the user questions about the product.
[0962] "Example 2 of combining an emotion engine"
[0963] (Claim 1)
[0964] A means of generating a scenario based on information entered by the user,
[0965] A means of conducting dialogue based on the generated scenario,
[0966] A means of recording the content of the conversation as audio data,
[0967] A means of analyzing and evaluating recorded audio data,
[0968] A means for generating feedback on the response based on the evaluation results,
[0969] Means of providing feedback to users,
[0970] An emotion recognition method that analyzes the user's emotional state in real time,
[0971] A means of adjusting feedback according to emotional state,
[0972] Means for securely managing information data,
[0973] A system that includes this.
[0974] (Claim 2)
[0975] The system according to claim 1, comprising means for receiving information about an item selected by the user and automatically generating a dialogue scenario based on that information.
[0976] (Claim 3)
[0977] The system according to claim 1, comprising means by which artificial intelligence plays the role of a customer in a generated dialogue scenario, and the artificial intelligence asks the user questions about the item.
[0978] "Application example 2 when combining with an emotional engine"
[0979] (Claim 1)
[0980] A means of generating a scenario based on information entered by the user,
[0981] A means of conducting dialogue based on the generated scenario,
[0982] A means of recording the content of the conversation as audio data,
[0983] A means of analyzing and evaluating recorded audio data,
[0984] A means for generating feedback on the response based on the evaluation results,
[0985] Means of providing feedback to users,
[0986] A method for estimating a user's emotional state in real time from their voice,
[0987] A means of adjusting feedback based on estimated emotional states,
[0988] Means for securely managing information data,
[0989] A system that includes this.
[0990] (Claim 2)
[0991] The system according to claim 1, comprising means for receiving information about a product selected by the user and automatically generating a dialogue scenario based on that information.
[0992] (Claim 3)
[0993] The system according to claim 1, comprising means by which the AI takes on the role of a customer in a generated dialogue scenario, and the AI asks the user questions about the product. [Explanation of Symbols]
[0994] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of generating a scenario based on information entered by the user, A means of conducting dialogue based on the generated scenario, A means of recording the content of the conversation as audio data, A means of analyzing and evaluating recorded audio data, A means for generating feedback on the response based on the evaluation results, Means of providing feedback to users, Means for securely managing information data, A system that includes this.
2. The system according to claim 1, which includes means for receiving information about a product selected by the user and automatically generating a dialogue scenario based on that information.
3. The system according to claim 1, comprising means by which the AI takes on the role of a customer in a generated dialogue scenario, and the AI asks the user questions about the product.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A