system

The integrated system addresses inefficiencies in customer service training by using virtual characters for real-time analysis and feedback, enhancing skills through diverse scenarios and multilingual support, achieving effective skill improvement.

JP2026074941APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing customer service training methods are inefficient, require significant time and resources, lack objective evaluation, and struggle with multilingual capabilities, making it difficult to improve customer service skills effectively.

Method used

An integrated system using a virtual dialogue method with real-time analysis and feedback, generating and staging virtual characters for diverse scenarios, and providing multilingual support to enhance customer service skills through continuous improvement.

Benefits of technology

Enables efficient and effective improvement of customer service skills by offering immediate feedback, tracking progress, and supporting multiple languages, allowing users to practice in realistic and varied training environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074941000001_ABST
    Figure 2026074941000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An information processing device that provides a virtual dialogue method for improving customer service skills, Means for generating and performing virtual characters, A means of analyzing and evaluating user conversations in real time, A means of generating feedback based on user performance, A means of providing the generated feedback to the user, A system that includes means for recording and managing user training progress data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern service industries, improving customer service skills is an important issue. However, there is a lack of efficient and practical training methods, and conventional training methods require a certain amount of time and human resources. Also, it is difficult to objectively evaluate the skills of individual employees while dealing with various scenarios. Furthermore, there is often a requirement for multilingual capabilities, but the appropriate environment for this is not well-established. In such a situation, there is a demand for providing an appropriate training platform for employees to effectively improve their customer service skills.

Means for Solving the Problems

[0005] This invention provides an integrated system that aims to improve customer service skills by providing virtual dialogue using an information processing device. Specifically, it provides a virtual customer service scenario to the user using means for generating and staging a virtual character. Furthermore, the system has means for analyzing and evaluating the user's dialogue content in real time, thereby enabling immediate feedback on the user's performance. The generated feedback is provided to the user to support continuous skill improvement. In addition, by recording and managing training progress data, long-term skill improvement is made visible. Furthermore, by providing a training environment with multiple scenarios and multilingual support, it is possible to meet the diverse training needs of users. In this way, this invention realizes efficient and effective improvement of customer service skills.

[0006] "Customer service skills" refer to the skills and methodologies for providing comfortable and satisfying service to customers.

[0007] A "virtual dialogue method" is a technique for conducting dialogue using a simulated conversational environment generated by a computer.

[0008] An "information processing device" is a general term for hardware and software used to collect and process data.

[0009] A "virtual character" refers to a simulated person or character created using digital technology.

[0010] A "user" is an individual or group that operates and utilizes a specific system or software.

[0011] "Real-time analysis and evaluation of dialogue content" refers to a process that instantly analyzes user input, such as their statements, and evaluates it according to established criteria.

[0012] "Feedback" refers to evaluations or suggestions for improvement given in response to specific behavior.

[0013] "Progress data" refers to information that shows the progress of a particular activity or process.

[0014] "Multilingual support" refers to the ability to support communication using multiple languages. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] A sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] To implement this invention, three parties—a server, a terminal, and a user—must work closely together. The server, as the central information processing device, is responsible for generating virtual characters, analyzing dialogue, and generating feedback. The terminal provides a user interface and functions as a device that enables interaction with the user.

[0037] First, the user logs into the system via a terminal and selects a virtual customer service scenario that suits their purpose. For example, they can choose scenarios such as "handling cross-cultural customer interactions" or "resolving complex customer complaints." After selection, the server generates a virtual character based on the specified scenario. This character simulates interaction with the user, providing a simulated training environment for honing customer service skills.

[0038] Once a conversation begins, the server analyzes the user's statements in real time. This analysis uses natural language processing to evaluate factors such as tone, appropriateness of responses, and the ability to understand customer needs. The analyzed information is immediately recorded and evaluated within the server.

[0039] Once the training is complete, the server generates feedback based on the user's performance. This feedback is provided to the user via their device, specifically highlighting areas for improvement and areas where they performed well. The user can use this feedback to further enhance their skills. The feedback is also stored as data in the user's training profile when it is generated.

[0040] This system allows users to improve their customer service skills through practical and diverse scenarios, often without requiring significant human intervention. Furthermore, its multilingual capabilities enhance its ability to serve international customers. This enables users to efficiently learn the skills required in real-world work environments, contributing to improved work performance.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user attempts to log in to the system using their device. Logging in requires personal authentication information, so the user enters their ID and password. The device then sends the entered information to the server.

[0044] Step 2:

[0045] The server compares the received login information with the registered information in the database and performs authentication. If authentication is successful, it prepares to display the user's dashboard on the terminal.

[0046] Step 3:

[0047] The user starts training from a dashboard on their device, selecting one of several scenarios provided. The selection is sent to the server, which retrieves information about the scenario.

[0048] Step 4:

[0049] The server generates a virtual character based on the scenario selected by the user. The generated character has its dialogue and situation pre-configured, and is ready to begin the scenario.

[0050] Step 5:

[0051] The terminal displays a virtual character to the user and initiates a virtual dialogue scenario. The user responds to the character using voice input or text input.

[0052] Step 6:

[0053] The server receives user input in real time and performs analysis using natural language processing. The analysis results take into account the appropriateness of the content of the statements, the wording, and the flow of the conversation.

[0054] Step 7:

[0055] Once the training is complete, the server evaluates the user's performance based on the analysis results. From the evaluation results, it extracts specific strengths and areas for improvement.

[0056] Step 8:

[0057] The server creates feedback based on the generated evaluation and sends it to the terminal. The terminal displays the feedback to the user, who then reviews it.

[0058] Step 9:

[0059] The server stores and records data about the user's training sessions in a format that can be referenced later. Based on this data, users can develop long-term skill improvement plans.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In today's business environment, customer service skills are crucial for improving customer satisfaction, but traditional training methods have struggled to efficiently provide practical and diverse scenario-based training. Furthermore, efficiently strengthening the ability to handle multilingual international customers has also been a challenging task. In response, there was a need for a feedback system tailored to individual skills, a means to visualize user capabilities, and a way to promote improvement.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating user dialogue in real time, and means for generating feedback based on user performance. This enables users to improve their customer service skills through practical and diverse scenarios. Furthermore, analysis using natural language processing can efficiently enhance multilingual communication skills. In addition, by combining a generative AI model with prompt sentences, it is possible to appropriately guide user responses and achieve higher quality dialogue training.

[0065] A "virtual character" is a digital humanoid entity created to simulate interaction with a user, and can have different personalities and backgrounds based on various scenarios.

[0066] A "user interface" is a general term for the screens and tools that users use to access and operate a system, and is used for inputting dialogues, selecting scenarios, and so on.

[0067] "Natural language processing" is a technology used to enable computers to understand and analyze human language. It is used to analyze user statements and provide appropriate responses and evaluations.

[0068] "Feedback" is evaluation information provided based on the results of user interaction training, indicating areas for improvement and strengths, and is useful for improving the user's skills.

[0069] A "training profile" is a collection of information that records a user's training progress and past evaluation results, and is referenced for future skill improvement.

[0070] A "generative AI model" is a type of artificial intelligence that generates appropriate information or characters in response to user requests, based on knowledge learned from large amounts of data.

[0071] A "prompt message" is a phrase provided to give instructions or draw attention to a user, and its role is to guide them to a specific response or action.

[0072] To implement this invention, a system is required in which a server, a terminal, and a user work in close cooperation. The server functions as a central information processing device responsible for generating virtual characters, analyzing user statements, and generating feedback. Specifically, the server utilizes natural language processing engines such as Google® Cloud Natural Language API and IBM Watson® for natural language processing, and platforms such as Unity and Unreal Engine for generating virtual characters.

[0073] The terminal acts as a user interface, creating an environment where users can participate in dialogue training. When a user logs into the system using the terminal, they are presented with selectable scenarios through a web application. By selecting from scenarios such as "Cross-Cultural Customer Service" or "Handling Complex Complaints," users can begin learning based on a specific scenario.

[0074] Based on the scenario selected by the user, the server uses a generative AI model to create a virtual character suitable for that scenario. The generative AI model enables diverse and realistic simulations, allowing the user to gain an experience close to actual work. For example, if the scenario is "handling cross-cultural customer service," the prompt message provided is "Your task is to handle a complaint from an international customer. The customer speaks only Spanish, and you need to resolve their issue using the virtual assistant," preparing an environment for the user to practice handling cross-cultural customer service.

[0075] As the conversation progresses, the server analyzes the user's statements in real time, evaluating their appropriateness and understanding of customer needs. Based on this analysis, feedback is generated at the end of the training session, reflecting the user's performance. This feedback is provided to the user via their device, allowing them to use the feedback to make improvements. The feedback is then stored in the training profile and used for future training sessions. This enables users to efficiently improve their customer service skills and strengthen their ability to serve international customers.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The user logs into the system using a terminal. The user accesses the user interface via a web application and authenticates by entering a specific ID and password. This information is sent to the server, and if authenticated, a success message is displayed, and the process proceeds to the next step.

[0079] Step 2:

[0080] The user selects a training scenario through an interface provided on the terminal. The user chooses one scenario from a list such as "Cross-Cultural Customer Service" or "Handling Difficult Complaints," and the selected scenario information is sent to the server. Based on this input information, the server outputs a virtual character profile that matches the scenario.

[0081] Step 3:

[0082] The server generates a virtual character using a generative AI model based on the selected scenario. The server uses the input scenario information to set the character's personality and dialogue patterns using the generative AI model, thereby generating a virtual customer. The generated character information is output to the terminal, allowing the user to begin interacting with the character.

[0083] Step 4:

[0084] The user initiates a conversation with a virtual character using a text field or voice input function on their device. The user's input data is sent to a server, which analyzes the utterance using natural language processing. The analysis results evaluate the intent and emotion of the utterance, and this information is stored. This analysis data forms the basis for providing real-time feedback to the user.

[0085] Step 5:

[0086] After the dialogue session ends, the server uses the accumulated analytical data to evaluate the user's performance and generate feedback. The server analyzes the appropriateness and areas for improvement of the user's responses and outputs feedback that includes specific improvement suggestions. This feedback data is sent to the terminal and becomes available for the user to view.

[0087] Step 6:

[0088] Feedback is provided to the user through the device. The user refers to the feedback provided and creates a self-improvement plan to enhance their skills. This feedback is recorded in the training profile by the server and used for future learning and evaluation. This allows the user's progress to be tracked and used to improve future training sessions.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] In recent years, there has been a growing demand for improved customer service skills in brick-and-mortar stores, but traditional methods make effective training difficult. In particular, developing the ability to respond appropriately and instantly to diverse customer needs in today's business environment is a challenge. Furthermore, with increasing globalization, multilingual support has become essential.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating human dialogue in real time, and means for generating feedback based on a person's work performance capabilities. This enables the rapid and effective improvement of customer service skills of staff in physical stores using a personal information display device.

[0094] A "virtual character" is a digital character created by a computer that simulates dialogue and actions.

[0095] "Real-time analysis" refers to a processing technology that instantly evaluates user input and provides immediate responses.

[0096] "Job performance ability" refers to the knowledge and skills necessary to effectively carry out tasks.

[0097] "Feedback" refers to information or advice that provides areas for improvement or evaluation based on user behavior and responses.

[0098] A "personal information display device" is an electronic device that can be worn or carried by an individual and used by that individual, and which has the function of visually displaying information.

[0099] To implement this invention, three parties—a server, a terminal, and a user—must work in cooperation. The server functions as a central information processing device for generating digital characters and simulating interactions with the user. The terminal utilizes personal information display devices or smart devices to provide an intuitive interface with the user.

[0100] The server analyzes user utterances using natural language processing technology and provides real-time evaluation. Specifically, this analysis process utilizes natural language processing libraries such as spaCy and BERT. Based on the evaluation results, feedback is generated to improve the user's work performance. This feedback is visually represented and provided to the user via a terminal. The provided feedback indicates specific areas for improvement that will help the user efficiently enhance their customer service skills.

[0101] As a concrete example, store staff could wear smart glasses and train their conversational skills by conversing with virtual characters representing customers from different cultures. The server would analyze the appropriateness of the words and responses used by the user in this conversation and visually present areas for improvement.

[0102] An example of a prompt message to input into a generative AI model is as follows:

[0103] "This session is for training customer service skills in a retail environment. The user will interact with a virtual customer character who speaks English. The system stitches the user's responses and provides real-time feedback on how to improve their service quality."

[0104] Through this system, users will be able to effectively improve their practical customer service skills.

[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0106] Step 1:

[0107] The user logs in via a terminal. The user selects a virtual customer service scenario that suits their purpose. Here, the user enters their ID information using an input terminal. As output, the selected scenario information is transmitted to the server.

[0108] Step 2:

[0109] The server generates a virtual character based on the selected scenario. In this process, the server uses a generation AI model and prompts to create a character suitable for the scenario. The output is the character's digital profile.

[0110] Step 3:

[0111] The server analyzes user speech in real time. The input is audio data sent from the user via their terminal. The server converts this data into text using natural language processing libraries (such as spaCy or BERT) and analyzes its content. The analysis results are output as a response quality evaluation score.

[0112] Step 4:

[0113] The server generates feedback based on the user's ability to perform their tasks. The input is the response quality evaluation score from the analysis results. Based on this, the server generates feedback that includes areas for improvement and strengths. The output is the text data of the feedback.

[0114] Step 5:

[0115] The terminal visualizes the generated feedback and provides it to the user. The input is text data of the feedback sent from the server, which the terminal displays on its screen. Based on the feedback, the user can strive to improve their customer service skills. The output is specific improvement guidelines provided to the user.

[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0117] As an embodiment of the present invention, a virtual dialogue system incorporating an emotion engine that recognizes emotions is introduced, with the cooperation of a server, a terminal, and a user. The server is primarily responsible for information processing, analyzing the user's statements and emotions and providing feedback. The terminal functions as a user interface, enabling interaction between the user and a virtual character.

[0118] First, the user logs into the system using a terminal and selects a virtual customer service scenario. The server generates a virtual character based on the selected scenario and prepares for the interaction. The user then begins interacting with the virtual character through voice or text input.

[0119] In this process, an emotion engine operates to recognize the user's emotional state during the conversation in real time. For example, if the user displays an unpleasant facial expression or tone of voice during a conversation, the emotion engine detects this and notifies the server. Based on the received emotional information, the server adjusts the virtual character's response to provide an interaction that is appropriate to the user's emotional state.

[0120] Once the training is complete, the server evaluates the user's performance by combining dialogue content and emotional data. The generated feedback takes into account the user's emotional responses and includes more personalized improvement suggestions. This feedback is provided to the user via their device, allowing them to review it and incorporate it into their next training session.

[0121] Furthermore, the server records training progress, including emotional data, and uses it for long-term analysis. This allows users to understand their skill development in detail, including changes in emotions. This invention enables training in advanced customer service techniques that incorporate emotion recognition, allowing users to provide emotionally sensitive customer service in actual work situations.

[0122] The following describes the processing flow.

[0123] Step 1:

[0124] The user enters their ID and password to log in to the system using their device. The device sends this information to the server and initiates the authentication process.

[0125] Step 2:

[0126] The server verifies the submitted login information against the database and performs the authentication process. If authentication is successful, it sends the data that makes up the user's dashboard to the device.

[0127] Step 3:

[0128] From a dashboard displayed on the device, the user selects a training scenario, such as one related to "stress management" or "nonverbal communication." The selected scenario is then notified to the server.

[0129] Step 4:

[0130] Based on the selected scenario information, the server generates an appropriate virtual character and simultaneously activates an emotion engine for emotion recognition.

[0131] Step 5:

[0132] The terminal displays a virtual character to the user and prepares to start a virtual dialogue scenario. The user begins interacting with the character using voice or text input.

[0133] Step 6:

[0134] The server analyzes the user's speech and visual data transmitted from the device in real time using an emotion engine to evaluate the user's emotional state.

[0135] Step 7:

[0136] Based on the analysis results of the emotion engine, the server adjusts the virtual character's response and generates feedback that is appropriate to the user's emotions.

[0137] Step 8:

[0138] Once the dialogue session ends, the server evaluates the user's training results based on all dialogue logs and sentiment data, and generates feedback that includes specific areas for improvement.

[0139] Step 9:

[0140] The feedback is sent to the device and presented to the user. Based on this, the user incorporates skill improvement measures that take emotional responses into account for the next training session.

[0141] Step 10:

[0142] The server stores training progress data, including collected emotional data, and records it as long-term analytical data to track user growth. This allows users to see how their emotional response skills have changed over time.

[0143] (Example 2)

[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0145] Traditional customer service techniques suffer from problems such as a lack of training that takes customer emotions into account and a lack of individualized feedback to users. As a result, it has been difficult for users to develop sufficient emotional response skills in actual customer service work. Furthermore, in multilingual situations, the loss of emotional nuances due to language translation has become a challenge.

[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0147] In this invention, the server includes means for generating and performing virtual characters, means for performing real-time emotion analysis of user speech or input data, and means for dynamically adjusting the virtual character's responses based on the emotion analysis results. This enables effective training that takes into account the user's emotional state and provides personalized feedback. Furthermore, it can improve the quality of communication by providing more natural, emotion-based translations even in multilingual environments.

[0148] "Means for generating and performing virtual characters" refers to a function in which a computer system uses its processing power to create a virtual entity for interaction with the user, and to control and display its responses and actions.

[0149] "Means for real-time sentiment analysis of user speech or input data" refers to technology that analyzes voice and text data provided by users and instantly identifies the emotions contained within it.

[0150] "Means for dynamically adjusting the virtual character's response based on emotion analysis results" refers to a function that changes the content of the virtual character's response and attitude in real time according to the detected emotions of the user.

[0151] "Means for generating personalized feedback" refers to a function that provides improvement suggestions and advice tailored to a specific user, based on the user's conversation history and sentiment data.

[0152] "Means for recording training progress data and analyzing long-term skill growth" refers to technologies that continuously save a user's learning history and conversation content, and measure skill improvement based on that data.

[0153] The "translation function that enables virtual dialogue in multiple languages ​​and further performs emotion-based translation" is a function that supports communication between different languages ​​and reflects emotional nuances when translating languages.

[0154] This invention provides a virtual dialogue system equipped with emotion analysis capabilities, through the collaboration of a server, a terminal, and a user. The system begins when the user logs in using a terminal and starts interacting with a virtual character.

[0155] First, the server receives login information from the user's device and performs user authentication. Once the user successfully logs in, the server receives the user's selection from several virtual customer service scenarios provided. Based on the selection, the server generates an appropriate virtual character using a generative AI model. The AI ​​model is based on natural language processing and machine learning techniques and has a mechanism to provide diverse dialogue patterns and responses that respond to emotions.

[0156] Next, the terminal functions as a user interface to enable interaction between the user and the virtual character. The user converses with the virtual character in real time through voice input or text input. The server uses an emotion analysis engine to analyze the emotional tone and keywords from the user's input data and identify the emotional state.

[0157] The server dynamically adjusts the virtual character's responses based on the sentiment analysis results. For example, if the user asks, "Could you tell me a little more?", the character will provide detailed information such as, "As a specific example of a recommendation, we have our special seafood pasta." If the user seems dissatisfied, the character's response tone can be adjusted, and suggestions for improvement can be offered.

[0158] Through this interaction, the server collects data to evaluate the user's performance and generates personalized feedback. This feedback is provided via the terminal so that the user can use it for future training sessions. The feedback specifically highlights areas where the user's skills can be improved and where their emotional response techniques can be enhanced.

[0159] Furthermore, this system supports multiple languages ​​and includes emotion-based translation capabilities, enabling natural dialogue even in different language environments. For example, by giving the AI ​​model instructions such as, "Explain how the virtual character should respond optimally based on user input," it becomes possible to explore more detailed ways to adjust the dialogue.

[0160] Through this format, users can effectively learn advanced customer service skills that take emotions into consideration.

[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0162] Step 1:

[0163] The user logs into the system using their device. The user enters their username and password, and the device sends this authentication information to the server. The server checks the user database to verify that the entered information is correct, and if successful, displays the home screen. As output, the user receives a screen showing the successful login.

[0164] Step 2:

[0165] The user selects their desired scenario from a list of virtual customer service scenarios displayed on their device. The device sends the selected scenario information to the server. Based on the received scenario information, the server uses a generation AI model to create the target virtual character. As output, the server sends the character and scenario settings to the device.

[0166] Step 3:

[0167] The user initiates a conversation with a virtual character by inputting voice or text through a terminal. The input data is sent from the terminal to the server. The server uses an emotion analysis engine to analyze the emotions contained in the user's input. As a result of the analysis, data is generated that determines the user's emotional state in real time. The analyzed emotion data is obtained as output.

[0168] Step 4:

[0169] The server dynamically adjusts the virtual character's responses based on the analyzed emotion data. Depending on the detected emotion, it changes the tone of the information and questions the virtual character provides. For example, if the user shows interest, the character will provide a more detailed explanation. The adapted response data is sent to the terminal as output and displayed to the user.

[0170] Step 5:

[0171] After a dialogue session ends, the server evaluates the user's performance based on their dialogue history and emotional data. Using a generative AI model, it performs analysis based on evaluation criteria and generates personalized feedback. This feedback includes suggestions for improving the user's emotional responses and skills. The feedback data is then sent to the terminal and presented to the user.

[0172] Step 6:

[0173] The server records user training progress data and analyzes long-term skill growth. The collected data is stored for later evaluation and improvement. This allows for a detailed understanding of how users improve their skills over time. As output, the recorded progress data is stored on the server.

[0174] (Application Example 2)

[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0176] Conventional customer service training systems have the problem of being unable to accurately understand user emotions and learn customer service methods based on them. The present invention aims to improve customer service skills by analyzing user emotions in real time and supporting appropriate customer service methods.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0178] In this invention, the server includes means for recognizing emotions, means for adjusting responses based on emotion analysis results, and means for analyzing emotional states in real time. This enables the provision of immediate feedback based on the user's emotions and improves personalized customer service techniques.

[0179] "Means of recognizing emotions" refers to technologies that analyze a user's facial expressions, tone of voice, etc., to identify their emotional state.

[0180] "Means for adjusting responses based on emotion analysis results" refers to techniques for appropriately modifying the responses of virtual dialogue characters or systems using analyzed emotion data.

[0181] "Methods for analyzing emotional states in real time" refer to technologies that instantly evaluate emotions during user interactions and provide immediate analysis results.

[0182] "Means for generating feedback to support customer service" refers to technologies that generate guidelines and suggestions for effective customer service based on the user's emotional state and the content of their conversations.

[0183] "Means for displaying generated feedback" refers to technologies for presenting feedback obtained from a system to the user visually or audibly.

[0184] "Means for recording and managing user progress data" refers to technologies for continuously saving and managing the progress of users' customer service training and the history of their emotional changes.

[0185] To implement this invention, a server, terminal, and user must cooperate and use the following configuration: The server is a system equipped with an emotion engine that recognizes emotions and an algorithm for generating feedback. The terminal is a device that functions as the user's interface, collecting dialogue content in real time and sending the data to the server for analysis. This system operates using software libraries specialized for image processing and speech analysis.

[0186] Specifically, a smartphone is used as the terminal, and customer facial expression data is captured using the OpenCV library, while emotions are analyzed from audio data using TENSORFLOW®. This data is sent to a server, which uses a generative AI model to generate feedback based on the analysis results. This feedback is then displayed on the terminal in real time, providing immediate advice to the user.

[0187] Furthermore, users can improve the quality of their customer service through their interactions with the system. For example, they can quickly detect changes in customer emotions during interactions, enabling them to respond flexibly and appropriately.

[0188] For example, if a user working in a bookstore shows interest in a customer's explanation of a particular product, feedback will appear stating, "You should continue with the detailed explanation." In this way, it is possible to efficiently improve customer service skills.

[0189] Furthermore, as an example of a prompt using a generative AI model, you can use a sentence like this: "Analyze the customer's facial expressions and voice, and suggest appropriate customer service methods. For example, I would like advice on the difference in how to respond when the customer is smiling and showing interest versus when they are showing confusion."

[0190] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0191] Step 1:

[0192] The device uses a camera and microphone to capture the facial expressions and voices of users and customers in real time. The input consists of visual and audio data. This data is passed to OpenCV, an image processing tool, and TensorFlow, an audio processing library, to perform facial feature extraction and audio data preprocessing.

[0193] Step 2:

[0194] The server receives pre-processed facial expression and audio data and begins emotion analysis. The input consists of feature-extracted visual and audio data. The emotion engine analyzes this data using a generated AI model to recognize the user's emotional state in real time. The output is data containing the emotion analysis results.

[0195] Step 3:

[0196] The server adjusts the virtual customer service scenario based on the analysis results. The input is the emotion analysis result, and the system determines the optimal character response pattern based on it. The server dynamically adjusts the dialogue scenario and generates data for feedback. The output is the adjusted response pattern.

[0197] Step 4:

[0198] The terminal presents information to the user visually or audibly based on feedback data received from the server. The input is feedback data. The presented information includes suggestions for specific customer service actions and approaches tailored to the customer's emotions. The user can use this information to adapt their behavior. The output is displaying the feedback on the user's screen.

[0199] Step 5:

[0200] The user adjusts their customer service behavior based on feedback and interacts with the customer. The user's responses become part of the data capture from the next processing step 1, helping to continuously improve the system. The input is feedback from the system, and the output is the user's adapted customer service behavior.

[0201] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0202] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0203] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0204] [Second Embodiment]

[0205] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0206] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0207] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0208] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0209] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0210] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0211] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0212] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0213] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0214] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0215] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0216] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0217] To implement this invention, three parties—a server, a terminal, and a user—must work closely together. The server, as the central information processing device, is responsible for generating virtual characters, analyzing dialogue, and generating feedback. The terminal provides a user interface and functions as a device that enables interaction with the user.

[0218] First, the user logs into the system via a terminal and selects a virtual customer service scenario that suits their purpose. For example, they can choose scenarios such as "handling cross-cultural customer interactions" or "resolving complex customer complaints." After selection, the server generates a virtual character based on the specified scenario. This character simulates interaction with the user, providing a simulated training environment for honing customer service skills.

[0219] Once a conversation begins, the server analyzes the user's statements in real time. This analysis uses natural language processing to evaluate factors such as tone, appropriateness of responses, and the ability to understand customer needs. The analyzed information is immediately recorded and evaluated within the server.

[0220] Once the training is complete, the server generates feedback based on the user's performance. This feedback is provided to the user via their device, specifically highlighting areas for improvement and areas where they performed well. The user can use this feedback to further enhance their skills. The feedback is also stored as data in the user's training profile when it is generated.

[0221] This system allows users to improve their customer service skills through practical and diverse scenarios, often without requiring significant human intervention. Furthermore, its multilingual capabilities enhance its ability to serve international customers. This enables users to efficiently learn the skills required in real-world work environments, contributing to improved work performance.

[0222] The following describes the processing flow.

[0223] Step 1:

[0224] The user attempts to log in to the system using their device. Logging in requires personal authentication information, so the user enters their ID and password. The device then sends the entered information to the server.

[0225] Step 2:

[0226] The server compares the received login information with the registered information in the database and performs authentication. If authentication is successful, it prepares to display the user's dashboard on the terminal.

[0227] Step 3:

[0228] The user starts training from a dashboard on their device, selecting one of several scenarios provided. The selection is sent to the server, which retrieves information about the scenario.

[0229] Step 4:

[0230] The server generates a virtual character based on the scenario selected by the user. The generated character has its dialogue and situation pre-configured, and is ready to begin the scenario.

[0231] Step 5:

[0232] The terminal displays a virtual character to the user and initiates a virtual dialogue scenario. The user responds to the character using voice input or text input.

[0233] Step 6:

[0234] The server receives user input in real time and performs analysis using natural language processing. The analysis results take into account the appropriateness of the content of the statements, the wording, and the flow of the conversation.

[0235] Step 7:

[0236] Once the training is complete, the server evaluates the user's performance based on the analysis results. From the evaluation results, it extracts specific strengths and areas for improvement.

[0237] Step 8:

[0238] The server creates feedback based on the generated evaluation and sends it to the terminal. The terminal displays the feedback to the user, who then reviews it.

[0239] Step 9:

[0240] The server stores and records data about the user's training sessions in a format that can be referenced later. Based on this data, users can develop long-term skill improvement plans.

[0241] (Example 1)

[0242] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0243] In today's business environment, customer service skills are crucial for improving customer satisfaction, but traditional training methods have struggled to efficiently provide practical and diverse scenario-based training. Furthermore, efficiently strengthening the ability to handle multilingual international customers has also been a challenging task. In response, there was a need for a feedback system tailored to individual skills, a means to visualize user capabilities, and a way to promote improvement.

[0244] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0245] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating user dialogue in real time, and means for generating feedback based on user performance. This enables users to improve their customer service skills through practical and diverse scenarios. Furthermore, analysis using natural language processing can efficiently enhance multilingual communication skills. In addition, by combining a generative AI model with prompt sentences, it is possible to appropriately guide user responses and achieve higher quality dialogue training.

[0246] A "virtual character" is a digital humanoid entity created to simulate interaction with a user, and can have different personalities and backgrounds based on various scenarios.

[0247] A "user interface" is a general term for the screens and tools that users use to access and operate a system, and is used for inputting dialogues, selecting scenarios, and so on.

[0248] "Natural language processing" is a technology used to enable computers to understand and analyze human language. It is used to analyze user statements and provide appropriate responses and evaluations.

[0249] "Feedback" is evaluation information provided based on the results of user interaction training, indicating areas for improvement and strengths, and is useful for improving the user's skills.

[0250] A "training profile" is a collection of information that records a user's training progress and past evaluation results, and is referenced for future skill improvement.

[0251] A "generative AI model" is a type of artificial intelligence that generates appropriate information or characters in response to user requests, based on knowledge learned from large amounts of data.

[0252] A "prompt message" is a phrase provided to give instructions or draw attention to a user, and its role is to guide them to a specific response or action.

[0253] To implement this invention, a system is required in which a server, a terminal, and a user work closely together. The server functions as a central information processing device responsible for generating virtual characters, analyzing user statements, and generating feedback. Specifically, the server utilizes natural language processing engines such as Google Cloud Natural Language API and IBM Watson for natural language processing, and platforms such as Unity and Unreal Engine for generating virtual characters.

[0254] The terminal acts as a user interface, creating an environment where users can participate in dialogue training. When a user logs into the system using the terminal, they are presented with selectable scenarios through a web application. By selecting from scenarios such as "Cross-Cultural Customer Service" or "Handling Complex Complaints," users can begin learning based on a specific scenario.

[0255] Based on the scenario selected by the user, the server uses a generative AI model to create a virtual character suitable for that scenario. The generative AI model enables diverse and realistic simulations, allowing the user to gain an experience close to actual work. For example, if the scenario is "handling cross-cultural customer service," the prompt message provided is "Your task is to handle a complaint from an international customer. The customer speaks only Spanish, and you need to resolve their issue using the virtual assistant," preparing an environment for the user to practice handling cross-cultural customer service.

[0256] As the conversation progresses, the server analyzes the user's statements in real time, evaluating their appropriateness and understanding of customer needs. Based on this analysis, feedback is generated at the end of the training session, reflecting the user's performance. This feedback is provided to the user via their device, allowing them to use the feedback to make improvements. The feedback is then stored in the training profile and used for future training sessions. This enables users to efficiently improve their customer service skills and strengthen their ability to serve international customers.

[0257] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0258] Step 1:

[0259] The user logs into the system using a terminal. The user accesses the user interface via a web application and authenticates by entering a specific ID and password. This information is sent to the server, and if authenticated, a success message is displayed, and the process proceeds to the next step.

[0260] Step 2:

[0261] The user selects a training scenario through an interface provided on the terminal. The user chooses one scenario from a list such as "Cross-Cultural Customer Service" or "Handling Difficult Complaints," and the selected scenario information is sent to the server. Based on this input information, the server outputs a virtual character profile that matches the scenario.

[0262] Step 3:

[0263] The server generates a virtual character using a generative AI model based on the selected scenario. The server uses the input scenario information to set the character's personality and dialogue patterns using the generative AI model, thereby generating a virtual customer. The generated character information is output to the terminal, allowing the user to begin interacting with the character.

[0264] Step 4:

[0265] The user initiates a conversation with a virtual character using a text field or voice input function on their device. The user's input data is sent to a server, which analyzes the utterance using natural language processing. The analysis results evaluate the intent and emotion of the utterance, and this information is stored. This analysis data forms the basis for providing real-time feedback to the user.

[0266] Step 5:

[0267] After the dialogue session ends, the server uses the accumulated analytical data to evaluate the user's performance and generate feedback. The server analyzes the appropriateness and areas for improvement of the user's responses and outputs feedback that includes specific improvement suggestions. This feedback data is sent to the terminal and becomes available for the user to view.

[0268] Step 6:

[0269] Feedback is provided to the user through the device. The user refers to the feedback provided and creates a self-improvement plan to enhance their skills. This feedback is recorded in the training profile by the server and used for future learning and evaluation. This allows the user's progress to be tracked and used to improve future training sessions.

[0270] (Application Example 1)

[0271] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0272] In recent years, there has been a growing demand for improved customer service skills in brick-and-mortar stores, but traditional methods make effective training difficult. In particular, developing the ability to respond appropriately and instantly to diverse customer needs in today's business environment is a challenge. Furthermore, with increasing globalization, multilingual support has become essential.

[0273] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0274] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating human dialogue in real time, and means for generating feedback based on a person's work performance capabilities. This enables the rapid and effective improvement of customer service skills of staff in physical stores using a personal information display device.

[0275] A "virtual character" is a digital character created by a computer that simulates dialogue and actions.

[0276] "Real-time analysis" refers to a processing technology that instantly evaluates user input and provides immediate responses.

[0277] "Job performance ability" refers to the knowledge and skills necessary to effectively carry out tasks.

[0278] "Feedback" refers to information or advice that provides areas for improvement or evaluation based on user behavior and responses.

[0279] A "personal information display device" is an electronic device that can be worn or carried by an individual and used by that individual, and which has the function of visually displaying information.

[0280] To implement this invention, three parties, namely the server, the terminal, and the user, need to cooperate and operate. The server functions as a central information processing device for generating digital characters and simulating interactions with users. The terminal utilizes personal information display devices and smart devices to provide an intuitive interface for users.

[0281] The server analyzes the user's speech using natural language processing technology and conducts real-time evaluations. Specifically, natural language processing libraries such as spaCy and BERT are used in this analysis process. Based on the evaluation results, feedback is generated to improve the user's business performance ability. This feedback is visually presented through the terminal and provided to the user. The provided feedback indicates specific areas for improvement for the user to efficiently enhance their customer service skills.

[0282] As a specific example, it is conceivable that a staff member in a physical store wears smart glasses and trains their dialogue skills through conversations with virtual customer characters for serving customers from different cultures. The server analyzes the words used by the user in this conversation and the appropriateness of their responses, and visually presents the points that need to be improved.

[0283] Examples of prompt texts input into the generative AI model are as follows.

[0284] "This session is for training customer service skills in a retail environment. The user will interact with a virtual customer character who speaks English. The system analyzes the user's responses and provides real-time feedback on how to improve their service quality."

[0285] Through this system, users can effectively improve their practical customer service skills.

[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0287] Step 1:

[0288] The user logs in through the terminal. The user selects a virtual customer service scenario according to the purpose. Here, the user uses the input terminal to enter their ID information. As output, the selected scenario information is transmitted to the server.

[0289] Step 2:

[0290] Based on the selected scenario, the server generates a virtual character. At this time, the server uses the generation AI model and uses the prompt text to generate a character suitable for the scenario. The output is the digital profile of the character.

[0291] Step 3:

[0292] The server analyzes the user's speech in real time. The input is the voice data transmitted from the user via the terminal. The server converts this data into text using a natural language processing library (such as spaCy or BERT) and analyzes the content. The analysis result is output as a response quality evaluation score.

[0293] Step 4:

[0294] Based on the user's business performance ability, the server generates feedback. The input is the response quality evaluation score of the analysis result. Based on this, the server generates feedback including points to be improved and excellent points. The output is the text data of the feedback.

[0295] Step 5:

[0296] The terminal visualizes the generated feedback and provides it to the user. The input is text data of the feedback sent from the server, which the terminal displays on its screen. Based on the feedback, the user can strive to improve their customer service skills. The output is specific improvement guidelines provided to the user.

[0297] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0298] As an embodiment of the present invention, a virtual dialogue system incorporating an emotion engine that recognizes emotions is introduced, with the cooperation of a server, a terminal, and a user. The server is primarily responsible for information processing, analyzing the user's statements and emotions and providing feedback. The terminal functions as a user interface, enabling interaction between the user and a virtual character.

[0299] First, the user logs into the system using a terminal and selects a virtual customer service scenario. The server generates a virtual character based on the selected scenario and prepares for the interaction. The user then begins interacting with the virtual character through voice or text input.

[0300] In this process, an emotion engine operates to recognize the user's emotional state during the conversation in real time. For example, if the user displays an unpleasant facial expression or tone of voice during a conversation, the emotion engine detects this and notifies the server. Based on the received emotional information, the server adjusts the virtual character's response to provide an interaction that is appropriate to the user's emotional state.

[0301] When the training is completed, the server evaluates the user's performance by integrating the conversation content and emotion data. The generated feedback takes into account the user's emotional reactions and includes more personalized improvement suggestions. This feedback is provided to the user through the terminal, and the user can review it and reflect it in the next training.

[0302] Furthermore, the server records the training progress including emotion data and utilizes it for long-term analysis. As a result, the user can grasp in detail the growth of skills including changes in emotions. With this invention, training in advanced customer service techniques incorporating emotion recognition is realized, and the user can provide customer service that takes emotions into account even in actual business.

[0303] The following describes the processing flow.

[0304] Step 1:

[0305] The user inputs an ID and password to log in to the system using the terminal. The terminal sends this information to the server and starts the authentication process.

[0306] Step 2:

[0307] The server collates the transmitted login information with the database and executes the authentication process. If the authentication is successful, it sends the data constituting the user's dashboard to the terminal.

[0308] Step 3:

[0309] The user selects a scenario for training, such as one related to "stress management" or "non-verbal communication", from the dashboard displayed on the terminal. The selected scenario is notified to the server.

[0310] Step 4:

[0311] Based on the selected scenario information, the server generates an appropriate virtual character and simultaneously activates an emotion engine for emotion recognition.

[0312] Step 5:

[0313] The terminal displays a virtual character to the user and prepares to start a virtual dialogue scenario. The user begins interacting with the character using voice or text input.

[0314] Step 6:

[0315] The server analyzes the user's speech and visual data transmitted from the device in real time using an emotion engine to evaluate the user's emotional state.

[0316] Step 7:

[0317] Based on the analysis results of the emotion engine, the server adjusts the virtual character's response and generates feedback that is appropriate to the user's emotions.

[0318] Step 8:

[0319] Once the dialogue session ends, the server evaluates the user's training results based on all dialogue logs and sentiment data, and generates feedback that includes specific areas for improvement.

[0320] Step 9:

[0321] The feedback is sent to the device and presented to the user. Based on this, the user incorporates skill improvement measures that take emotional responses into account for the next training session.

[0322] Step 10:

[0323] The server stores training progress data, including collected emotional data, and records it as long-term analytical data to track user growth. This allows users to see how their emotional response skills have changed over time.

[0324] (Example 2)

[0325] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0326] Traditional customer service techniques suffer from problems such as a lack of training that takes customer emotions into account and a lack of individualized feedback to users. As a result, it has been difficult for users to develop sufficient emotional response skills in actual customer service work. Furthermore, in multilingual situations, the loss of emotional nuances due to language translation has become a challenge.

[0327] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0328] In this invention, the server includes means for generating and performing virtual characters, means for performing real-time emotion analysis of user speech or input data, and means for dynamically adjusting the virtual character's responses based on the emotion analysis results. This enables effective training that takes into account the user's emotional state and provides personalized feedback. Furthermore, it can improve the quality of communication by providing more natural, emotion-based translations even in multilingual environments.

[0329] "Means for generating and performing virtual characters" refers to a function in which a computer system uses its processing power to create a virtual entity for interaction with the user, and to control and display its responses and actions.

[0330] "Means for real-time sentiment analysis of user speech or input data" refers to technology that analyzes voice and text data provided by users and instantly identifies the emotions contained within it.

[0331] "Means for dynamically adjusting the virtual character's response based on emotion analysis results" refers to a function that changes the content of the virtual character's response and attitude in real time according to the detected emotions of the user.

[0332] "Means for generating personalized feedback" refers to a function that provides improvement suggestions and advice tailored to a specific user, based on the user's conversation history and sentiment data.

[0333] "Means for recording training progress data and analyzing long-term skill growth" refers to technologies that continuously save a user's learning history and conversation content, and measure skill improvement based on that data.

[0334] The "translation function that enables virtual dialogue in multiple languages ​​and further performs emotion-based translation" is a function that supports communication between different languages ​​and reflects emotional nuances when translating languages.

[0335] This invention provides a virtual dialogue system equipped with emotion analysis capabilities, through the collaboration of a server, a terminal, and a user. The system begins when the user logs in using a terminal and starts interacting with a virtual character.

[0336] First, the server receives login information from the user's device and performs user authentication. Once the user successfully logs in, the server receives the user's selection from several virtual customer service scenarios provided. Based on the selection, the server generates an appropriate virtual character using a generative AI model. The AI ​​model is based on natural language processing and machine learning techniques and has a mechanism to provide diverse dialogue patterns and responses that respond to emotions.

[0337] Next, the terminal functions as a user interface to enable interaction between the user and the virtual character. The user converses with the virtual character in real time through voice input or text input. The server uses an emotion analysis engine to analyze the emotional tone and keywords from the user's input data and identify the emotional state.

[0338] The server dynamically adjusts the virtual character's responses based on the sentiment analysis results. For example, if the user asks, "Could you tell me a little more?", the character will provide detailed information such as, "As a specific example of a recommendation, we have our special seafood pasta." If the user seems dissatisfied, the character's response tone can be adjusted, and suggestions for improvement can be offered.

[0339] Through this interaction, the server collects data to evaluate the user's performance and generates personalized feedback. This feedback is provided via the terminal so that the user can use it for future training sessions. The feedback specifically highlights areas where the user's skills can be improved and where their emotional response techniques can be enhanced.

[0340] Furthermore, this system supports multiple languages ​​and includes emotion-based translation capabilities, enabling natural dialogue even in different language environments. For example, by giving the AI ​​model instructions such as, "Explain how the virtual character should respond optimally based on user input," it becomes possible to explore more detailed ways to adjust the dialogue.

[0341] Through this format, users can effectively learn advanced customer service skills that take emotions into consideration.

[0342] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0343] Step 1:

[0344] The user logs into the system using their device. The user enters their username and password, and the device sends this authentication information to the server. The server checks the user database to verify that the entered information is correct, and if successful, displays the home screen. As output, the user receives a screen showing the successful login.

[0345] Step 2:

[0346] The user selects their desired scenario from a list of virtual customer service scenarios displayed on their device. The device sends the selected scenario information to the server. Based on the received scenario information, the server uses a generation AI model to create the target virtual character. As output, the server sends the character and scenario settings to the device.

[0347] Step 3:

[0348] The user initiates a conversation with a virtual character by inputting voice or text through a terminal. The input data is sent from the terminal to the server. The server uses an emotion analysis engine to analyze the emotions contained in the user's input. As a result of the analysis, data is generated that determines the user's emotional state in real time. The analyzed emotion data is obtained as output.

[0349] Step 4:

[0350] The server dynamically adjusts the virtual character's responses based on the analyzed emotion data. Depending on the detected emotion, it changes the tone of the information and questions the virtual character provides. For example, if the user shows interest, the character will provide a more detailed explanation. The adapted response data is sent to the terminal as output and displayed to the user.

[0351] Step 5:

[0352] After a dialogue session ends, the server evaluates the user's performance based on their dialogue history and emotional data. Using a generative AI model, it performs analysis based on evaluation criteria and generates personalized feedback. This feedback includes suggestions for improving the user's emotional responses and skills. The feedback data is then sent to the terminal and presented to the user.

[0353] Step 6:

[0354] The server records user training progress data and analyzes long-term skill growth. The collected data is stored for later evaluation and improvement. This allows for a detailed understanding of how users improve their skills over time. As output, the recorded progress data is stored on the server.

[0355] (Application Example 2)

[0356] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0357] Conventional customer service training systems have the problem of being unable to accurately understand user emotions and learn customer service methods based on them. The present invention aims to improve customer service skills by analyzing user emotions in real time and supporting appropriate customer service methods.

[0358] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0359] In this invention, the server includes means for recognizing emotions, means for adjusting responses based on emotion analysis results, and means for analyzing emotional states in real time. This enables the provision of immediate feedback based on the user's emotions and improves personalized customer service techniques.

[0360] "Means of recognizing emotions" refers to technologies that analyze a user's facial expressions, tone of voice, etc., to identify their emotional state.

[0361] "Means for adjusting responses based on emotion analysis results" refers to techniques for appropriately modifying the responses of virtual dialogue characters or systems using analyzed emotion data.

[0362] "Methods for analyzing emotional states in real time" refer to technologies that instantly evaluate emotions during user interactions and provide immediate analysis results.

[0363] "Means for generating feedback to support customer service" refers to technologies that generate guidelines and suggestions for effective customer service based on the user's emotional state and the content of their conversations.

[0364] "Means for displaying generated feedback" refers to technologies for presenting feedback obtained from a system to the user visually or audibly.

[0365] "Means for recording and managing user progress data" refers to technologies for continuously saving and managing the progress of users' customer service training and the history of their emotional changes.

[0366] To implement this invention, a server, terminal, and user must cooperate and use the following configuration: The server is a system equipped with an emotion engine that recognizes emotions and an algorithm for generating feedback. The terminal is a device that functions as the user's interface, collecting dialogue content in real time and sending the data to the server for analysis. This system operates using software libraries specialized for image processing and speech analysis.

[0367] Specifically, a smartphone is used as the terminal, and customer facial expression data is captured using the OpenCV library, while emotions are analyzed from audio data using TensorFlow. This data is sent to a server, which uses a generative AI model to generate feedback based on the analysis results. This feedback is then displayed on the terminal in real time, providing immediate advice to the user.

[0368] Furthermore, users can improve the quality of their customer service through their interactions with the system. For example, they can quickly detect changes in customer emotions during interactions, enabling them to respond flexibly and appropriately.

[0369] For example, if a user working in a bookstore shows interest in a customer's explanation of a particular product, feedback will appear stating, "You should continue with the detailed explanation." In this way, it is possible to efficiently improve customer service skills.

[0370] Furthermore, as an example of a prompt using a generative AI model, you can use a sentence like this: "Analyze the customer's facial expressions and voice, and suggest appropriate customer service methods. For example, I would like advice on the difference in how to respond when the customer is smiling and showing interest versus when they are showing confusion."

[0371] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0372] Step 1:

[0373] The device uses a camera and microphone to capture the facial expressions and voices of users and customers in real time. The input consists of visual and audio data. This data is passed to OpenCV, an image processing tool, and TensorFlow, an audio processing library, to perform facial feature extraction and audio data preprocessing.

[0374] Step 2:

[0375] The server receives pre-processed facial expression and audio data and begins emotion analysis. The input consists of feature-extracted visual and audio data. The emotion engine analyzes this data using a generated AI model to recognize the user's emotional state in real time. The output is data containing the emotion analysis results.

[0376] Step 3:

[0377] The server adjusts the virtual customer service scenario based on the analysis results. The input is the emotion analysis result, and the system determines the optimal character response pattern based on it. The server dynamically adjusts the dialogue scenario and generates data for feedback. The output is the adjusted response pattern.

[0378] Step 4:

[0379] The terminal presents information to the user visually or audibly based on feedback data received from the server. The input is feedback data. The presented information includes suggestions for specific customer service actions and approaches tailored to the customer's emotions. The user can use this information to adapt their behavior. The output is displaying the feedback on the user's screen.

[0380] Step 5:

[0381] The user adjusts their customer service behavior based on feedback and interacts with the customer. The user's responses become part of the data capture from the next processing step 1, helping to continuously improve the system. The input is feedback from the system, and the output is the user's adapted customer service behavior.

[0382] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0383] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0384] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0385] [Third Embodiment]

[0386] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0387] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0388] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0389] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0390] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0392] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0393] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0394] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0395] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0396] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0397] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0398] To implement this invention, three parties—a server, a terminal, and a user—must work closely together. The server, as the central information processing device, is responsible for generating virtual characters, analyzing dialogue, and generating feedback. The terminal provides a user interface and functions as a device that enables interaction with the user.

[0399] First, the user logs into the system via a terminal and selects a virtual customer service scenario that suits their purpose. For example, they can choose scenarios such as "handling cross-cultural customer interactions" or "resolving complex customer complaints." After selection, the server generates a virtual character based on the specified scenario. This character simulates interaction with the user, providing a simulated training environment for honing customer service skills.

[0400] Once a conversation begins, the server analyzes the user's statements in real time. This analysis uses natural language processing to evaluate factors such as tone, appropriateness of responses, and the ability to understand customer needs. The analyzed information is immediately recorded and evaluated within the server.

[0401] Once the training is complete, the server generates feedback based on the user's performance. This feedback is provided to the user via their device, specifically highlighting areas for improvement and areas where they performed well. The user can use this feedback to further enhance their skills. The feedback is also stored as data in the user's training profile when it is generated.

[0402] This system allows users to improve their customer service skills through practical and diverse scenarios, often without requiring significant human intervention. Furthermore, its multilingual capabilities enhance its ability to serve international customers. This enables users to efficiently learn the skills required in real-world work environments, contributing to improved work performance.

[0403] The following describes the processing flow.

[0404] Step 1:

[0405] The user attempts to log in to the system using their device. Logging in requires personal authentication information, so the user enters their ID and password. The device then sends the entered information to the server.

[0406] Step 2:

[0407] The server compares the received login information with the registered information in the database and performs authentication. If authentication is successful, it prepares to display the user's dashboard on the terminal.

[0408] Step 3:

[0409] The user starts training from a dashboard on their device, selecting one of several scenarios provided. The selection is sent to the server, which retrieves information about the scenario.

[0410] Step 4:

[0411] The server generates a virtual character based on the scenario selected by the user. The generated character has its dialogue and situation pre-configured, and is ready to begin the scenario.

[0412] Step 5:

[0413] The terminal displays a virtual character to the user and initiates a virtual dialogue scenario. The user responds to the character using voice input or text input.

[0414] Step 6:

[0415] The server receives user input in real time and performs analysis using natural language processing. The analysis results take into account the appropriateness of the content of the statements, the wording, and the flow of the conversation.

[0416] Step 7:

[0417] Once the training is complete, the server evaluates the user's performance based on the analysis results. From the evaluation results, it extracts specific strengths and areas for improvement.

[0418] Step 8:

[0419] The server creates feedback based on the generated evaluation and sends it to the terminal. The terminal displays the feedback to the user, who then reviews it.

[0420] Step 9:

[0421] The server stores and records data about the user's training sessions in a format that can be referenced later. Based on this data, users can develop long-term skill improvement plans.

[0422] (Example 1)

[0423] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0424] In today's business environment, customer service skills are crucial for improving customer satisfaction, but traditional training methods have struggled to efficiently provide practical and diverse scenario-based training. Furthermore, efficiently strengthening the ability to handle multilingual international customers has also been a challenging task. In response, there was a need for a feedback system tailored to individual skills, a means to visualize user capabilities, and a way to promote improvement.

[0425] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0426] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating user dialogue in real time, and means for generating feedback based on user performance. This enables users to improve their customer service skills through practical and diverse scenarios. Furthermore, analysis using natural language processing can efficiently enhance multilingual communication skills. In addition, by combining a generative AI model with prompt sentences, it is possible to appropriately guide user responses and achieve higher quality dialogue training.

[0427] A "virtual character" is a digital humanoid entity created to simulate interaction with a user, and can have different personalities and backgrounds based on various scenarios.

[0428] A "user interface" is a general term for the screens and tools that users use to access and operate a system, and is used for inputting dialogues, selecting scenarios, and so on.

[0429] "Natural language processing" is a technology used to enable computers to understand and analyze human language. It is used to analyze user statements and provide appropriate responses and evaluations.

[0430] "Feedback" is evaluation information provided based on the results of user interaction training, indicating areas for improvement and strengths, and is useful for improving the user's skills.

[0431] A "training profile" is a collection of information that records a user's training progress and past evaluation results, and is referenced for future skill improvement.

[0432] A "generative AI model" is a type of artificial intelligence that generates appropriate information or characters in response to user requests, based on knowledge learned from large amounts of data.

[0433] A "prompt message" is a phrase provided to give instructions or draw attention to a user, and its role is to guide them to a specific response or action.

[0434] To implement this invention, a system is required in which a server, a terminal, and a user work closely together. The server functions as a central information processing device responsible for generating virtual characters, analyzing user statements, and generating feedback. Specifically, the server utilizes natural language processing engines such as Google Cloud Natural Language API and IBM Watson for natural language processing, and platforms such as Unity and Unreal Engine for generating virtual characters.

[0435] The terminal acts as a user interface, creating an environment where users can participate in dialogue training. When a user logs into the system using the terminal, they are presented with selectable scenarios through a web application. By selecting from scenarios such as "Cross-Cultural Customer Service" or "Handling Complex Complaints," users can begin learning based on a specific scenario.

[0436] Based on the scenario selected by the user, the server uses a generative AI model to create a virtual character suitable for that scenario. The generative AI model enables diverse and realistic simulations, allowing the user to gain an experience close to actual work. For example, if the scenario is "handling cross-cultural customer service," the prompt message provided is "Your task is to handle a complaint from an international customer. The customer speaks only Spanish, and you need to resolve their issue using the virtual assistant," preparing an environment for the user to practice handling cross-cultural customer service.

[0437] As the conversation progresses, the server analyzes the user's statements in real time, evaluating their appropriateness and understanding of customer needs. Based on this analysis, feedback is generated at the end of the training session, reflecting the user's performance. This feedback is provided to the user via their device, allowing them to use the feedback to make improvements. The feedback is then stored in the training profile and used for future training sessions. This enables users to efficiently improve their customer service skills and strengthen their ability to serve international customers.

[0438] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0439] Step 1:

[0440] The user logs into the system using a terminal. The user accesses the user interface via a web application and authenticates by entering a specific ID and password. This information is sent to the server, and if authenticated, a success message is displayed, and the process proceeds to the next step.

[0441] Step 2:

[0442] The user selects a training scenario through an interface provided on the terminal. The user chooses one scenario from a list such as "Cross-Cultural Customer Service" or "Handling Difficult Complaints," and the selected scenario information is sent to the server. Based on this input information, the server outputs a virtual character profile that matches the scenario.

[0443] Step 3:

[0444] The server generates a virtual character using a generative AI model based on the selected scenario. The server uses the input scenario information to set the character's personality and dialogue patterns using the generative AI model, thereby generating a virtual customer. The generated character information is output to the terminal, allowing the user to begin interacting with the character.

[0445] Step 4:

[0446] The user initiates a conversation with a virtual character using a text field or voice input function on their device. The user's input data is sent to a server, which analyzes the utterance using natural language processing. The analysis results evaluate the intent and emotion of the utterance, and this information is stored. This analysis data forms the basis for providing real-time feedback to the user.

[0447] Step 5:

[0448] After the dialogue session ends, the server uses the accumulated analytical data to evaluate the user's performance and generate feedback. The server analyzes the appropriateness and areas for improvement of the user's responses and outputs feedback that includes specific improvement suggestions. This feedback data is sent to the terminal and becomes available for the user to view.

[0449] Step 6:

[0450] Feedback is provided to the user through the device. The user refers to the feedback provided and creates a self-improvement plan to enhance their skills. This feedback is recorded in the training profile by the server and used for future learning and evaluation. This allows the user's progress to be tracked and used to improve future training sessions.

[0451] (Application Example 1)

[0452] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0453] In recent years, there has been a growing demand for improved customer service skills in brick-and-mortar stores, but traditional methods make effective training difficult. In particular, developing the ability to respond appropriately and instantly to diverse customer needs in today's business environment is a challenge. Furthermore, with increasing globalization, multilingual support has become essential.

[0454] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0455] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating human dialogue in real time, and means for generating feedback based on a person's work performance capabilities. This enables the rapid and effective improvement of customer service skills of staff in physical stores using a personal information display device.

[0456] A "virtual character" is a digital character created by a computer that simulates dialogue and actions.

[0457] "Real-time analysis" refers to a processing technology that instantly evaluates user input and provides immediate responses.

[0458] "Job performance ability" refers to the knowledge and skills necessary to effectively carry out tasks.

[0459] "Feedback" refers to information or advice that provides areas for improvement or evaluation based on user behavior and responses.

[0460] A "personal information display device" is an electronic device that can be worn or carried by an individual and used by that individual, and which has the function of visually displaying information.

[0461] To implement this invention, three parties—a server, a terminal, and a user—must work in cooperation. The server functions as a central information processing device for generating digital characters and simulating interactions with the user. The terminal utilizes personal information display devices or smart devices to provide an intuitive interface with the user.

[0462] The server analyzes user utterances using natural language processing technology and provides real-time evaluation. Specifically, this analysis process utilizes natural language processing libraries such as spaCy and BERT. Based on the evaluation results, feedback is generated to improve the user's work performance. This feedback is visually represented and provided to the user via a terminal. The provided feedback indicates specific areas for improvement that will help the user efficiently enhance their customer service skills.

[0463] As a concrete example, store staff could wear smart glasses and train their conversational skills by conversing with virtual characters representing customers from different cultures. The server would analyze the appropriateness of the words and responses used by the user in this conversation and visually present areas for improvement.

[0464] An example of a prompt message to input into a generative AI model is as follows:

[0465] "This session is for training customer service skills in a retail environment. The user will interact with a virtual customer character who speaks English. The system stitches the user's responses and provides real-time feedback on how to improve their service quality."

[0466] Through this system, users will be able to effectively improve their practical customer service skills.

[0467] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0468] Step 1:

[0469] The user logs in via a terminal. The user selects a virtual customer service scenario that suits their purpose. Here, the user enters their ID information using an input terminal. As output, the selected scenario information is transmitted to the server.

[0470] Step 2:

[0471] The server generates a virtual character based on the selected scenario. In this process, the server uses a generation AI model and prompts to create a character suitable for the scenario. The output is the character's digital profile.

[0472] Step 3:

[0473] The server analyzes user speech in real time. The input is audio data sent from the user via their terminal. The server converts this data into text using natural language processing libraries (such as spaCy or BERT) and analyzes its content. The analysis results are output as a response quality evaluation score.

[0474] Step 4:

[0475] The server generates feedback based on the user's ability to perform their tasks. The input is the response quality evaluation score from the analysis results. Based on this, the server generates feedback that includes areas for improvement and strengths. The output is the text data of the feedback.

[0476] Step 5:

[0477] The terminal visualizes the generated feedback and provides it to the user. The input is text data of the feedback sent from the server, which the terminal displays on its screen. Based on the feedback, the user can strive to improve their customer service skills. The output is specific improvement guidelines provided to the user.

[0478] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0479] As an embodiment of the present invention, a virtual dialogue system incorporating an emotion engine that recognizes emotions is introduced, with the cooperation of a server, a terminal, and a user. The server is primarily responsible for information processing, analyzing the user's statements and emotions and providing feedback. The terminal functions as a user interface, enabling interaction between the user and a virtual character.

[0480] First, the user logs into the system using a terminal and selects a virtual customer service scenario. The server generates a virtual character based on the selected scenario and prepares for the interaction. The user then begins interacting with the virtual character through voice or text input.

[0481] In this process, an emotion engine operates to recognize the user's emotional state during the conversation in real time. For example, if the user displays an unpleasant facial expression or tone of voice during a conversation, the emotion engine detects this and notifies the server. Based on the received emotional information, the server adjusts the virtual character's response to provide an interaction that is appropriate to the user's emotional state.

[0482] Once the training is complete, the server evaluates the user's performance by combining dialogue content and emotional data. The generated feedback takes into account the user's emotional responses and includes more personalized improvement suggestions. This feedback is provided to the user via their device, allowing them to review it and incorporate it into their next training session.

[0483] Furthermore, the server records training progress, including emotional data, and uses it for long-term analysis. This allows users to understand their skill development in detail, including changes in emotions. This invention enables training in advanced customer service techniques that incorporate emotion recognition, allowing users to provide emotionally sensitive customer service in actual work situations.

[0484] The following describes the processing flow.

[0485] Step 1:

[0486] The user enters their ID and password to log in to the system using their device. The device sends this information to the server and initiates the authentication process.

[0487] Step 2:

[0488] The server verifies the submitted login information against the database and performs the authentication process. If authentication is successful, it sends the data that makes up the user's dashboard to the device.

[0489] Step 3:

[0490] From a dashboard displayed on the device, the user selects a training scenario, such as one related to "stress management" or "nonverbal communication." The selected scenario is then notified to the server.

[0491] Step 4:

[0492] Based on the selected scenario information, the server generates an appropriate virtual character and simultaneously activates an emotion engine for emotion recognition.

[0493] Step 5:

[0494] The terminal displays a virtual character to the user and prepares to start a virtual dialogue scenario. The user begins interacting with the character using voice or text input.

[0495] Step 6:

[0496] The server analyzes the user's speech and visual data transmitted from the device in real time using an emotion engine to evaluate the user's emotional state.

[0497] Step 7:

[0498] Based on the analysis results of the emotion engine, the server adjusts the virtual character's response and generates feedback that is appropriate to the user's emotions.

[0499] Step 8:

[0500] Once the dialogue session ends, the server evaluates the user's training results based on all dialogue logs and sentiment data, and generates feedback that includes specific areas for improvement.

[0501] Step 9:

[0502] The feedback is sent to the device and presented to the user. Based on this, the user incorporates skill improvement measures that take emotional responses into account for the next training session.

[0503] Step 10:

[0504] The server stores training progress data, including collected emotional data, and records it as long-term analytical data to track user growth. This allows users to see how their emotional response skills have changed over time.

[0505] (Example 2)

[0506] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0507] Traditional customer service techniques suffer from problems such as a lack of training that takes customer emotions into account and a lack of individualized feedback to users. As a result, it has been difficult for users to develop sufficient emotional response skills in actual customer service work. Furthermore, in multilingual situations, the loss of emotional nuances due to language translation has become a challenge.

[0508] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0509] In this invention, the server includes means for generating and performing virtual characters, means for performing real-time emotion analysis of user speech or input data, and means for dynamically adjusting the virtual character's responses based on the emotion analysis results. This enables effective training that takes into account the user's emotional state and provides personalized feedback. Furthermore, it can improve the quality of communication by providing more natural, emotion-based translations even in multilingual environments.

[0510] "Means for generating and performing virtual characters" refers to a function in which a computer system uses its processing power to create a virtual entity for interaction with the user, and to control and display its responses and actions.

[0511] "Means for real-time sentiment analysis of user speech or input data" refers to technology that analyzes voice and text data provided by users and instantly identifies the emotions contained within it.

[0512] "Means for dynamically adjusting the virtual character's response based on emotion analysis results" refers to a function that changes the content of the virtual character's response and attitude in real time according to the detected emotions of the user.

[0513] "Means for generating personalized feedback" refers to a function that provides improvement suggestions and advice tailored to a specific user, based on the user's conversation history and sentiment data.

[0514] "Means for recording training progress data and analyzing long-term skill growth" refers to technologies that continuously save a user's learning history and conversation content, and measure skill improvement based on that data.

[0515] The "translation function that enables virtual dialogue in multiple languages ​​and further performs emotion-based translation" is a function that supports communication between different languages ​​and reflects emotional nuances when translating languages.

[0516] This invention provides a virtual dialogue system equipped with emotion analysis capabilities, through the collaboration of a server, a terminal, and a user. The system begins when the user logs in using a terminal and starts interacting with a virtual character.

[0517] First, the server receives login information from the user's device and performs user authentication. Once the user successfully logs in, the server receives the user's selection from several virtual customer service scenarios provided. Based on the selection, the server generates an appropriate virtual character using a generative AI model. The AI ​​model is based on natural language processing and machine learning techniques and has a mechanism to provide diverse dialogue patterns and responses that respond to emotions.

[0518] Next, the terminal functions as a user interface to enable interaction between the user and the virtual character. The user converses with the virtual character in real time through voice input or text input. The server uses an emotion analysis engine to analyze the emotional tone and keywords from the user's input data and identify the emotional state.

[0519] The server dynamically adjusts the virtual character's responses based on the sentiment analysis results. For example, if the user asks, "Could you tell me a little more?", the character will provide detailed information such as, "As a specific example of a recommendation, we have our special seafood pasta." If the user seems dissatisfied, the character's response tone can be adjusted, and suggestions for improvement can be offered.

[0520] Through this interaction, the server collects data to evaluate the user's performance and generates personalized feedback. This feedback is provided via the terminal so that the user can use it for future training sessions. The feedback specifically highlights areas where the user's skills can be improved and where their emotional response techniques can be enhanced.

[0521] Furthermore, this system supports multiple languages ​​and includes emotion-based translation capabilities, enabling natural dialogue even in different language environments. For example, by giving the AI ​​model instructions such as, "Explain how the virtual character should respond optimally based on user input," it becomes possible to explore more detailed ways to adjust the dialogue.

[0522] Through this format, users can effectively learn advanced customer service skills that take emotions into consideration.

[0523] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0524] Step 1:

[0525] The user logs into the system using their device. The user enters their username and password, and the device sends this authentication information to the server. The server checks the user database to verify that the entered information is correct, and if successful, displays the home screen. As output, the user receives a screen showing the successful login.

[0526] Step 2:

[0527] The user selects their desired scenario from a list of virtual customer service scenarios displayed on their device. The device sends the selected scenario information to the server. Based on the received scenario information, the server uses a generation AI model to create the target virtual character. As output, the server sends the character and scenario settings to the device.

[0528] Step 3:

[0529] The user initiates a conversation with a virtual character by inputting voice or text through a terminal. The input data is sent from the terminal to the server. The server uses an emotion analysis engine to analyze the emotions contained in the user's input. As a result of the analysis, data is generated that determines the user's emotional state in real time. The analyzed emotion data is obtained as output.

[0530] Step 4:

[0531] The server dynamically adjusts the virtual character's responses based on the analyzed emotion data. Depending on the detected emotion, it changes the tone of the information and questions the virtual character provides. For example, if the user shows interest, the character will provide a more detailed explanation. The adapted response data is sent to the terminal as output and displayed to the user.

[0532] Step 5:

[0533] After a dialogue session ends, the server evaluates the user's performance based on their dialogue history and emotional data. Using a generative AI model, it performs analysis based on evaluation criteria and generates personalized feedback. This feedback includes suggestions for improving the user's emotional responses and skills. The feedback data is then sent to the terminal and presented to the user.

[0534] Step 6:

[0535] The server records user training progress data and analyzes long-term skill growth. The collected data is stored for later evaluation and improvement. This allows for a detailed understanding of how users improve their skills over time. As output, the recorded progress data is stored on the server.

[0536] (Application Example 2)

[0537] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0538] Conventional customer service training systems have the problem of being unable to accurately understand user emotions and learn customer service methods based on them. The present invention aims to improve customer service skills by analyzing user emotions in real time and supporting appropriate customer service methods.

[0539] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0540] In this invention, the server includes means for recognizing emotions, means for adjusting responses based on emotion analysis results, and means for analyzing emotional states in real time. This enables the provision of immediate feedback based on the user's emotions and improves personalized customer service techniques.

[0541] "Means of recognizing emotions" refers to technologies that analyze a user's facial expressions, tone of voice, etc., to identify their emotional state.

[0542] "Means for adjusting responses based on emotion analysis results" refers to techniques for appropriately modifying the responses of virtual dialogue characters or systems using analyzed emotion data.

[0543] "Methods for analyzing emotional states in real time" refer to technologies that instantly evaluate emotions during user interactions and provide immediate analysis results.

[0544] "Means for generating feedback to support customer service" refers to technologies that generate guidelines and suggestions for effective customer service based on the user's emotional state and the content of their conversations.

[0545] "Means for displaying generated feedback" refers to technologies for presenting feedback obtained from a system to the user visually or audibly.

[0546] "Means for recording and managing user progress data" refers to technologies for continuously saving and managing the progress of users' customer service training and the history of their emotional changes.

[0547] To implement this invention, a server, terminal, and user must cooperate and use the following configuration: The server is a system equipped with an emotion engine that recognizes emotions and an algorithm for generating feedback. The terminal is a device that functions as the user's interface, collecting dialogue content in real time and sending the data to the server for analysis. This system operates using software libraries specialized for image processing and speech analysis.

[0548] Specifically, a smartphone is used as the terminal, and customer facial expression data is captured using the OpenCV library, while emotions are analyzed from audio data using TensorFlow. This data is sent to a server, which uses a generative AI model to generate feedback based on the analysis results. This feedback is then displayed on the terminal in real time, providing immediate advice to the user.

[0549] Furthermore, users can improve the quality of their customer service through their interactions with the system. For example, they can quickly detect changes in customer emotions during interactions, enabling them to respond flexibly and appropriately.

[0550] For example, if a user working in a bookstore shows interest in a customer's explanation of a particular product, feedback will appear stating, "You should continue with the detailed explanation." In this way, it is possible to efficiently improve customer service skills.

[0551] Furthermore, as an example of a prompt using a generative AI model, you can use a sentence like this: "Analyze the customer's facial expressions and voice, and suggest appropriate customer service methods. For example, I would like advice on the difference in how to respond when the customer is smiling and showing interest versus when they are showing confusion."

[0552] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0553] Step 1:

[0554] The device uses a camera and microphone to capture the facial expressions and voices of users and customers in real time. The input consists of visual and audio data. This data is passed to OpenCV, an image processing tool, and TensorFlow, an audio processing library, to perform facial feature extraction and audio data preprocessing.

[0555] Step 2:

[0556] The server receives pre-processed facial expression and audio data and begins emotion analysis. The input consists of feature-extracted visual and audio data. The emotion engine analyzes this data using a generated AI model to recognize the user's emotional state in real time. The output is data containing the emotion analysis results.

[0557] Step 3:

[0558] The server adjusts the virtual customer service scenario based on the analysis results. The input is the emotion analysis result, and the system determines the optimal character response pattern based on it. The server dynamically adjusts the dialogue scenario and generates data for feedback. The output is the adjusted response pattern.

[0559] Step 4:

[0560] The terminal presents information to the user visually or audibly based on feedback data received from the server. The input is feedback data. The presented information includes suggestions for specific customer service actions and approaches tailored to the customer's emotions. The user can use this information to adapt their behavior. The output is displaying the feedback on the user's screen.

[0561] Step 5:

[0562] The user adjusts their customer service behavior based on feedback and interacts with the customer. The user's responses become part of the data capture from the next processing step 1, helping to continuously improve the system. The input is feedback from the system, and the output is the user's adapted customer service behavior.

[0563] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0564] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0565] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0566] [Fourth Embodiment]

[0567] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0568] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0569] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0570] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0571] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0572] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0573] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0574] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0575] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0576] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0577] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0578] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0579] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0580] To implement this invention, three parties—a server, a terminal, and a user—must work closely together. The server, as the central information processing device, is responsible for generating virtual characters, analyzing dialogue, and generating feedback. The terminal provides a user interface and functions as a device that enables interaction with the user.

[0581] First, the user logs into the system via a terminal and selects a virtual customer service scenario that suits their purpose. For example, they can choose scenarios such as "handling cross-cultural customer interactions" or "resolving complex customer complaints." After selection, the server generates a virtual character based on the specified scenario. This character simulates interaction with the user, providing a simulated training environment for honing customer service skills.

[0582] Once a conversation begins, the server analyzes the user's statements in real time. This analysis uses natural language processing to evaluate factors such as tone, appropriateness of responses, and the ability to understand customer needs. The analyzed information is immediately recorded and evaluated within the server.

[0583] Once the training is complete, the server generates feedback based on the user's performance. This feedback is provided to the user via their device, specifically highlighting areas for improvement and areas where they performed well. The user can use this feedback to further enhance their skills. The feedback is also stored as data in the user's training profile when it is generated.

[0584] This system allows users to improve their customer service skills through practical and diverse scenarios, often without requiring significant human intervention. Furthermore, its multilingual capabilities enhance its ability to serve international customers. This enables users to efficiently learn the skills required in real-world work environments, contributing to improved work performance.

[0585] The following describes the processing flow.

[0586] Step 1:

[0587] The user attempts to log in to the system using their device. Logging in requires personal authentication information, so the user enters their ID and password. The device then sends the entered information to the server.

[0588] Step 2:

[0589] The server compares the received login information with the registered information in the database and performs authentication. If authentication is successful, it prepares to display the user's dashboard on the terminal.

[0590] Step 3:

[0591] The user starts training from a dashboard on their device, selecting one of several scenarios provided. The selection is sent to the server, which retrieves information about the scenario.

[0592] Step 4:

[0593] The server generates a virtual character based on the scenario selected by the user. The generated character has its dialogue and situation pre-configured, and is ready to begin the scenario.

[0594] Step 5:

[0595] The terminal displays a virtual character to the user and initiates a virtual dialogue scenario. The user responds to the character using voice input or text input.

[0596] Step 6:

[0597] The server receives user input in real time and performs analysis using natural language processing. The analysis results take into account the appropriateness of the content of the statements, the wording, and the flow of the conversation.

[0598] Step 7:

[0599] Once the training is complete, the server evaluates the user's performance based on the analysis results. From the evaluation results, it extracts specific strengths and areas for improvement.

[0600] Step 8:

[0601] The server creates feedback based on the generated evaluation and sends it to the terminal. The terminal displays the feedback to the user, who then reviews it.

[0602] Step 9:

[0603] The server stores and records data about the user's training sessions in a format that can be referenced later. Based on this data, users can develop long-term skill improvement plans.

[0604] (Example 1)

[0605] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0606] In today's business environment, customer service skills are crucial for improving customer satisfaction, but traditional training methods have struggled to efficiently provide practical and diverse scenario-based training. Furthermore, efficiently strengthening the ability to handle multilingual international customers has also been a challenging task. In response, there was a need for a feedback system tailored to individual skills, a means to visualize user capabilities, and a way to promote improvement.

[0607] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0608] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating user dialogue in real time, and means for generating feedback based on user performance. This enables users to improve their customer service skills through practical and diverse scenarios. Furthermore, analysis using natural language processing can efficiently enhance multilingual communication skills. In addition, by combining a generative AI model with prompt sentences, it is possible to appropriately guide user responses and achieve higher quality dialogue training.

[0609] A "virtual character" is a digital humanoid entity created to simulate interaction with a user, and can have different personalities and backgrounds based on various scenarios.

[0610] A "user interface" is a general term for the screens and tools that users use to access and operate a system, and is used for inputting dialogues, selecting scenarios, and so on.

[0611] "Natural language processing" is a technology used to enable computers to understand and analyze human language. It is used to analyze user statements and provide appropriate responses and evaluations.

[0612] "Feedback" is evaluation information provided based on the results of user interaction training, indicating areas for improvement and strengths, and is useful for improving the user's skills.

[0613] A "training profile" is a collection of information that records a user's training progress and past evaluation results, and is referenced for future skill improvement.

[0614] A "generative AI model" is a type of artificial intelligence that generates appropriate information or characters in response to user requests, based on knowledge learned from large amounts of data.

[0615] A "prompt message" is a phrase provided to give instructions or draw attention to a user, and its role is to guide them to a specific response or action.

[0616] To implement this invention, a system is required in which a server, a terminal, and a user work closely together. The server functions as a central information processing device responsible for generating virtual characters, analyzing user statements, and generating feedback. Specifically, the server utilizes natural language processing engines such as Google Cloud Natural Language API and IBM Watson for natural language processing, and platforms such as Unity and Unreal Engine for generating virtual characters.

[0617] The terminal acts as a user interface, creating an environment where users can participate in dialogue training. When a user logs into the system using the terminal, they are presented with selectable scenarios through a web application. By selecting from scenarios such as "Cross-Cultural Customer Service" or "Handling Complex Complaints," users can begin learning based on a specific scenario.

[0618] Based on the scenario selected by the user, the server uses a generative AI model to create a virtual character suitable for that scenario. The generative AI model enables diverse and realistic simulations, allowing the user to gain an experience close to actual work. For example, if the scenario is "handling cross-cultural customer service," the prompt message provided is "Your task is to handle a complaint from an international customer. The customer speaks only Spanish, and you need to resolve their issue using the virtual assistant," preparing an environment for the user to practice handling cross-cultural customer service.

[0619] As the conversation progresses, the server analyzes the user's statements in real time, evaluating their appropriateness and understanding of customer needs. Based on this analysis, feedback is generated at the end of the training session, reflecting the user's performance. This feedback is provided to the user via their device, allowing them to use the feedback to make improvements. The feedback is then stored in the training profile and used for future training sessions. This enables users to efficiently improve their customer service skills and strengthen their ability to serve international customers.

[0620] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0621] Step 1:

[0622] The user logs into the system using a terminal. The user accesses the user interface via a web application and authenticates by entering a specific ID and password. This information is sent to the server, and if authenticated, a success message is displayed, and the process proceeds to the next step.

[0623] Step 2:

[0624] The user selects a training scenario through an interface provided on the terminal. The user chooses one scenario from a list such as "Cross-Cultural Customer Service" or "Handling Difficult Complaints," and the selected scenario information is sent to the server. Based on this input information, the server outputs a virtual character profile that matches the scenario.

[0625] Step 3:

[0626] The server generates a virtual character using a generative AI model based on the selected scenario. The server uses the input scenario information to set the character's personality and dialogue patterns using the generative AI model, thereby generating a virtual customer. The generated character information is output to the terminal, allowing the user to begin interacting with the character.

[0627] Step 4:

[0628] The user initiates a conversation with a virtual character using a text field or voice input function on their device. The user's input data is sent to a server, which analyzes the utterance using natural language processing. The analysis results evaluate the intent and emotion of the utterance, and this information is stored. This analysis data forms the basis for providing real-time feedback to the user.

[0629] Step 5:

[0630] After the dialogue session ends, the server uses the accumulated analytical data to evaluate the user's performance and generate feedback. The server analyzes the appropriateness and areas for improvement of the user's responses and outputs feedback that includes specific improvement suggestions. This feedback data is sent to the terminal and becomes available for the user to view.

[0631] Step 6:

[0632] Feedback is provided to the user through the device. The user refers to the feedback provided and creates a self-improvement plan to enhance their skills. This feedback is recorded in the training profile by the server and used for future learning and evaluation. This allows the user's progress to be tracked and used to improve future training sessions.

[0633] (Application Example 1)

[0634] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0635] In recent years, there has been a growing demand for improved customer service skills in brick-and-mortar stores, but traditional methods make effective training difficult. In particular, developing the ability to respond appropriately and instantly to diverse customer needs in today's business environment is a challenge. Furthermore, with increasing globalization, multilingual support has become essential.

[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0637] In this invention, the server includes means for generating and staging virtual characters, means for analyzing and evaluating human dialogue in real time, and means for generating feedback based on a person's work performance capabilities. This enables the rapid and effective improvement of customer service skills of staff in physical stores using a personal information display device.

[0638] A "virtual character" is a digital character created by a computer that simulates dialogue and actions.

[0639] "Real-time analysis" refers to a processing technology that instantly evaluates user input and provides immediate responses.

[0640] "Job performance ability" refers to the knowledge and skills necessary to effectively carry out tasks.

[0641] "Feedback" refers to information or advice that provides areas for improvement or evaluation based on user behavior and responses.

[0642] A "personal information display device" is an electronic device that can be worn or carried by an individual and used by that individual, and which has the function of visually displaying information.

[0643] To implement this invention, three parties—a server, a terminal, and a user—must work in cooperation. The server functions as a central information processing device for generating digital characters and simulating interactions with the user. The terminal utilizes personal information display devices or smart devices to provide an intuitive interface with the user.

[0644] The server analyzes user utterances using natural language processing technology and provides real-time evaluation. Specifically, this analysis process utilizes natural language processing libraries such as spaCy and BERT. Based on the evaluation results, feedback is generated to improve the user's work performance. This feedback is visually represented and provided to the user via a terminal. The provided feedback indicates specific areas for improvement that will help the user efficiently enhance their customer service skills.

[0645] As a concrete example, store staff could wear smart glasses and train their conversational skills by conversing with virtual characters representing customers from different cultures. The server would analyze the appropriateness of the words and responses used by the user in this conversation and visually present areas for improvement.

[0646] An example of a prompt message to input into a generative AI model is as follows:

[0647] "This session is for training customer service skills in a retail environment. The user will interact with a virtual customer character who speaks English. The system stitches the user's responses and provides real-time feedback on how to improve their service quality."

[0648] Through this system, users will be able to effectively improve their practical customer service skills.

[0649] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0650] Step 1:

[0651] The user logs in via a terminal. The user selects a virtual customer service scenario that suits their purpose. Here, the user enters their ID information using an input terminal. As output, the selected scenario information is transmitted to the server.

[0652] Step 2:

[0653] The server generates a virtual character based on the selected scenario. In this process, the server uses a generation AI model and prompts to create a character suitable for the scenario. The output is the character's digital profile.

[0654] Step 3:

[0655] The server analyzes user speech in real time. The input is audio data sent from the user via their terminal. The server converts this data into text using natural language processing libraries (such as spaCy or BERT) and analyzes its content. The analysis results are output as a response quality evaluation score.

[0656] Step 4:

[0657] The server generates feedback based on the user's ability to perform their tasks. The input is the response quality evaluation score from the analysis results. Based on this, the server generates feedback that includes areas for improvement and strengths. The output is the text data of the feedback.

[0658] Step 5:

[0659] The terminal visualizes the generated feedback and provides it to the user. The input is text data of the feedback sent from the server, which the terminal displays on its screen. Based on the feedback, the user can strive to improve their customer service skills. The output is specific improvement guidelines provided to the user.

[0660] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0661] As an embodiment of the present invention, a virtual dialogue system incorporating an emotion engine that recognizes emotions is introduced, with the cooperation of a server, a terminal, and a user. The server is primarily responsible for information processing, analyzing the user's statements and emotions and providing feedback. The terminal functions as a user interface, enabling interaction between the user and a virtual character.

[0662] First, the user logs into the system using a terminal and selects a virtual customer service scenario. The server generates a virtual character based on the selected scenario and prepares for the interaction. The user then begins interacting with the virtual character through voice or text input.

[0663] In this process, an emotion engine operates to recognize the user's emotional state during the conversation in real time. For example, if the user displays an unpleasant facial expression or tone of voice during a conversation, the emotion engine detects this and notifies the server. Based on the received emotional information, the server adjusts the virtual character's response to provide an interaction that is appropriate to the user's emotional state.

[0664] Once the training is complete, the server evaluates the user's performance by combining dialogue content and emotional data. The generated feedback takes into account the user's emotional responses and includes more personalized improvement suggestions. This feedback is provided to the user via their device, allowing them to review it and incorporate it into their next training session.

[0665] Furthermore, the server records training progress, including emotional data, and uses it for long-term analysis. This allows users to understand their skill development in detail, including changes in emotions. This invention enables training in advanced customer service techniques that incorporate emotion recognition, allowing users to provide emotionally sensitive customer service in actual work situations.

[0666] The following describes the processing flow.

[0667] Step 1:

[0668] The user enters their ID and password to log in to the system using their device. The device sends this information to the server and initiates the authentication process.

[0669] Step 2:

[0670] The server verifies the submitted login information against the database and performs the authentication process. If authentication is successful, it sends the data that makes up the user's dashboard to the device.

[0671] Step 3:

[0672] From a dashboard displayed on the device, the user selects a training scenario, such as one related to "stress management" or "nonverbal communication." The selected scenario is then notified to the server.

[0673] Step 4:

[0674] Based on the selected scenario information, the server generates an appropriate virtual character and simultaneously activates an emotion engine for emotion recognition.

[0675] Step 5:

[0676] The terminal displays a virtual character to the user and prepares to start a virtual dialogue scenario. The user begins interacting with the character using voice or text input.

[0677] Step 6:

[0678] The server analyzes the user's speech and visual data transmitted from the device in real time using an emotion engine to evaluate the user's emotional state.

[0679] Step 7:

[0680] Based on the analysis results of the emotion engine, the server adjusts the virtual character's response and generates feedback that is appropriate to the user's emotions.

[0681] Step 8:

[0682] Once the dialogue session ends, the server evaluates the user's training results based on all dialogue logs and sentiment data, and generates feedback that includes specific areas for improvement.

[0683] Step 9:

[0684] The feedback is sent to the device and presented to the user. Based on this, the user incorporates skill improvement measures that take emotional responses into account for the next training session.

[0685] Step 10:

[0686] The server stores training progress data, including collected emotional data, and records it as long-term analytical data to track user growth. This allows users to see how their emotional response skills have changed over time.

[0687] (Example 2)

[0688] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0689] Traditional customer service techniques suffer from problems such as a lack of training that takes customer emotions into account and a lack of individualized feedback to users. As a result, it has been difficult for users to develop sufficient emotional response skills in actual customer service work. Furthermore, in multilingual situations, the loss of emotional nuances due to language translation has become a challenge.

[0690] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0691] In this invention, the server includes means for generating and performing virtual characters, means for performing real-time emotion analysis of user speech or input data, and means for dynamically adjusting the virtual character's responses based on the emotion analysis results. This enables effective training that takes into account the user's emotional state and provides personalized feedback. Furthermore, it can improve the quality of communication by providing more natural, emotion-based translations even in multilingual environments.

[0692] "Means for generating and performing virtual characters" refers to a function in which a computer system uses its processing power to create a virtual entity for interaction with the user, and to control and display its responses and actions.

[0693] "Means for real-time sentiment analysis of user speech or input data" refers to technology that analyzes voice and text data provided by users and instantly identifies the emotions contained within it.

[0694] "Means for dynamically adjusting the virtual character's response based on emotion analysis results" refers to a function that changes the content of the virtual character's response and attitude in real time according to the detected emotions of the user.

[0695] "Means for generating personalized feedback" refers to a function that provides improvement suggestions and advice tailored to a specific user, based on the user's conversation history and sentiment data.

[0696] "Means for recording training progress data and analyzing long-term skill growth" refers to technologies that continuously save a user's learning history and conversation content, and measure skill improvement based on that data.

[0697] The "translation function that enables virtual dialogue in multiple languages ​​and further performs emotion-based translation" is a function that supports communication between different languages ​​and reflects emotional nuances when translating languages.

[0698] This invention provides a virtual dialogue system equipped with emotion analysis capabilities, through the collaboration of a server, a terminal, and a user. The system begins when the user logs in using a terminal and starts interacting with a virtual character.

[0699] First, the server receives login information from the user's device and performs user authentication. Once the user successfully logs in, the server receives the user's selection from several virtual customer service scenarios provided. Based on the selection, the server generates an appropriate virtual character using a generative AI model. The AI ​​model is based on natural language processing and machine learning techniques and has a mechanism to provide diverse dialogue patterns and responses that respond to emotions.

[0700] Next, the terminal functions as a user interface to enable interaction between the user and the virtual character. The user converses with the virtual character in real time through voice input or text input. The server uses an emotion analysis engine to analyze the emotional tone and keywords from the user's input data and identify the emotional state.

[0701] The server dynamically adjusts the virtual character's responses based on the sentiment analysis results. For example, if the user asks, "Could you tell me a little more?", the character will provide detailed information such as, "As a specific example of a recommendation, we have our special seafood pasta." If the user seems dissatisfied, the character's response tone can be adjusted, and suggestions for improvement can be offered.

[0702] Through this interaction, the server collects data to evaluate the user's performance and generates personalized feedback. This feedback is provided via the terminal so that the user can use it for future training sessions. The feedback specifically highlights areas where the user's skills can be improved and where their emotional response techniques can be enhanced.

[0703] Furthermore, this system supports multiple languages ​​and includes emotion-based translation capabilities, enabling natural dialogue even in different language environments. For example, by giving the AI ​​model instructions such as, "Explain how the virtual character should respond optimally based on user input," it becomes possible to explore more detailed ways to adjust the dialogue.

[0704] Through this format, users can effectively learn advanced customer service skills that take emotions into consideration.

[0705] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0706] Step 1:

[0707] The user logs into the system using their device. The user enters their username and password, and the device sends this authentication information to the server. The server checks the user database to verify that the entered information is correct, and if successful, displays the home screen. As output, the user receives a screen showing the successful login.

[0708] Step 2:

[0709] The user selects their desired scenario from a list of virtual customer service scenarios displayed on their device. The device sends the selected scenario information to the server. Based on the received scenario information, the server uses a generation AI model to create the target virtual character. As output, the server sends the character and scenario settings to the device.

[0710] Step 3:

[0711] The user initiates a conversation with a virtual character by inputting voice or text through a terminal. The input data is sent from the terminal to the server. The server uses an emotion analysis engine to analyze the emotions contained in the user's input. As a result of the analysis, data is generated that determines the user's emotional state in real time. The analyzed emotion data is obtained as output.

[0712] Step 4:

[0713] The server dynamically adjusts the virtual character's responses based on the analyzed emotion data. Depending on the detected emotion, it changes the tone of the information and questions the virtual character provides. For example, if the user shows interest, the character will provide a more detailed explanation. The adapted response data is sent to the terminal as output and displayed to the user.

[0714] Step 5:

[0715] After a dialogue session ends, the server evaluates the user's performance based on their dialogue history and emotional data. Using a generative AI model, it performs analysis based on evaluation criteria and generates personalized feedback. This feedback includes suggestions for improving the user's emotional responses and skills. The feedback data is then sent to the terminal and presented to the user.

[0716] Step 6:

[0717] The server records user training progress data and analyzes long-term skill growth. The collected data is stored for later evaluation and improvement. This allows for a detailed understanding of how users improve their skills over time. As output, the recorded progress data is stored on the server.

[0718] (Application Example 2)

[0719] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0720] Conventional customer service training systems have the problem of being unable to accurately understand user emotions and learn customer service methods based on them. The present invention aims to improve customer service skills by analyzing user emotions in real time and supporting appropriate customer service methods.

[0721] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0722] In this invention, the server includes means for recognizing emotions, means for adjusting responses based on emotion analysis results, and means for analyzing emotional states in real time. This enables the provision of immediate feedback based on the user's emotions and improves personalized customer service techniques.

[0723] "Means of recognizing emotions" refers to technologies that analyze a user's facial expressions, tone of voice, etc., to identify their emotional state.

[0724] "Means for adjusting responses based on emotion analysis results" refers to techniques for appropriately modifying the responses of virtual dialogue characters or systems using analyzed emotion data.

[0725] "Methods for analyzing emotional states in real time" refer to technologies that instantly evaluate emotions during user interactions and provide immediate analysis results.

[0726] "Means for generating feedback to support customer service" refers to technologies that generate guidelines and suggestions for effective customer service based on the user's emotional state and the content of their conversations.

[0727] "Means for displaying generated feedback" refers to technologies for presenting feedback obtained from a system to the user visually or audibly.

[0728] "Means for recording and managing user progress data" refers to technologies for continuously saving and managing the progress of users' customer service training and the history of their emotional changes.

[0729] To implement this invention, a server, terminal, and user must cooperate and use the following configuration: The server is a system equipped with an emotion engine that recognizes emotions and an algorithm for generating feedback. The terminal is a device that functions as the user's interface, collecting dialogue content in real time and sending the data to the server for analysis. This system operates using software libraries specialized for image processing and speech analysis.

[0730] Specifically, a smartphone is used as the terminal, and customer facial expression data is captured using the OpenCV library, while emotions are analyzed from audio data using TensorFlow. This data is sent to a server, which uses a generative AI model to generate feedback based on the analysis results. This feedback is then displayed on the terminal in real time, providing immediate advice to the user.

[0731] Furthermore, users can improve the quality of their customer service through their interactions with the system. For example, they can quickly detect changes in customer emotions during interactions, enabling them to respond flexibly and appropriately.

[0732] For example, if a user working in a bookstore shows interest in a customer's explanation of a particular product, feedback will appear stating, "You should continue with the detailed explanation." In this way, it is possible to efficiently improve customer service skills.

[0733] Furthermore, as an example of a prompt using a generative AI model, you can use a sentence like this: "Analyze the customer's facial expressions and voice, and suggest appropriate customer service methods. For example, I would like advice on the difference in how to respond when the customer is smiling and showing interest versus when they are showing confusion."

[0734] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0735] Step 1:

[0736] The device uses a camera and microphone to capture the facial expressions and voices of users and customers in real time. The input consists of visual and audio data. This data is passed to OpenCV, an image processing tool, and TensorFlow, an audio processing library, to perform facial feature extraction and audio data preprocessing.

[0737] Step 2:

[0738] The server receives pre-processed facial expression and audio data and begins emotion analysis. The input consists of feature-extracted visual and audio data. The emotion engine analyzes this data using a generated AI model to recognize the user's emotional state in real time. The output is data containing the emotion analysis results.

[0739] Step 3:

[0740] The server adjusts the virtual customer service scenario based on the analysis results. The input is the emotion analysis result, and the system determines the optimal character response pattern based on it. The server dynamically adjusts the dialogue scenario and generates data for feedback. The output is the adjusted response pattern.

[0741] Step 4:

[0742] The terminal presents information to the user visually or audibly based on feedback data received from the server. The input is feedback data. The presented information includes suggestions for specific customer service actions and approaches tailored to the customer's emotions. The user can use this information to adapt their behavior. The output is displaying the feedback on the user's screen.

[0743] Step 5:

[0744] The user adjusts their customer service behavior based on feedback and interacts with the customer. The user's responses become part of the data capture from the next processing step 1, helping to continuously improve the system. The input is feedback from the system, and the output is the user's adapted customer service behavior.

[0745] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0746] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0747] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0748] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0749] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0750] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0751] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0752] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0753] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0754] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0755] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0756] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0757] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0758] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0759] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0760] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0761] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0762] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0763] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0764] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0765] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0766] The following is further disclosed regarding the embodiments described above.

[0767] (Claim 1)

[0768] An information processing device that provides a virtual dialogue method for improving customer service skills,

[0769] Means for generating and performing virtual characters,

[0770] A means of analyzing and evaluating user conversations in real time,

[0771] A means of generating feedback based on user performance,

[0772] A means of providing the generated feedback to the user,

[0773] A system that includes means for recording and managing user training progress data.

[0774] (Claim 2)

[0775] It has a configuration for conducting virtual interactions based on multiple scenarios selected by the user.

[0776] The system according to claim 1.

[0777] (Claim 3)

[0778] It has a translation function to enable virtual dialogue in multiple languages.

[0779] The system according to claim 1.

[0780] "Example 1"

[0781] (Claim 1)

[0782] Means for generating and performing virtual characters,

[0783] A means of analyzing and evaluating user conversations in real time,

[0784] A means of generating feedback based on user performance,

[0785] A means of providing the generated feedback to the user,

[0786] A means of recording and managing user training progress data,

[0787] A method for analyzing user utterances using natural language processing and evaluating their appropriateness and understanding of customer needs,

[0788] A means to evaluate the quality of user responses based on the analysis results and to accumulate that evaluation in a training profile,

[0789] A means to enable users to select from a variety of virtual customer service scenarios through the user interface,

[0790] A method for creating appropriate characters from scenario data using a generative AI model,

[0791] A system that includes this.

[0792] (Claim 2)

[0793] The system according to claim 1, having a configuration for conducting virtual interactions based on multiple scenarios selected by the user.

[0794] (Claim 3)

[0795] The system according to claim 1, comprising a translation function for enabling multilingual virtual dialogue and means for guiding the user's response using prompt sentences.

[0796] "Application Example 1"

[0797] (Claim 1)

[0798] A means of creating and portraying a virtual character,

[0799] A means of analyzing and evaluating the content of human conversations in real time,

[0800] A means of generating feedback based on a person's ability to perform their job,

[0801] A means of providing the generated feedback to people,

[0802] A means of recording and managing human training progress data,

[0803] A system characterized in that the above means is incorporated into a personal information display device and has a configuration that enables dialogue training.

[0804] (Claim 2)

[0805] The system according to claim 1, having a configuration for conducting virtual interactions based on multiple scenarios selected by the user.

[0806] (Claim 3)

[0807] The system according to claim 1, having a translation function to enable virtual dialogue in multiple languages.

[0808] "Example 2 of combining an emotion engine"

[0809] (Claim 1)

[0810] Means for generating and performing virtual characters,

[0811] A means for performing real-time emotion analysis on user speech or input data,

[0812] A means for dynamically adjusting the response of a virtual character based on the results of emotion analysis,

[0813] A means of generating personalized feedback based on user performance,

[0814] A means of providing the generated feedback to the user,

[0815] A system that includes means for recording user training progress data and analyzing long-term skill growth.

[0816] (Claim 2)

[0817] The system conducts virtual interactions based on multiple scenarios selected by the user.

[0818] The system according to claim 1, having a configuration that provides an emotion-responsive response.

[0819] (Claim 3)

[0820] The system according to claim 1, which enables virtual dialogue in multiple languages ​​and further has a translation function that performs emotion-based translation.

[0821] "Application example 2 when combining with an emotional engine"

[0822] (Claim 1)

[0823] Means of recognizing emotions,

[0824] A means of adjusting responses based on emotion analysis results,

[0825] A means of analyzing emotional states in real time,

[0826] A means of generating feedback to support customer service,

[0827] A means of displaying the generated feedback,

[0828] A system that includes means for recording and managing user progress data.

[0829] (Claim 2)

[0830] The system according to claim 1, having a configuration for conducting a virtual dialogue that is emotionally appropriate based on a scenario selected by the user.

[0831] (Claim 3)

[0832] The system according to claim 1, having a translation function to enable multilingual virtual dialogue and sentiment analysis. [Explanation of Symbols]

[0833] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An information processing device that provides a virtual dialogue method for improving customer service skills, Means for generating and performing virtual characters, A means of analyzing and evaluating user conversations in real time, A means of generating feedback based on user performance, A means of providing the generated feedback to the user, A system that includes means for recording and managing user training progress data.

2. It has a configuration for conducting virtual interactions based on multiple scenarios selected by the user. The system according to claim 1.

3. It has a translation function to enable virtual dialogue in multiple languages. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A