system
The system addresses limitations in conventional training by integrating advanced speech recognition and natural language processing to provide realistic and objective customer service training with real-time feedback, enhancing user performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
Smart Images

Figure 2026063875000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional customer service training system, there are problems that the scenarios are limited and the actual customer service scenarios cannot be fully reproduced. Also, since the evaluation of training depends on subjectivity, it has been a problem that it is difficult to obtain objective feedback. Furthermore, due to the low responsiveness in real time, it has been difficult for users to obtain a sense of presence as if they are actually providing customer service.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides a system that includes means for receiving and saving customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback on the evaluation results to the user. According to the present invention, since voice responses are generated in real time based on the user's profile information and scenario information, it is possible to conduct training that is close to actual customer service, and furthermore, objective evaluation and feedback can be provided.
[0006] "Customer settings" refers to the customer profile information and scenario information that users register in the system.
[0007] "Voice input" refers to the voice data that a user provides to the system through their device.
[0008] "Means of converting to text" refers to a function that analyzes voice input and generates corresponding text data.
[0009] "Means for generating natural responses" refers to a function that creates appropriate and natural response sentences based on converted text data.
[0010] "Means of converting to speech" refers to the function of converting the generated response sentence into audio data for auditory communication.
[0011] "Voice response" refers to the content of a response provided as voice data converted from text data.
[0012] "User" refers to a person who uses the system to conduct customer service training.
[0013] "Means for analyzing conversation content" refers to a function that analyzes conversation data exchanged between the user and customer settings.
[0014] The "means for evaluating the user's performance" refers to a function for evaluating the user's response ability and skills based on the analyzed conversation content.
[0015] The "evaluation result" refers to the score and comment of the evaluation regarding the user's performance.
[0016] The "means for providing feedback" refers to a function for providing the evaluation result to the user and conveying the improvement points and achievements.
Brief Explanation of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Modes for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user. Specific embodiments are described below.
[0039] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[0040] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as Google® Speech-to-Text API and IBM Watson® Speech to Text can be used for this speech recognition.
[0041] The converted text is then passed to the server's natural language processing (NLP) engine. The NLP engine takes into account the user's settings and the context of the conversation to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," the server will generate a response such as, "We apologize for the wait, customer."
[0042] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0043] Once the role-playing session ends, the server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, providing information on what went well and where there is room for improvement.
[0044] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal.
[0045] This allows users to conduct realistic customer service training and identify specific areas for improvement. This is a specific embodiment of the system according to the present invention.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The user accesses the customer settings registration screen from their device.
[0049] Step 2:
[0050] The user enters their profile information and scenario information into the device.
[0051] Step 3:
[0052] The terminal sends the entered customer settings information to the server.
[0053] Step 4:
[0054] The server saves the received information to the database.
[0055] Step 5:
[0056] The user presses the "Start Conversation" button on their device.
[0057] Step 6:
[0058] The device records the user's voice in real time and sends that audio data to the server.
[0059] Step 7:
[0060] The server's speech recognition module converts the speech data into text.
[0061] Step 8:
[0062] The server passes the converted text data to the natural language processing (NLP) engine.
[0063] Step 9:
[0064] The server's NLP engine generates appropriate responses based on customer settings and the context of the conversation.
[0065] Step 10:
[0066] The server sends the generated response text to the speech synthesis module.
[0067] Step 11:
[0068] The server's speech synthesis module converts text into speech data.
[0069] Step 12:
[0070] The server sends the audio data to the terminal.
[0071] Step 13:
[0072] The device plays the received audio data to the user.
[0073] Step 14:
[0074] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0075] Step 15:
[0076] The server analyzes all recorded conversation data.
[0077] Step 16:
[0078] The server generates a score and detailed feedback based on the analysis data to evaluate the user's responsiveness and skills.
[0079] Step 17:
[0080] The server sends the evaluation results and feedback to the terminal.
[0081] Step 18:
[0082] The device displays the evaluation results to the user and provides feedback.
[0083] (Example 1)
[0084] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0085] Existing speech recognition systems have the challenge of not being able to quickly provide appropriate responses based on specific scenarios and profile information set by the user, and not being able to evaluate user performance from multiple perspectives in real time. This invention solves these problems and provides a system that enables users to conduct more realistic and effective training.
[0086] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0087] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to speech, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, a terminal for inputting user setting information, a terminal for recording speech in real time and transmitting it to the server, a response generation means using a natural language processing engine, a response speech conversion means using a speech synthesis module, and means for evaluating the user's performance using multiple indicators and generating detailed feedback. This enables highly accurate response generation based on the user's settings and the provision of multifaceted evaluation results.
[0088] "Customer settings" refers to settings that include user profile information and scenario information.
[0089] "Voice input" refers to voice data that a user produces using a device.
[0090] "Text conversion" refers to the process of converting voice input into text data.
[0091] "Generating natural responses" refers to the process of generating appropriate and natural responses based on converted text data.
[0092] "Speech synthesis" is the process of converting generated text responses into speech data.
[0093] "Providing voice response" refers to playing back the converted voice data to the user.
[0094] "Conversation content analysis" refers to the process of analyzing recorded conversation data to evaluate user performance.
[0095] "Performance evaluation" refers to a comprehensive assessment of factors such as the appropriateness, speed, and clarity of user responses.
[0096] "Feedback" refers to providing users with evaluation results to offer suggestions for improvement and highlight strengths in their training.
[0097] A "terminal" is a device used by a user to access a system and perform tasks such as entering settings or using voice input.
[0098] A "server" is a computer system that centrally stores, processes, and analyzes various types of data.
[0099] A "natural language processing engine" is a software module that analyzes text data and generates appropriate responses.
[0100] A "speech synthesis module" is a software module used to convert text data into speech data.
[0101] "User configuration information" refers to profile information and scenario information entered by the user.
[0102] "Real-time recording" refers to the instantaneous recording of the user's voice.
[0103] The present invention is a system that includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user.
[0104] First, the user accesses the terminal and enters customer settings, including profile information and scenario information. For example, they might enter information such as "male in his 50s, bank employee, customer complaint handling scenario." The terminal sends this information to the server, which stores the received data in a database.
[0105] Next, when the user presses the "Start Conversation" button, the device records the user's voice in real time and sends the recording data to the server. The server uses speech recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text to convert the voice into text data. For example, if the user says, "Excuse me, sorry to have kept you waiting," the speech recognition technology will convert this into text format.
[0106] The converted text data is passed to the server's natural language processing engine. This engine considers the user's scenario settings and conversational context to generate an appropriate response. For example, it might generate a response such as "We apologize for the wait, sir / madam" in response to a user's utterance. This natural language processing engine may utilize existing AI services (e.g., GPT-4®).
[0107] The generated text response is sent to a speech synthesis module and converted into audio data. Google Text-to-Speech API or Amazon Polly are used for speech synthesis. The generated audio data is sent from the server to the device, which then plays it back to the user.
[0108] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. Based on this evaluation data, the server calculates a score and generates detailed feedback. The evaluation results are provided to the user via the terminal. For example, an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points" might be displayed.
[0109] As a concrete example, the user enters the following prompt:
[0110] "The user sets the information as 'a man in his 50s, a bank employee, handling a customer complaint scenario,' and begins role-playing. When the user says, 'Could you explain that matter again?', how should the system respond?"
[0111] In response to this prompt, the system generates an appropriate response and provides it in voice.
[0112] Therefore, the present invention provides a systematic system that comprehensively handles everything from inputting user configuration information to speech recognition, response generation, speech synthesis, and performance evaluation, and provides high-quality feedback to the user. This enables the user to receive realistic and effective training.
[0113] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0114] Step 1:
[0115] The user accesses the terminal and enters profile information and scenario information. For example, the user enters "50s male, banker, complaint handling scenario". This input data is sent from the terminal to the server as user settings information. The server saves the received data to a database. Specifically, the server executes and saves the data by executing the SQL query "INSERT INTO customer_settings (age, occupation, scenario) VALUES (50, 'banker', 'complaint handling')".
[0116] Input: User profile information and scenario information
[0117] Output: User settings information stored in the database
[0118] Step 2:
[0119] When the user presses the "Start Conversation" button, the device activates the microphone and records the user's voice in real time. The recorded audio data is sent from the device to the server via WebSocket. The server receives this audio data.
[0120] Input: User voice input
[0121] Output: Audio data sent to the server
[0122] Step 3:
[0123] The server receives the audio data and passes it to the Google Speech-to-Text API. The API converts the audio data into text data and returns that text data to the server. For example, the user's utterance "Excuse me, sorry to keep you waiting" is converted into the text data "Sorry to keep you waiting". The server receives this converted result.
[0124] Input: Audio data
[0125] Output: Text data
[0126] Step 4:
[0127] The server passes text data to a natural language processing (NLP) engine. The NLP engine automatically generates an appropriate response, taking into account the user's scenario setting and the context of the conversation. For example, if the user says, "Excuse me, sorry to have kept you waiting," the NLP engine will generate a response such as, "We apologize for the wait, customer."
[0128] Input: Text data, user settings information
[0129] Output: Generated text response
[0130] Step 5:
[0131] The server passes the generated text response to a text-to-speech module. The text-to-speech module (e.g., Google Text-to-Speech API) converts this into speech data and returns the speech data to the server. The server receives this speech data and sends it to the device. The device plays the received speech data and provides the user with a speech response.
[0132] Input: Generated text response
[0133] Output: Audio data
[0134] Step 6:
[0135] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. Specifically, the server analyzes metrics such as the appropriateness and speed of responses and the clarity of pronunciation, and calculates a score. For example, it might generate an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points." The generated evaluation result is sent to the terminal, which then displays it to the user.
[0136] Input: Conversation data
[0137] Output: Evaluation results
[0138] Through these steps, the system provides users with a consistent experience, from setting up their configuration to engaging in actual role-playing and providing detailed feedback.
[0139] (Application Example 1)
[0140] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0141] In modern customer service, staff need to improve their customer service skills quickly and effectively. However, traditional training methods often involve manual scenario setting and feedback, resulting in inefficiency and a lack of real-time response. Furthermore, the naturalness of voice responses and the reflection of conversational context are insufficient, making it difficult to provide a practical training environment. To solve these problems, there is a need for a system that utilizes more advanced speech recognition and natural response generation technologies.
[0142] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0143] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for initiating a training session based on profile information and scenario information, means for generating prompt sentences using a generative AI model, means for generating a response that reflects the conversation context, and means for generating a natural response based on the generated prompt sentences. This enables real-time voice recognition and response generation, resulting in high effectiveness and immediate results in customer service training.
[0144] "Customer settings" refer to configuration data that includes user profile information and scenario information.
[0145] "Voice input" is the process of acquiring a user's speech in digital format.
[0146] "Text conversion" is the process of converting voice input into written text.
[0147] "Natural response generation" is the process of generating natural conversational responses based on converted text and configured profile information.
[0148] "Speech synthesis" is the process of converting generated text responses into speech data.
[0149] "Voice response" is a method of providing users with synthesized speech responses.
[0150] "Performance evaluation" is a process that analyzes conversation content and evaluates user response and other performance aspects.
[0151] "Feedback" is the process of returning evaluation results to the user, highlighting areas for improvement and strengths.
[0152] "Profile information" refers to personal information such as the user's age, occupation, and job title.
[0153] "Scenario information" refers to scenario data for conversations that are set with specific situations or conditions.
[0154] "Prompt generation" is the process of generating input sentences that are appropriate to the conversational context using a generative AI model.
[0155] "Conversation context" refers to contextual information that includes the current situation and background information of the conversation.
[0156] The system according to the present invention is a voice-response-based role-playing system intended for staff training and improving customer service skills in physical stores.
[0157] First, the user accesses the system and registers profile information (age, occupation, etc.) and scenario information (specific training scenarios) as customer settings. This setting information is entered from a smartphone or other device and sent to the server. The server stores this information in a database.
[0158] Next, the user starts a role-playing session by pressing the "Start Conversation" button. The user's voice input is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the speech to text.
[0159] The converted text is passed to the server's natural language processing (NLP) engine. The NLP engine uses a generative AI model to generate a prompt based on the user's profile information and the context of the conversation. This prompt will look like this:
[0160] Please generate a natural response considering the following context.
[0161] Context: Bank customer complaint handling
[0162] User utterance: I'm sorry, sir / madam, but...
[0163] Response: "We apologize for the wait."
[0164] A natural-sounding response is generated based on this generated prompt. The generated text response is converted into audio data using the Google Text-to-Speech API and sent to the device. The device plays this audio data and provides the user with an audio response.
[0165] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes evaluation criteria such as the appropriateness, speed, and clarity of responses. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, highlighting strengths and areas for improvement.
[0166] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal. This evaluation and feedback allows users to improve their practical customer service skills while simulating real-world scenarios.
[0167] Thus, the present invention achieves real-time performance and high training effectiveness by using advanced speech recognition technology and natural response generation technology.
[0168] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0169] Step 1:
[0170] The user accesses the system and enters profile information and scenario information. Specifically, they enter information such as "female in her 30s, bank employee, customer complaint handling scenario" from a smartphone or other device, and the device sends this information to the server. The server stores the received customer settings information in a database.
[0171] Input: User profile information and scenario information
[0172] Output: Customer settings information stored in the database
[0173] Operation: Sending information from the terminal, receiving information on the server, and saving it to the database.
[0174] Step 2:
[0175] The user presses the "Start Conversation" button to begin a role-playing session. The user's speech is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the received audio data into text.
[0176] Input: User voice input
[0177] Output: User utterance converted to text
[0178] Function: Audio recording on the device, transmission of audio data, speech recognition and text conversion on the server.
[0179] Step 3:
[0180] The server passes the converted text to a natural language processing (NLP) engine, which generates prompt sentences based on profile information and conversational context. Using a generative AI model, it generates prompt sentences such as, "Please generate a natural response considering the following context. Context: Bank customer service, User utterance: I'm sorry, sir / madam, but..., Response: I'm sorry to have kept you waiting."
[0181] Input: User utterances converted to text, profile information, scenario information
[0182] Output: Generated prompt message
[0183] Operation: Uses a natural language processing engine and a generative AI model to generate prompt sentences.
[0184] Step 4:
[0185] Based on the generated prompt, the server produces a natural-sounding response. This response is converted into speech data using the Google Text-to-Speech API. The synthesized response is then sent from the server to the terminal.
[0186] Input: Generated prompt message
[0187] Output: Response converted into audio data
[0188] Function: Generate text responses, convert them to audio data, and send the audio data.
[0189] Step 5:
[0190] The device plays the received audio data and provides the user with an audio response. The user listens to the audio response and then speaks again.
[0191] Input: Response converted into voice data
[0192] Output: Providing voice responses to the user
[0193] Function: Receiving audio data, playing audio
[0194] Step 6:
[0195] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes aspects such as the appropriateness of responses, speed, and clarity of pronunciation. The server generates a score and detailed feedback from the evaluation data and sends the results to the terminal.
[0196] Input: Recorded conversation data
[0197] Output: User ratings and feedback
[0198] Function: Analyzes conversation data, generates evaluation scores, and creates and sends feedback.
[0199] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0200] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing conversation content and evaluating user performance, means for providing feedback of the evaluation results to the user, and an emotion engine for recognizing the user's emotions.
[0201] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in the database.
[0202] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for speech recognition.
[0203] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0204] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0205] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion.
[0206] The server generates a score and detailed feedback from this data and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements. For example, the server generates evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points" and displays them to the user through the terminal.
[0207] This allows users to receive multifaceted feedback, including on their own emotional management, while undergoing realistic customer service training. The above is a specific embodiment of the system according to the present invention.
[0208] The following describes the processing flow.
[0209] Step 1:
[0210] The user accesses the customer settings registration screen from their device.
[0211] Step 2:
[0212] The user enters their profile information and scenario information into the device.
[0213] Step 3:
[0214] The terminal sends the entered customer settings information to the server.
[0215] Step 4:
[0216] The server saves the received information to the database.
[0217] Step 5:
[0218] The user presses the "Start Conversation" button on their device.
[0219] Step 6:
[0220] The device records the user's voice in real time and sends that audio data to the server.
[0221] Step 7:
[0222] The server's speech recognition module converts the speech data into text.
[0223] Step 8:
[0224] The server passes the converted text data to the natural language processing (NLP) engine.
[0225] Step 9:
[0226] The server's NLP engine analyzes the converted text and generates an appropriate response text.
[0227] Step 10:
[0228] The server passes the voice data to the emotion engine, which then analyzes the user's emotions.
[0229] Step 11:
[0230] The emotion engine feeds back the detected emotion information to the NLP engine.
[0231] Step 12:
[0232] The server's NLP engine generates more natural response text based on the user's emotions.
[0233] Step 13:
[0234] The server sends the generated response text to the speech synthesis module.
[0235] Step 14:
[0236] The server's speech synthesis module converts text into speech data.
[0237] Step 15:
[0238] The server sends the audio data to the terminal.
[0239] Step 16:
[0240] The device plays the received audio data to the user.
[0241] Step 17:
[0242] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0243] Step 18:
[0244] The server analyzes the recorded conversation data.
[0245] Step 19:
[0246] The server generates scores and detailed feedback based on analytical data to evaluate the user's responsiveness, skills, and emotional response.
[0247] Step 20:
[0248] The server sends the evaluation results and feedback to the terminal.
[0249] Step 21:
[0250] The device displays evaluation results to the user, providing specific areas for improvement and highlighting achievements.
[0251] (Example 2)
[0252] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0253] Traditional customer service training systems struggled to generate responses that took user emotions into account and to evaluate performance. As a result, training was ineffective, and improvements in actual customer service skills were not achieved. In particular, they were unable to provide appropriate feedback when users were experiencing emotions such as tension or fear.
[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0255] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, and means for recognizing the user's emotions and feeding the results back into generating a natural response. This makes it possible to provide appropriate responses that take the user's emotions into account in real time, as well as to perform detailed performance evaluations and provide feedback.
[0256] "Customer settings" refer to the user's profile information and conversation scenario information, which the system uses to customize responses according to the user's individual needs.
[0257] "Voice input" refers to audio data supplied to the system by the user using a microphone or other voice acquisition device.
[0258] "Methods for converting to text" refers to technologies that analyze voice input and convert its content into text data.
[0259] "Means for generating natural responses" refers to technologies that generate appropriate responses based on converted text data, taking into account the flow, context, and emotions of the user interaction.
[0260] "Means of converting to speech" refers to technology that converts the generated text response back into speech data and provides it to the user as speech.
[0261] "Means of providing voice responses to users" refers to technology that transmits converted voice data to the user's device, allowing the user to listen to the voice.
[0262] "Means for analyzing conversation content" refers to technologies that analyze recorded conversation data and evaluate the content and quality of the conversation.
[0263] "Means for evaluating user performance" refers to technologies that evaluate a user's responsiveness and communication skills based on the content of the conversation and the user's responses at that time.
[0264] "Means of providing feedback on evaluation results to users" refers to technologies that communicate the results of the analyzed performance evaluation to users and provide detailed information on areas for improvement and achievements.
[0265] "Means of recognizing user emotions" refers to technologies that extract and analyze a user's emotional state from audio data and text data.
[0266] An "emotion engine" refers to an engine that identifies emotions from a user's voice or text and feeds the results back to other system components.
[0267] The system according to the present invention is for users to perform interactive training and for which their performance is evaluated and feedback is provided. This system has the following means:
[0268] 1. Register user settings
[0269] First, the user accesses the system using a terminal and enters profile information and scenario information on the settings screen. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario." This information is sent from the terminal to the server, which stores it in a database. The database used is a relational database such as MySQL® or PostgreSQL.
[0270] 2. Start of role-playing
[0271] When the user presses the "Start Conversation" button on the system, the device records the user's voice and sends it to the server in real time. At this time, the device uses its built-in microphone to capture the voice.
[0272] 3. Text conversion of voice input
[0273] The server receives the audio data sent from the terminal. This data is then converted into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text.
[0274] 4. Generating a response
[0275] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used here include SpaCy and NLTK. The audio data is also input to an emotion engine, which uses the Emotion API or a generative AI model (e.g., GPT-3®) to extract the user's emotions. The emotion engine feeds its results back to the NLP engine, which then generates a natural response considering the customer's settings, the context of the conversation, and the user's emotions. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0276] 5. Providing voice response
[0277] The generated text response is passed to a speech synthesis module. This speech synthesis module converts the text into speech data using the Google Text-to-Speech API or Amazon Polly. The server sends the generated speech data to the device, and the device plays the audio for the user.
[0278] 6. Performance Evaluation and Feedback
[0279] Once the role-playing session ends, the server collects and stores all recorded conversation data. Next, the server analyzes the stored conversation data and evaluates the user's performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression. Based on these evaluation results, the server generates detailed feedback and a score. The evaluation results are sent to the terminal and displayed to the user. For example, it might be displayed in the format: "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[0280] As a concrete example, consider a scenario where a user enters information such as "25-year-old female, call center representative, product return scenario" on the settings screen. The terminal sends this information to the server, which stores it in a database. The user presses the "Start Conversation" button, and the terminal records the audio and sends it to the server. The server converts the audio data into text using the Google Speech-to-Text API. The NLP engine generates a response such as "Excuse me, sir / madam. I will handle this matter." The speech synthesis module uses Amazon Polly to convert this response into audio data, which the terminal plays. After the session ends, the server generates an evaluation result such as "Response Speed: 93 points, Customer Satisfaction Response: 85 points, Pronunciation: 88 points, Emotion Recognition: 80 points" and sends it to the terminal. The terminal displays this to the user.
[0281] Examples of prompt messages include the following:
[0282] "Please generate a conversation regarding a customer being made to wait in a claim handling scenario for a bank clerk modeled on a 50-year-old male."
[0283] "Based on the following conversation, analyze the user's emotions and provide feedback: 'Sorry to have made you wait.'"
[0284] With this system, the user can conduct customer service training similar to actual business scenarios and improve skills through detailed feedback."
[0285] The flow of the specific process in Example 2 will be described using FIG. 13."
[0286] Step 1:
[0287] Registration of user settings
[0288] The user accesses the system using a terminal and inputs profile information and scenario information. For example, information such as '50-year-old male, bank clerk, claim handling scenario' is included."
[0289] Input: Profile information and scenario information input by the user via the terminal."
[0290] The terminal sends the input information to the server. At this time, the terminal uses an HTTP request to send the information to the server."
[0291] The server receives the transmitted information and saves it in the database. MySQL or PostgreSQL is used as the database."
[0292] Output: User setting information saved in the database."
[0293] Step 2:
[0294] Start of role-playing
[0295] The user presses the "Start Conversation" button on the terminal.
[0296] Input: Pressing operation of the "Start Conversation" button.
[0297] The terminal uses the built-in microphone to record the user's voice in real time. This voice data is sent to the server using WebSocket or HTTP POST request. <00009३९>
[0298] Output: The recorded voice data is sent to the server.
[0299] Step 3:
[0300] Text conversion of voice input
[0301] The server receives the voice data sent from the terminal.
[0302] Input: Voice data sent from the terminal.
[0303] The server converts the received voice data into text using a speech recognition module such as Google Speech-to-Text API or IBM Watson Speech to Text. This module analyzes the voice waveform and converts its content into text format.
[0304] Output: Voice data converted into text.
[0305] Step 4:
[0306] Response generation
[0307] The server passes the converted text data to a natural language processing (NLP) engine. As the NLP engine, for example, SpaCy or NLTK is used.
[0308] Input: Voice data converted into text.
[0309] The server simultaneously inputs voice data into the emotion engine to detect the user's emotions. The emotion engine uses the Emotion API or a generative AI model (e.g., GPT-3).
[0310] The NLP engine generates natural responses based on text data, customer settings, conversation context, and user emotions. For example, if a user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, it will generate a response such as, "You seem a little nervous, is there something you're worried about?"
[0311] Output: Text data of the generated natural response.
[0312] Step 5:
[0313] Providing voice response
[0314] The speech synthesis module receives the generated text response.
[0315] Input: Text data of the generated natural response.
[0316] The speech synthesis module uses the Google Text-to-Speech API or Amazon Polly to convert this text data into speech data.
[0317] The server sends the generated audio data to the terminal.
[0318] The device plays this audio data, and the user can hear the voice response.
[0319] Output: Audio data played by the device.
[0320] Step 6:
[0321] Performance evaluation and feedback
[0322] Once the role-playing session ends, the server collects and saves all conversation data.
[0323] Input: All recorded conversation data.
[0324] The server analyzes stored conversation data to evaluate user performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression.
[0325] The server generates a score and detailed feedback based on these evaluation results. For example, it might be displayed in the format of "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[0326] Output: Performance evaluation results and feedback provided to the user.
[0327] (Application Example 2)
[0328] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0329] Conventional customer service training systems have limitations in how they evaluate user voice input and the appropriateness of responses. They lack the ability to perform real-time sentiment analysis and provide feedback based on that analysis, resulting in insufficient improvement in the performance of customer service staff. Furthermore, they are unable to generate flexible responses based on emotions, making effective training in real-world customer interactions at physical stores difficult.
[0330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and saving customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for analyzing emotions, and means for adjusting the response based on emotions. This enables more practical and effective customer service training by analyzing the user's emotions in real time and generating flexible responses based on those emotions.
[0331] "Customer settings" refer to configuration information that includes user profile information and scenario information, and are customized data that the user provides to the system.
[0332] "Voice input" refers to the voice data spoken by the user, which is the input information used by the system to convert into text.
[0333] "Text conversion" is the process of converting voice input into text information using speech recognition technology.
[0334] "Natural response generation" is the process of generating appropriate and natural responses by considering the context of the converted text and conversation, as well as the user's sentiment data.
[0335] "Speech conversion" is the process of converting generated text responses into speech data.
[0336] "Voice response provision" refers to the process of playing back converted voice data to the user.
[0337] "Conversation content analysis" is the process of analyzing recorded conversation data to evaluate user performance.
[0338] "Evaluation result feedback" is a process that provides evaluation results, such as the appropriateness and speed of user responses, based on the analyzed data.
[0339] "Emotion analysis" is the process of extracting and identifying emotions from a user's statements, tone of voice, and other factors.
[0340] "Response adjustment" is the process of appropriately modifying and adjusting the generated responses based on analyzed emotional data.
[0341] A "prompt" is text input into a generative AI model, and it is an instruction to generate a response based on the user's situation and emotions.
[0342] In the system according to this invention, the user first accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information such as "30s, store clerk, customer complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[0343] Next, when the user presses the "Start Conversation" button, the device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Common speech recognition technologies can be used for speech recognition.
[0344] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0345] The generated text response is sent to the server's speech synthesis module and converted into audio data. Common speech synthesis technologies can be used for this conversion. The server then sends this audio data to the terminal, which plays the response back to the user. This allows the user to receive real-time, emotion-based feedback.
[0346] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. The evaluation includes appropriateness of responses, speed, clarity of pronunciation, and emotional changes. From this data, the server generates a score and detailed feedback, and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements.
[0347] This process uses the following specific hardware and software:
[0348] Hardware: Smartphone, microphone
[0349] Software: Speech recognition library (speech_recognition), text-to-speech conversion library (pyttsx3), natural language processing library (transformers)
[0350] As a concrete example, a response is generated using the following prompt:
[0351] The user is 30 years old, working as a shop assistant, handling a customer complaint scenario. The user said: "Excuse me, sorry to have kept you waiting." Detected emotion: "Nervous." Respond appropriately.
[0352] In this way, users can receive practical customer service training and have their performance evaluated from multiple perspectives.
[0353] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0354] Step 1:
[0355] The user enters their customer settings on the connected device. The user uses a smartphone or computer to enter profile information (age, occupation, etc.) and scenario information (e.g., a customer complaint handling scenario). The entered customer settings information is sent from the device to the server. The server receives this information and stores it in a database. This allows the system to prepare a customized training state for each user.
[0356] Step 2:
[0357] When the user presses the "Start Conversation" button, the device records the user's voice through the microphone and sends the audio data to the server in real time. The server uses a speech recognition library to convert this audio data into text. Specifically, it analyzes the audio data and converts phonemes and syllables into text. Audio data is input, and the converted text is output.
[0358] Step 3:
[0359] The server passes the converted text to a natural language processing (NLP) engine and simultaneously to an emotion analysis engine. The emotion analysis engine analyzes the audio data and text to detect the user's emotions. This process identifies the user's emotions (e.g., tension, joy, anger) from the tone of voice and word choice. The input is audio data and text data, and the output is the emotion analysis result.
[0360] Step 4:
[0361] Once the sentiment analysis results are output, they are fed back to the NLP engine. The NLP engine generates an appropriate and natural response, taking into account the user's emotions, profile information, and the context of the conversation. For example, if tension is detected, the NLP engine will generate a response such as, "You seem a little nervous, is there anything you're worried about?" The input is text data and sentiment analysis results, and the output is the generated text response.
[0362] Step 5:
[0363] The server sends the generated text response to a speech synthesis module, which converts it into speech data. The speech synthesis module uses a technique to convert text data into speech data (e.g., text-to-speech software). The generated speech data is sent to the terminal, which plays this speech data and provides the user with a voice response. The input is the generated text response, and the output is the speech data.
[0364] Step 6:
[0365] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation process uses criteria such as the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion. The input is the conversation data, and the output is the user's evaluation result.
[0366] Step 7:
[0367] The server generates a score and detailed feedback based on the evaluation results and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user and provides specific areas for improvement and achievements. For example, the user is provided with detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points." The input is the evaluation result, and the output is the feedback presented to the user.
[0368] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0369] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0370] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0371] [Second Embodiment]
[0372] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0373] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0374] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0375] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0376] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0378] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0379] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0380] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0381] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0382] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0383] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0384] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user. Specific embodiments are described below.
[0385] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[0386] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for this speech recognition.
[0387] The converted text is then passed to the server's natural language processing (NLP) engine. The NLP engine takes into account the user's settings and the context of the conversation to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," the server will generate a response such as, "We apologize for the wait, customer."
[0388] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0389] Once the role-playing session ends, the server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, providing information on what went well and where there is room for improvement.
[0390] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal.
[0391] This allows users to conduct realistic customer service training and identify specific areas for improvement. This is a specific embodiment of the system according to the present invention.
[0392] The following describes the processing flow.
[0393] Step 1:
[0394] The user accesses the customer settings registration screen from their device.
[0395] Step 2:
[0396] The user enters their profile information and scenario information into the device.
[0397] Step 3:
[0398] The terminal sends the entered customer settings information to the server.
[0399] Step 4:
[0400] The server saves the received information to the database.
[0401] Step 5:
[0402] The user presses the "Start Conversation" button on their device.
[0403] Step 6:
[0404] The device records the user's voice in real time and sends that audio data to the server.
[0405] Step 7:
[0406] The server's speech recognition module converts the speech data into text.
[0407] Step 8:
[0408] The server passes the converted text data to the natural language processing (NLP) engine.
[0409] Step 9:
[0410] The server's NLP engine generates appropriate responses based on customer settings and the context of the conversation.
[0411] Step 10:
[0412] The server sends the generated response text to the speech synthesis module.
[0413] Step 11:
[0414] The server's speech synthesis module converts text into speech data.
[0415] Step 12:
[0416] The server sends the audio data to the terminal.
[0417] Step 13:
[0418] The device plays the received audio data to the user.
[0419] Step 14:
[0420] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0421] Step 15:
[0422] The server analyzes all recorded conversation data.
[0423] Step 16:
[0424] The server generates a score and detailed feedback based on the analysis data to evaluate the user's responsiveness and skills.
[0425] Step 17:
[0426] The server sends the evaluation results and feedback to the terminal.
[0427] Step 18:
[0428] The device displays the evaluation results to the user and provides feedback.
[0429] (Example 1)
[0430] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0431] Existing speech recognition systems have the challenge of not being able to quickly provide appropriate responses based on specific scenarios and profile information set by the user, and not being able to evaluate user performance from multiple perspectives in real time. This invention solves these problems and provides a system that enables users to conduct more realistic and effective training.
[0432] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0433] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to speech, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, a terminal for inputting user setting information, a terminal for recording speech in real time and transmitting it to the server, a response generation means using a natural language processing engine, a response speech conversion means using a speech synthesis module, and means for evaluating the user's performance using multiple indicators and generating detailed feedback. This enables highly accurate response generation based on the user's settings and the provision of multifaceted evaluation results.
[0434] "Customer settings" refers to settings that include user profile information and scenario information.
[0435] "Voice input" refers to voice data that a user produces using a device.
[0436] "Text conversion" refers to the process of converting voice input into text data.
[0437] "Generating natural responses" refers to the process of generating appropriate and natural responses based on converted text data.
[0438] "Speech synthesis" is the process of converting generated text responses into speech data.
[0439] "Providing voice response" refers to playing back the converted voice data to the user.
[0440] "Conversation content analysis" refers to the process of analyzing recorded conversation data to evaluate user performance.
[0441] "Performance evaluation" refers to a comprehensive assessment of factors such as the appropriateness, speed, and clarity of user responses.
[0442] "Feedback" refers to providing users with evaluation results to offer suggestions for improvement and highlight strengths in their training.
[0443] A "terminal" is a device used by a user to access a system and perform tasks such as entering settings or using voice input.
[0444] A "server" is a computer system that centrally stores, processes, and analyzes various types of data.
[0445] A "natural language processing engine" is a software module that analyzes text data and generates appropriate responses.
[0446] A "speech synthesis module" is a software module used to convert text data into speech data.
[0447] "User configuration information" refers to profile information and scenario information entered by the user.
[0448] "Real-time recording" refers to the instantaneous recording of the user's voice.
[0449] The present invention is a system that includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user.
[0450] First, the user accesses the terminal and enters customer settings, including profile information and scenario information. For example, they might enter information such as "male in his 50s, bank employee, customer complaint handling scenario." The terminal sends this information to the server, which stores the received data in a database.
[0451] Next, when the user presses the "Start Conversation" button, the device records the user's voice in real time and sends the recording data to the server. The server uses speech recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text to convert the voice into text data. For example, if the user says, "Excuse me, sorry to have kept you waiting," the speech recognition technology will convert this into text format.
[0452] The converted text data is passed to the server's natural language processing engine. This engine considers the user's scenario settings and conversational context to generate an appropriate response. For example, it might generate a response such as "We apologize for the wait, customer" in response to a user's utterance. This natural language processing engine may utilize existing AI services (e.g., GPT-4).
[0453] The generated text response is sent to a speech synthesis module and converted into audio data. Google Text-to-Speech API or Amazon Polly are used for speech synthesis. The generated audio data is sent from the server to the device, which then plays it back to the user.
[0454] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. Based on this evaluation data, the server calculates a score and generates detailed feedback. The evaluation results are provided to the user via the terminal. For example, an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points" might be displayed.
[0455] As a concrete example, the user enters the following prompt:
[0456] "The user sets the information as 'a man in his 50s, a bank employee, handling a customer complaint scenario,' and begins role-playing. When the user says, 'Could you explain that matter again?', how should the system respond?"
[0457] In response to this prompt, the system generates an appropriate response and provides it in voice.
[0458] Therefore, the present invention provides a systematic system that comprehensively handles everything from inputting user configuration information to speech recognition, response generation, speech synthesis, and performance evaluation, and provides high-quality feedback to the user. This enables the user to receive realistic and effective training.
[0459] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0460] Step 1:
[0461] The user accesses the terminal and enters profile information and scenario information. For example, the user enters "50s male, banker, complaint handling scenario". This input data is sent from the terminal to the server as user settings information. The server saves the received data to a database. Specifically, the server executes and saves the data by executing the SQL query "INSERT INTO customer_settings (age, occupation, scenario) VALUES (50, 'banker', 'complaint handling')".
[0462] Input: User profile information and scenario information
[0463] Output: User settings information stored in the database
[0464] Step 2:
[0465] When the user presses the "Start Conversation" button, the device activates the microphone and records the user's voice in real time. The recorded audio data is sent from the device to the server via WebSocket. The server receives this audio data.
[0466] Input: User voice input
[0467] Output: Audio data sent to the server
[0468] Step 3:
[0469] The server receives the audio data and passes it to the Google Speech-to-Text API. The API converts the audio data into text data and returns that text data to the server. For example, the user's utterance "Excuse me, sorry to keep you waiting" is converted into the text data "Sorry to keep you waiting". The server receives this converted result.
[0470] Input: Audio data
[0471] Output: Text data
[0472] Step 4:
[0473] The server passes text data to a natural language processing (NLP) engine. The NLP engine automatically generates an appropriate response, taking into account the user's scenario setting and the context of the conversation. For example, if the user says, "Excuse me, sorry to have kept you waiting," the NLP engine will generate a response such as, "We apologize for the wait, customer."
[0474] Input: Text data, user settings information
[0475] Output: Generated text response
[0476] Step 5:
[0477] The server passes the generated text response to a text-to-speech module. The text-to-speech module (e.g., Google Text-to-Speech API) converts this into speech data and returns the speech data to the server. The server receives this speech data and sends it to the device. The device plays the received speech data and provides the user with a speech response.
[0478] Input: Generated text response
[0479] Output: Audio data
[0480] Step 6:
[0481] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. Specifically, the server analyzes metrics such as the appropriateness and speed of responses and the clarity of pronunciation, and calculates a score. For example, it might generate an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points." The generated evaluation result is sent to the terminal, which then displays it to the user.
[0482] Input: Conversation data
[0483] Output: Evaluation results
[0484] Through these steps, the system provides users with a consistent experience, from setting up their configuration to engaging in actual role-playing and providing detailed feedback.
[0485] (Application Example 1)
[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0487] In modern customer service, staff need to improve their customer service skills quickly and effectively. However, traditional training methods often involve manual scenario setting and feedback, resulting in inefficiency and a lack of real-time response. Furthermore, the naturalness of voice responses and the reflection of conversational context are insufficient, making it difficult to provide a practical training environment. To solve these problems, there is a need for a system that utilizes more advanced speech recognition and natural response generation technologies.
[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0489] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for initiating a training session based on profile information and scenario information, means for generating prompt sentences using a generative AI model, means for generating a response that reflects the conversation context, and means for generating a natural response based on the generated prompt sentences. This enables real-time voice recognition and response generation, resulting in high effectiveness and immediate results in customer service training.
[0490] "Customer settings" refer to configuration data that includes user profile information and scenario information.
[0491] "Voice input" is the process of acquiring a user's speech in digital format.
[0492] "Text conversion" is the process of converting voice input into written text.
[0493] "Natural response generation" is the process of generating natural conversational responses based on converted text and configured profile information.
[0494] "Speech synthesis" is the process of converting generated text responses into speech data.
[0495] "Voice response" is a method of providing users with synthesized speech responses.
[0496] "Performance evaluation" is a process that analyzes conversation content and evaluates user response and other performance aspects.
[0497] "Feedback" is the process of returning evaluation results to the user, highlighting areas for improvement and strengths.
[0498] "Profile information" refers to personal information such as the user's age, occupation, and job title.
[0499] "Scenario information" refers to scenario data for conversations that are set with specific situations or conditions.
[0500] "Prompt generation" is the process of generating input sentences that are appropriate to the conversational context using a generative AI model.
[0501] "Conversation context" refers to contextual information that includes the current situation and background information of the conversation.
[0502] The system according to the present invention is a voice-response-based role-playing system intended for staff training and improving customer service skills in physical stores.
[0503] First, the user accesses the system and registers profile information (age, occupation, etc.) and scenario information (specific training scenarios) as customer settings. This setting information is entered from a smartphone or other device and sent to the server. The server stores this information in a database.
[0504] Next, the user starts a role-playing session by pressing the "Start Conversation" button. The user's voice input is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the speech to text.
[0505] The converted text is passed to the server's natural language processing (NLP) engine. The NLP engine uses a generative AI model to generate a prompt based on the user's profile information and the context of the conversation. This prompt will look like this:
[0506] Please generate a natural response considering the following context.
[0507] Context: Bank customer complaint handling
[0508] User utterance: I'm sorry, sir / madam, but...
[0509] Response: "We apologize for the wait."
[0510] A natural-sounding response is generated based on this generated prompt. The generated text response is converted into audio data using the Google Text-to-Speech API and sent to the device. The device plays this audio data and provides the user with an audio response.
[0511] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes evaluation criteria such as the appropriateness, speed, and clarity of responses. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, highlighting strengths and areas for improvement.
[0512] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal. This evaluation and feedback allows users to improve their practical customer service skills while simulating real-world scenarios.
[0513] Thus, the present invention achieves real-time performance and high training effectiveness by using advanced speech recognition technology and natural response generation technology.
[0514] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0515] Step 1:
[0516] The user accesses the system and enters profile information and scenario information. Specifically, they enter information such as "female in her 30s, bank employee, customer complaint handling scenario" from a smartphone or other device, and the device sends this information to the server. The server stores the received customer settings information in a database.
[0517] Input: User profile information and scenario information
[0518] Output: Customer settings information stored in the database
[0519] Operation: Sending information from the terminal, receiving information on the server, and saving it to the database.
[0520] Step 2:
[0521] The user presses the "Start Conversation" button to begin a role-playing session. The user's speech is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the received audio data into text.
[0522] Input: User voice input
[0523] Output: User utterance converted to text
[0524] Function: Audio recording on the device, transmission of audio data, speech recognition and text conversion on the server.
[0525] Step 3:
[0526] The server passes the converted text to a natural language processing (NLP) engine, which generates prompt sentences based on profile information and conversational context. Using a generative AI model, it generates prompt sentences such as, "Please generate a natural response considering the following context. Context: Bank customer service, User utterance: I'm sorry, sir / madam, but..., Response: I'm sorry to have kept you waiting."
[0527] Input: User utterances converted to text, profile information, scenario information
[0528] Output: Generated prompt message
[0529] Operation: Uses a natural language processing engine and a generative AI model to generate prompt sentences.
[0530] Step 4:
[0531] Based on the generated prompt, the server produces a natural-sounding response. This response is converted into speech data using the Google Text-to-Speech API. The synthesized response is then sent from the server to the terminal.
[0532] Input: Generated prompt message
[0533] Output: Response converted into audio data
[0534] Function: Generate text responses, convert them to audio data, and send the audio data.
[0535] Step 5:
[0536] The device plays the received audio data and provides the user with an audio response. The user listens to the audio response and then speaks again.
[0537] Input: Response converted into voice data
[0538] Output: Providing voice responses to the user
[0539] Function: Receiving audio data, playing audio
[0540] Step 6:
[0541] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes aspects such as the appropriateness of responses, speed, and clarity of pronunciation. The server generates a score and detailed feedback from the evaluation data and sends the results to the terminal.
[0542] Input: Recorded conversation data
[0543] Output: User ratings and feedback
[0544] Function: Analyzes conversation data, generates evaluation scores, and creates and sends feedback.
[0545] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0546] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing conversation content and evaluating user performance, means for providing feedback of the evaluation results to the user, and an emotion engine for recognizing the user's emotions.
[0547] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in the database.
[0548] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for speech recognition.
[0549] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0550] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0551] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion.
[0552] The server generates a score and detailed feedback from this data and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements. For example, the server generates evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points" and displays them to the user through the terminal.
[0553] This allows users to receive multifaceted feedback, including on their own emotional management, while undergoing realistic customer service training. The above is a specific embodiment of the system according to the present invention.
[0554] The following describes the processing flow.
[0555] Step 1:
[0556] The user accesses the customer settings registration screen from their device.
[0557] Step 2:
[0558] The user enters their profile information and scenario information into the device.
[0559] Step 3:
[0560] The terminal sends the entered customer settings information to the server.
[0561] Step 4:
[0562] The server saves the received information to the database.
[0563] Step 5:
[0564] The user presses the "Start Conversation" button on their device.
[0565] Step 6:
[0566] The device records the user's voice in real time and sends that audio data to the server.
[0567] Step 7:
[0568] The server's speech recognition module converts the speech data into text.
[0569] Step 8:
[0570] The server passes the converted text data to the natural language processing (NLP) engine.
[0571] Step 9:
[0572] The server's NLP engine analyzes the converted text and generates an appropriate response text.
[0573] Step 10:
[0574] The server passes the voice data to the emotion engine, which then analyzes the user's emotions.
[0575] Step 11:
[0576] The emotion engine feeds back the detected emotion information to the NLP engine.
[0577] Step 12:
[0578] The server's NLP engine generates more natural response text based on the user's emotions.
[0579] Step 13:
[0580] The server sends the generated response text to the speech synthesis module.
[0581] Step 14:
[0582] The server's speech synthesis module converts text into speech data.
[0583] Step 15:
[0584] The server sends the audio data to the terminal.
[0585] Step 16:
[0586] The device plays the received audio data to the user.
[0587] Step 17:
[0588] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0589] Step 18:
[0590] The server analyzes the recorded conversation data.
[0591] Step 19:
[0592] The server generates scores and detailed feedback based on analytical data to evaluate the user's responsiveness, skills, and emotional response.
[0593] Step 20:
[0594] The server sends the evaluation results and feedback to the terminal.
[0595] Step 21:
[0596] The device displays evaluation results to the user, providing specific areas for improvement and highlighting achievements.
[0597] (Example 2)
[0598] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0599] Traditional customer service training systems struggled to generate responses that took user emotions into account and to evaluate performance. As a result, training was ineffective, and improvements in actual customer service skills were not achieved. In particular, they were unable to provide appropriate feedback when users were experiencing emotions such as tension or fear.
[0600] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0601] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, and means for recognizing the user's emotions and feeding the results back into generating a natural response. This makes it possible to provide appropriate responses that take the user's emotions into account in real time, as well as to perform detailed performance evaluations and provide feedback.
[0602] "Customer settings" refer to the user's profile information and conversation scenario information, which the system uses to customize responses according to the user's individual needs.
[0603] "Voice input" refers to audio data supplied to the system by the user using a microphone or other voice acquisition device.
[0604] "Methods for converting to text" refers to technologies that analyze voice input and convert its content into text data.
[0605] "Means for generating natural responses" refers to technologies that generate appropriate responses based on converted text data, taking into account the flow, context, and emotions of the user interaction.
[0606] "Means of converting to speech" refers to technology that converts the generated text response back into speech data and provides it to the user as speech.
[0607] "Means of providing voice responses to users" refers to technology that transmits converted voice data to the user's device, allowing the user to listen to the voice.
[0608] "Means for analyzing conversation content" refers to technologies that analyze recorded conversation data and evaluate the content and quality of the conversation.
[0609] "Means for evaluating user performance" refers to technologies that evaluate a user's responsiveness and communication skills based on the content of the conversation and the user's responses at that time.
[0610] "Means of providing feedback on evaluation results to users" refers to technologies that communicate the results of the analyzed performance evaluation to users and provide detailed information on areas for improvement and achievements.
[0611] "Means of recognizing user emotions" refers to technologies that extract and analyze a user's emotional state from audio data and text data.
[0612] An "emotion engine" refers to an engine that identifies emotions from a user's voice or text and feeds the results back to other system components.
[0613] The system according to the present invention is for users to perform interactive training and for which their performance is evaluated and feedback is provided. This system has the following means:
[0614] 1. Register user settings
[0615] First, the user accesses the system using a terminal and enters profile information and scenario information on the settings screen. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario." This information is sent from the terminal to the server, which stores it in a database. Relational databases such as MySQL or PostgreSQL are used as the database.
[0616] 2. Start of role-playing
[0617] When the user presses the "Start Conversation" button on the system, the device records the user's voice and sends it to the server in real time. At this time, the device uses its built-in microphone to capture the voice.
[0618] 3. Text conversion of voice input
[0619] The server receives the audio data sent from the terminal. This data is then converted into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text.
[0620] 4. Generating a response
[0621] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used here include SpaCy and NLTK. The audio data is also input to an emotion engine, which uses the Emotion API or a generative AI model (such as GPT-3) to extract the user's emotions. The emotion engine feeds its results back to the NLP engine, which then generates a natural response considering the customer's settings, the context of the conversation, and the user's emotions. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0622] 5. Providing voice response
[0623] The generated text response is passed to a speech synthesis module. This speech synthesis module converts the text into speech data using the Google Text-to-Speech API or Amazon Polly. The server sends the generated speech data to the device, and the device plays the audio for the user.
[0624] 6. Performance Evaluation and Feedback
[0625] Once the role-playing session ends, the server collects and stores all recorded conversation data. Next, the server analyzes the stored conversation data and evaluates the user's performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression. Based on these evaluation results, the server generates detailed feedback and a score. The evaluation results are sent to the terminal and displayed to the user. For example, it might be displayed in the format: "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[0626] As a concrete example, consider a scenario where a user enters information such as "25-year-old female, call center representative, product return scenario" on the settings screen. The terminal sends this information to the server, which stores it in a database. The user presses the "Start Conversation" button, and the terminal records the audio and sends it to the server. The server converts the audio data into text using the Google Speech-to-Text API. The NLP engine generates a response such as "Excuse me, sir / madam. I will handle this matter." The speech synthesis module uses Amazon Polly to convert this response into audio data, which the terminal plays. After the session ends, the server generates an evaluation result such as "Response Speed: 93 points, Customer Satisfaction Response: 85 points, Pronunciation: 88 points, Emotion Recognition: 80 points" and sends it to the terminal. The terminal displays this to the user.
[0627] Examples of prompt messages include the following:
[0628] "Generate a conversation about a customer being kept waiting, using a bank employee modeled after a man in his 50s, as the scenario for handling a customer complaint."
[0629] "Based on the following conversation, we analyze the user's emotions and provide feedback: 'Sorry to keep you waiting.'"
[0630] This system allows users to conduct customer service training that simulates real-world work situations and improve their skills through detailed feedback.
[0631] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0632] Step 1:
[0633] User settings registration
[0634] Users access the system using a terminal and input profile information and scenario information. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario."
[0635] Input: Profile information and scenario information entered by the user via the terminal.
[0636] The terminal sends the input information to the server. In this process, the terminal sends the information to the server using an HTTP request.
[0637] The server receives the transmitted information and stores it in a database. MySQL or PostgreSQL are used as the database.
[0638] Output: User settings information stored in the database.
[0639] Step 2:
[0640] Start of role-playing
[0641] The user presses the "Start Conversation" button on their device.
[0642] Input: Pressing the "Start Conversation" button.
[0643] The device uses its built-in microphone to record the user's voice in real time. This audio data is sent to the server using WebSocket or HTTP POST requests.
[0644] Output: The recorded audio data is sent to the server.
[0645] Step 3:
[0646] Voice input to text conversion
[0647] The server receives the audio data sent from the terminal.
[0648] Input: Audio data transmitted from the device.
[0649] The server converts the received audio data into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text. These modules analyze the audio waveform and convert its content into text format.
[0650] Output: Audio data converted to text.
[0651] Step 4:
[0652] Response generation
[0653] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used include SpaCy and NLTK.
[0654] Input: Audio data converted to text.
[0655] The server simultaneously inputs voice data into the emotion engine to detect the user's emotions. The emotion engine uses the Emotion API or a generative AI model (e.g., GPT-3).
[0656] The NLP engine generates natural responses based on text data, customer settings, conversation context, and user emotions. For example, if a user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, it will generate a response such as, "You seem a little nervous, is there something you're worried about?"
[0657] Output: Text data of the generated natural response.
[0658] Step 5:
[0659] Providing voice response
[0660] The speech synthesis module receives the generated text response.
[0661] Input: Text data of the generated natural response.
[0662] The speech synthesis module uses the Google Text-to-Speech API or Amazon Polly to convert this text data into speech data.
[0663] The server sends the generated audio data to the terminal.
[0664] The device plays this audio data, and the user can hear the voice response.
[0665] Output: Audio data played by the device.
[0666] Step 6:
[0667] Performance evaluation and feedback
[0668] Once the role-playing session ends, the server collects and saves all conversation data.
[0669] Input: All recorded conversation data.
[0670] The server analyzes stored conversation data to evaluate user performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression.
[0671] The server generates a score and detailed feedback based on these evaluation results. For example, it might be displayed in the format of "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[0672] Output: Performance evaluation results and feedback provided to the user.
[0673] (Application Example 2)
[0674] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0675] Conventional customer service training systems have limitations in how they evaluate user voice input and the appropriateness of responses. They lack the ability to perform real-time sentiment analysis and provide feedback based on that analysis, resulting in insufficient improvement in the performance of customer service staff. Furthermore, they are unable to generate flexible responses based on emotions, making effective training in real-world customer interactions at physical stores difficult.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and saving customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for analyzing emotions, and means for adjusting the response based on emotions. This enables more practical and effective customer service training by analyzing the user's emotions in real time and generating flexible responses based on those emotions.
[0677] "Customer settings" refer to configuration information that includes user profile information and scenario information, and are customized data that the user provides to the system.
[0678] "Voice input" refers to the voice data spoken by the user, which is the input information used by the system to convert into text.
[0679] "Text conversion" is the process of converting voice input into text information using speech recognition technology.
[0680] "Natural response generation" is the process of generating appropriate and natural responses by considering the context of the converted text and conversation, as well as the user's sentiment data.
[0681] "Speech conversion" is the process of converting generated text responses into speech data.
[0682] "Voice response provision" refers to the process of playing back converted voice data to the user.
[0683] "Conversation content analysis" is the process of analyzing recorded conversation data to evaluate user performance.
[0684] "Evaluation result feedback" is a process that provides evaluation results, such as the appropriateness and speed of user responses, based on the analyzed data.
[0685] "Emotion analysis" is the process of extracting and identifying emotions from a user's statements, tone of voice, and other factors.
[0686] "Response adjustment" is the process of appropriately modifying and adjusting the generated responses based on analyzed emotional data.
[0687] A "prompt" is text input into a generative AI model, and it is an instruction to generate a response based on the user's situation and emotions.
[0688] In the system according to this invention, the user first accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information such as "30s, store clerk, customer complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[0689] Next, when the user presses the "Start Conversation" button, the device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Common speech recognition technologies can be used for speech recognition.
[0690] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0691] The generated text response is sent to the server's speech synthesis module and converted into audio data. Common speech synthesis technologies can be used for this conversion. The server then sends this audio data to the terminal, which plays the response back to the user. This allows the user to receive real-time, emotion-based feedback.
[0692] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. The evaluation includes appropriateness of responses, speed, clarity of pronunciation, and emotional changes. From this data, the server generates a score and detailed feedback, and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements.
[0693] This process uses the following specific hardware and software:
[0694] Hardware: Smartphone, microphone
[0695] Software: Speech recognition library (speech_recognition), text-to-speech conversion library (pyttsx3), natural language processing library (transformers)
[0696] As a concrete example, a response is generated using the following prompt:
[0697] The user is 30 years old, working as a shop assistant, handling a customer complaint scenario. The user said: "Excuse me, sorry to have kept you waiting." Detected emotion: "Nervous." Respond appropriately.
[0698] In this way, users can receive practical customer service training and have their performance evaluated from multiple perspectives.
[0699] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0700] Step 1:
[0701] The user enters their customer settings on the connected device. The user uses a smartphone or computer to enter profile information (age, occupation, etc.) and scenario information (e.g., a customer complaint handling scenario). The entered customer settings information is sent from the device to the server. The server receives this information and stores it in a database. This allows the system to prepare a customized training state for each user.
[0702] Step 2:
[0703] When the user presses the "Start Conversation" button, the device records the user's voice through the microphone and sends the audio data to the server in real time. The server uses a speech recognition library to convert this audio data into text. Specifically, it analyzes the audio data and converts phonemes and syllables into text. Audio data is input, and the converted text is output.
[0704] Step 3:
[0705] The server passes the converted text to a natural language processing (NLP) engine and simultaneously to an emotion analysis engine. The emotion analysis engine analyzes the audio data and text to detect the user's emotions. This process identifies the user's emotions (e.g., tension, joy, anger) from the tone of voice and word choice. The input is audio data and text data, and the output is the emotion analysis result.
[0706] Step 4:
[0707] Once the sentiment analysis results are output, they are fed back to the NLP engine. The NLP engine generates an appropriate and natural response, taking into account the user's emotions, profile information, and the context of the conversation. For example, if tension is detected, the NLP engine will generate a response such as, "You seem a little nervous, is there anything you're worried about?" The input is text data and sentiment analysis results, and the output is the generated text response.
[0708] Step 5:
[0709] The server sends the generated text response to a speech synthesis module, which converts it into speech data. The speech synthesis module uses a technique to convert text data into speech data (e.g., text-to-speech software). The generated speech data is sent to the terminal, which plays this speech data and provides the user with a voice response. The input is the generated text response, and the output is the speech data.
[0710] Step 6:
[0711] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation process uses criteria such as the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion. The input is the conversation data, and the output is the user's evaluation result.
[0712] Step 7:
[0713] The server generates a score and detailed feedback based on the evaluation results and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user and provides specific areas for improvement and achievements. For example, the user is provided with detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points." The input is the evaluation result, and the output is the feedback presented to the user.
[0714] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0715] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0716] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0717] [Third Embodiment]
[0718] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0719] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0720] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0721] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0722] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0723] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0724] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0725] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0726] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0727] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0728] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0729] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0730] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user. Specific embodiments are described below.
[0731] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[0732] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for this speech recognition.
[0733] The converted text is then passed to the server's natural language processing (NLP) engine. The NLP engine takes into account the user's settings and the context of the conversation to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," the server will generate a response such as, "We apologize for the wait, customer."
[0734] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0735] Once the role-playing session ends, the server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, providing information on what went well and where there is room for improvement.
[0736] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal.
[0737] This allows users to conduct realistic customer service training and identify specific areas for improvement. This is a specific embodiment of the system according to the present invention.
[0738] The following describes the processing flow.
[0739] Step 1:
[0740] The user accesses the customer settings registration screen from their device.
[0741] Step 2:
[0742] The user enters their profile information and scenario information into the device.
[0743] Step 3:
[0744] The terminal sends the entered customer settings information to the server.
[0745] Step 4:
[0746] The server saves the received information to the database.
[0747] Step 5:
[0748] The user presses the "Start Conversation" button on their device.
[0749] Step 6:
[0750] The device records the user's voice in real time and sends that audio data to the server.
[0751] Step 7:
[0752] The server's speech recognition module converts the speech data into text.
[0753] Step 8:
[0754] The server passes the converted text data to the natural language processing (NLP) engine.
[0755] Step 9:
[0756] The server's NLP engine generates appropriate responses based on customer settings and the context of the conversation.
[0757] Step 10:
[0758] The server sends the generated response text to the speech synthesis module.
[0759] Step 11:
[0760] The server's speech synthesis module converts text into speech data.
[0761] Step 12:
[0762] The server sends the audio data to the terminal.
[0763] Step 13:
[0764] The device plays the received audio data to the user.
[0765] Step 14:
[0766] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0767] Step 15:
[0768] The server analyzes all recorded conversation data.
[0769] Step 16:
[0770] The server generates a score and detailed feedback based on the analysis data to evaluate the user's responsiveness and skills.
[0771] Step 17:
[0772] The server sends the evaluation results and feedback to the terminal.
[0773] Step 18:
[0774] The device displays the evaluation results to the user and provides feedback.
[0775] (Example 1)
[0776] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0777] Existing speech recognition systems have the challenge of not being able to quickly provide appropriate responses based on specific scenarios and profile information set by the user, and not being able to evaluate user performance from multiple perspectives in real time. This invention solves these problems and provides a system that enables users to conduct more realistic and effective training.
[0778] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0779] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to speech, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, a terminal for inputting user setting information, a terminal for recording speech in real time and transmitting it to the server, a response generation means using a natural language processing engine, a response speech conversion means using a speech synthesis module, and means for evaluating the user's performance using multiple indicators and generating detailed feedback. This enables highly accurate response generation based on the user's settings and the provision of multifaceted evaluation results.
[0780] "Customer settings" refers to settings that include user profile information and scenario information.
[0781] "Voice input" refers to voice data that a user produces using a device.
[0782] "Text conversion" refers to the process of converting voice input into text data.
[0783] "Generating natural responses" refers to the process of generating appropriate and natural responses based on converted text data.
[0784] "Speech synthesis" is the process of converting generated text responses into speech data.
[0785] "Providing voice response" refers to playing back the converted voice data to the user.
[0786] "Conversation content analysis" refers to the process of analyzing recorded conversation data to evaluate user performance.
[0787] "Performance evaluation" refers to a comprehensive assessment of factors such as the appropriateness, speed, and clarity of user responses.
[0788] "Feedback" refers to providing users with evaluation results to offer suggestions for improvement and highlight strengths in their training.
[0789] A "terminal" is a device used by a user to access a system and perform tasks such as entering settings or using voice input.
[0790] A "server" is a computer system that centrally stores, processes, and analyzes various types of data.
[0791] A "natural language processing engine" is a software module that analyzes text data and generates appropriate responses.
[0792] A "speech synthesis module" is a software module used to convert text data into speech data.
[0793] "User configuration information" refers to profile information and scenario information entered by the user.
[0794] "Real-time recording" refers to the instantaneous recording of the user's voice.
[0795] The present invention is a system that includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user.
[0796] First, the user accesses the terminal and enters customer settings, including profile information and scenario information. For example, they might enter information such as "male in his 50s, bank employee, customer complaint handling scenario." The terminal sends this information to the server, which stores the received data in a database.
[0797] Next, when the user presses the "Start Conversation" button, the device records the user's voice in real time and sends the recording data to the server. The server uses speech recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text to convert the voice into text data. For example, if the user says, "Excuse me, sorry to have kept you waiting," the speech recognition technology will convert this into text format.
[0798] The converted text data is passed to the server's natural language processing engine. This engine considers the user's scenario settings and conversational context to generate an appropriate response. For example, it might generate a response such as "We apologize for the wait, customer" in response to a user's utterance. This natural language processing engine may utilize existing AI services (e.g., GPT-4).
[0799] The generated text response is sent to a speech synthesis module and converted into audio data. Google Text-to-Speech API or Amazon Polly are used for speech synthesis. The generated audio data is sent from the server to the device, which then plays it back to the user.
[0800] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. Based on this evaluation data, the server calculates a score and generates detailed feedback. The evaluation results are provided to the user via the terminal. For example, an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points" might be displayed.
[0801] As a concrete example, the user enters the following prompt:
[0802] "The user sets the information as 'a man in his 50s, a bank employee, handling a customer complaint scenario,' and begins role-playing. When the user says, 'Could you explain that matter again?', how should the system respond?"
[0803] In response to this prompt, the system generates an appropriate response and provides it in voice.
[0804] Therefore, the present invention provides a systematic system that comprehensively handles everything from inputting user configuration information to speech recognition, response generation, speech synthesis, and performance evaluation, and provides high-quality feedback to the user. This enables the user to receive realistic and effective training.
[0805] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0806] Step 1:
[0807] The user accesses the terminal and enters profile information and scenario information. For example, the user enters "50s male, banker, complaint handling scenario". This input data is sent from the terminal to the server as user settings information. The server saves the received data to a database. Specifically, the server executes and saves the data by executing the SQL query "INSERT INTO customer_settings (age, occupation, scenario) VALUES (50, 'banker', 'complaint handling')".
[0808] Input: User profile information and scenario information
[0809] Output: User settings information stored in the database
[0810] Step 2:
[0811] When the user presses the "Start Conversation" button, the device activates the microphone and records the user's voice in real time. The recorded audio data is sent from the device to the server via WebSocket. The server receives this audio data.
[0812] Input: User voice input
[0813] Output: Audio data sent to the server
[0814] Step 3:
[0815] The server receives the audio data and passes it to the Google Speech-to-Text API. The API converts the audio data into text data and returns that text data to the server. For example, the user's utterance "Excuse me, sorry to keep you waiting" is converted into the text data "Sorry to keep you waiting". The server receives this converted result.
[0816] Input: Audio data
[0817] Output: Text data
[0818] Step 4:
[0819] The server passes text data to a natural language processing (NLP) engine. The NLP engine automatically generates an appropriate response, taking into account the user's scenario setting and the context of the conversation. For example, if the user says, "Excuse me, sorry to have kept you waiting," the NLP engine will generate a response such as, "We apologize for the wait, customer."
[0820] Input: Text data, user settings information
[0821] Output: Generated text response
[0822] Step 5:
[0823] The server passes the generated text response to a text-to-speech module. The text-to-speech module (e.g., Google Text-to-Speech API) converts this into speech data and returns the speech data to the server. The server receives this speech data and sends it to the device. The device plays the received speech data and provides the user with a speech response.
[0824] Input: Generated text response
[0825] Output: Audio data
[0826] Step 6:
[0827] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. Specifically, the server analyzes metrics such as the appropriateness and speed of responses and the clarity of pronunciation, and calculates a score. For example, it might generate an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points." The generated evaluation result is sent to the terminal, which then displays it to the user.
[0828] Input: Conversation data
[0829] Output: Evaluation results
[0830] Through these steps, the system provides users with a consistent experience, from setting up their configuration to engaging in actual role-playing and providing detailed feedback.
[0831] (Application Example 1)
[0832] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0833] In modern customer service, staff need to improve their customer service skills quickly and effectively. However, traditional training methods often involve manual scenario setting and feedback, resulting in inefficiency and a lack of real-time response. Furthermore, the naturalness of voice responses and the reflection of conversational context are insufficient, making it difficult to provide a practical training environment. To solve these problems, there is a need for a system that utilizes more advanced speech recognition and natural response generation technologies.
[0834] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0835] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for initiating a training session based on profile information and scenario information, means for generating prompt sentences using a generative AI model, means for generating a response that reflects the conversation context, and means for generating a natural response based on the generated prompt sentences. This enables real-time voice recognition and response generation, resulting in high effectiveness and immediate results in customer service training.
[0836] "Customer settings" refer to configuration data that includes user profile information and scenario information.
[0837] "Voice input" is the process of acquiring a user's speech in digital format.
[0838] "Text conversion" is the process of converting voice input into written text.
[0839] "Natural response generation" is the process of generating natural conversational responses based on converted text and configured profile information.
[0840] "Speech synthesis" is the process of converting generated text responses into speech data.
[0841] "Voice response" is a method of providing users with synthesized speech responses.
[0842] "Performance evaluation" is a process that analyzes conversation content and evaluates user response and other performance aspects.
[0843] "Feedback" is the process of returning evaluation results to the user, highlighting areas for improvement and strengths.
[0844] "Profile information" refers to personal information such as the user's age, occupation, and job title.
[0845] "Scenario information" refers to scenario data for conversations that are set with specific situations or conditions.
[0846] "Prompt generation" is the process of generating input sentences that are appropriate to the conversational context using a generative AI model.
[0847] "Conversation context" refers to contextual information that includes the current situation and background information of the conversation.
[0848] The system according to the present invention is a voice-response-based role-playing system intended for staff training and improving customer service skills in physical stores.
[0849] First, the user accesses the system and registers profile information (age, occupation, etc.) and scenario information (specific training scenarios) as customer settings. This setting information is entered from a smartphone or other device and sent to the server. The server stores this information in a database.
[0850] Next, the user starts a role-playing session by pressing the "Start Conversation" button. The user's voice input is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the speech to text.
[0851] The converted text is passed to the server's natural language processing (NLP) engine. The NLP engine uses a generative AI model to generate a prompt based on the user's profile information and the context of the conversation. This prompt will look like this:
[0852] Please generate a natural response considering the following context.
[0853] Context: Bank customer complaint handling
[0854] User utterance: I'm sorry, sir / madam, but...
[0855] Response: "We apologize for the wait."
[0856] A natural-sounding response is generated based on this generated prompt. The generated text response is converted into audio data using the Google Text-to-Speech API and sent to the device. The device plays this audio data and provides the user with an audio response.
[0857] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes evaluation criteria such as the appropriateness, speed, and clarity of responses. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, highlighting strengths and areas for improvement.
[0858] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal. This evaluation and feedback allows users to improve their practical customer service skills while simulating real-world scenarios.
[0859] Thus, the present invention achieves real-time performance and high training effectiveness by using advanced speech recognition technology and natural response generation technology.
[0860] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0861] Step 1:
[0862] The user accesses the system and enters profile information and scenario information. Specifically, they enter information such as "female in her 30s, bank employee, customer complaint handling scenario" from a smartphone or other device, and the device sends this information to the server. The server stores the received customer settings information in a database.
[0863] Input: User profile information and scenario information
[0864] Output: Customer settings information stored in the database
[0865] Operation: Sending information from the terminal, receiving information on the server, and saving it to the database.
[0866] Step 2:
[0867] The user presses the "Start Conversation" button to begin a role-playing session. The user's speech is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the received audio data into text.
[0868] Input: User voice input
[0869] Output: User utterance converted to text
[0870] Function: Audio recording on the device, transmission of audio data, speech recognition and text conversion on the server.
[0871] Step 3:
[0872] The server passes the converted text to a natural language processing (NLP) engine, which generates prompt sentences based on profile information and conversational context. Using a generative AI model, it generates prompt sentences such as, "Please generate a natural response considering the following context. Context: Bank customer service, User utterance: I'm sorry, sir / madam, but..., Response: I'm sorry to have kept you waiting."
[0873] Input: User utterances converted to text, profile information, scenario information
[0874] Output: Generated prompt message
[0875] Operation: Uses a natural language processing engine and a generative AI model to generate prompt sentences.
[0876] Step 4:
[0877] Based on the generated prompt, the server produces a natural-sounding response. This response is converted into speech data using the Google Text-to-Speech API. The synthesized response is then sent from the server to the terminal.
[0878] Input: Generated prompt message
[0879] Output: Response converted into audio data
[0880] Function: Generate text responses, convert them to audio data, and send the audio data.
[0881] Step 5:
[0882] The device plays the received audio data and provides the user with an audio response. The user listens to the audio response and then speaks again.
[0883] Input: Response converted into voice data
[0884] Output: Providing voice responses to the user
[0885] Function: Receiving audio data, playing audio
[0886] Step 6:
[0887] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes aspects such as the appropriateness of responses, speed, and clarity of pronunciation. The server generates a score and detailed feedback from the evaluation data and sends the results to the terminal.
[0888] Input: Recorded conversation data
[0889] Output: User ratings and feedback
[0890] Function: Analyzes conversation data, generates evaluation scores, and creates and sends feedback.
[0891] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0892] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing conversation content and evaluating user performance, means for providing feedback of the evaluation results to the user, and an emotion engine for recognizing the user's emotions.
[0893] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in the database.
[0894] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for speech recognition.
[0895] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0896] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[0897] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion.
[0898] The server generates a score and detailed feedback from this data and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements. For example, the server generates evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points" and displays them to the user through the terminal.
[0899] This allows users to receive multifaceted feedback, including on their own emotional management, while undergoing realistic customer service training. The above is a specific embodiment of the system according to the present invention.
[0900] The following describes the processing flow.
[0901] Step 1:
[0902] The user accesses the customer settings registration screen from their device.
[0903] Step 2:
[0904] The user enters their profile information and scenario information into the device.
[0905] Step 3:
[0906] The terminal sends the entered customer settings information to the server.
[0907] Step 4:
[0908] The server saves the received information to the database.
[0909] Step 5:
[0910] The user presses the "Start Conversation" button on their device.
[0911] Step 6:
[0912] The device records the user's voice in real time and sends that audio data to the server.
[0913] Step 7:
[0914] The server's speech recognition module converts the speech data into text.
[0915] Step 8:
[0916] The server passes the converted text data to the natural language processing (NLP) engine.
[0917] Step 9:
[0918] The server's NLP engine analyzes the converted text and generates an appropriate response text.
[0919] Step 10:
[0920] The server passes the voice data to the emotion engine, which then analyzes the user's emotions.
[0921] Step 11:
[0922] The emotion engine feeds back the detected emotion information to the NLP engine.
[0923] Step 12:
[0924] The server's NLP engine generates more natural response text based on the user's emotions.
[0925] Step 13:
[0926] The server sends the generated response text to the speech synthesis module.
[0927] Step 14:
[0928] The server's speech synthesis module converts text into speech data.
[0929] Step 15:
[0930] The server sends the audio data to the terminal.
[0931] Step 16:
[0932] The device plays the received audio data to the user.
[0933] Step 17:
[0934] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[0935] Step 18:
[0936] The server analyzes the recorded conversation data.
[0937] Step 19:
[0938] The server generates scores and detailed feedback based on analytical data to evaluate the user's responsiveness, skills, and emotional response.
[0939] Step 20:
[0940] The server sends the evaluation results and feedback to the terminal.
[0941] Step 21:
[0942] The device displays evaluation results to the user, providing specific areas for improvement and highlighting achievements.
[0943] (Example 2)
[0944] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0945] Traditional customer service training systems struggled to generate responses that took user emotions into account and to evaluate performance. As a result, training was ineffective, and improvements in actual customer service skills were not achieved. In particular, they were unable to provide appropriate feedback when users were experiencing emotions such as tension or fear.
[0946] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0947] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, and means for recognizing the user's emotions and feeding the results back into generating a natural response. This makes it possible to provide appropriate responses that take the user's emotions into account in real time, as well as to perform detailed performance evaluations and provide feedback.
[0948] "Customer settings" refer to the user's profile information and conversation scenario information, which the system uses to customize responses according to the user's individual needs.
[0949] "Voice input" refers to audio data supplied to the system by the user using a microphone or other voice acquisition device.
[0950] "Methods for converting to text" refers to technologies that analyze voice input and convert its content into text data.
[0951] "Means for generating natural responses" refers to technologies that generate appropriate responses based on converted text data, taking into account the flow, context, and emotions of the user interaction.
[0952] "Means of converting to speech" refers to technology that converts the generated text response back into speech data and provides it to the user as speech.
[0953] "Means of providing voice responses to users" refers to technology that transmits converted voice data to the user's device, allowing the user to listen to the voice.
[0954] "Means for analyzing conversation content" refers to technologies that analyze recorded conversation data and evaluate the content and quality of the conversation.
[0955] "Means for evaluating user performance" refers to technologies that evaluate a user's responsiveness and communication skills based on the content of the conversation and the user's responses at that time.
[0956] "Means of providing feedback on evaluation results to users" refers to technologies that communicate the results of the analyzed performance evaluation to users and provide detailed information on areas for improvement and achievements.
[0957] "Means of recognizing user emotions" refers to technologies that extract and analyze a user's emotional state from audio data and text data.
[0958] An "emotion engine" refers to an engine that identifies emotions from a user's voice or text and feeds the results back to other system components.
[0959] The system according to the present invention is for users to perform interactive training and for which their performance is evaluated and feedback is provided. This system has the following means:
[0960] 1. Register user settings
[0961] First, the user accesses the system using a terminal and enters profile information and scenario information on the settings screen. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario." This information is sent from the terminal to the server, which stores it in a database. Relational databases such as MySQL or PostgreSQL are used as the database.
[0962] 2. Start of role-playing
[0963] When the user presses the "Start Conversation" button on the system, the device records the user's voice and sends it to the server in real time. At this time, the device uses its built-in microphone to capture the voice.
[0964] 3. Text conversion of voice input
[0965] The server receives the audio data sent from the terminal. This data is then converted into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text.
[0966] 4. Generating a response
[0967] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used here include SpaCy and NLTK. The audio data is also input to an emotion engine, which uses the Emotion API or a generative AI model (such as GPT-3) to extract the user's emotions. The emotion engine feeds its results back to the NLP engine, which then generates a natural response considering the customer's settings, the context of the conversation, and the user's emotions. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[0968] 5. Providing voice response
[0969] The generated text response is passed to a speech synthesis module. This speech synthesis module converts the text into speech data using the Google Text-to-Speech API or Amazon Polly. The server sends the generated speech data to the device, and the device plays the audio for the user.
[0970] 6. Performance Evaluation and Feedback
[0971] Once the role-playing session ends, the server collects and stores all recorded conversation data. Next, the server analyzes the stored conversation data and evaluates the user's performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression. Based on these evaluation results, the server generates detailed feedback and a score. The evaluation results are sent to the terminal and displayed to the user. For example, it might be displayed in the format: "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[0972] As a concrete example, consider a scenario where a user enters information such as "25-year-old female, call center representative, product return scenario" on the settings screen. The terminal sends this information to the server, which stores it in a database. The user presses the "Start Conversation" button, and the terminal records the audio and sends it to the server. The server converts the audio data into text using the Google Speech-to-Text API. The NLP engine generates a response such as "Excuse me, sir / madam. I will handle this matter." The speech synthesis module uses Amazon Polly to convert this response into audio data, which the terminal plays. After the session ends, the server generates an evaluation result such as "Response Speed: 93 points, Customer Satisfaction Response: 85 points, Pronunciation: 88 points, Emotion Recognition: 80 points" and sends it to the terminal. The terminal displays this to the user.
[0973] Examples of prompt messages include the following:
[0974] "Generate a conversation about a customer being kept waiting, using a bank employee modeled after a man in his 50s, as the scenario for handling a customer complaint."
[0975] "Based on the following conversation, we analyze the user's emotions and provide feedback: 'Sorry to keep you waiting.'"
[0976] This system allows users to conduct customer service training that simulates real-world work situations and improve their skills through detailed feedback.
[0977] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0978] Step 1:
[0979] User settings registration
[0980] Users access the system using a terminal and input profile information and scenario information. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario."
[0981] Input: Profile information and scenario information entered by the user via the terminal.
[0982] The terminal sends the input information to the server. In this process, the terminal sends the information to the server using an HTTP request.
[0983] The server receives the transmitted information and stores it in a database. MySQL or PostgreSQL are used as the database.
[0984] Output: User settings information stored in the database.
[0985] Step 2:
[0986] Start of role-playing
[0987] The user presses the "Start Conversation" button on their device.
[0988] Input: Pressing the "Start Conversation" button.
[0989] The device uses its built-in microphone to record the user's voice in real time. This audio data is sent to the server using WebSocket or HTTP POST requests.
[0990] Output: The recorded audio data is sent to the server.
[0991] Step 3:
[0992] Voice input to text conversion
[0993] The server receives the audio data sent from the terminal.
[0994] Input: Audio data transmitted from the device.
[0995] The server converts the received audio data into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text. These modules analyze the audio waveform and convert its content into text format.
[0996] Output: Audio data converted to text.
[0997] Step 4:
[0998] Response generation
[0999] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used include SpaCy and NLTK.
[1000] Input: Audio data converted to text.
[1001] The server simultaneously inputs voice data into the emotion engine to detect the user's emotions. The emotion engine uses the Emotion API or a generative AI model (e.g., GPT-3).
[1002] The NLP engine generates natural responses based on text data, customer settings, conversation context, and user emotions. For example, if a user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, it will generate a response such as, "You seem a little nervous, is there something you're worried about?"
[1003] Output: Text data of the generated natural response.
[1004] Step 5:
[1005] Providing voice response
[1006] The speech synthesis module receives the generated text response.
[1007] Input: Text data of the generated natural response.
[1008] The speech synthesis module uses the Google Text-to-Speech API or Amazon Polly to convert this text data into speech data.
[1009] The server sends the generated audio data to the terminal.
[1010] The device plays this audio data, and the user can hear the voice response.
[1011] Output: Audio data played by the device.
[1012] Step 6:
[1013] Performance evaluation and feedback
[1014] Once the role-playing session ends, the server collects and saves all conversation data.
[1015] Input: All recorded conversation data.
[1016] The server analyzes stored conversation data to evaluate user performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression.
[1017] The server generates a score and detailed feedback based on these evaluation results. For example, it might be displayed in the format of "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[1018] Output: Performance evaluation results and feedback provided to the user.
[1019] (Application Example 2)
[1020] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1021] Conventional customer service training systems have limitations in how they evaluate user voice input and the appropriateness of responses. They lack the ability to perform real-time sentiment analysis and provide feedback based on that analysis, resulting in insufficient improvement in the performance of customer service staff. Furthermore, they are unable to generate flexible responses based on emotions, making effective training in real-world customer interactions at physical stores difficult.
[1022] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and saving customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for analyzing emotions, and means for adjusting the response based on emotions. This enables more practical and effective customer service training by analyzing the user's emotions in real time and generating flexible responses based on those emotions.
[1023] "Customer settings" refer to configuration information that includes user profile information and scenario information, and are customized data that the user provides to the system.
[1024] "Voice input" refers to the voice data spoken by the user, which is the input information used by the system to convert into text.
[1025] "Text conversion" is the process of converting voice input into text information using speech recognition technology.
[1026] "Natural response generation" is the process of generating appropriate and natural responses by considering the context of the converted text and conversation, as well as the user's sentiment data.
[1027] "Speech conversion" is the process of converting generated text responses into speech data.
[1028] "Voice response provision" refers to the process of playing back converted voice data to the user.
[1029] "Conversation content analysis" is the process of analyzing recorded conversation data to evaluate user performance.
[1030] "Evaluation result feedback" is a process that provides evaluation results, such as the appropriateness and speed of user responses, based on the analyzed data.
[1031] "Emotion analysis" is the process of extracting and identifying emotions from a user's statements, tone of voice, and other factors.
[1032] "Response adjustment" is the process of appropriately modifying and adjusting the generated responses based on analyzed emotional data.
[1033] A "prompt" is text input into a generative AI model, and it is an instruction to generate a response based on the user's situation and emotions.
[1034] In the system according to this invention, the user first accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information such as "30s, store clerk, customer complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[1035] Next, when the user presses the "Start Conversation" button, the device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Common speech recognition technologies can be used for speech recognition.
[1036] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[1037] The generated text response is sent to the server's speech synthesis module and converted into audio data. Common speech synthesis technologies can be used for this conversion. The server then sends this audio data to the terminal, which plays the response back to the user. This allows the user to receive real-time, emotion-based feedback.
[1038] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. The evaluation includes appropriateness of responses, speed, clarity of pronunciation, and emotional changes. From this data, the server generates a score and detailed feedback, and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements.
[1039] This process uses the following specific hardware and software:
[1040] Hardware: Smartphone, microphone
[1041] Software: Speech recognition library (speech_recognition), text-to-speech conversion library (pyttsx3), natural language processing library (transformers)
[1042] As a concrete example, a response is generated using the following prompt:
[1043] The user is 30 years old, working as a shop assistant, handling a customer complaint scenario. The user said: "Excuse me, sorry to have kept you waiting." Detected emotion: "Nervous." Respond appropriately.
[1044] In this way, users can receive practical customer service training and have their performance evaluated from multiple perspectives.
[1045] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1046] Step 1:
[1047] The user enters their customer settings on the connected device. The user uses a smartphone or computer to enter profile information (age, occupation, etc.) and scenario information (e.g., a customer complaint handling scenario). The entered customer settings information is sent from the device to the server. The server receives this information and stores it in a database. This allows the system to prepare a customized training state for each user.
[1048] Step 2:
[1049] When the user presses the "Start Conversation" button, the device records the user's voice through the microphone and sends the audio data to the server in real time. The server uses a speech recognition library to convert this audio data into text. Specifically, it analyzes the audio data and converts phonemes and syllables into text. Audio data is input, and the converted text is output.
[1050] Step 3:
[1051] The server passes the converted text to a natural language processing (NLP) engine and simultaneously to an emotion analysis engine. The emotion analysis engine analyzes the audio data and text to detect the user's emotions. This process identifies the user's emotions (e.g., tension, joy, anger) from the tone of voice and word choice. The input is audio data and text data, and the output is the emotion analysis result.
[1052] Step 4:
[1053] Once the sentiment analysis results are output, they are fed back to the NLP engine. The NLP engine generates an appropriate and natural response, taking into account the user's emotions, profile information, and the context of the conversation. For example, if tension is detected, the NLP engine will generate a response such as, "You seem a little nervous, is there anything you're worried about?" The input is text data and sentiment analysis results, and the output is the generated text response.
[1054] Step 5:
[1055] The server sends the generated text response to a speech synthesis module, which converts it into speech data. The speech synthesis module uses a technique to convert text data into speech data (e.g., text-to-speech software). The generated speech data is sent to the terminal, which plays this speech data and provides the user with a voice response. The input is the generated text response, and the output is the speech data.
[1056] Step 6:
[1057] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation process uses criteria such as the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion. The input is the conversation data, and the output is the user's evaluation result.
[1058] Step 7:
[1059] The server generates a score and detailed feedback based on the evaluation results and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user and provides specific areas for improvement and achievements. For example, the user is provided with detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points." The input is the evaluation result, and the output is the feedback presented to the user.
[1060] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1061] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1062] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1063] [Fourth Embodiment]
[1064] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1065] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1066] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1067] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1068] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1069] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1070] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1071] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1072] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1073] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1074] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1075] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1076] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1077] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user. Specific embodiments are described below.
[1078] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[1079] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for this speech recognition.
[1080] The converted text is then passed to the server's natural language processing (NLP) engine. The NLP engine takes into account the user's settings and the context of the conversation to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," the server will generate a response such as, "We apologize for the wait, customer."
[1081] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[1082] Once the role-playing session ends, the server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, providing information on what went well and where there is room for improvement.
[1083] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal.
[1084] This allows users to conduct realistic customer service training and identify specific areas for improvement. This is a specific embodiment of the system according to the present invention.
[1085] The following describes the processing flow.
[1086] Step 1:
[1087] The user accesses the customer settings registration screen from their device.
[1088] Step 2:
[1089] The user enters their profile information and scenario information into the device.
[1090] Step 3:
[1091] The terminal sends the entered customer settings information to the server.
[1092] Step 4:
[1093] The server saves the received information to the database.
[1094] Step 5:
[1095] The user presses the "Start Conversation" button on their device.
[1096] Step 6:
[1097] The device records the user's voice in real time and sends that audio data to the server.
[1098] Step 7:
[1099] The server's speech recognition module converts the speech data into text.
[1100] Step 8:
[1101] The server passes the converted text data to the natural language processing (NLP) engine.
[1102] Step 9:
[1103] The server's NLP engine generates appropriate responses based on customer settings and the context of the conversation.
[1104] Step 10:
[1105] The server sends the generated response text to the speech synthesis module.
[1106] Step 11:
[1107] The server's speech synthesis module converts text into speech data.
[1108] Step 12:
[1109] The server sends the audio data to the terminal.
[1110] Step 13:
[1111] The device plays the received audio data to the user.
[1112] Step 14:
[1113] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[1114] Step 15:
[1115] The server analyzes all recorded conversation data.
[1116] Step 16:
[1117] The server generates a score and detailed feedback based on the analysis data to evaluate the user's responsiveness and skills.
[1118] Step 17:
[1119] The server sends the evaluation results and feedback to the terminal.
[1120] Step 18:
[1121] The device displays the evaluation results to the user and provides feedback.
[1122] (Example 1)
[1123] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1124] Existing speech recognition systems have the challenge of not being able to quickly provide appropriate responses based on specific scenarios and profile information set by the user, and not being able to evaluate user performance from multiple perspectives in real time. This invention solves these problems and provides a system that enables users to conduct more realistic and effective training.
[1125] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1126] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to speech, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, a terminal for inputting user setting information, a terminal for recording speech in real time and transmitting it to the server, a response generation means using a natural language processing engine, a response speech conversion means using a speech synthesis module, and means for evaluating the user's performance using multiple indicators and generating detailed feedback. This enables highly accurate response generation based on the user's settings and the provision of multifaceted evaluation results.
[1127] "Customer settings" refers to settings that include user profile information and scenario information.
[1128] "Voice input" refers to voice data that a user produces using a device.
[1129] "Text conversion" refers to the process of converting voice input into text data.
[1130] "Generating natural responses" refers to the process of generating appropriate and natural responses based on converted text data.
[1131] "Speech synthesis" is the process of converting generated text responses into speech data.
[1132] "Providing voice response" refers to playing back the converted voice data to the user.
[1133] "Conversation content analysis" refers to the process of analyzing recorded conversation data to evaluate user performance.
[1134] "Performance evaluation" refers to a comprehensive assessment of factors such as the appropriateness, speed, and clarity of user responses.
[1135] "Feedback" refers to providing users with evaluation results to offer suggestions for improvement and highlight strengths in their training.
[1136] A "terminal" is a device used by a user to access a system and perform tasks such as entering settings or using voice input.
[1137] A "server" is a computer system that centrally stores, processes, and analyzes various types of data.
[1138] A "natural language processing engine" is a software module that analyzes text data and generates appropriate responses.
[1139] A "speech synthesis module" is a software module used to convert text data into speech data.
[1140] "User configuration information" refers to profile information and scenario information entered by the user.
[1141] "Real-time recording" refers to the instantaneous recording of the user's voice.
[1142] The present invention is a system that includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing the conversation content and evaluating the user's performance, and means for providing feedback of the evaluation results to the user.
[1143] First, the user accesses the terminal and enters customer settings, including profile information and scenario information. For example, they might enter information such as "male in his 50s, bank employee, customer complaint handling scenario." The terminal sends this information to the server, which stores the received data in a database.
[1144] Next, when the user presses the "Start Conversation" button, the device records the user's voice in real time and sends the recording data to the server. The server uses speech recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text to convert the voice into text data. For example, if the user says, "Excuse me, sorry to have kept you waiting," the speech recognition technology will convert this into text format.
[1145] The converted text data is passed to the server's natural language processing engine. This engine considers the user's scenario settings and conversational context to generate an appropriate response. For example, it might generate a response such as "We apologize for the wait, customer" in response to a user's utterance. This natural language processing engine may utilize existing AI services (e.g., GPT-4).
[1146] The generated text response is sent to a speech synthesis module and converted into audio data. Google Text-to-Speech API or Amazon Polly are used for speech synthesis. The generated audio data is sent from the server to the device, which then plays it back to the user.
[1147] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. This evaluation includes aspects such as the appropriateness and speed of responses and the clarity of pronunciation. Based on this evaluation data, the server calculates a score and generates detailed feedback. The evaluation results are provided to the user via the terminal. For example, an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points" might be displayed.
[1148] As a concrete example, the user enters the following prompt:
[1149] "The user sets the information as 'a man in his 50s, a bank employee, handling a customer complaint scenario,' and begins role-playing. When the user says, 'Could you explain that matter again?', how should the system respond?"
[1150] In response to this prompt, the system generates an appropriate response and provides it in voice.
[1151] Therefore, the present invention provides a systematic system that comprehensively handles everything from inputting user configuration information to speech recognition, response generation, speech synthesis, and performance evaluation, and provides high-quality feedback to the user. This enables the user to receive realistic and effective training.
[1152] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1153] Step 1:
[1154] The user accesses the terminal and enters profile information and scenario information. For example, the user enters "50s male, banker, complaint handling scenario". This input data is sent from the terminal to the server as user settings information. The server saves the received data to a database. Specifically, the server executes and saves the data by executing the SQL query "INSERT INTO customer_settings (age, occupation, scenario) VALUES (50, 'banker', 'complaint handling')".
[1155] Input: User profile information and scenario information
[1156] Output: User settings information stored in the database
[1157] Step 2:
[1158] When the user presses the "Start Conversation" button, the device activates the microphone and records the user's voice in real time. The recorded audio data is sent from the device to the server via WebSocket. The server receives this audio data.
[1159] Input: User voice input
[1160] Output: Audio data sent to the server
[1161] Step 3:
[1162] The server receives the audio data and passes it to the Google Speech-to-Text API. The API converts the audio data into text data and returns that text data to the server. For example, the user's utterance "Excuse me, sorry to keep you waiting" is converted into the text data "Sorry to keep you waiting". The server receives this converted result.
[1163] Input: Audio data
[1164] Output: Text data
[1165] Step 4:
[1166] The server passes text data to a natural language processing (NLP) engine. The NLP engine automatically generates an appropriate response, taking into account the user's scenario setting and the context of the conversation. For example, if the user says, "Excuse me, sorry to have kept you waiting," the NLP engine will generate a response such as, "We apologize for the wait, customer."
[1167] Input: Text data, user settings information
[1168] Output: Generated text response
[1169] Step 5:
[1170] The server passes the generated text response to a text-to-speech module. The text-to-speech module (e.g., Google Text-to-Speech API) converts this into speech data and returns the speech data to the server. The server receives this speech data and sends it to the device. The device plays the received speech data and provides the user with a speech response.
[1171] Input: Generated text response
[1172] Output: Audio data
[1173] Step 6:
[1174] Once the role-playing session ends, the server analyzes all conversation data and evaluates the user's performance. Specifically, the server analyzes metrics such as the appropriateness and speed of responses and the clarity of pronunciation, and calculates a score. For example, it might generate an evaluation result such as "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points." The generated evaluation result is sent to the terminal, which then displays it to the user.
[1175] Input: Conversation data
[1176] Output: Evaluation results
[1177] Through these steps, the system provides users with a consistent experience, from setting up their configuration to engaging in actual role-playing and providing detailed feedback.
[1178] (Application Example 1)
[1179] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1180] In modern customer service, staff need to improve their customer service skills quickly and effectively. However, traditional training methods often involve manual scenario setting and feedback, resulting in inefficiency and a lack of real-time response. Furthermore, the naturalness of voice responses and the reflection of conversational context are insufficient, making it difficult to provide a practical training environment. To solve these problems, there is a need for a system that utilizes more advanced speech recognition and natural response generation technologies.
[1181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1182] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for initiating a training session based on profile information and scenario information, means for generating prompt sentences using a generative AI model, means for generating a response that reflects the conversation context, and means for generating a natural response based on the generated prompt sentences. This enables real-time voice recognition and response generation, resulting in high effectiveness and immediate results in customer service training.
[1183] "Customer settings" refer to configuration data that includes user profile information and scenario information.
[1184] "Voice input" is the process of acquiring a user's speech in digital format.
[1185] "Text conversion" is the process of converting voice input into written text.
[1186] "Natural response generation" is the process of generating natural conversational responses based on converted text and configured profile information.
[1187] "Speech synthesis" is the process of converting generated text responses into speech data.
[1188] "Voice response" is a method of providing users with synthesized speech responses.
[1189] "Performance evaluation" is a process that analyzes conversation content and evaluates user response and other performance aspects.
[1190] "Feedback" is the process of returning evaluation results to the user, highlighting areas for improvement and strengths.
[1191] "Profile information" refers to personal information such as the user's age, occupation, and job title.
[1192] "Scenario information" refers to scenario data for conversations that are set with specific situations or conditions.
[1193] "Prompt generation" is the process of generating input sentences that are appropriate to the conversational context using a generative AI model.
[1194] "Conversation context" refers to contextual information that includes the current situation and background information of the conversation.
[1195] The system according to the present invention is a voice-response-based role-playing system intended for staff training and improving customer service skills in physical stores.
[1196] First, the user accesses the system and registers profile information (age, occupation, etc.) and scenario information (specific training scenarios) as customer settings. This setting information is entered from a smartphone or other device and sent to the server. The server stores this information in a database.
[1197] Next, the user starts a role-playing session by pressing the "Start Conversation" button. The user's voice input is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the speech to text.
[1198] The converted text is passed to the server's natural language processing (NLP) engine. The NLP engine uses a generative AI model to generate a prompt based on the user's profile information and the context of the conversation. This prompt will look like this:
[1199] Please generate a natural response considering the following context.
[1200] Context: Bank customer complaint handling
[1201] User utterance: I'm sorry, sir / madam, but...
[1202] Response: "We apologize for the wait."
[1203] A natural-sounding response is generated based on this generated prompt. The generated text response is converted into audio data using the Google Text-to-Speech API and sent to the device. The device plays this audio data and provides the user with an audio response.
[1204] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes evaluation criteria such as the appropriateness, speed, and clarity of responses. From this data, the server generates a score and detailed feedback, and sends the results to the terminal. The terminal displays the evaluation results to the user, highlighting strengths and areas for improvement.
[1205] For example, the server generates detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points" and provides them to the user via the terminal. This evaluation and feedback allows users to improve their practical customer service skills while simulating real-world scenarios.
[1206] Thus, the present invention achieves real-time performance and high training effectiveness by using advanced speech recognition technology and natural response generation technology.
[1207] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1208] Step 1:
[1209] The user accesses the system and enters profile information and scenario information. Specifically, they enter information such as "female in her 30s, bank employee, customer complaint handling scenario" from a smartphone or other device, and the device sends this information to the server. The server stores the received customer settings information in a database.
[1210] Input: User profile information and scenario information
[1211] Output: Customer settings information stored in the database
[1212] Operation: Sending information from the terminal, receiving information on the server, and saving it to the database.
[1213] Step 2:
[1214] The user presses the "Start Conversation" button to begin a role-playing session. The user's speech is recorded on a device with a microphone and sent to the server in real time. The server uses the Google Cloud Speech-to-Text API to convert the received audio data into text.
[1215] Input: User voice input
[1216] Output: User utterance converted to text
[1217] Function: Audio recording on the device, transmission of audio data, speech recognition and text conversion on the server.
[1218] Step 3:
[1219] The server passes the converted text to a natural language processing (NLP) engine, which generates prompt sentences based on profile information and conversational context. Using a generative AI model, it generates prompt sentences such as, "Please generate a natural response considering the following context. Context: Bank customer service, User utterance: I'm sorry, sir / madam, but..., Response: I'm sorry to have kept you waiting."
[1220] Input: User utterances converted to text, profile information, scenario information
[1221] Output: Generated prompt message
[1222] Operation: Uses a natural language processing engine and a generative AI model to generate prompt sentences.
[1223] Step 4:
[1224] Based on the generated prompt, the server produces a natural-sounding response. This response is converted into speech data using the Google Text-to-Speech API. The synthesized response is then sent from the server to the terminal.
[1225] Input: Generated prompt message
[1226] Output: Response converted into audio data
[1227] Function: Generate text responses, convert them to audio data, and send the audio data.
[1228] Step 5:
[1229] The device plays the received audio data and provides the user with an audio response. The user listens to the audio response and then speaks again.
[1230] Input: Response converted into voice data
[1231] Output: Providing voice responses to the user
[1232] Function: Receiving audio data, playing audio
[1233] Step 6:
[1234] Once the role-playing session ends, the server analyzes all recorded conversation data. This analysis includes aspects such as the appropriateness of responses, speed, and clarity of pronunciation. The server generates a score and detailed feedback from the evaluation data and sends the results to the terminal.
[1235] Input: Recorded conversation data
[1236] Output: User ratings and feedback
[1237] Function: Analyzes conversation data, generates evaluation scores, and creates and sends feedback.
[1238] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1239] The system according to the present invention includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing voice responses to the user, means for analyzing conversation content and evaluating user performance, means for providing feedback of the evaluation results to the user, and an emotion engine for recognizing the user's emotions.
[1240] First, the user accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information, such as "male in his 50s, bank employee, complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in the database.
[1241] Next, the user presses the "Start Conversation" button to begin role-playing. The device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Existing technologies such as the Google Speech-to-Text API or IBM Watson Speech to Text can be used for speech recognition.
[1242] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[1243] The generated text response is sent to a speech synthesis module and converted into audio data. Speech synthesis can utilize APIs such as Google Text-to-Speech or Amazon Polly. The server sends this audio data to the device, which then plays the response back to the user.
[1244] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation includes the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion.
[1245] The server generates a score and detailed feedback from this data and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements. For example, the server generates evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points" and displays them to the user through the terminal.
[1246] This allows users to receive multifaceted feedback, including on their own emotional management, while undergoing realistic customer service training. The above is a specific embodiment of the system according to the present invention.
[1247] The following describes the processing flow.
[1248] Step 1:
[1249] The user accesses the customer settings registration screen from their device.
[1250] Step 2:
[1251] The user enters their profile information and scenario information into the device.
[1252] Step 3:
[1253] The terminal sends the entered customer settings information to the server.
[1254] Step 4:
[1255] The server saves the received information to the database.
[1256] Step 5:
[1257] The user presses the "Start Conversation" button on their device.
[1258] Step 6:
[1259] The device records the user's voice in real time and sends that audio data to the server.
[1260] Step 7:
[1261] The server's speech recognition module converts the speech data into text.
[1262] Step 8:
[1263] The server passes the converted text data to the natural language processing (NLP) engine.
[1264] Step 9:
[1265] The server's NLP engine analyzes the converted text and generates an appropriate response text.
[1266] Step 10:
[1267] The server passes the voice data to the emotion engine, which then analyzes the user's emotions.
[1268] Step 11:
[1269] The emotion engine feeds back the detected emotion information to the NLP engine.
[1270] Step 12:
[1271] The server's NLP engine generates more natural response text based on the user's emotions.
[1272] Step 13:
[1273] The server sends the generated response text to the speech synthesis module.
[1274] Step 14:
[1275] The server's speech synthesis module converts text into speech data.
[1276] Step 15:
[1277] The server sends the audio data to the terminal.
[1278] Step 16:
[1279] The device plays the received audio data to the user.
[1280] Step 17:
[1281] When a role-playing session ends, it will either end when the user presses the "End Conversation" button or automatically after a certain period of time has elapsed.
[1282] Step 18:
[1283] The server analyzes the recorded conversation data.
[1284] Step 19:
[1285] The server generates scores and detailed feedback based on analytical data to evaluate the user's responsiveness, skills, and emotional response.
[1286] Step 20:
[1287] The server sends the evaluation results and feedback to the terminal.
[1288] Step 21:
[1289] The device displays evaluation results to the user, providing specific areas for improvement and highlighting achievements.
[1290] (Example 2)
[1291] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1292] Traditional customer service training systems struggled to generate responses that took user emotions into account and to evaluate performance. As a result, training was ineffective, and improvements in actual customer service skills were not achieved. In particular, they were unable to provide appropriate feedback when users were experiencing emotions such as tension or fear.
[1293] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1294] In this invention, the server includes means for receiving and storing customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback on the evaluation results to the user, and means for recognizing the user's emotions and feeding the results back into generating a natural response. This makes it possible to provide appropriate responses that take the user's emotions into account in real time, as well as to perform detailed performance evaluations and provide feedback.
[1295] "Customer settings" refer to the user's profile information and conversation scenario information, which the system uses to customize responses according to the user's individual needs.
[1296] "Voice input" refers to audio data supplied to the system by the user using a microphone or other voice acquisition device.
[1297] "Methods for converting to text" refers to technologies that analyze voice input and convert its content into text data.
[1298] "Means for generating natural responses" refers to technologies that generate appropriate responses based on converted text data, taking into account the flow, context, and emotions of the user interaction.
[1299] "Means of converting to speech" refers to technology that converts the generated text response back into speech data and provides it to the user as speech.
[1300] "Means of providing voice responses to users" refers to technology that transmits converted voice data to the user's device, allowing the user to listen to the voice.
[1301] "Means for analyzing conversation content" refers to technologies that analyze recorded conversation data and evaluate the content and quality of the conversation.
[1302] "Means for evaluating user performance" refers to technologies that evaluate a user's responsiveness and communication skills based on the content of the conversation and the user's responses at that time.
[1303] "Means of providing feedback on evaluation results to users" refers to technologies that communicate the results of the analyzed performance evaluation to users and provide detailed information on areas for improvement and achievements.
[1304] "Means of recognizing user emotions" refers to technologies that extract and analyze a user's emotional state from audio data and text data.
[1305] An "emotion engine" refers to an engine that identifies emotions from a user's voice or text and feeds the results back to other system components.
[1306] The system according to the present invention is for users to perform interactive training and for which their performance is evaluated and feedback is provided. This system has the following means:
[1307] 1. Register user settings
[1308] First, the user accesses the system using a terminal and enters profile information and scenario information on the settings screen. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario." This information is sent from the terminal to the server, which stores it in a database. Relational databases such as MySQL or PostgreSQL are used as the database.
[1309] 2. Start of role-playing
[1310] When the user presses the "Start Conversation" button on the system, the device records the user's voice and sends it to the server in real time. At this time, the device uses its built-in microphone to capture the voice.
[1311] 3. Text conversion of voice input
[1312] The server receives the audio data sent from the terminal. This data is then converted into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text.
[1313] 4. Generating a response
[1314] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used here include SpaCy and NLTK. The audio data is also input to an emotion engine, which uses the Emotion API or a generative AI model (such as GPT-3) to extract the user's emotions. The emotion engine feeds its results back to the NLP engine, which then generates a natural response considering the customer's settings, the context of the conversation, and the user's emotions. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[1315] 5. Providing voice response
[1316] The generated text response is passed to a speech synthesis module. This speech synthesis module converts the text into speech data using the Google Text-to-Speech API or Amazon Polly. The server sends the generated speech data to the device, and the device plays the audio for the user.
[1317] 6. Performance Evaluation and Feedback
[1318] Once the role-playing session ends, the server collects and stores all recorded conversation data. Next, the server analyzes the stored conversation data and evaluates the user's performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression. Based on these evaluation results, the server generates detailed feedback and a score. The evaluation results are sent to the terminal and displayed to the user. For example, it might be displayed in the format: "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[1319] As a concrete example, consider a scenario where a user enters information such as "25-year-old female, call center representative, product return scenario" on the settings screen. The terminal sends this information to the server, which stores it in a database. The user presses the "Start Conversation" button, and the terminal records the audio and sends it to the server. The server converts the audio data into text using the Google Speech-to-Text API. The NLP engine generates a response such as "Excuse me, sir / madam. I will handle this matter." The speech synthesis module uses Amazon Polly to convert this response into audio data, which the terminal plays. After the session ends, the server generates an evaluation result such as "Response Speed: 93 points, Customer Satisfaction Response: 85 points, Pronunciation: 88 points, Emotion Recognition: 80 points" and sends it to the terminal. The terminal displays this to the user.
[1320] Examples of prompt messages include the following:
[1321] "Generate a conversation about a customer being kept waiting, using a bank employee modeled after a man in his 50s, as the scenario for handling a customer complaint."
[1322] "Based on the following conversation, we analyze the user's emotions and provide feedback: 'Sorry to keep you waiting.'"
[1323] This system allows users to conduct customer service training that simulates real-world work situations and improve their skills through detailed feedback.
[1324] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1325] Step 1:
[1326] User settings registration
[1327] Users access the system using a terminal and input profile information and scenario information. For example, this information may include "male in his 50s, bank employee, customer complaint handling scenario."
[1328] Input: Profile information and scenario information entered by the user via the terminal.
[1329] The terminal sends the input information to the server. In this process, the terminal sends the information to the server using an HTTP request.
[1330] The server receives the transmitted information and stores it in a database. MySQL or PostgreSQL are used as the database.
[1331] Output: User settings information stored in the database.
[1332] Step 2:
[1333] Start of role-playing
[1334] The user presses the "Start Conversation" button on their device.
[1335] Input: Pressing the "Start Conversation" button.
[1336] The device uses its built-in microphone to record the user's voice in real time. This audio data is sent to the server using WebSocket or HTTP POST requests.
[1337] Output: The recorded audio data is sent to the server.
[1338] Step 3:
[1339] Voice input to text conversion
[1340] The server receives the audio data sent from the terminal.
[1341] Input: Audio data transmitted from the device.
[1342] The server converts the received audio data into text using speech recognition modules such as the Google Speech-to-Text API or IBM Watson Speech to Text. These modules analyze the audio waveform and convert its content into text format.
[1343] Output: Audio data converted to text.
[1344] Step 4:
[1345] Response generation
[1346] The server passes the converted text data to a natural language processing (NLP) engine. Examples of NLP engines used include SpaCy and NLTK.
[1347] Input: Audio data converted to text.
[1348] The server simultaneously inputs voice data into the emotion engine to detect the user's emotions. The emotion engine uses the Emotion API or a generative AI model (e.g., GPT-3).
[1349] The NLP engine generates natural responses based on text data, customer settings, conversation context, and user emotions. For example, if a user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension, it will generate a response such as, "You seem a little nervous, is there something you're worried about?"
[1350] Output: Text data of the generated natural response.
[1351] Step 5:
[1352] Providing voice response
[1353] The speech synthesis module receives the generated text response.
[1354] Input: Text data of the generated natural response.
[1355] The speech synthesis module uses the Google Text-to-Speech API or Amazon Polly to convert this text data into speech data.
[1356] The server sends the generated audio data to the terminal.
[1357] The device plays this audio data, and the user can hear the voice response.
[1358] Output: Audio data played by the device.
[1359] Step 6:
[1360] Performance evaluation and feedback
[1361] Once the role-playing session ends, the server collects and saves all conversation data.
[1362] Input: All recorded conversation data.
[1363] The server analyzes stored conversation data to evaluate user performance. Evaluation criteria include appropriateness of responses, speed, clarity of pronunciation, and emotional expression.
[1364] The server generates a score and detailed feedback based on these evaluation results. For example, it might be displayed in the format of "Response Speed: 95 points, Customer Satisfaction Response: 88 points, Pronunciation: 90 points, Emotion Recognition: 85 points."
[1365] Output: Performance evaluation results and feedback provided to the user.
[1366] (Application Example 2)
[1367] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1368] Conventional customer service training systems have limitations in how they evaluate user voice input and the appropriateness of responses. They lack the ability to perform real-time sentiment analysis and provide feedback based on that analysis, resulting in insufficient improvement in the performance of customer service staff. Furthermore, they are unable to generate flexible responses based on emotions, making effective training in real-world customer interactions at physical stores difficult.
[1369] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving and saving customer settings, means for recognizing voice input and converting it to text, means for generating a natural response based on the converted text, means for converting the generated response to voice, means for providing the voice response to the user, means for analyzing the conversation content and evaluating the user's performance, means for providing feedback of the evaluation results to the user, means for analyzing emotions, and means for adjusting the response based on emotions. This enables more practical and effective customer service training by analyzing the user's emotions in real time and generating flexible responses based on those emotions.
[1370] "Customer settings" refer to configuration information that includes user profile information and scenario information, and are customized data that the user provides to the system.
[1371] "Voice input" refers to the voice data spoken by the user, which is the input information used by the system to convert into text.
[1372] "Text conversion" is the process of converting voice input into text information using speech recognition technology.
[1373] "Natural response generation" is the process of generating appropriate and natural responses by considering the context of the converted text and conversation, as well as the user's sentiment data.
[1374] "Speech conversion" is the process of converting generated text responses into speech data.
[1375] "Voice response provision" refers to the process of playing back converted voice data to the user.
[1376] "Conversation content analysis" is the process of analyzing recorded conversation data to evaluate user performance.
[1377] "Evaluation result feedback" is a process that provides evaluation results, such as the appropriateness and speed of user responses, based on the analyzed data.
[1378] "Emotion analysis" is the process of extracting and identifying emotions from a user's statements, tone of voice, and other factors.
[1379] "Response adjustment" is the process of appropriately modifying and adjusting the generated responses based on analyzed emotional data.
[1380] A "prompt" is text input into a generative AI model, and it is an instruction to generate a response based on the user's situation and emotions.
[1381] In the system according to this invention, the user first accesses the system and registers their customer settings. The user uses a terminal to input profile information and scenario information such as "30s, store clerk, customer complaint handling scenario." The terminal sends this information to the server, which receives the information and stores it in a database.
[1382] Next, when the user presses the "Start Conversation" button, the device records the user's voice and sends it to the server in real time. The server uses a speech recognition module to convert this voice into text. Common speech recognition technologies can be used for speech recognition.
[1383] The converted text is passed to the server's natural language processing (NLP) engine, which in turn is input to the emotion engine. The emotion engine extracts the user's emotions from the audio data and feeds the results back to the NLP engine. The NLP engine considers the customer's settings, the context of the conversation, and the user's emotions to generate an appropriate and natural response. For example, if the user says, "Excuse me, sorry to have kept you waiting," and the emotion engine detects tension in that statement, the server will generate a response such as, "You seem a little nervous, is there anything you're worried about?"
[1384] The generated text response is sent to the server's speech synthesis module and converted into audio data. Common speech synthesis technologies can be used for this conversion. The server then sends this audio data to the terminal, which plays the response back to the user. This allows the user to receive real-time, emotion-based feedback.
[1385] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. The evaluation includes appropriateness of responses, speed, clarity of pronunciation, and emotional changes. From this data, the server generates a score and detailed feedback, and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user, providing specific areas for improvement and achievements.
[1386] This process uses the following specific hardware and software:
[1387] Hardware: Smartphone, microphone
[1388] Software: Speech recognition library (speech_recognition), text-to-speech conversion library (pyttsx3), natural language processing library (transformers)
[1389] As a concrete example, a response is generated using the following prompt:
[1390] The user is 30 years old, working as a shop assistant, handling a customer complaint scenario. The user said: "Excuse me, sorry to have kept you waiting." Detected emotion: "Nervous." Respond appropriately.
[1391] In this way, users can receive practical customer service training and have their performance evaluated from multiple perspectives.
[1392] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1393] Step 1:
[1394] The user enters their customer settings on the connected device. The user uses a smartphone or computer to enter profile information (age, occupation, etc.) and scenario information (e.g., a customer complaint handling scenario). The entered customer settings information is sent from the device to the server. The server receives this information and stores it in a database. This allows the system to prepare a customized training state for each user.
[1395] Step 2:
[1396] When the user presses the "Start Conversation" button, the device records the user's voice through the microphone and sends the audio data to the server in real time. The server uses a speech recognition library to convert this audio data into text. Specifically, it analyzes the audio data and converts phonemes and syllables into text. Audio data is input, and the converted text is output.
[1397] Step 3:
[1398] The server passes the converted text to a natural language processing (NLP) engine and simultaneously to an emotion analysis engine. The emotion analysis engine analyzes the audio data and text to detect the user's emotions. This process identifies the user's emotions (e.g., tension, joy, anger) from the tone of voice and word choice. The input is audio data and text data, and the output is the emotion analysis result.
[1399] Step 4:
[1400] Once the sentiment analysis results are output, they are fed back to the NLP engine. The NLP engine generates an appropriate and natural response, taking into account the user's emotions, profile information, and the context of the conversation. For example, if tension is detected, the NLP engine will generate a response such as, "You seem a little nervous, is there anything you're worried about?" The input is text data and sentiment analysis results, and the output is the generated text response.
[1401] Step 5:
[1402] The server sends the generated text response to a speech synthesis module, which converts it into speech data. The speech synthesis module uses a technique to convert text data into speech data (e.g., text-to-speech software). The generated speech data is sent to the terminal, which plays this speech data and provides the user with a voice response. The input is the generated text response, and the output is the speech data.
[1403] Step 6:
[1404] When a role-playing session ends, the user can press the "End Conversation" button, or the session will automatically end after a certain period of time. The server analyzes all recorded conversation data and evaluates the user's performance. This evaluation process uses criteria such as the appropriateness and speed of responses, clarity of pronunciation, and changes in emotion. The input is the conversation data, and the output is the user's evaluation result.
[1405] Step 7:
[1406] The server generates a score and detailed feedback based on the evaluation results and sends the evaluation results and feedback to the terminal. The terminal displays the evaluation results to the user and provides specific areas for improvement and achievements. For example, the user is provided with detailed evaluation results such as "Response speed: 95 points, Customer satisfaction response: 88 points, Pronunciation: 90 points, Emotion recognition: 85 points." The input is the evaluation result, and the output is the feedback presented to the user.
[1407] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1410] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1411] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1412] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1413] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1414] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1415] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1416] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1417] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1418] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1419] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1420] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1421] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1422] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1423] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1424] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1425] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1426] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1427] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1428] The following is further disclosed regarding the embodiments described above.
[1429] (Claim 1)
[1430] A means of receiving and saving customer settings,
[1431] A means of recognizing voice input and converting it to text,
[1432] A means for generating a natural response based on the converted text,
[1433] A means for converting the generated response into speech,
[1434] A means of providing voice response to the user,
[1435] A means of analyzing conversation content and evaluating user performance,
[1436] A means of providing feedback on evaluation results to users,
[1437] A system that includes this.
[1438] (Claim 2)
[1439] The system according to claim 1, wherein the customer settings include user profile information and scenario information.
[1440] (Claim 3)
[1441] The system according to claim 1, which recognizes voice input in real time and generates a response in real time.
[1442] "Example 1"
[1443] (Claim 1)
[1444] A means of receiving and saving customer settings,
[1445] A means of recognizing voice input and converting it to text,
[1446] A means for generating a natural response based on the converted text,
[1447] A means for converting the generated response into speech,
[1448] A means of providing voice response to the user,
[1449] A means of analyzing conversation content and evaluating user performance,
[1450] A means of providing feedback on evaluation results to users,
[1451] A terminal for entering user settings information,
[1452] A terminal that records audio in real time and sends it to a server,
[1453] A response generation means using a natural language processing engine,
[1454] A response speech conversion means using a speech synthesis module,
[1455] A means of evaluating user performance using multiple metrics and generating detailed feedback,
[1456] A system that includes this.
[1457] (Claim 2)
[1458] The system according to claim 1, wherein the customer settings include user profile information and scenario information.
[1459] (Claim 3)
[1460] The system according to claim 1, which recognizes voice input in real time and generates a response in real time.
[1461] "Application Example 1"
[1462] (Claim 1)
[1463] A means of receiving and saving customer settings,
[1464] A means of recognizing voice input and converting it to text,
[1465] A means for generating a natural response based on the converted text,
[1466] A means for converting the generated response into speech,
[1467] A means of providing voice response to the user,
[1468] A means of analyzing conversation content and evaluating user performance,
[1469] A means of providing feedback on evaluation results to users,
[1470] A means of initiating a training session based on profile information and scenario information,
[1471] A means for generating prompt sentences using a generative AI model,
[1472] A means for generating a response that reflects the conversational context,
[1473] A means for generating a natural response based on the generated prompt sentence,
[1474] A system that includes this.
[1475] (Claim 2)
[1476] The system according to claim 1, wherein the customer settings include user profile information and scenario information.
[1477] (Claim 3)
[1478] The system according to claim 1, which recognizes voice input in real time and generates a response in real time.
[1479] "Example 2 of combining an emotion engine"
[1480] (Claim 1)
[1481] A means of receiving and saving customer settings,
[1482] A means of recognizing voice input and converting it to text,
[1483] A means for generating a natural response based on the converted text,
[1484] A means for converting the generated response into speech,
[1485] A means of providing voice response to the user,
[1486] A means of analyzing conversation content and evaluating user performance,
[1487] A means of providing feedback on evaluation results to users,
[1488] A means of recognizing user emotions and feeding the results back into generating natural responses,
[1489] A system that includes this.
[1490] (Claim 2)
[1491] The system according to claim 1, wherein the customer settings include user profile information and scenario information.
[1492] (Claim 3)
[1493] The system according to claim 1, which recognizes voice input in real time and generates a response in real time based on the converted text and the user's emotions.
[1494] "Application example 2 when combining with an emotional engine"
[1495] (Claim 1)
[1496] A means of receiving and saving customer settings,
[1497] A means of recognizing voice input and converting it to text,
[1498] A means for generating a natural response based on the converted text,
[1499] A means for converting the generated response into speech,
[1500] A means of providing voice response to the user,
[1501] A means of analyzing conversation content and evaluating user performance,
[1502] A means of providing feedback on evaluation results to users,
[1503] A means of analyzing emotions,
[1504] Means of adjusting responses based on emotions,
[1505] A system that includes this.
[1506] (Claim 2)
[1507] The system according to claim 1, wherein the customer settings include user profile information and scenario information.
[1508] (Claim 3)
[1509] The system according to claim 1, which recognizes voice input in real time and generates a response in real time.
[1510] (Claim 4)
[1511] The system according to claim 1, which includes a prompt statement that generates a response based on emotion. [Explanation of symbols]
[1512] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving and saving customer settings, A means of recognizing voice input and converting it to text, A means for generating a natural response based on the converted text, A means for converting the generated response into speech, A means of providing voice response to the user, A means of analyzing conversation content and evaluating user performance, A means of providing feedback on evaluation results to users, A system that includes this.
2. The system according to claim 1, wherein the customer settings include user profile information and scenario information.
3. The system according to claim 1, which recognizes voice input in real time and generates a response in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A