System
The system addresses the lack of real-time feedback and progress tracking in communication training by authenticating users, determining skill levels, providing tailored scenarios, and analyzing speech in real-time, resulting in effective skill improvement.
Patent Information
- Application Number
- JP2024116410
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Existing communication skill training systems lack real-time feedback and effective progress tracking, making it difficult for users to improve their communication skills sustainably.
A system that authenticates users, determines their skill level, provides tailored training scenarios, analyzes speech in real-time, and stores progress data for self-evaluation, enabling real-time feedback and long-term growth tracking.
Provides practical and effective communication skill development with real-time feedback and long-term progress tracking, allowing users to improve their communication skills effectively.
Smart Images

Figure 2026014936000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the lack of effective communication skills is a major problem for many people. Even as face-to-face communication declines, there are still many situations in business and everyday life that require high communication skills. However, opportunities to hone skills while receiving appropriate feedback in real time are limited. Furthermore, self-evaluation methods are inadequate, making it difficult to properly track individual progress. Given these circumstances, there is a need for a system that allows users to effectively and sustainably improve their communication skills. [Means for solving the problem]
[0005] The present invention provides a means for authenticating a user's login information and determining the user's skill level based on past training data. It then implements a means for providing an appropriate training scenario based on the user's skill level, and provides a means for generating a communication scene based on the training scenario selected by the user. It also provides a means for analyzing the user's speech in real time, generating feedback, and presenting it to the user. It also provides a means for saving individual user progress data and tracking and analyzing growth records, providing information for users to self-evaluate and promote their growth. This series of means allows users to improve their communication skills while receiving real-time feedback, promoting long-term growth.
[0006] "User" refers to an individual who intends to use this system to improve their communication skills.
[0007] "Login Information" means the authentication information (e.g., username and password) used by a User to access a System.
[0008] "Training Data" refers to data (e.g., speech content, feedback, progress information) from a user's previous training sessions.
[0009] "Skill level" refers to an index that indicates the degree of a user's communication ability.
[0010] "Training scenario" refers to a specific task or situation provided to a user to train their communication skills.
[0011] A "communication scene" refers to a specific situation in which a user should speak or take action, which is generated based on a training scenario.
[0012] "Speech data" refers to voice data of speech made by a user using a microphone.
[0013] "Natural language processing" refers to the technology that allows computers to analyze, understand, and automatically process human language.
[0014] "Feedback" refers to evaluations and instructions for improvement presented to a user in response to their speech or actions.
[0015] "Database" means a repository of information where a user's training and progress data is stored.
[0016] "Progress Data" refers to a record of the progress and growth a User has achieved through training.
[0017] "Self-evaluation" refers to the process in which the user evaluates the degree of improvement in his or her own communication skills.
[0018] "Growth record" refers to the history of skill improvement as the user continues to train. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to a communication skill improvement platform, which is designed to provide a set of functions for users to effectively improve their communication skills. The system is implemented as a cross-platform application and can be used through a web browser or a mobile application.
[0041] server
[0042] The server plays a central role in this system and provides the following functions:
[0043] 1. User authentication and information management
[0044] The server verifies the user's credentials when they submit a login request, and if authentication is successful, retrieves the user's past training data from a database to determine their current skill level.
[0045] 2. Providing training scenarios
[0046] The server generates appropriate training scenarios based on the user's skill level and sends the information to the user's device. The scenarios contain various communication scenes, and the user can select a scenario that suits their level.
[0047] 3. Communication Scene Generation
[0048] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[0049] 4. Real-time analysis and feedback
[0050] The server receives the user's voice data in real time and converts it into text using natural language processing (NLP) technology. The converted text data is then analyzed, and appropriate responses and feedback are generated and sent to the device.
[0051] 5. Saving and analyzing progress data
[0052] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. Later, this data is analyzed and used to provide information for visualizing the user's progress.
[0053] Terminal
[0054] The user's device works in conjunction with the server to provide the following functions:
[0055] 1. User Interface
[0056] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[0057] 2. Audio capture and transmission
[0058] The device captures the user's speech with a microphone and transmits the audio data to the server in real time.
[0059] 3. Displaying feedback and responses
[0060] Feedback and responses received from the server are displayed as voice or text, allowing the user to understand what needs to be improved and take the next action.
[0061] User
[0062] The user performs the following actions through this system:
[0063] 1. Log in and check your profile
[0064] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[0065] 2. Selecting a training scenario
[0066] The user selects a scenario that suits their level from a list of training scenarios provided.
[0067] 3. Practice in communication situations
[0068] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[0069] 4. Receive feedback and improve
[0070] See real-time feedback and act on it to improve your skills.
[0071] 5. Track your progress over the long term
[0072] Users can view progress data on a dashboard and perform self-assessments to promote their growth.
[0073] As a concrete example, consider the case where a user selects a job interview scenario. The user experiences self-introductions and question-and-answer sessions, and receives real-time feedback from the server. For example, the server may provide feedback such as, "You speak too quickly, so you should speak more slowly," or "You need to improve your explanation, including your specific work experience." By practicing based on this feedback, the user will be able to respond confidently in an actual interview.
[0074] Thus, the present invention provides users with an environment for practical and effective communication skill development, with real-time feedback and long-term progress tracking.
[0075] The processing flow will be explained below.
[0076] Step 1:
[0077] User
[0078] The user enters their account information on the login screen and clicks the login button.
[0079] server
[0080] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[0081] Step 2:
[0082] server
[0083] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[0084] Terminal
[0085] The terminal displays the list of training scenarios received from the server on the screen.
[0086] Step 3:
[0087] User
[0088] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[0089] server
[0090] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[0091] Step 4:
[0092] Terminal
[0093] The terminal displays the scene information received from the server on the screen and prepares for audio playback.
[0094] User
[0095] The user checks the scene information and clicks the button to start training.
[0096] Step 5:
[0097] Terminal
[0098] The device captures the user's speech with a microphone and transmits the voice data to the server in real time.
[0099] server
[0100] The server receives the voice data and converts it into text data via a natural language processing (NLP) module.
[0101] Step 6:
[0102] server
[0103] The server analyzes the converted text data, generates appropriate responses and feedback, converts the generated responses and feedback into audio data, and sends it to the user's device.
[0104] Terminal
[0105] The terminal plays back the audio data received from the server and provides feedback to the user.
[0106] Step 7:
[0107] User
[0108] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[0109] Step 8:
[0110] server
[0111] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. The stored data is analyzed and information is generated to visualize the progress.
[0112] User
[0113] After the session, users can view their progress data on a dashboard and self-evaluate.
[0114] Example 1
[0115] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0116] Current communication skill training systems lack the functionality to provide appropriate feedback in real time based on the user's skill level. They also have limited functionality to effectively track the user's progress and suggest specific ways to improve. This makes it difficult for existing systems to effectively improve skills.
[0117] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0118] In this invention, the server includes means for verifying user authentication information, means for determining the user's skill level based on past learning data, means for generating appropriate training scenarios based on the user's skill level, means for constructing dialogue scenes based on the training scenario selected by the user, means for analyzing the user's voice data in real time using natural language processing technology and generating feedback, and means for storing session data in a database and analyzing and visualizing the user's progress, thereby enabling the user to receive effective feedback in real time, thereby promoting self-evaluation and growth.
[0119] "Means for verifying user authentication information" refers to a means for checking the authentication information (user name and password) entered by the user and comparing it with a database to see if it is correct.
[0120] The "means for determining the user's skill level based on past learning data" refers to a means for analyzing the user's past training data and evaluation results to determine the user's current skill level.
[0121] The "means for generating an appropriate training scenario based on the skill level of the user" is a means for creating an optimal training scenario for the user in accordance with the determined skill level.
[0122] The "means for constructing a dialogue scene based on a training scenario selected by the user" refers to a means for constructing a scene including specific dialogue content and questions based on a scenario selected by the user.
[0123] "Means for analyzing user voice data in real time using natural language processing technology and generating feedback" refers to means for capturing a user's speech as voice data, converting it into text, analyzing it, and generating the necessary feedback in real time.
[0124] "Means for storing session data in a database and analyzing and visualizing the user's progress" refers to a means for storing data after a training session in a database and subsequently analyzing the data to visualize the user's progress in graphs or other formats.
[0125] The communication skill improvement system of the present invention can be effectively implemented by cooperation between the server, the terminal, and the user.
[0126] Server Roles
[0127] The server plays a central role in this system and provides the following functions:
[0128] 1. Verify user credentials
[0129] The server receives the authentication information entered by the user on the login screen and authenticates the user by comparing it with the database. The user can only use the system if authentication is successful.
[0130] 2. Determine the user's skill level
[0131] The server analyzes past learning data to determine the user's current skill level, including stored training session results and evaluations.
[0132] 3. Generating training scenarios
[0133] The server automatically generates an appropriate training scenario based on the user's skill level. The generated scenario is sent to the user's device. For example, a "basic self-introduction" scenario for beginners and a "job interview" scenario for intermediate users are generated.
[0134] 4. Creating a communication scene
[0135] Based on the training scenario selected by the user, the system constructs a specific dialogue scene and transmits the scene information to the device. The scene includes the actual utterances and questions the user will ask.
[0136] 5. Real-time analysis and feedback
[0137] The server receives the user's voice data in real time, converts it into text using natural language processing technology, and generates feedback, which is then immediately sent to the device and presented to the user.
[0138] 6. Session data storage and progress analysis
[0139] After each session, the server stores all relevant data in a database. By analyzing the stored data, the user's progress is visualized and displayed on a dashboard, allowing users to self-evaluate and track their progress.
[0140] Device Role
[0141] The user's device works in conjunction with the server to provide the following functions:
[0142] 1. Providing a user interface
[0143] The terminal allows the user to interact with the system through interfaces such as a login screen, scenario selection screen, and training start screen.
[0144] 2. Capture and transmit audio data
[0145] The device captures the user's speech with a microphone and transmits the audio data to a server in real time, allowing for immediate analysis and feedback.
[0146] 3. Viewing Feedback
[0147] The terminal displays the feedback or response received from the server in voice or text form and provides it to the user.
[0148] User Behavior
[0149] The user performs the following actions through this system:
[0150] 1. Log in and check your profile
[0151] Users enter their login information to access the system. After logging in, they can check their skill level and past training data on the dashboard.
[0152] 2. Selecting a training scenario
[0153] From the list of training scenarios provided, choose one that matches your skill level, for example, the intermediate "Job Interview" scenario.
[0154] 3. Practice in communication situations
[0155] The user will practice speaking in a dialogue scene generated based on the selected scenario. Once the user has finished speaking, they can move on to the next step.
[0156] 4. Receive feedback and improve
[0157] See real-time feedback and use it to improve your skills, for example, "You speak too fast; you should speak a little slower."
[0158] 5. Track your progress over the long term
[0159] Users can view progress data and self-assess on the dashboard, allowing them to understand their progress and set goals for further improvement.
[0160] Examples and prompts
[0161] Specific examples
[0162] If a user selects the "Job Interview" scenario for intermediate level learners, they will experience self-introductions and question-and-answer sessions, and receive real-time feedback from the server. For example, they may receive feedback such as, "You speak too quickly; you should speak a little more slowly," or "You need to improve your explanation, including your specific work experience." Users can then practice based on this feedback, gaining confidence in the actual interview.
[0163] Prompt Sentence Examples
[0164] The user has selected the job interview scenario. Please simulate the scene where the user introduces themselves and provide feedback on the following points:
[0165] 1. Speaking Speed
[0166] 2. Description of specific work experience
[0167] 3. Language and attitude
[0168] By including these specific details, the present invention provides users with an environment for practical and effective communication skill development, allowing for real-time feedback and long-term growth tracking.
[0169] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0170] Step 1:
[0171] Entering and validating user credentials
[0172] Input: The username and password the user enters on the login screen.
[0173] Specific behavior:
[0174] The user enters the username and password on the login screen of the device.
[0175] The terminal transmits the entered authentication information to the server.
[0176] The server checks the received authentication information against a database.
[0177] Data processing / calculation:
[0178] The server performs a calculation to match the username and password in its database.
[0179] output:
[0180] If the authentication is successful, the server sends a "user authentication successful" response to the terminal and starts a session.
[0181] If the authentication fails, the server sends a "user authentication failed" response to the terminal, prompting it to try again.
[0182] Step 2:
[0183] Obtaining user history data and determining skill level
[0184] Input: User ID and previous learning data request.
[0185] Specific behavior:
[0186] The server uses the user ID to load the user's past training data from a database.
[0187] Data processing / calculation:
[0188] The server analyzes the acquired past training data and performs calculations to evaluate the user's skill level.
[0189] output:
[0190] The user's current skill level information is transmitted to the terminal.
[0191] Step 3:
[0192] Training scenario generation and distribution
[0193] Input: User skill level information.
[0194] Specific behavior:
[0195] The server generates appropriate training scenarios based on skill level.
[0196] Data processing / calculation:
[0197] The server selects the most suitable scenario for the user's skill level from among multiple scenario candidates and constructs scenario data.
[0198] output:
[0199] The generated training scenario is sent to the terminal.
[0200] Step 4:
[0201] Building and sending communication scenes
[0202] Input: A training scenario selected by the user.
[0203] Specific behavior:
[0204] The user selects one of the scenarios provided on the terminal.
[0205] The terminal transmits the selected scenario information to the server.
[0206] The server constructs a specific dialogue scene based on the selected scenario.
[0207] output:
[0208] The constructed dialogue scene is sent to the device.
[0209] Step 5:
[0210] Capture and transmit audio data
[0211] Input: User speech.
[0212] Specific behavior:
[0213] The user speaks into the microphone of the terminal.
[0214] The device captures audio data through a microphone.
[0215] The device transmits the captured audio data to the server in real time.
[0216] output:
[0217] Send the audio data to the server.
[0218] Step 6:
[0219] Providing real-time analysis and feedback
[0220] Input: User's voice data sent from the device.
[0221] Specific behavior:
[0222] The server converts the received voice data into text using natural language processing technology.
[0223] The server analyzes the converted text data and generates appropriate feedback.
[0224] Data processing / calculation:
[0225] The system converts voice data into text and performs feedback generation calculations based on text analysis.
[0226] output:
[0227] The generated feedback is sent to the device.
[0228] Step 7:
[0229] Session data storage and progress visualization
[0230] Input: Data from each session (utterances, feedback, and evaluation results).
[0231] Specific behavior:
[0232] The server stores all relevant data in a database after the session ends.
[0233] Data processing / calculation:
[0234] The stored data is analyzed and calculations are performed to visualize the progress data.
[0235] output:
[0236] Progress data is visualized and presented to users via a dashboard.
[0237] Through the above steps, this system can effectively improve the user's communication skills.
[0238] (Application example 1)
[0239] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0240] In modern brick-and-mortar stores, improving staff communication skills is a high priority. High skills are especially required for dealing with new customers and handling complaints, as these skills are directly linked to customer satisfaction. However, in brick-and-mortar stores, staff have limited opportunities to effectively improve their skills while performing their daily tasks. Furthermore, traditional training methods make it difficult to train many staff members at once, and individual feedback is lacking. As a result, there is a problem of inefficient improvement of individual skills.
[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0242] In this invention, the server includes means for authenticating a user's login information, means for determining a user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating communication scenes based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for saving the user's progress data and tracking and analyzing the growth record, means for evaluating the user's speech speed and content and suggesting specific areas for improvement, and means for creating scenarios for staff to respond to in-store situations and conducting training. This enables store staff to effectively and efficiently improve the communication skills required in-store while performing their daily work.
[0243] The means for "authenticating user login information" is a means for verifying authentication information such as a user name and password entered when a user accesses the system, and confirming that the user is a legitimate user.
[0244] The means for "determining the user's skill level based on past training data" is a means for analyzing data on training previously performed by the user and evaluating the current level of the user's skill based on the data.
[0245] The means for "providing an appropriate training scenario based on the user's skill level" is a means for automatically selecting training according to the determined skill level and presenting it to the user.
[0246] The means for "generating a communication scene based on a training scenario selected by the user" is a means for reproducing a specific communication situation or condition based on a training scenario selected by the user.
[0247] The method of "analyzing user speech in real time and generating feedback" involves collecting the user's speech with a microphone, instantly analyzing it using natural language processing technology, and presenting appropriate suggestions and areas for improvement based on the analysis results.
[0248] The means for "storing user progress data and tracking and analyzing growth records" refers to storing the user's training details and results in a database, and analyzing them over the long term to track skill improvement.
[0249] The method of "evaluating the user's speech speed and content and suggesting specific areas for improvement" involves analyzing the speed and quality of the content of the user's speech and, based on the results, providing specific instructions to the user on how to improve.
[0250] The method of "creating scenarios for staff to respond to in-store situations and conducting training" involves creating virtual scenarios for specific situations such as customer service and handling complaints in a real store, and then having staff practice using those scenarios.
[0251] This invention is a system for improving communication skills for store staff. This system mainly consists of three components: a server, a terminal, and a user. The following is a detailed description of the role of each component and its processing.
[0252] server
[0253] The server plays a central role in the system, providing the following functions:
[0254] User authentication:
[0255] When a user logs in to a system, the server authenticates them based on their username and password, ensuring that only authorized users can access the system.
[0256] Skill Level Determination:
[0257] The server analyzes past training data to assess the user's current skill level, which is stored in a database and updated in real time.
[0258] Training scenarios provided:
[0259] Based on the user's skill level, appropriate training scenarios are selected and provided to the user, including scenarios for greeting new customers and handling complaints.
[0260] Communication Scene Generation:
[0261] Based on the scenario selected by the user, specific communication scenes are generated and sent to the device, allowing the user to receive training tailored to their actual work.
[0262] Real-time analysis and feedback:
[0263] It collects user utterances in real time, analyzes them using natural language processing (NLP) technology, and generates feedback based on the analysis results and sends it to the device.
[0264] Save and analyze progress data:
[0265] Session data is stored in a database and progress is visualized for each user, allowing users to self-assess and foster long-term growth.
[0266] Terminal
[0267] The user's terminal provides the following functions in cooperation with the server.
[0268] User Interface:
[0269] The terminal provides an interface for interacting with the user, such as a login screen, a scenario selection screen, and a training scene.
[0270] Audio Capture:
[0271] The device's microphone is used to capture the user's speech in real time and send the data to the server.
[0272] Show feedback:
[0273] The analysis results and feedback received from the server are provided to the user via text or voice, allowing the user to make immediate corrections and improve their skills.
[0274] User
[0275] The users are staff members of a physical store, and by utilizing this system, they perform the following actions:
[0276] Login:
[0277] The user logs in to the terminal and checks their account information.
[0278] Select a training scenario:
[0279] Choose from the list of scenarios provided that fit your skill level.
[0280] Communication practice:
[0281] Participants will experience and speak in a series of real-life communication situations, including greeting new customers and handling complex complaints.
[0282] Receiving feedback and making corrections:
[0283] Improve your skills with real-time feedback.
[0284] Long-term progress check:
[0285] View your training history and progress data on the dashboard and self-assess.
[0286] Hardware and software used
[0287] Hardware: Smartphones, tablets
[0288] software:
[0289] Web browser (cross-platform compatible)
[0290] Python (Flask framework)
[0291] Database (MySQL or PostgreSQL)
[0292] Natural Language Processing (NLP) libraries (spaCy and nltk)
[0293] Frontend (ReactJS)
[0294] Examples of concrete examples and prompts
[0295] For example, if the user selects a scenario for handling a complaint, the following dialogue is generated:
[0296] 1. User: "Sorry, but we can't accept your return at this time."
[0297] 2. System: Performs real-time analysis and provides feedback such as "The text is too long. Make it concise" or "Intonation is important."
[0298] As described above, by using this system, store staff can efficiently improve the communication skills required in physical stores.
[0299] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0300] Step 1:
[0301] Authenticate your login details
[0302] The server receives the login information (username and password) entered by the user. It compares this information with existing data in the database, and if it matches, it completes the authentication. Based on the information entered, it determines whether the user is legitimate, and if successful, it proceeds to the next step. This process allows access to the system.
[0303] Input: Username, Password
[0304] Output: Authentication success / failure status
[0305] Step 2:
[0306] Skill level determination
[0307] The server retrieves past training data for successfully authenticated users from the database and analyzes it. It uses Python data analysis libraries (such as pandas and NumPy) to evaluate the user's current skill level. Based on this, it selects a training scenario of the appropriate level.
[0308] Input: Past training data
[0309] Output: Skill level judgment result
[0310] Step 3:
[0311] Providing training scenarios
[0312] The server selects multiple training scenarios based on the determined skill level and sends them to the user's device. The scenarios include many scenes that can be used in real stores, and the user can choose the scenario they want to train in.
[0313] Input: Skill Level
[0314] Output: List of training scenarios provided
[0315] Step 4:
[0316] Communication Scene Generation
[0317] Based on the training scenario selected by the user, the server generates specific communication scenes, including scripts and dialogue flows, to recreate real-life store situations. The generated scenes are then sent to the device for the user to practice.
[0318] Input: The selected training scenario
[0319] Output: Communication scene details
[0320] Step 5:
[0321] Audio capture and transmission
[0322] When the user presses the start button, the device uses the built-in microphone to capture the user's speech in real time. The speech data is immediately sent to the server and prepared for analysis, particularly for speech recognition and natural language processing (NLP), where it is converted into an appropriate data format (e.g., text).
[0323] Input: User utterance
[0324] Output: Sends audio data to the server
[0325] Step 6:
[0326] Real-time analysis and feedback generation
[0327] The server analyzes the received voice data using a natural language processing (NLP) engine. It analyzes the text of the speech and prosody (speech characteristics) such as speed and intonation, and generates appropriate feedback. The generated feedback is sent to the user's device in real time.
[0328] Input: Audio data
[0329] Output: Feedback
[0330] Step 7:
[0331] View Feedback
[0332] The device receives feedback from the server and displays it to the user in text or audio, allowing the user to review the feedback and learn how to improve their skills.
[0333] Input: Feedback data
[0334] Output: Display feedback
[0335] Step 8:
[0336] Save and analyze progress data
[0337] At the end of the session, the device sends the session data (such as what was said and feedback) to the server, which stores the data in a database and visualizes each user's progress. This data is analyzed to provide insights to drive long-term growth.
[0338] Input: Session data
[0339] Output: Progress data visualization and analysis results
[0340] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0341] This invention is a system that further enhances the effectiveness of training by incorporating an emotion engine into a platform for users to effectively improve their communication skills. The system has the function of recognizing the user's emotions in real time and adjusting feedback and training scenarios based on that information.
[0342] server
[0343] The server plays a central role in this system and provides the following functions:
[0344] 1. User authentication and information management
[0345] The server verifies the user's credentials when they submit a login request, and if authentication is successful, determines the user's past training data and skill level.
[0346] 2. Providing training scenarios
[0347] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[0348] 3. Communication Scene Generation
[0349] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[0350] 4. Real-time analysis and feedback
[0351] The server receives and analyzes the user's voice and facial expression data in real time, and then uses a natural language processing (NLP) module to convert the voice data into text and generate appropriate responses and feedback.
[0352] 5. Emotion recognition and feedback regulation
[0353] The server uses an emotion engine to analyze the user's emotional state and adjusts the feedback accordingly. For example, if the user is nervous, it provides feedback to calm them down.
[0354] 6. Saving and analyzing progress data
[0355] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The server analyzes the stored data and provides users with information that visualizes their progress.
[0356] Terminal
[0357] The user's device works in conjunction with the server to provide the following functions:
[0358] 1. User Interface
[0359] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[0360] 2. Capture and transmit voice and facial expressions
[0361] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the data to a server in real time.
[0362] 3. Displaying feedback and responses
[0363] It displays feedback and responses received from the server in audio or text format, and also displays emotion-based feedback.
[0364] User
[0365] The user performs the following actions through this system:
[0366] 1. Log in and check your profile
[0367] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[0368] 2. Selecting a training scenario
[0369] The user selects the scenario that suits them from a list of training scenarios provided.
[0370] 3. Practice in communication situations
[0371] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[0372] 4. Receiving and improving feedback based on emotion recognition
[0373] See and improve your skills based on real-time feedback, including emotional feedback recognized by the emotion engine.
[0374] 5. Track your progress over the long term
[0375] Users can view their progress and self-assess on a dashboard, which also includes a record of their emotions, allowing them to see their mental progress.
[0376] As a concrete example, consider the case where a user selects a job interview scenario and experiences self-introduction and question-and-answer sessions. If the emotion engine recognizes that the user is nervous, the server will provide feedback such as "Take a deep breath to relieve tension." It can also provide specific advice such as "You're speaking too fast, so speak a little more slowly." By practicing based on this feedback, the user can improve their actual performance in the interview.
[0377] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[0378] The processing flow will be explained below.
[0379] Step 1:
[0380] User
[0381] The user enters their account information on the login screen and clicks the login button.
[0382] server
[0383] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[0384] Step 2:
[0385] server
[0386] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[0387] Terminal
[0388] The terminal displays the list of training scenarios received from the server on the screen.
[0389] Step 3:
[0390] User
[0391] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[0392] server
[0393] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[0394] Step 4:
[0395] Terminal
[0396] The terminal displays the scene information received from the server on the screen and prepares to play audio and capture facial expressions.
[0397] User
[0398] The user checks the scene information and clicks the button to start training.
[0399] Step 5:
[0400] Terminal
[0401] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the audio and video data to a server in real time.
[0402] server
[0403] The server receives the voice data and facial expression data and converts the voice data into text via a natural language processing (NLP) module.
[0404] Step 6:
[0405] server
[0406] The server analyzes the converted text data and facial expression data and uses an emotion engine to recognize the user's emotional state. Based on the analysis results, it generates appropriate responses and feedback. The generated responses and feedback are converted into voice data and sent to the user's device.
[0407] Terminal
[0408] The terminal plays back the voice data and emotional feedback received from the server and provides them to the user.
[0409] Step 7:
[0410] User
[0411] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[0412] Step 8:
[0413] server
[0414] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The stored data is analyzed and information is generated to visualize the progress.
[0415] User
[0416] After the session, users can check their progress data on the dashboard and self-evaluate. They can also review their progress, including changes in their emotions, and use this information to improve their next training session.
[0417] Example 2
[0418] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0419] Conventional communication skill improvement systems struggle to accurately recognize a user's emotional state and provide feedback based on that, making effective training difficult. Furthermore, they often lack real-time speech analysis and feedback generation, which often results in ineffective improvement of the user's skills. Furthermore, they lack sufficient visualization of progress data and long-term growth records, leaving users with a lack of reference information for self-evaluation.
[0420] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for authenticating the user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on the training scenario selected by the user, means for analyzing the user's voice data and facial expression data in real time and generating feedback, means for analyzing the user's emotional state and adjusting the feedback based on the results, and means for saving the user's progress data and tracking and analyzing the growth record. This makes it possible to analyze the user's emotional state in real time and provide appropriate feedback based on the analysis results. Furthermore, by visualizing the progress data and providing a long-term growth record, it is possible to enhance the reference information for the user when performing self-evaluation.
[0421] "User Login Information" means the authentication data (e.g., user ID and password) provided by a User to access a System.
[0422] "Training data" refers to historical information such as the results and evaluations of past training sessions conducted by the user, as well as the content of utterances.
[0423] "Skill level" is an index that evaluates a user's communication ability and proficiency.
[0424] A "training scenario" is an exercise designed to allow a user to practice in a specific situation.
[0425] A "communication scene" is a specific situation or setting that a user simulates, which is generated based on a training scenario.
[0426] "Voice data" refers to data that digitally captures a user's speech.
[0427] "Facial Expression Data" refers to data that captures a user's facial expressions and represents them in digital form.
[0428] "Feedback" refers to the evaluation and advice provided by the system in response to the user's actions and utterances.
[0429] An "emotional state" is the emotion (e.g., tension, joy, anger, etc.) that a user is feeling at a particular moment.
[0430] "Analysis" is the process of processing collected data to extract useful information and patterns.
[0431] "Progress Data" means the performance and evaluation recorded for each of a User's training sessions.
[0432] A "growth record" is information that records the history and results of a user's training sessions over a long period of time and shows the trajectory of their growth.
[0433] "Visualization" is the process of transforming and displaying complex data into visual formats such as graphs and charts.
[0434] The present invention is a training system for improving a user's communication skills. This system recognizes the user's emotions in real time and provides feedback based on the results, thereby achieving more effective training.
[0435] System Configuration
[0436] server
[0437] The server plays a central role in the system and includes the following hardware and software components:
[0438] Authentication system: Validates user login information and performs authentication. The software used is a common authentication framework.
[0439] Database Management System (DBMS): Stores and manages users' past training data and skill levels.
[0440] Training scenario generation module: Selects an appropriate training scenario based on the user's skill level.
[0441] Communication scene generation module: Generates specific communication scenes based on the scenario selected by the user.
[0442] Real-time analysis engine: Using an NLP module, the user's voice is converted into text, and the content is then analyzed to generate an appropriate response.
[0443] Emotion Engine: Analyzes the user's facial expression data and evaluates the user's emotional state.
[0444] Feedback generation module: Generates feedback for the user based on the results of real-time analysis and sentiment analysis.
[0445] Progress Data Storage and Analysis Module: Saves and visualizes progress data after a user's training session.
[0446] Terminal
[0447] The user's device works in conjunction with the server to provide the following functions:
[0448] User Interface: Provides interfaces for users to interact with the system, such as a login screen, scenario selection screen, and training screen.
[0449] Voice capture device: Uses a microphone to capture the user's speech and sends it to the server as voice data.
[0450] Facial Expression Capture Device: Uses a camera to capture the user's facial expressions and transmits the data to a server.
[0451] Feedback display device: Displays feedback or responses received from the server in audio or text format.
[0452] User
[0453] The user performs the following actions:
[0454] 1. Log in and check your profile: Log in to the system using your device and check your skill level and past training data on the dashboard.
[0455] 2. Select a training scenario: Select an appropriate scenario from the list of training scenarios provided by the server.
[0456] 3. Practice in communication scenes: Practice your speech and actions in scenes generated based on the scenario you select.
[0457] 4. Receive real-time feedback: See feedback displayed on your device and improve your training.
[0458] 5. Track your progress: Check your progress data on the dashboard and self-assess.
[0459] Specific examples
[0460] As a concrete example, consider the case where a user selects a "job interview" scenario. The user actually introduces themselves and engages in a question-and-answer session with an AI acting as an interviewer. The device's camera and microphone are used to capture the user's facial expressions and voice. The server analyzes this data in real time and provides feedback such as "You seem nervous. Take a deep breath." It also displays specific advice such as "You speak too quickly. Please speak more slowly."
[0461] Prompt Sentence Examples
[0462] Prompt: Start a practice job interview scenario. First, introduce yourself. Then move on to questions and answers. If you're nervous, the system will provide feedback. It will also give you specific advice on how to speak.
[0463] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[0464] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0465] Step 1:
[0466] Enter your login information
[0467] Input: The user enters their user ID and password on the login screen.
[0468] Operation: The device sends the authentication information entered by the user to the server.
[0469] Output: The server receives the authentication information.
[0470] Step 2:
[0471] User authentication and information management
[0472] Input: The server checks the user's login information against a database.
[0473] How it works: After successful authentication, the server retrieves past training data and skill level from the database.
[0474] Output: Send the acquired training data and skill level to the device.
[0475] Step 3:
[0476] Providing training scenarios
[0477] Input: The server selects appropriate training scenarios based on the user's skill level.
[0478] Operation: The server generates a list of selected training scenarios and sends it to the terminal.
[0479] Output: A list of scenarios will be displayed on the terminal.
[0480] Step 4:
[0481] Scenario Selection
[0482] Input: The user selects one from the list of scenarios provided.
[0483] Operation: The terminal transmits the selected scenario information to the server.
[0484] Output: The server receives the scenario selection information.
[0485] Step 5:
[0486] Communication Scene Generation
[0487] Input: The server generates a scene based on the selected scenario.
[0488] Operation: The server creates detailed communication scenes based on the scenario and sends the information to the terminal.
[0489] Output: A specific training scene is displayed on the device.
[0490] Step 6:
[0491] Voice and facial expression capture
[0492] Input: The user makes utterances and expressions in the simulation.
[0493] How it works: The device uses a microphone and camera to capture the user's voice and facial expression data in real time.
[0494] Output: The captured data is sent to the server.
[0495] Step 7:
[0496] Real-time analytics
[0497] Input: The server receives the voice data and facial expression data sent from the device.
[0498] How it works: The server uses an NLP module to convert voice data into text and analyze its content, as well as analyze facial expression data to assess emotional state.
[0499] Output: Text data and emotion evaluation results are obtained as the analysis results.
[0500] Step 8:
[0501] Feedback Generation
[0502] Input: The server generates feedback based on the real-time analysis results and emotion evaluation results.
[0503] How it works: The server generates and adjusts the feedback based on the user's emotional state. For example, if the user is nervous, the server generates the feedback "Take a deep breath."
[0504] Output: The generated feedback information is sent to the terminal.
[0505] Step 9:
[0506] View Feedback
[0507] Input: The terminal obtains the feedback information received from the server.
[0508] Action: The device displays feedback in the form of audio or text.
[0509] Output: User sees feedback.
[0510] Step 10:
[0511] Save and visualize progress data
[0512] Input: The server receives all the data from the training session.
[0513] How it works: After a session ends, the server stores progress data in a database and analyzes it to generate visualizations, such as graphs and charts.
[0514] Output: Visualized progress information is sent to the device and made available to the user on a dashboard.
[0515] Through this series of processes, users receive real-time feedback to improve their communication skills, enabling effective training.
[0516] (Application example 2)
[0517] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0518] Providing effective customer service skill training is crucial in today's virtual environments. However, conventional training systems struggle to provide users with real-time feedback or feedback based on emotion recognition, resulting in delayed or insufficient improvement. They also lack a means to effectively track and visualize user progress. It is necessary to provide a system that can solve these issues and efficiently support the improvement of customer service skills.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0520] In this invention, the server includes means for authenticating a user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for capturing the user's facial expressions and recognizing their emotions, means for providing feedback based on the recognized emotions and adjusting the training scenario, means for saving the user's progress data and tracking and analyzing the growth record, means for the user to select from a plurality of customer service scenarios, and means for presenting prompt sentences for the user to train their customer service skills. This enables the user to efficiently improve their customer service skills while receiving feedback based on emotion recognition in real time.
[0521] "User Login Information" means the authentication information for a User to access the System.
[0522] "Past Training Data" means records and performance of a User's previous training sessions.
[0523] "Skill level" refers to a user's ability or proficiency in a particular skill.
[0524] A "training scenario" is a simulation or challenge that allows a user to practice a particular skill.
[0525] A "communication scene" is a virtual interaction or situation that a user experiences during training.
[0526] "User utterance" refers to the sounds and language that a user makes during training.
[0527] "Feedback" refers to evaluation of the user's actions and utterances and advice for improvement.
[0528] "Facial Expression" refers to the facial expression of the user.
[0529] "Means for recognizing emotions" refers to techniques or devices for analyzing a user's emotional state.
[0530] "Progress Data" means a record showing how much a User has improved through training.
[0531] A "customer service scenario" is a hypothetical situation that users use to train their customer service skills.
[0532] A "prompt" is an instruction or question that prompts the user to take a specific action or speak a certain word.
[0533] This invention is a system for efficiently training customer service skills in a virtual store. The system consists of three main elements: a server, a terminal, and a user.
[0534] server
[0535] The server plays a central role in this system and provides the following functions:
[0536] 1. User authentication and information management
[0537] The server verifies the user's credentials when they submit a login request, and if the login is successful, determines the user's skill level based on past training data.
[0538] 2. Providing training scenarios
[0539] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[0540] 3. Communication Scene Generation
[0541] When a user selects a specific training scenario, the server constructs a specific communication scene based on that scenario and transmits the information to the terminal.
[0542] 4. Real-time analysis and feedback
[0543] The server receives and analyzes the user's speech data and facial expression data in real time using OpenCV (for capturing facial expressions), Keras (for loading emotion recognition models and analyzing emotions), and NLP modules (for natural language processing).
[0544] 5. Emotion recognition and feedback regulation
[0545] The server analyzes the user's emotions using an emotion engine, provides feedback based on the results, and adjusts the training scenario. For example, if the server determines that the user is tense, it provides feedback encouraging the user to relax.
[0546] 6. Saving and analyzing progress data
[0547] After the session ends, the server stores all the user's session data in a database, analyzes the stored data, and provides the user with a visualization of their progress.
[0548] Terminal
[0549] The user's device works in conjunction with the server to provide the following functions:
[0550] 1. User Interface
[0551] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[0552] 2. Capture and transmit voice and facial expressions
[0553] The device captures the user's speech using a voice input device, recognizes facial expressions using a camera, and transmits the data to a server in real time.
[0554] 3. Displaying feedback and responses
[0555] The device displays feedback and responses received from the server in audio or text format, as well as emotion-based feedback.
[0556] User
[0557] Through this system, users can carry out the following training activities:
[0558] 1. Log in and check your profile
[0559] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[0560] 2. Selecting a training scenario
[0561] The user selects the scenario that suits them from a list of training scenarios provided.
[0562] 3. Practice in communication situations
[0563] The user trains by actually speaking in a communication scene generated based on the selected scenario.
[0564] 4. Receiving and improving feedback based on emotion recognition
[0565] See real-time feedback and act on it to improve your skills.
[0566] 5. Track your progress over the long term
[0567] Users can view their progress and self-assess on a dashboard, which also includes tracking their emotions and showing their mental progress.
[0568] Specific examples
[0569] For example, if a user selects the customer service scenario "Welcome Customers," the prompt message will be displayed: "As a salesperson in a virtual store, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." The emotion recognition engine will detect tension in the user's facial expressions, and the server will provide feedback such as "Relax and smile more." In this way, the user can improve their customer service skills while receiving specific advice.
[0570] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0571] Step 1:
[0572] The server receives the user's login information and performs authentication.
[0573] Input: User login information (user ID and password)
[0574] Data processing: The server checks the data against the authentication database
[0575] Output: Authentication result (success or failure)
[0576] Specific operation: The server checks the login information against the authentication database, and if authentication is successful, retrieves the user's past training data.
[0577] Step 2:
[0578] The server determines the user's skill level based on past training data.
[0579] Input: Past training data
[0580] Data processing: Analysis of training data and assessment of skill level
[0581] Output: User's skill level
[0582] Specific operation: The server analyzes past training data and applies a skill assessment algorithm to calculate the user's skill level.
[0583] Step 3:
[0584] The server generates an appropriate training scenario based on the user's skill level and transmits it to the terminal.
[0585] Input: User's skill level
[0586] Data processing: Selection of appropriate scenarios from the scenario database
[0587] Output: List of training scenarios
[0588] Specific operation: The server extracts training scenarios that match the user's skill level from the scenario database and sends the list to the terminal.
[0589] Step 4:
[0590] The user selects a training scenario on the terminal.
[0591] Input: List of training scenarios
[0592] Data Calculation: Scenario Selection
[0593] Output: Selected training scenario
[0594] Specific operation: The user selects the desired scenario from the list of training scenarios displayed on the terminal.
[0595] Step 5:
[0596] The server generates a communication scene based on the selected scenario and transmits it to the terminal.
[0597] Input: Selected training scenario
[0598] Data processing: Generation of communication scenes based on scenarios
[0599] Output: Communication scene data
[0600] Specific operation: The server constructs a virtual communication scene based on the selected scenario and sends the data to the terminal.
[0601] Step 6:
[0602] The device captures the user's voice and facial expressions and transmits them to the server.
[0603] Input: User utterances and facial expressions
[0604] Data processing: Audio input device and camera capture
[0605] Output: Capture data
[0606] Specific operation: The device captures the user's voice with a microphone, recognizes facial expressions with a camera, and sends the data to a server.
[0607] Step 7:
[0608] The server analyzes the user's speech data and facial expression data in real time, generates feedback, and sends it to the device.
[0609] Input: User speech data and facial expression data
[0610] Data processing: speech recognition, natural language processing, sentiment analysis
[0611] Output: Real-time feedback
[0612] Specific operation: The server analyzes the user's speech using speech recognition and natural language processing, and analyzes the facial expressions using an emotion engine. Based on the obtained information, it generates feedback and sends it to the device.
[0613] Step 8:
[0614] The terminal presents the feedback received from the server to the user in real time.
[0615] Input: Feedback data
[0616] Data Calculation: Feedback Display
[0617] Output: Feedback that is displayed to the user
[0618] Specific operation: The device displays the feedback received from the server to the user in voice or text.
[0619] Step 9:
[0620] The server stores the progress data in a database and generates visualization data.
[0621] Input: Training session data
[0622] Data processing: saving to database, analyzing and visualizing progress data
[0623] Output: Progress report
[0624] What it does: After the session ends, the server stores all data in a database and generates a report to visualize the progress.
[0625] Step 10:
[0626] Users can check their progress and self-assess through a dashboard.
[0627] Input: Visualized progress data
[0628] Data calculation: Check progress data and self-evaluation
[0629] Output: User self-assessment results
[0630] Specific operation: The user uses the device's dashboard function to check their training progress and perform self-evaluation.
[0631] As a specific example, if a user selects the customer service scenario "welcoming a customer," the prompt message will be displayed: "As a virtual store clerk, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." If tension is detected in the user's facial expression, the feedback will be "Relax and smile more."
[0632] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0633] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0634] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0635] [Second embodiment]
[0636] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0637] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0638] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0639] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0640] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0641] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0642] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0643] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0644] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0645] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0646] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0647] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0648] The present invention relates to a communication skill improvement platform, which is designed to provide a set of functions for users to effectively improve their communication skills. The system is implemented as a cross-platform application and can be used through a web browser or a mobile application.
[0649] server
[0650] The server plays a central role in this system and provides the following functions:
[0651] 1. User authentication and information management
[0652] The server verifies the user's credentials when they submit a login request, and if authentication is successful, retrieves the user's past training data from a database to determine their current skill level.
[0653] 2. Providing training scenarios
[0654] The server generates appropriate training scenarios based on the user's skill level and sends the information to the user's device. The scenarios contain various communication scenes, and the user can select a scenario that suits their level.
[0655] 3. Communication Scene Generation
[0656] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[0657] 4. Real-time analysis and feedback
[0658] The server receives the user's voice data in real time and converts it into text using natural language processing (NLP) technology. The converted text data is then analyzed, and appropriate responses and feedback are generated and sent to the device.
[0659] 5. Saving and analyzing progress data
[0660] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. Later, this data is analyzed and used to provide information for visualizing the user's progress.
[0661] Terminal
[0662] The user's device works in conjunction with the server to provide the following functions:
[0663] 1. User Interface
[0664] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[0665] 2. Audio capture and transmission
[0666] The device captures the user's speech with a microphone and transmits the audio data to the server in real time.
[0667] 3. Displaying feedback and responses
[0668] Feedback and responses received from the server are displayed as voice or text, allowing the user to understand what needs to be improved and take the next action.
[0669] User
[0670] The user performs the following actions through this system:
[0671] 1. Log in and check your profile
[0672] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[0673] 2. Selecting a training scenario
[0674] The user selects a scenario that suits their level from a list of training scenarios provided.
[0675] 3. Practice in communication situations
[0676] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[0677] 4. Receive feedback and improve
[0678] See real-time feedback and act on it to improve your skills.
[0679] 5. Track your progress over the long term
[0680] Users can view progress data on a dashboard and perform self-assessments to promote their growth.
[0681] As a concrete example, consider the case where a user selects a job interview scenario. The user experiences self-introductions and question-and-answer sessions, and receives real-time feedback from the server. For example, the server may provide feedback such as, "You speak too quickly, so you should speak more slowly," or "You need to improve your explanation, including your specific work experience." By practicing based on this feedback, the user will be able to respond confidently in an actual interview.
[0682] Thus, the present invention provides users with an environment for practical and effective communication skill development, with real-time feedback and long-term progress tracking.
[0683] The processing flow will be explained below.
[0684] Step 1:
[0685] User
[0686] The user enters their account information on the login screen and clicks the login button.
[0687] server
[0688] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[0689] Step 2:
[0690] server
[0691] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[0692] Terminal
[0693] The terminal displays the list of training scenarios received from the server on the screen.
[0694] Step 3:
[0695] User
[0696] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[0697] server
[0698] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[0699] Step 4:
[0700] Terminal
[0701] The terminal displays the scene information received from the server on the screen and prepares for audio playback.
[0702] User
[0703] The user checks the scene information and clicks the button to start training.
[0704] Step 5:
[0705] Terminal
[0706] The device captures the user's speech with a microphone and transmits the voice data to the server in real time.
[0707] server
[0708] The server receives the voice data and converts it into text data via a natural language processing (NLP) module.
[0709] Step 6:
[0710] server
[0711] The server analyzes the converted text data, generates appropriate responses and feedback, converts the generated responses and feedback into audio data, and sends it to the user's device.
[0712] Terminal
[0713] The terminal plays back the audio data received from the server and provides feedback to the user.
[0714] Step 7:
[0715] User
[0716] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[0717] Step 8:
[0718] server
[0719] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. The stored data is analyzed and information is generated to visualize the progress.
[0720] User
[0721] After the session, users can view their progress data on a dashboard and self-evaluate.
[0722] Example 1
[0723] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0724] Current communication skill training systems lack the functionality to provide appropriate feedback in real time based on the user's skill level. They also have limited functionality to effectively track the user's progress and suggest specific ways to improve. This makes it difficult for existing systems to effectively improve skills.
[0725] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0726] In this invention, the server includes means for verifying user authentication information, means for determining the user's skill level based on past learning data, means for generating appropriate training scenarios based on the user's skill level, means for constructing dialogue scenes based on the training scenario selected by the user, means for analyzing the user's voice data in real time using natural language processing technology and generating feedback, and means for storing session data in a database and analyzing and visualizing the user's progress, thereby enabling the user to receive effective feedback in real time, thereby promoting self-evaluation and growth.
[0727] "Means for verifying user authentication information" refers to a means for checking the authentication information (user name and password) entered by the user and comparing it with a database to see if it is correct.
[0728] The "means for determining the user's skill level based on past learning data" refers to a means for analyzing the user's past training data and evaluation results to determine the user's current skill level.
[0729] The "means for generating an appropriate training scenario based on the skill level of the user" is a means for creating an optimal training scenario for the user in accordance with the determined skill level.
[0730] The "means for constructing a dialogue scene based on a training scenario selected by the user" refers to a means for constructing a scene including specific dialogue content and questions based on a scenario selected by the user.
[0731] "Means for analyzing user voice data in real time using natural language processing technology and generating feedback" refers to means for capturing a user's speech as voice data, converting it into text, analyzing it, and generating the necessary feedback in real time.
[0732] "Means for storing session data in a database and analyzing and visualizing the user's progress" refers to a means for storing data after a training session in a database and subsequently analyzing the data to visualize the user's progress in graphs or other formats.
[0733] The communication skill improvement system of the present invention can be effectively implemented by cooperation between the server, the terminal, and the user.
[0734] Server Roles
[0735] The server plays a central role in this system and provides the following functions:
[0736] 1. Verify user credentials
[0737] The server receives the authentication information entered by the user on the login screen and authenticates the user by comparing it with the database. The user can only use the system if authentication is successful.
[0738] 2. Determine the user's skill level
[0739] The server analyzes past learning data to determine the user's current skill level, including stored training session results and evaluations.
[0740] 3. Generating training scenarios
[0741] The server automatically generates an appropriate training scenario based on the user's skill level. The generated scenario is sent to the user's device. For example, a "basic self-introduction" scenario for beginners and a "job interview" scenario for intermediate users are generated.
[0742] 4. Creating a communication scene
[0743] Based on the training scenario selected by the user, the system constructs a specific dialogue scene and transmits the scene information to the device. The scene includes the actual utterances and questions the user will ask.
[0744] 5. Real-time analysis and feedback
[0745] The server receives the user's voice data in real time, converts it into text using natural language processing technology, and generates feedback, which is then immediately sent to the device and presented to the user.
[0746] 6. Session data storage and progress analysis
[0747] After each session, the server stores all relevant data in a database. By analyzing the stored data, the user's progress is visualized and displayed on a dashboard, allowing users to self-evaluate and track their progress.
[0748] Device Role
[0749] The user's device works in conjunction with the server to provide the following functions:
[0750] 1. Providing a user interface
[0751] The terminal allows the user to interact with the system through interfaces such as a login screen, scenario selection screen, and training start screen.
[0752] 2. Capture and transmit audio data
[0753] The device captures the user's speech with a microphone and transmits the audio data to a server in real time, allowing for immediate analysis and feedback.
[0754] 3. Viewing Feedback
[0755] The terminal displays the feedback or response received from the server in voice or text form and provides it to the user.
[0756] User Behavior
[0757] The user performs the following actions through this system:
[0758] 1. Log in and check your profile
[0759] Users enter their login information to access the system. After logging in, they can check their skill level and past training data on the dashboard.
[0760] 2. Selecting a training scenario
[0761] From the list of training scenarios provided, choose one that matches your skill level, for example, the intermediate "Job Interview" scenario.
[0762] 3. Practice in communication situations
[0763] The user will practice speaking in a dialogue scene generated based on the selected scenario. Once the user has finished speaking, they can move on to the next step.
[0764] 4. Receive feedback and improve
[0765] See real-time feedback and use it to improve your skills, for example, "You speak too fast; you should speak a little slower."
[0766] 5. Track your progress over the long term
[0767] Users can view progress data and self-assess on the dashboard, allowing them to understand their progress and set goals for further improvement.
[0768] Examples and prompts
[0769] Specific examples
[0770] If a user selects the "Job Interview" scenario for intermediate level learners, they will experience self-introductions and question-and-answer sessions, and receive real-time feedback from the server. For example, they may receive feedback such as, "You speak too quickly; you should speak a little more slowly," or "You need to improve your explanation, including your specific work experience." Users can then practice based on this feedback, gaining confidence in the actual interview.
[0771] Prompt Sentence Examples
[0772] The user has selected the job interview scenario. Please simulate the scene where the user introduces themselves and provide feedback on the following points:
[0773] 1. Speaking Speed
[0774] 2. Description of specific work experience
[0775] 3. Language and attitude
[0776] By including these specific details, the present invention provides users with an environment for practical and effective communication skill development, allowing for real-time feedback and long-term growth tracking.
[0777] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0778] Step 1:
[0779] Entering and validating user credentials
[0780] Input: The username and password the user enters on the login screen.
[0781] Specific behavior:
[0782] The user enters the username and password on the login screen of the device.
[0783] The terminal transmits the entered authentication information to the server.
[0784] The server checks the received authentication information against a database.
[0785] Data processing / calculation:
[0786] The server performs a calculation to match the username and password in its database.
[0787] output:
[0788] If the authentication is successful, the server sends a "user authentication successful" response to the terminal and starts a session.
[0789] If the authentication fails, the server sends a "user authentication failed" response to the terminal, prompting it to try again.
[0790] Step 2:
[0791] Obtaining user history data and determining skill level
[0792] Input: User ID and previous learning data request.
[0793] Specific behavior:
[0794] The server uses the user ID to load the user's past training data from a database.
[0795] Data processing / calculation:
[0796] The server analyzes the acquired past training data and performs calculations to evaluate the user's skill level.
[0797] output:
[0798] The user's current skill level information is transmitted to the terminal.
[0799] Step 3:
[0800] Training scenario generation and distribution
[0801] Input: User skill level information.
[0802] Specific behavior:
[0803] The server generates appropriate training scenarios based on skill level.
[0804] Data processing / calculation:
[0805] The server selects the most suitable scenario for the user's skill level from among multiple scenario candidates and constructs scenario data.
[0806] output:
[0807] The generated training scenario is sent to the terminal.
[0808] Step 4:
[0809] Building and sending communication scenes
[0810] Input: A training scenario selected by the user.
[0811] Specific behavior:
[0812] The user selects one of the scenarios provided on the terminal.
[0813] The terminal transmits the selected scenario information to the server.
[0814] The server constructs a specific dialogue scene based on the selected scenario.
[0815] output:
[0816] The constructed dialogue scene is sent to the device.
[0817] Step 5:
[0818] Capture and transmit audio data
[0819] Input: User speech.
[0820] Specific behavior:
[0821] The user speaks into the microphone of the terminal.
[0822] The device captures audio data through a microphone.
[0823] The device transmits the captured audio data to the server in real time.
[0824] output:
[0825] Send the audio data to the server.
[0826] Step 6:
[0827] Providing real-time analysis and feedback
[0828] Input: User's voice data sent from the device.
[0829] Specific behavior:
[0830] The server converts the received voice data into text using natural language processing technology.
[0831] The server analyzes the converted text data and generates appropriate feedback.
[0832] Data processing / calculation:
[0833] The system converts voice data into text and performs feedback generation calculations based on text analysis.
[0834] output:
[0835] The generated feedback is sent to the device.
[0836] Step 7:
[0837] Session data storage and progress visualization
[0838] Input: Data from each session (utterances, feedback, and evaluation results).
[0839] Specific behavior:
[0840] The server stores all relevant data in a database after the session ends.
[0841] Data processing / calculation:
[0842] The stored data is analyzed and calculations are performed to visualize the progress data.
[0843] output:
[0844] Progress data is visualized and presented to users via a dashboard.
[0845] Through the above steps, this system can effectively improve the user's communication skills.
[0846] (Application example 1)
[0847] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0848] In modern brick-and-mortar stores, improving staff communication skills is a high priority. High skills are especially required for dealing with new customers and handling complaints, as these skills are directly linked to customer satisfaction. However, in brick-and-mortar stores, staff have limited opportunities to effectively improve their skills while performing their daily tasks. Furthermore, traditional training methods make it difficult to train many staff members at once, and individual feedback is lacking. As a result, there is a problem of inefficient improvement of individual skills.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0850] In this invention, the server includes means for authenticating a user's login information, means for determining a user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating communication scenes based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for saving the user's progress data and tracking and analyzing the growth record, means for evaluating the user's speech speed and content and suggesting specific areas for improvement, and means for creating scenarios for staff to respond to in-store situations and conducting training. This enables store staff to effectively and efficiently improve the communication skills required in-store while performing their daily work.
[0851] The means for "authenticating user login information" is a means for verifying authentication information such as a user name and password entered when a user accesses the system, and confirming that the user is a legitimate user.
[0852] The means for "determining the user's skill level based on past training data" is a means for analyzing data on training previously performed by the user and evaluating the current level of the user's skill based on the data.
[0853] The means for "providing an appropriate training scenario based on the user's skill level" is a means for automatically selecting training according to the determined skill level and presenting it to the user.
[0854] The means for "generating a communication scene based on a training scenario selected by the user" is a means for reproducing a specific communication situation or condition based on a training scenario selected by the user.
[0855] The method of "analyzing user speech in real time and generating feedback" involves collecting the user's speech with a microphone, instantly analyzing it using natural language processing technology, and presenting appropriate suggestions and areas for improvement based on the analysis results.
[0856] The means for "storing user progress data and tracking and analyzing growth records" refers to storing the user's training details and results in a database, and analyzing them over the long term to track skill improvement.
[0857] The method of "evaluating the user's speech speed and content and suggesting specific areas for improvement" involves analyzing the speed and quality of the content of the user's speech and, based on the results, providing specific instructions to the user on how to improve.
[0858] The method of "creating scenarios for staff to respond to in-store situations and conducting training" involves creating virtual scenarios for specific situations such as customer service and handling complaints in a real store, and then having staff practice using those scenarios.
[0859] This invention is a system for improving communication skills for store staff. This system mainly consists of three components: a server, a terminal, and a user. The following is a detailed description of the role of each component and its processing.
[0860] server
[0861] The server plays a central role in the system, providing the following functions:
[0862] User authentication:
[0863] When a user logs in to a system, the server authenticates them based on their username and password, ensuring that only authorized users can access the system.
[0864] Skill Level Determination:
[0865] The server analyzes past training data to assess the user's current skill level, which is stored in a database and updated in real time.
[0866] Training scenarios provided:
[0867] Based on the user's skill level, appropriate training scenarios are selected and provided to the user, including scenarios for greeting new customers and handling complaints.
[0868] Communication Scene Generation:
[0869] Based on the scenario selected by the user, specific communication scenes are generated and sent to the device, allowing the user to receive training tailored to their actual work.
[0870] Real-time analysis and feedback:
[0871] It collects user utterances in real time, analyzes them using natural language processing (NLP) technology, and generates feedback based on the analysis results and sends it to the device.
[0872] Save and analyze progress data:
[0873] Session data is stored in a database and progress is visualized for each user, allowing users to self-assess and foster long-term growth.
[0874] Terminal
[0875] The user's terminal provides the following functions in cooperation with the server.
[0876] User Interface:
[0877] The terminal provides an interface for interacting with the user, such as a login screen, a scenario selection screen, and a training scene.
[0878] Audio Capture:
[0879] The device's microphone is used to capture the user's speech in real time and send the data to the server.
[0880] Show feedback:
[0881] The analysis results and feedback received from the server are provided to the user via text or voice, allowing the user to make immediate corrections and improve their skills.
[0882] User
[0883] The users are staff members of a physical store, and by utilizing this system, they perform the following actions:
[0884] Login:
[0885] The user logs in to the terminal and checks their account information.
[0886] Select a training scenario:
[0887] Choose from the list of scenarios provided that fit your skill level.
[0888] Communication practice:
[0889] Participants will experience and speak in a series of real-life communication situations, including greeting new customers and handling complex complaints.
[0890] Receiving feedback and making corrections:
[0891] Improve your skills with real-time feedback.
[0892] Long-term progress check:
[0893] View your training history and progress data on the dashboard and self-assess.
[0894] Hardware and software used
[0895] Hardware: Smartphones, tablets
[0896] software:
[0897] Web browser (cross-platform compatible)
[0898] Python (Flask framework)
[0899] Database (MySQL or PostgreSQL)
[0900] Natural Language Processing (NLP) libraries (spaCy and nltk)
[0901] Frontend (ReactJS)
[0902] Examples of concrete examples and prompts
[0903] For example, if the user selects a scenario for handling a complaint, the following dialogue is generated:
[0904] 1. User: "Sorry, but we can't accept your return at this time."
[0905] 2. System: Performs real-time analysis and provides feedback such as "The text is too long. Make it concise" or "Intonation is important."
[0906] As described above, by using this system, store staff can efficiently improve the communication skills required in physical stores.
[0907] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0908] Step 1:
[0909] Authenticate your login details
[0910] The server receives the login information (username and password) entered by the user. It compares this information with existing data in the database, and if it matches, it completes the authentication. Based on the information entered, it determines whether the user is legitimate, and if successful, it proceeds to the next step. This process allows access to the system.
[0911] Input: Username, Password
[0912] Output: Authentication success / failure status
[0913] Step 2:
[0914] Skill level determination
[0915] The server retrieves past training data for successfully authenticated users from the database and analyzes it. It uses Python data analysis libraries (such as pandas and NumPy) to evaluate the user's current skill level. Based on this, it selects a training scenario of the appropriate level.
[0916] Input: Past training data
[0917] Output: Skill level judgment result
[0918] Step 3:
[0919] Providing training scenarios
[0920] The server selects multiple training scenarios based on the determined skill level and sends them to the user's device. The scenarios include many scenes that can be used in real stores, and the user can choose the scenario they want to train in.
[0921] Input: Skill Level
[0922] Output: List of training scenarios provided
[0923] Step 4:
[0924] Communication Scene Generation
[0925] Based on the training scenario selected by the user, the server generates specific communication scenes, including scripts and dialogue flows, to recreate real-life store situations. The generated scenes are then sent to the device for the user to practice.
[0926] Input: The selected training scenario
[0927] Output: Communication scene details
[0928] Step 5:
[0929] Audio capture and transmission
[0930] When the user presses the start button, the device uses the built-in microphone to capture the user's speech in real time. The speech data is immediately sent to the server and prepared for analysis, particularly for speech recognition and natural language processing (NLP), where it is converted into an appropriate data format (e.g., text).
[0931] Input: User utterance
[0932] Output: Sends audio data to the server
[0933] Step 6:
[0934] Real-time analysis and feedback generation
[0935] The server analyzes the received voice data using a natural language processing (NLP) engine. It analyzes the text of the speech and prosody (speech characteristics) such as speed and intonation, and generates appropriate feedback. The generated feedback is sent to the user's device in real time.
[0936] Input: Audio data
[0937] Output: Feedback
[0938] Step 7:
[0939] View Feedback
[0940] The device receives feedback from the server and displays it to the user in text or audio, allowing the user to review the feedback and learn how to improve their skills.
[0941] Input: Feedback data
[0942] Output: Display feedback
[0943] Step 8:
[0944] Save and analyze progress data
[0945] At the end of the session, the device sends the session data (such as what was said and feedback) to the server, which stores the data in a database and visualizes each user's progress. This data is analyzed to provide insights to drive long-term growth.
[0946] Input: Session data
[0947] Output: Progress data visualization and analysis results
[0948] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0949] This invention is a system that further enhances the effectiveness of training by incorporating an emotion engine into a platform for users to effectively improve their communication skills. The system has the function of recognizing the user's emotions in real time and adjusting feedback and training scenarios based on that information.
[0950] server
[0951] The server plays a central role in this system and provides the following functions:
[0952] 1. User authentication and information management
[0953] The server verifies the user's credentials when they submit a login request, and if authentication is successful, determines the user's past training data and skill level.
[0954] 2. Providing training scenarios
[0955] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[0956] 3. Communication Scene Generation
[0957] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[0958] 4. Real-time analysis and feedback
[0959] The server receives and analyzes the user's voice and facial expression data in real time, and then uses a natural language processing (NLP) module to convert the voice data into text and generate appropriate responses and feedback.
[0960] 5. Emotion recognition and feedback regulation
[0961] The server uses an emotion engine to analyze the user's emotional state and adjusts the feedback accordingly. For example, if the user is nervous, it provides feedback to calm them down.
[0962] 6. Saving and analyzing progress data
[0963] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The server analyzes the stored data and provides users with information that visualizes their progress.
[0964] Terminal
[0965] The user's device works in conjunction with the server to provide the following functions:
[0966] 1. User Interface
[0967] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[0968] 2. Capture and transmit voice and facial expressions
[0969] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the data to a server in real time.
[0970] 3. Displaying feedback and responses
[0971] It displays feedback and responses received from the server in audio or text format, and also displays emotion-based feedback.
[0972] User
[0973] The user performs the following actions through this system:
[0974] 1. Log in and check your profile
[0975] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[0976] 2. Selecting a training scenario
[0977] The user selects the scenario that suits them from a list of training scenarios provided.
[0978] 3. Practice in communication situations
[0979] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[0980] 4. Receiving and improving feedback based on emotion recognition
[0981] See and improve your skills based on real-time feedback, including emotional feedback recognized by the emotion engine.
[0982] 5. Track your progress over the long term
[0983] Users can view their progress and self-assess on a dashboard, which also includes a record of their emotions, allowing them to see their mental progress.
[0984] As a concrete example, consider the case where a user selects a job interview scenario and experiences self-introduction and question-and-answer sessions. If the emotion engine recognizes that the user is nervous, the server will provide feedback such as "Take a deep breath to relieve tension." It can also provide specific advice such as "You're speaking too fast, so speak a little more slowly." By practicing based on this feedback, the user can improve their actual performance in the interview.
[0985] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[0986] The processing flow will be explained below.
[0987] Step 1:
[0988] User
[0989] The user enters their account information on the login screen and clicks the login button.
[0990] server
[0991] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[0992] Step 2:
[0993] server
[0994] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[0995] Terminal
[0996] The terminal displays the list of training scenarios received from the server on the screen.
[0997] Step 3:
[0998] User
[0999] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[1000] server
[1001] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[1002] Step 4:
[1003] Terminal
[1004] The terminal displays the scene information received from the server on the screen and prepares to play audio and capture facial expressions.
[1005] User
[1006] The user checks the scene information and clicks the button to start training.
[1007] Step 5:
[1008] Terminal
[1009] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the audio and video data to a server in real time.
[1010] server
[1011] The server receives the voice data and facial expression data and converts the voice data into text via a natural language processing (NLP) module.
[1012] Step 6:
[1013] server
[1014] The server analyzes the converted text data and facial expression data and uses an emotion engine to recognize the user's emotional state. Based on the analysis results, it generates appropriate responses and feedback. The generated responses and feedback are converted into voice data and sent to the user's device.
[1015] Terminal
[1016] The terminal plays back the voice data and emotional feedback received from the server and provides them to the user.
[1017] Step 7:
[1018] User
[1019] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[1020] Step 8:
[1021] server
[1022] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The stored data is analyzed and information is generated to visualize the progress.
[1023] User
[1024] After the session, users can check their progress data on the dashboard and self-evaluate. They can also review their progress, including changes in their emotions, and use this information to improve their next training session.
[1025] Example 2
[1026] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1027] Conventional communication skill improvement systems struggle to accurately recognize a user's emotional state and provide feedback based on that, making effective training difficult. Furthermore, they often lack real-time speech analysis and feedback generation, which often results in ineffective improvement of the user's skills. Furthermore, they lack sufficient visualization of progress data and long-term growth records, leaving users with a lack of reference information for self-evaluation.
[1028] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for authenticating the user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on the training scenario selected by the user, means for analyzing the user's voice data and facial expression data in real time and generating feedback, means for analyzing the user's emotional state and adjusting the feedback based on the results, and means for saving the user's progress data and tracking and analyzing the growth record. This makes it possible to analyze the user's emotional state in real time and provide appropriate feedback based on the analysis results. Furthermore, by visualizing the progress data and providing a long-term growth record, it is possible to enhance the reference information for the user when performing self-evaluation.
[1029] "User Login Information" means the authentication data (e.g., user ID and password) provided by a User to access a System.
[1030] "Training data" refers to historical information such as the results and evaluations of past training sessions conducted by the user, as well as the content of utterances.
[1031] "Skill level" is an index that evaluates a user's communication ability and proficiency.
[1032] A "training scenario" is an exercise designed to allow a user to practice in a specific situation.
[1033] A "communication scene" is a specific situation or setting that a user simulates, which is generated based on a training scenario.
[1034] "Voice data" refers to data that digitally captures a user's speech.
[1035] "Facial Expression Data" refers to data that captures a user's facial expressions and represents them in digital form.
[1036] "Feedback" refers to the evaluation and advice provided by the system in response to the user's actions and utterances.
[1037] An "emotional state" is the emotion (e.g., tension, joy, anger, etc.) that a user is feeling at a particular moment.
[1038] "Analysis" is the process of processing collected data to extract useful information and patterns.
[1039] "Progress Data" means the performance and evaluation recorded for each of a User's training sessions.
[1040] A "growth record" is information that records the history and results of a user's training sessions over a long period of time and shows the trajectory of their growth.
[1041] "Visualization" is the process of transforming and displaying complex data into visual formats such as graphs and charts.
[1042] The present invention is a training system for improving a user's communication skills. This system recognizes the user's emotions in real time and provides feedback based on the results, thereby achieving more effective training.
[1043] System Configuration
[1044] server
[1045] The server plays a central role in the system and includes the following hardware and software components:
[1046] Authentication system: Validates user login information and performs authentication. The software used is a common authentication framework.
[1047] Database Management System (DBMS): Stores and manages users' past training data and skill levels.
[1048] Training scenario generation module: Selects an appropriate training scenario based on the user's skill level.
[1049] Communication scene generation module: Generates specific communication scenes based on the scenario selected by the user.
[1050] Real-time analysis engine: Using an NLP module, the user's voice is converted into text, and the content is then analyzed to generate an appropriate response.
[1051] Emotion Engine: Analyzes the user's facial expression data and evaluates the user's emotional state.
[1052] Feedback generation module: Generates feedback for the user based on the results of real-time analysis and sentiment analysis.
[1053] Progress Data Storage and Analysis Module: Saves and visualizes progress data after a user's training session.
[1054] Terminal
[1055] The user's device works in conjunction with the server to provide the following functions:
[1056] User Interface: Provides interfaces for users to interact with the system, such as a login screen, scenario selection screen, and training screen.
[1057] Voice capture device: Uses a microphone to capture the user's speech and sends it to the server as voice data.
[1058] Facial Expression Capture Device: Uses a camera to capture the user's facial expressions and transmits the data to a server.
[1059] Feedback display device: Displays feedback or responses received from the server in audio or text format.
[1060] User
[1061] The user performs the following actions:
[1062] 1. Log in and check your profile: Log in to the system using your device and check your skill level and past training data on the dashboard.
[1063] 2. Select a training scenario: Select an appropriate scenario from the list of training scenarios provided by the server.
[1064] 3. Practice in communication scenes: Practice your speech and actions in scenes generated based on the scenario you select.
[1065] 4. Receive real-time feedback: See feedback displayed on your device and improve your training.
[1066] 5. Track your progress: Check your progress data on the dashboard and self-assess.
[1067] Specific examples
[1068] As a concrete example, consider the case where a user selects a "job interview" scenario. The user actually introduces themselves and engages in a question-and-answer session with an AI acting as an interviewer. The device's camera and microphone are used to capture the user's facial expressions and voice. The server analyzes this data in real time and provides feedback such as "You seem nervous. Take a deep breath." It also displays specific advice such as "You speak too quickly. Please speak more slowly."
[1069] Prompt Sentence Examples
[1070] Prompt: Start a practice job interview scenario. First, introduce yourself. Then move on to questions and answers. If you're nervous, the system will provide feedback. It will also give you specific advice on how to speak.
[1071] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[1072] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1073] Step 1:
[1074] Enter your login information
[1075] Input: The user enters their user ID and password on the login screen.
[1076] Operation: The device sends the authentication information entered by the user to the server.
[1077] Output: The server receives the authentication information.
[1078] Step 2:
[1079] User authentication and information management
[1080] Input: The server checks the user's login information against a database.
[1081] How it works: After successful authentication, the server retrieves past training data and skill level from the database.
[1082] Output: Send the acquired training data and skill level to the device.
[1083] Step 3:
[1084] Providing training scenarios
[1085] Input: The server selects appropriate training scenarios based on the user's skill level.
[1086] Operation: The server generates a list of selected training scenarios and sends it to the terminal.
[1087] Output: A list of scenarios will be displayed on the terminal.
[1088] Step 4:
[1089] Scenario Selection
[1090] Input: The user selects one from the list of scenarios provided.
[1091] Operation: The terminal transmits the selected scenario information to the server.
[1092] Output: The server receives the scenario selection information.
[1093] Step 5:
[1094] Communication Scene Generation
[1095] Input: The server generates a scene based on the selected scenario.
[1096] Operation: The server creates detailed communication scenes based on the scenario and sends the information to the terminal.
[1097] Output: A specific training scene is displayed on the device.
[1098] Step 6:
[1099] Voice and facial expression capture
[1100] Input: The user makes utterances and expressions in the simulation.
[1101] How it works: The device uses a microphone and camera to capture the user's voice and facial expression data in real time.
[1102] Output: The captured data is sent to the server.
[1103] Step 7:
[1104] Real-time analytics
[1105] Input: The server receives the voice data and facial expression data sent from the device.
[1106] How it works: The server uses an NLP module to convert voice data into text and analyze its content, as well as analyze facial expression data to assess emotional state.
[1107] Output: Text data and emotion evaluation results are obtained as the analysis results.
[1108] Step 8:
[1109] Feedback Generation
[1110] Input: The server generates feedback based on the real-time analysis results and emotion evaluation results.
[1111] How it works: The server generates and adjusts the feedback based on the user's emotional state. For example, if the user is nervous, the server generates the feedback "Take a deep breath."
[1112] Output: The generated feedback information is sent to the terminal.
[1113] Step 9:
[1114] View Feedback
[1115] Input: The terminal obtains the feedback information received from the server.
[1116] Action: The device displays feedback in the form of audio or text.
[1117] Output: User sees feedback.
[1118] Step 10:
[1119] Save and visualize progress data
[1120] Input: The server receives all the data from the training session.
[1121] How it works: After a session ends, the server stores progress data in a database and analyzes it to generate visualizations, such as graphs and charts.
[1122] Output: Visualized progress information is sent to the device and made available to the user on a dashboard.
[1123] Through this series of processes, users receive real-time feedback to improve their communication skills, enabling effective training.
[1124] (Application example 2)
[1125] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1126] Providing effective customer service skill training is crucial in today's virtual environments. However, conventional training systems struggle to provide users with real-time feedback or feedback based on emotion recognition, resulting in delayed or insufficient improvement. They also lack a means to effectively track and visualize user progress. It is necessary to provide a system that can solve these issues and efficiently support the improvement of customer service skills.
[1127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1128] In this invention, the server includes means for authenticating a user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for capturing the user's facial expressions and recognizing their emotions, means for providing feedback based on the recognized emotions and adjusting the training scenario, means for saving the user's progress data and tracking and analyzing the growth record, means for the user to select from a plurality of customer service scenarios, and means for presenting prompt sentences for the user to train their customer service skills. This enables the user to efficiently improve their customer service skills while receiving feedback based on emotion recognition in real time.
[1129] "User Login Information" means the authentication information for a User to access the System.
[1130] "Past Training Data" means records and performance of a User's previous training sessions.
[1131] "Skill level" refers to a user's ability or proficiency in a particular skill.
[1132] A "training scenario" is a simulation or challenge that allows a user to practice a particular skill.
[1133] A "communication scene" is a virtual interaction or situation that a user experiences during training.
[1134] "User utterance" refers to the sounds and language that a user makes during training.
[1135] "Feedback" refers to evaluation of the user's actions and utterances and advice for improvement.
[1136] "Facial Expression" refers to the facial expression of the user.
[1137] "Means for recognizing emotions" refers to techniques or devices for analyzing a user's emotional state.
[1138] "Progress Data" means a record showing how much a User has improved through training.
[1139] A "customer service scenario" is a hypothetical situation that users use to train their customer service skills.
[1140] A "prompt" is an instruction or question that prompts the user to take a specific action or speak a certain word.
[1141] This invention is a system for efficiently training customer service skills in a virtual store. The system consists of three main elements: a server, a terminal, and a user.
[1142] server
[1143] The server plays a central role in this system and provides the following functions:
[1144] 1. User authentication and information management
[1145] The server verifies the user's credentials when they submit a login request, and if the login is successful, determines the user's skill level based on past training data.
[1146] 2. Providing training scenarios
[1147] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[1148] 3. Communication Scene Generation
[1149] When a user selects a specific training scenario, the server constructs a specific communication scene based on that scenario and transmits the information to the terminal.
[1150] 4. Real-time analysis and feedback
[1151] The server receives and analyzes the user's speech data and facial expression data in real time using OpenCV (for capturing facial expressions), Keras (for loading emotion recognition models and analyzing emotions), and NLP modules (for natural language processing).
[1152] 5. Emotion recognition and feedback regulation
[1153] The server analyzes the user's emotions using an emotion engine, provides feedback based on the results, and adjusts the training scenario. For example, if the server determines that the user is tense, it provides feedback encouraging the user to relax.
[1154] 6. Saving and analyzing progress data
[1155] After the session ends, the server stores all the user's session data in a database, analyzes the stored data, and provides the user with a visualization of their progress.
[1156] Terminal
[1157] The user's device works in conjunction with the server to provide the following functions:
[1158] 1. User Interface
[1159] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[1160] 2. Capture and transmit voice and facial expressions
[1161] The device captures the user's speech using a voice input device, recognizes facial expressions using a camera, and transmits the data to a server in real time.
[1162] 3. Displaying feedback and responses
[1163] The device displays feedback and responses received from the server in audio or text format, as well as emotion-based feedback.
[1164] User
[1165] Through this system, users can carry out the following training activities:
[1166] 1. Log in and check your profile
[1167] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[1168] 2. Selecting a training scenario
[1169] The user selects the scenario that suits them from a list of training scenarios provided.
[1170] 3. Practice in communication situations
[1171] The user trains by actually speaking in a communication scene generated based on the selected scenario.
[1172] 4. Receiving and improving feedback based on emotion recognition
[1173] See real-time feedback and act on it to improve your skills.
[1174] 5. Track your progress over the long term
[1175] Users can view their progress and self-assess on a dashboard, which also includes tracking their emotions and showing their mental progress.
[1176] Specific examples
[1177] For example, if a user selects the customer service scenario "Welcome Customers," the prompt message will be displayed: "As a salesperson in a virtual store, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." The emotion recognition engine will detect tension in the user's facial expressions, and the server will provide feedback such as "Relax and smile more." In this way, the user can improve their customer service skills while receiving specific advice.
[1178] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1179] Step 1:
[1180] The server receives the user's login information and performs authentication.
[1181] Input: User login information (user ID and password)
[1182] Data processing: The server checks the data against the authentication database
[1183] Output: Authentication result (success or failure)
[1184] Specific operation: The server checks the login information against the authentication database, and if authentication is successful, retrieves the user's past training data.
[1185] Step 2:
[1186] The server determines the user's skill level based on past training data.
[1187] Input: Past training data
[1188] Data processing: Analysis of training data and assessment of skill level
[1189] Output: User's skill level
[1190] Specific operation: The server analyzes past training data and applies a skill assessment algorithm to calculate the user's skill level.
[1191] Step 3:
[1192] The server generates an appropriate training scenario based on the user's skill level and transmits it to the terminal.
[1193] Input: User's skill level
[1194] Data processing: Selection of appropriate scenarios from the scenario database
[1195] Output: List of training scenarios
[1196] Specific operation: The server extracts training scenarios that match the user's skill level from the scenario database and sends the list to the terminal.
[1197] Step 4:
[1198] The user selects a training scenario on the terminal.
[1199] Input: List of training scenarios
[1200] Data Calculation: Scenario Selection
[1201] Output: Selected training scenario
[1202] Specific operation: The user selects the desired scenario from the list of training scenarios displayed on the terminal.
[1203] Step 5:
[1204] The server generates a communication scene based on the selected scenario and transmits it to the terminal.
[1205] Input: Selected training scenario
[1206] Data processing: Generation of communication scenes based on scenarios
[1207] Output: Communication scene data
[1208] Specific operation: The server constructs a virtual communication scene based on the selected scenario and sends the data to the terminal.
[1209] Step 6:
[1210] The device captures the user's voice and facial expressions and transmits them to the server.
[1211] Input: User utterances and facial expressions
[1212] Data processing: Audio input device and camera capture
[1213] Output: Capture data
[1214] Specific operation: The device captures the user's voice with a microphone, recognizes facial expressions with a camera, and sends the data to a server.
[1215] Step 7:
[1216] The server analyzes the user's speech data and facial expression data in real time, generates feedback, and sends it to the device.
[1217] Input: User speech data and facial expression data
[1218] Data processing: speech recognition, natural language processing, sentiment analysis
[1219] Output: Real-time feedback
[1220] Specific operation: The server analyzes the user's speech using speech recognition and natural language processing, and analyzes the facial expressions using an emotion engine. Based on the obtained information, it generates feedback and sends it to the device.
[1221] Step 8:
[1222] The terminal presents the feedback received from the server to the user in real time.
[1223] Input: Feedback data
[1224] Data Calculation: Feedback Display
[1225] Output: Feedback that is displayed to the user
[1226] Specific operation: The device displays the feedback received from the server to the user in voice or text.
[1227] Step 9:
[1228] The server stores the progress data in a database and generates visualization data.
[1229] Input: Training session data
[1230] Data processing: saving to database, analyzing and visualizing progress data
[1231] Output: Progress report
[1232] What it does: After the session ends, the server stores all data in a database and generates a report to visualize the progress.
[1233] Step 10:
[1234] Users can check their progress and self-assess through a dashboard.
[1235] Input: Visualized progress data
[1236] Data calculation: Check progress data and self-evaluation
[1237] Output: User self-assessment results
[1238] Specific operation: The user uses the device's dashboard function to check their training progress and perform self-evaluation.
[1239] As a specific example, if a user selects the customer service scenario "welcoming a customer," the prompt message will be displayed: "As a virtual store clerk, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." If tension is detected in the user's facial expression, the feedback will be "Relax and smile more."
[1240] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1241] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1242] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1243] [Third embodiment]
[1244] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1245] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1246] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1247] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1248] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1249] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1250] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1251] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1252] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1253] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1254] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1255] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1256] The present invention relates to a communication skill improvement platform, which is designed to provide a set of functions for users to effectively improve their communication skills. The system is implemented as a cross-platform application and can be used through a web browser or a mobile application.
[1257] server
[1258] The server plays a central role in this system and provides the following functions:
[1259] 1. User authentication and information management
[1260] The server verifies the user's credentials when they submit a login request, and if authentication is successful, retrieves the user's past training data from a database to determine their current skill level.
[1261] 2. Providing training scenarios
[1262] The server generates appropriate training scenarios based on the user's skill level and sends the information to the user's device. The scenarios contain various communication scenes, and the user can select a scenario that suits their level.
[1263] 3. Communication Scene Generation
[1264] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[1265] 4. Real-time analysis and feedback
[1266] The server receives the user's voice data in real time and converts it into text using natural language processing (NLP) technology. The converted text data is then analyzed, and appropriate responses and feedback are generated and sent to the device.
[1267] 5. Saving and analyzing progress data
[1268] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. Later, this data is analyzed and used to provide information for visualizing the user's progress.
[1269] Terminal
[1270] The user's device works in conjunction with the server to provide the following functions:
[1271] 1. User Interface
[1272] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[1273] 2. Audio capture and transmission
[1274] The device captures the user's speech with a microphone and transmits the audio data to the server in real time.
[1275] 3. Displaying feedback and responses
[1276] Feedback and responses received from the server are displayed as voice or text, allowing the user to understand what needs to be improved and take the next action.
[1277] User
[1278] The user performs the following actions through this system:
[1279] 1. Log in and check your profile
[1280] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[1281] 2. Selecting a training scenario
[1282] The user selects a scenario that suits their level from a list of training scenarios provided.
[1283] 3. Practice in communication situations
[1284] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[1285] 4. Receive feedback and improve
[1286] See real-time feedback and act on it to improve your skills.
[1287] 5. Track your progress over the long term
[1288] Users can view progress data on a dashboard and perform self-assessments to promote their growth.
[1289] As a concrete example, consider the case where a user selects a job interview scenario. The user experiences self-introductions and question-and-answer sessions, and receives real-time feedback from the server. For example, the server may provide feedback such as, "You speak too quickly, so you should speak more slowly," or "You need to improve your explanation, including your specific work experience." By practicing based on this feedback, the user will be able to respond confidently in an actual interview.
[1290] Thus, the present invention provides users with an environment for practical and effective communication skill development, with real-time feedback and long-term progress tracking.
[1291] The processing flow will be explained below.
[1292] Step 1:
[1293] User
[1294] The user enters their account information on the login screen and clicks the login button.
[1295] server
[1296] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[1297] Step 2:
[1298] server
[1299] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[1300] Terminal
[1301] The terminal displays the list of training scenarios received from the server on the screen.
[1302] Step 3:
[1303] User
[1304] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[1305] server
[1306] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[1307] Step 4:
[1308] Terminal
[1309] The terminal displays the scene information received from the server on the screen and prepares for audio playback.
[1310] User
[1311] The user checks the scene information and clicks the button to start training.
[1312] Step 5:
[1313] Terminal
[1314] The device captures the user's speech with a microphone and transmits the voice data to the server in real time.
[1315] server
[1316] The server receives the voice data and converts it into text data via a natural language processing (NLP) module.
[1317] Step 6:
[1318] server
[1319] The server analyzes the converted text data, generates appropriate responses and feedback, converts the generated responses and feedback into audio data, and sends it to the user's device.
[1320] Terminal
[1321] The terminal plays back the audio data received from the server and provides feedback to the user.
[1322] Step 7:
[1323] User
[1324] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[1325] Step 8:
[1326] server
[1327] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. The stored data is analyzed and information is generated to visualize the progress.
[1328] User
[1329] After the session, users can view their progress data on a dashboard and self-evaluate.
[1330] Example 1
[1331] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1332] Current communication skill training systems lack the functionality to provide appropriate feedback in real time based on the user's skill level. They also have limited functionality to effectively track the user's progress and suggest specific ways to improve. This makes it difficult for existing systems to effectively improve skills.
[1333] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1334] In this invention, the server includes means for verifying user authentication information, means for determining the user's skill level based on past learning data, means for generating appropriate training scenarios based on the user's skill level, means for constructing dialogue scenes based on the training scenario selected by the user, means for analyzing the user's voice data in real time using natural language processing technology and generating feedback, and means for storing session data in a database and analyzing and visualizing the user's progress, thereby enabling the user to receive effective feedback in real time, thereby promoting self-evaluation and growth.
[1335] "Means for verifying user authentication information" refers to a means for checking the authentication information (user name and password) entered by the user and comparing it with a database to see if it is correct.
[1336] The "means for determining the user's skill level based on past learning data" refers to a means for analyzing the user's past training data and evaluation results to determine the user's current skill level.
[1337] The "means for generating an appropriate training scenario based on the skill level of the user" is a means for creating an optimal training scenario for the user in accordance with the determined skill level.
[1338] The "means for constructing a dialogue scene based on a training scenario selected by the user" refers to a means for constructing a scene including specific dialogue content and questions based on a scenario selected by the user.
[1339] "Means for analyzing user voice data in real time using natural language processing technology and generating feedback" refers to means for capturing a user's speech as voice data, converting it into text, analyzing it, and generating the necessary feedback in real time.
[1340] "Means for storing session data in a database and analyzing and visualizing the user's progress" refers to a means for storing data after a training session in a database and subsequently analyzing the data to visualize the user's progress in graphs or other formats.
[1341] The communication skill improvement system of the present invention can be effectively implemented by cooperation between the server, the terminal, and the user.
[1342] Server Roles
[1343] The server plays a central role in this system and provides the following functions:
[1344] 1. Verify user credentials
[1345] The server receives the authentication information entered by the user on the login screen and authenticates the user by comparing it with the database. The user can only use the system if authentication is successful.
[1346] 2. Determine the user's skill level
[1347] The server analyzes past learning data to determine the user's current skill level, including stored training session results and evaluations.
[1348] 3. Generating training scenarios
[1349] The server automatically generates an appropriate training scenario based on the user's skill level. The generated scenario is sent to the user's device. For example, a "basic self-introduction" scenario for beginners and a "job interview" scenario for intermediate users are generated.
[1350] 4. Creating a communication scene
[1351] Based on the training scenario selected by the user, the system constructs a specific dialogue scene and transmits the scene information to the device. The scene includes the actual utterances and questions the user will ask.
[1352] 5. Real-time analysis and feedback
[1353] The server receives the user's voice data in real time, converts it into text using natural language processing technology, and generates feedback, which is then immediately sent to the device and presented to the user.
[1354] 6. Session data storage and progress analysis
[1355] After each session, the server stores all relevant data in a database. By analyzing the stored data, the user's progress is visualized and displayed on a dashboard, allowing users to self-evaluate and track their progress.
[1356] Device Role
[1357] The user's device works in conjunction with the server to provide the following functions:
[1358] 1. Providing a user interface
[1359] The terminal allows the user to interact with the system through interfaces such as a login screen, scenario selection screen, and training start screen.
[1360] 2. Capture and transmit audio data
[1361] The device captures the user's speech with a microphone and transmits the audio data to a server in real time, allowing for immediate analysis and feedback.
[1362] 3. Viewing Feedback
[1363] The terminal displays the feedback or response received from the server in voice or text form and provides it to the user.
[1364] User Behavior
[1365] The user performs the following actions through this system:
[1366] 1. Log in and check your profile
[1367] Users enter their login information to access the system. After logging in, they can check their skill level and past training data on the dashboard.
[1368] 2. Selecting a training scenario
[1369] From the list of training scenarios provided, choose one that matches your skill level, for example, the intermediate "Job Interview" scenario.
[1370] 3. Practice in communication situations
[1371] The user will practice speaking in a dialogue scene generated based on the selected scenario. Once the user has finished speaking, they can move on to the next step.
[1372] 4. Receive feedback and improve
[1373] See real-time feedback and use it to improve your skills, for example, "You speak too fast; you should speak a little slower."
[1374] 5. Track your progress over the long term
[1375] Users can view progress data and self-assess on the dashboard, allowing them to understand their progress and set goals for further improvement.
[1376] Examples and prompts
[1377] Specific examples
[1378] If a user selects the "Job Interview" scenario for intermediate level learners, they will experience self-introductions and question-and-answer sessions, and receive real-time feedback from the server. For example, they may receive feedback such as, "You speak too quickly; you should speak a little more slowly," or "You need to improve your explanation, including your specific work experience." Users can then practice based on this feedback, gaining confidence in the actual interview.
[1379] Prompt Sentence Examples
[1380] The user has selected the job interview scenario. Please simulate the scene where the user introduces themselves and provide feedback on the following points:
[1381] 1. Speaking Speed
[1382] 2. Description of specific work experience
[1383] 3. Language and attitude
[1384] By including these specific details, the present invention provides users with an environment for practical and effective communication skill development, allowing for real-time feedback and long-term growth tracking.
[1385] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1386] Step 1:
[1387] Entering and validating user credentials
[1388] Input: The username and password the user enters on the login screen.
[1389] Specific behavior:
[1390] The user enters the username and password on the login screen of the device.
[1391] The terminal transmits the entered authentication information to the server.
[1392] The server checks the received authentication information against a database.
[1393] Data processing / calculation:
[1394] The server performs a calculation to match the username and password in its database.
[1395] output:
[1396] If the authentication is successful, the server sends a "user authentication successful" response to the terminal and starts a session.
[1397] If the authentication fails, the server sends a "user authentication failed" response to the terminal, prompting it to try again.
[1398] Step 2:
[1399] Obtaining user history data and determining skill level
[1400] Input: User ID and previous learning data request.
[1401] Specific behavior:
[1402] The server uses the user ID to load the user's past training data from a database.
[1403] Data processing / calculation:
[1404] The server analyzes the acquired past training data and performs calculations to evaluate the user's skill level.
[1405] output:
[1406] The user's current skill level information is transmitted to the terminal.
[1407] Step 3:
[1408] Training scenario generation and distribution
[1409] Input: User skill level information.
[1410] Specific behavior:
[1411] The server generates appropriate training scenarios based on skill level.
[1412] Data processing / calculation:
[1413] The server selects the most suitable scenario for the user's skill level from among multiple scenario candidates and constructs scenario data.
[1414] output:
[1415] The generated training scenario is sent to the terminal.
[1416] Step 4:
[1417] Building and sending communication scenes
[1418] Input: A training scenario selected by the user.
[1419] Specific behavior:
[1420] The user selects one of the scenarios provided on the terminal.
[1421] The terminal transmits the selected scenario information to the server.
[1422] The server constructs a specific dialogue scene based on the selected scenario.
[1423] output:
[1424] The constructed dialogue scene is sent to the device.
[1425] Step 5:
[1426] Capture and transmit audio data
[1427] Input: User speech.
[1428] Specific behavior:
[1429] The user speaks into the microphone of the terminal.
[1430] The device captures audio data through a microphone.
[1431] The device transmits the captured audio data to the server in real time.
[1432] output:
[1433] Send the audio data to the server.
[1434] Step 6:
[1435] Providing real-time analysis and feedback
[1436] Input: User's voice data sent from the device.
[1437] Specific behavior:
[1438] The server converts the received voice data into text using natural language processing technology.
[1439] The server analyzes the converted text data and generates appropriate feedback.
[1440] Data processing / calculation:
[1441] The system converts voice data into text and performs feedback generation calculations based on text analysis.
[1442] output:
[1443] The generated feedback is sent to the device.
[1444] Step 7:
[1445] Session data storage and progress visualization
[1446] Input: Data from each session (utterances, feedback, and evaluation results).
[1447] Specific behavior:
[1448] The server stores all relevant data in a database after the session ends.
[1449] Data processing / calculation:
[1450] The stored data is analyzed and calculations are performed to visualize the progress data.
[1451] output:
[1452] Progress data is visualized and presented to users via a dashboard.
[1453] Through the above steps, this system can effectively improve the user's communication skills.
[1454] (Application example 1)
[1455] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1456] In modern brick-and-mortar stores, improving staff communication skills is a high priority. High skills are especially required for dealing with new customers and handling complaints, as these skills are directly linked to customer satisfaction. However, in brick-and-mortar stores, staff have limited opportunities to effectively improve their skills while performing their daily tasks. Furthermore, traditional training methods make it difficult to train many staff members at once, and individual feedback is lacking. As a result, there is a problem of inefficient improvement of individual skills.
[1457] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1458] In this invention, the server includes means for authenticating a user's login information, means for determining a user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating communication scenes based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for saving the user's progress data and tracking and analyzing the growth record, means for evaluating the user's speech speed and content and suggesting specific areas for improvement, and means for creating scenarios for staff to respond to in-store situations and conducting training. This enables store staff to effectively and efficiently improve the communication skills required in-store while performing their daily work.
[1459] The means for "authenticating user login information" is a means for verifying authentication information such as a user name and password entered when a user accesses the system, and confirming that the user is a legitimate user.
[1460] The means for "determining the user's skill level based on past training data" is a means for analyzing data on training previously performed by the user and evaluating the current level of the user's skill based on the data.
[1461] The means for "providing an appropriate training scenario based on the user's skill level" is a means for automatically selecting training according to the determined skill level and presenting it to the user.
[1462] The means for "generating a communication scene based on a training scenario selected by the user" is a means for reproducing a specific communication situation or condition based on a training scenario selected by the user.
[1463] The method of "analyzing user speech in real time and generating feedback" involves collecting the user's speech with a microphone, instantly analyzing it using natural language processing technology, and presenting appropriate suggestions and areas for improvement based on the analysis results.
[1464] The means for "storing user progress data and tracking and analyzing growth records" refers to storing the user's training details and results in a database, and analyzing them over the long term to track skill improvement.
[1465] The method of "evaluating the user's speech speed and content and suggesting specific areas for improvement" involves analyzing the speed and quality of the content of the user's speech and, based on the results, providing specific instructions to the user on how to improve.
[1466] The method of "creating scenarios for staff to respond to in-store situations and conducting training" involves creating virtual scenarios for specific situations such as customer service and handling complaints in a real store, and then having staff practice using those scenarios.
[1467] This invention is a system for improving communication skills for store staff. This system mainly consists of three components: a server, a terminal, and a user. The following is a detailed description of the role of each component and its processing.
[1468] server
[1469] The server plays a central role in the system, providing the following functions:
[1470] User authentication:
[1471] When a user logs in to a system, the server authenticates them based on their username and password, ensuring that only authorized users can access the system.
[1472] Skill Level Determination:
[1473] The server analyzes past training data to assess the user's current skill level, which is stored in a database and updated in real time.
[1474] Training scenarios provided:
[1475] Based on the user's skill level, appropriate training scenarios are selected and provided to the user, including scenarios for greeting new customers and handling complaints.
[1476] Communication Scene Generation:
[1477] Based on the scenario selected by the user, specific communication scenes are generated and sent to the device, allowing the user to receive training tailored to their actual work.
[1478] Real-time analysis and feedback:
[1479] It collects user utterances in real time, analyzes them using natural language processing (NLP) technology, and generates feedback based on the analysis results and sends it to the device.
[1480] Save and analyze progress data:
[1481] Session data is stored in a database and progress is visualized for each user, allowing users to self-assess and foster long-term growth.
[1482] Terminal
[1483] The user's terminal provides the following functions in cooperation with the server.
[1484] User Interface:
[1485] The terminal provides an interface for interacting with the user, such as a login screen, a scenario selection screen, and a training scene.
[1486] Audio Capture:
[1487] The device's microphone is used to capture the user's speech in real time and send the data to the server.
[1488] Show feedback:
[1489] The analysis results and feedback received from the server are provided to the user via text or voice, allowing the user to make immediate corrections and improve their skills.
[1490] User
[1491] The users are staff members of a physical store, and by utilizing this system, they perform the following actions:
[1492] Login:
[1493] The user logs in to the terminal and checks their account information.
[1494] Select a training scenario:
[1495] Choose from the list of scenarios provided that fit your skill level.
[1496] Communication practice:
[1497] Participants will experience and speak in a series of real-life communication situations, including greeting new customers and handling complex complaints.
[1498] Receiving feedback and making corrections:
[1499] Improve your skills with real-time feedback.
[1500] Long-term progress check:
[1501] View your training history and progress data on the dashboard and self-assess.
[1502] Hardware and software used
[1503] Hardware: Smartphones, tablets
[1504] software:
[1505] Web browser (cross-platform compatible)
[1506] Python (Flask framework)
[1507] Database (MySQL or PostgreSQL)
[1508] Natural Language Processing (NLP) libraries (spaCy and nltk)
[1509] Frontend (ReactJS)
[1510] Examples of concrete examples and prompts
[1511] For example, if the user selects a scenario for handling a complaint, the following dialogue is generated:
[1512] 1. User: "Sorry, but we can't accept your return at this time."
[1513] 2. System: Performs real-time analysis and provides feedback such as "The text is too long. Make it concise" or "Intonation is important."
[1514] As described above, by using this system, store staff can efficiently improve the communication skills required in physical stores.
[1515] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1516] Step 1:
[1517] Authenticate your login details
[1518] The server receives the login information (username and password) entered by the user. It compares this information with existing data in the database, and if it matches, it completes the authentication. Based on the information entered, it determines whether the user is legitimate, and if successful, it proceeds to the next step. This process allows access to the system.
[1519] Input: Username, Password
[1520] Output: Authentication success / failure status
[1521] Step 2:
[1522] Skill level determination
[1523] The server retrieves past training data for successfully authenticated users from the database and analyzes it. It uses Python data analysis libraries (such as pandas and NumPy) to evaluate the user's current skill level. Based on this, it selects a training scenario of the appropriate level.
[1524] Input: Past training data
[1525] Output: Skill level judgment result
[1526] Step 3:
[1527] Providing training scenarios
[1528] The server selects multiple training scenarios based on the determined skill level and sends them to the user's device. The scenarios include many scenes that can be used in real stores, and the user can choose the scenario they want to train in.
[1529] Input: Skill Level
[1530] Output: List of training scenarios provided
[1531] Step 4:
[1532] Communication Scene Generation
[1533] Based on the training scenario selected by the user, the server generates specific communication scenes, including scripts and dialogue flows, to recreate real-life store situations. The generated scenes are then sent to the device for the user to practice.
[1534] Input: The selected training scenario
[1535] Output: Communication scene details
[1536] Step 5:
[1537] Audio capture and transmission
[1538] When the user presses the start button, the device uses the built-in microphone to capture the user's speech in real time. The speech data is immediately sent to the server and prepared for analysis, particularly for speech recognition and natural language processing (NLP), where it is converted into an appropriate data format (e.g., text).
[1539] Input: User utterance
[1540] Output: Sends audio data to the server
[1541] Step 6:
[1542] Real-time analysis and feedback generation
[1543] The server analyzes the received voice data using a natural language processing (NLP) engine. It analyzes the text of the speech and prosody (speech characteristics) such as speed and intonation, and generates appropriate feedback. The generated feedback is sent to the user's device in real time.
[1544] Input: Audio data
[1545] Output: Feedback
[1546] Step 7:
[1547] View Feedback
[1548] The device receives feedback from the server and displays it to the user in text or audio, allowing the user to review the feedback and learn how to improve their skills.
[1549] Input: Feedback data
[1550] Output: Display feedback
[1551] Step 8:
[1552] Save and analyze progress data
[1553] At the end of the session, the device sends the session data (such as what was said and feedback) to the server, which stores the data in a database and visualizes each user's progress. This data is analyzed to provide insights to drive long-term growth.
[1554] Input: Session data
[1555] Output: Progress data visualization and analysis results
[1556] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1557] This invention is a system that further enhances the effectiveness of training by incorporating an emotion engine into a platform for users to effectively improve their communication skills. The system has the function of recognizing the user's emotions in real time and adjusting feedback and training scenarios based on that information.
[1558] server
[1559] The server plays a central role in this system and provides the following functions:
[1560] 1. User authentication and information management
[1561] The server verifies the user's credentials when they submit a login request, and if authentication is successful, determines the user's past training data and skill level.
[1562] 2. Providing training scenarios
[1563] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[1564] 3. Communication Scene Generation
[1565] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[1566] 4. Real-time analysis and feedback
[1567] The server receives and analyzes the user's voice and facial expression data in real time, and then uses a natural language processing (NLP) module to convert the voice data into text and generate appropriate responses and feedback.
[1568] 5. Emotion recognition and feedback regulation
[1569] The server uses an emotion engine to analyze the user's emotional state and adjusts the feedback accordingly. For example, if the user is nervous, it provides feedback to calm them down.
[1570] 6. Saving and analyzing progress data
[1571] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The server analyzes the stored data and provides users with information that visualizes their progress.
[1572] Terminal
[1573] The user's device works in conjunction with the server to provide the following functions:
[1574] 1. User Interface
[1575] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[1576] 2. Capture and transmit voice and facial expressions
[1577] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the data to a server in real time.
[1578] 3. Displaying feedback and responses
[1579] It displays feedback and responses received from the server in audio or text format, and also displays emotion-based feedback.
[1580] User
[1581] The user performs the following actions through this system:
[1582] 1. Log in and check your profile
[1583] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[1584] 2. Selecting a training scenario
[1585] The user selects the scenario that suits them from a list of training scenarios provided.
[1586] 3. Practice in communication situations
[1587] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[1588] 4. Receiving and improving feedback based on emotion recognition
[1589] See and improve your skills based on real-time feedback, including emotional feedback recognized by the emotion engine.
[1590] 5. Track your progress over the long term
[1591] Users can view their progress and self-assess on a dashboard, which also includes a record of their emotions, allowing them to see their mental progress.
[1592] As a concrete example, consider the case where a user selects a job interview scenario and experiences self-introduction and question-and-answer sessions. If the emotion engine recognizes that the user is nervous, the server will provide feedback such as "Take a deep breath to relieve tension." It can also provide specific advice such as "You're speaking too fast, so speak a little more slowly." By practicing based on this feedback, the user can improve their actual performance in the interview.
[1593] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[1594] The processing flow will be explained below.
[1595] Step 1:
[1596] User
[1597] The user enters their account information on the login screen and clicks the login button.
[1598] server
[1599] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[1600] Step 2:
[1601] server
[1602] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[1603] Terminal
[1604] The terminal displays the list of training scenarios received from the server on the screen.
[1605] Step 3:
[1606] User
[1607] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[1608] server
[1609] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[1610] Step 4:
[1611] Terminal
[1612] The terminal displays the scene information received from the server on the screen and prepares to play audio and capture facial expressions.
[1613] User
[1614] The user checks the scene information and clicks the button to start training.
[1615] Step 5:
[1616] Terminal
[1617] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the audio and video data to a server in real time.
[1618] server
[1619] The server receives the voice data and facial expression data and converts the voice data into text via a natural language processing (NLP) module.
[1620] Step 6:
[1621] server
[1622] The server analyzes the converted text data and facial expression data and uses an emotion engine to recognize the user's emotional state. Based on the analysis results, it generates appropriate responses and feedback. The generated responses and feedback are converted into voice data and sent to the user's device.
[1623] Terminal
[1624] The terminal plays back the voice data and emotional feedback received from the server and provides them to the user.
[1625] Step 7:
[1626] User
[1627] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[1628] Step 8:
[1629] server
[1630] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The stored data is analyzed and information is generated to visualize the progress.
[1631] User
[1632] After the session, users can check their progress data on the dashboard and self-evaluate. They can also review their progress, including changes in their emotions, and use this information to improve their next training session.
[1633] Example 2
[1634] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1635] Conventional communication skill improvement systems struggle to accurately recognize a user's emotional state and provide feedback based on that, making effective training difficult. Furthermore, they often lack real-time speech analysis and feedback generation, which often results in ineffective improvement of the user's skills. Furthermore, they lack sufficient visualization of progress data and long-term growth records, leaving users with a lack of reference information for self-evaluation.
[1636] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for authenticating the user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on the training scenario selected by the user, means for analyzing the user's voice data and facial expression data in real time and generating feedback, means for analyzing the user's emotional state and adjusting the feedback based on the results, and means for saving the user's progress data and tracking and analyzing the growth record. This makes it possible to analyze the user's emotional state in real time and provide appropriate feedback based on the analysis results. Furthermore, by visualizing the progress data and providing a long-term growth record, it is possible to enhance the reference information for the user when performing self-evaluation.
[1637] "User Login Information" means the authentication data (e.g., user ID and password) provided by a User to access a System.
[1638] "Training data" refers to historical information such as the results and evaluations of past training sessions conducted by the user, as well as the content of utterances.
[1639] "Skill level" is an index that evaluates a user's communication ability and proficiency.
[1640] A "training scenario" is an exercise designed to allow a user to practice in a specific situation.
[1641] A "communication scene" is a specific situation or setting that a user simulates, which is generated based on a training scenario.
[1642] "Voice data" refers to data that digitally captures a user's speech.
[1643] "Facial Expression Data" refers to data that captures a user's facial expressions and represents them in digital form.
[1644] "Feedback" refers to the evaluation and advice provided by the system in response to the user's actions and utterances.
[1645] An "emotional state" is the emotion (e.g., tension, joy, anger, etc.) that a user is feeling at a particular moment.
[1646] "Analysis" is the process of processing collected data to extract useful information and patterns.
[1647] "Progress Data" means the performance and evaluation recorded for each of a User's training sessions.
[1648] A "growth record" is information that records the history and results of a user's training sessions over a long period of time and shows the trajectory of their growth.
[1649] "Visualization" is the process of transforming and displaying complex data into visual formats such as graphs and charts.
[1650] The present invention is a training system for improving a user's communication skills. This system recognizes the user's emotions in real time and provides feedback based on the results, thereby achieving more effective training.
[1651] System Configuration
[1652] server
[1653] The server plays a central role in the system and includes the following hardware and software components:
[1654] Authentication system: Validates user login information and performs authentication. The software used is a common authentication framework.
[1655] Database Management System (DBMS): Stores and manages users' past training data and skill levels.
[1656] Training scenario generation module: Selects an appropriate training scenario based on the user's skill level.
[1657] Communication scene generation module: Generates specific communication scenes based on the scenario selected by the user.
[1658] Real-time analysis engine: Using an NLP module, the user's voice is converted into text, and the content is then analyzed to generate an appropriate response.
[1659] Emotion Engine: Analyzes the user's facial expression data and evaluates the user's emotional state.
[1660] Feedback generation module: Generates feedback for the user based on the results of real-time analysis and sentiment analysis.
[1661] Progress Data Storage and Analysis Module: Saves and visualizes progress data after a user's training session.
[1662] Terminal
[1663] The user's device works in conjunction with the server to provide the following functions:
[1664] User Interface: Provides interfaces for users to interact with the system, such as a login screen, scenario selection screen, and training screen.
[1665] Voice capture device: Uses a microphone to capture the user's speech and sends it to the server as voice data.
[1666] Facial Expression Capture Device: Uses a camera to capture the user's facial expressions and transmits the data to a server.
[1667] Feedback display device: Displays feedback or responses received from the server in audio or text format.
[1668] User
[1669] The user performs the following actions:
[1670] 1. Log in and check your profile: Log in to the system using your device and check your skill level and past training data on the dashboard.
[1671] 2. Select a training scenario: Select an appropriate scenario from the list of training scenarios provided by the server.
[1672] 3. Practice in communication scenes: Practice your speech and actions in scenes generated based on the scenario you select.
[1673] 4. Receive real-time feedback: See feedback displayed on your device and improve your training.
[1674] 5. Track your progress: Check your progress data on the dashboard and self-assess.
[1675] Specific examples
[1676] As a concrete example, consider the case where a user selects a "job interview" scenario. The user actually introduces themselves and engages in a question-and-answer session with an AI acting as an interviewer. The device's camera and microphone are used to capture the user's facial expressions and voice. The server analyzes this data in real time and provides feedback such as "You seem nervous. Take a deep breath." It also displays specific advice such as "You speak too quickly. Please speak more slowly."
[1677] Prompt Sentence Examples
[1678] Prompt: Start a practice job interview scenario. First, introduce yourself. Then move on to questions and answers. If you're nervous, the system will provide feedback. It will also give you specific advice on how to speak.
[1679] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[1680] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1681] Step 1:
[1682] Enter your login information
[1683] Input: The user enters their user ID and password on the login screen.
[1684] Operation: The device sends the authentication information entered by the user to the server.
[1685] Output: The server receives the authentication information.
[1686] Step 2:
[1687] User authentication and information management
[1688] Input: The server checks the user's login information against a database.
[1689] How it works: After successful authentication, the server retrieves past training data and skill level from the database.
[1690] Output: Send the acquired training data and skill level to the device.
[1691] Step 3:
[1692] Providing training scenarios
[1693] Input: The server selects appropriate training scenarios based on the user's skill level.
[1694] Operation: The server generates a list of selected training scenarios and sends it to the terminal.
[1695] Output: A list of scenarios will be displayed on the terminal.
[1696] Step 4:
[1697] Scenario Selection
[1698] Input: The user selects one from the list of scenarios provided.
[1699] Operation: The terminal transmits the selected scenario information to the server.
[1700] Output: The server receives the scenario selection information.
[1701] Step 5:
[1702] Communication Scene Generation
[1703] Input: The server generates a scene based on the selected scenario.
[1704] Operation: The server creates detailed communication scenes based on the scenario and sends the information to the terminal.
[1705] Output: A specific training scene is displayed on the device.
[1706] Step 6:
[1707] Voice and facial expression capture
[1708] Input: The user makes utterances and expressions in the simulation.
[1709] How it works: The device uses a microphone and camera to capture the user's voice and facial expression data in real time.
[1710] Output: The captured data is sent to the server.
[1711] Step 7:
[1712] Real-time analytics
[1713] Input: The server receives the voice data and facial expression data sent from the device.
[1714] How it works: The server uses an NLP module to convert voice data into text and analyze its content, as well as analyze facial expression data to assess emotional state.
[1715] Output: Text data and emotion evaluation results are obtained as the analysis results.
[1716] Step 8:
[1717] Feedback Generation
[1718] Input: The server generates feedback based on the real-time analysis results and emotion evaluation results.
[1719] How it works: The server generates and adjusts the feedback based on the user's emotional state. For example, if the user is nervous, the server generates the feedback "Take a deep breath."
[1720] Output: The generated feedback information is sent to the terminal.
[1721] Step 9:
[1722] View Feedback
[1723] Input: The terminal obtains the feedback information received from the server.
[1724] Action: The device displays feedback in the form of audio or text.
[1725] Output: User sees feedback.
[1726] Step 10:
[1727] Save and visualize progress data
[1728] Input: The server receives all the data from the training session.
[1729] How it works: After a session ends, the server stores progress data in a database and analyzes it to generate visualizations, such as graphs and charts.
[1730] Output: Visualized progress information is sent to the device and made available to the user on a dashboard.
[1731] Through this series of processes, users receive real-time feedback to improve their communication skills, enabling effective training.
[1732] (Application example 2)
[1733] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1734] Providing effective customer service skill training is crucial in today's virtual environments. However, conventional training systems struggle to provide users with real-time feedback or feedback based on emotion recognition, resulting in delayed or insufficient improvement. They also lack a means to effectively track and visualize user progress. It is necessary to provide a system that can solve these issues and efficiently support the improvement of customer service skills.
[1735] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1736] In this invention, the server includes means for authenticating a user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for capturing the user's facial expressions and recognizing their emotions, means for providing feedback based on the recognized emotions and adjusting the training scenario, means for saving the user's progress data and tracking and analyzing the growth record, means for the user to select from a plurality of customer service scenarios, and means for presenting prompt sentences for the user to train their customer service skills. This enables the user to efficiently improve their customer service skills while receiving feedback based on emotion recognition in real time.
[1737] "User Login Information" means the authentication information for a User to access the System.
[1738] "Past Training Data" means records and performance of a User's previous training sessions.
[1739] "Skill level" refers to a user's ability or proficiency in a particular skill.
[1740] A "training scenario" is a simulation or challenge that allows a user to practice a particular skill.
[1741] A "communication scene" is a virtual interaction or situation that a user experiences during training.
[1742] "User utterance" refers to the sounds and language that a user makes during training.
[1743] "Feedback" refers to evaluation of the user's actions and utterances and advice for improvement.
[1744] "Facial Expression" refers to the facial expression of the user.
[1745] "Means for recognizing emotions" refers to techniques or devices for analyzing a user's emotional state.
[1746] "Progress Data" means a record showing how much a User has improved through training.
[1747] A "customer service scenario" is a hypothetical situation that users use to train their customer service skills.
[1748] A "prompt" is an instruction or question that prompts the user to take a specific action or speak a certain word.
[1749] This invention is a system for efficiently training customer service skills in a virtual store. The system consists of three main elements: a server, a terminal, and a user.
[1750] server
[1751] The server plays a central role in this system and provides the following functions:
[1752] 1. User authentication and information management
[1753] The server verifies the user's credentials when they submit a login request, and if the login is successful, determines the user's skill level based on past training data.
[1754] 2. Providing training scenarios
[1755] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[1756] 3. Communication Scene Generation
[1757] When a user selects a specific training scenario, the server constructs a specific communication scene based on that scenario and transmits the information to the terminal.
[1758] 4. Real-time analysis and feedback
[1759] The server receives and analyzes the user's speech data and facial expression data in real time using OpenCV (for capturing facial expressions), Keras (for loading emotion recognition models and analyzing emotions), and NLP modules (for natural language processing).
[1760] 5. Emotion recognition and feedback regulation
[1761] The server analyzes the user's emotions using an emotion engine, provides feedback based on the results, and adjusts the training scenario. For example, if the server determines that the user is tense, it provides feedback encouraging the user to relax.
[1762] 6. Saving and analyzing progress data
[1763] After the session ends, the server stores all the user's session data in a database, analyzes the stored data, and provides the user with a visualization of their progress.
[1764] Terminal
[1765] The user's device works in conjunction with the server to provide the following functions:
[1766] 1. User Interface
[1767] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[1768] 2. Capture and transmit voice and facial expressions
[1769] The device captures the user's speech using a voice input device, recognizes facial expressions using a camera, and transmits the data to a server in real time.
[1770] 3. Displaying feedback and responses
[1771] The device displays feedback and responses received from the server in audio or text format, as well as emotion-based feedback.
[1772] User
[1773] Through this system, users can carry out the following training activities:
[1774] 1. Log in and check your profile
[1775] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[1776] 2. Selecting a training scenario
[1777] The user selects the scenario that suits them from a list of training scenarios provided.
[1778] 3. Practice in communication situations
[1779] The user trains by actually speaking in a communication scene generated based on the selected scenario.
[1780] 4. Receiving and improving feedback based on emotion recognition
[1781] See real-time feedback and act on it to improve your skills.
[1782] 5. Track your progress over the long term
[1783] Users can view their progress and self-assess on a dashboard, which also includes tracking their emotions and showing their mental progress.
[1784] Specific examples
[1785] For example, if a user selects the customer service scenario "Welcome Customers," the prompt message will be displayed: "As a salesperson in a virtual store, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." The emotion recognition engine will detect tension in the user's facial expressions, and the server will provide feedback such as "Relax and smile more." In this way, the user can improve their customer service skills while receiving specific advice.
[1786] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1787] Step 1:
[1788] The server receives the user's login information and performs authentication.
[1789] Input: User login information (user ID and password)
[1790] Data processing: The server checks the data against the authentication database
[1791] Output: Authentication result (success or failure)
[1792] Specific operation: The server checks the login information against the authentication database, and if authentication is successful, retrieves the user's past training data.
[1793] Step 2:
[1794] The server determines the user's skill level based on past training data.
[1795] Input: Past training data
[1796] Data processing: Analysis of training data and assessment of skill level
[1797] Output: User's skill level
[1798] Specific operation: The server analyzes past training data and applies a skill assessment algorithm to calculate the user's skill level.
[1799] Step 3:
[1800] The server generates an appropriate training scenario based on the user's skill level and transmits it to the terminal.
[1801] Input: User's skill level
[1802] Data processing: Selection of appropriate scenarios from the scenario database
[1803] Output: List of training scenarios
[1804] Specific operation: The server extracts training scenarios that match the user's skill level from the scenario database and sends the list to the terminal.
[1805] Step 4:
[1806] The user selects a training scenario on the terminal.
[1807] Input: List of training scenarios
[1808] Data Calculation: Scenario Selection
[1809] Output: Selected training scenario
[1810] Specific operation: The user selects the desired scenario from the list of training scenarios displayed on the terminal.
[1811] Step 5:
[1812] The server generates a communication scene based on the selected scenario and transmits it to the terminal.
[1813] Input: Selected training scenario
[1814] Data processing: Generation of communication scenes based on scenarios
[1815] Output: Communication scene data
[1816] Specific operation: The server constructs a virtual communication scene based on the selected scenario and sends the data to the terminal.
[1817] Step 6:
[1818] The device captures the user's voice and facial expressions and transmits them to the server.
[1819] Input: User utterances and facial expressions
[1820] Data processing: Audio input device and camera capture
[1821] Output: Capture data
[1822] Specific operation: The device captures the user's voice with a microphone, recognizes facial expressions with a camera, and sends the data to a server.
[1823] Step 7:
[1824] The server analyzes the user's speech data and facial expression data in real time, generates feedback, and sends it to the device.
[1825] Input: User speech data and facial expression data
[1826] Data processing: speech recognition, natural language processing, sentiment analysis
[1827] Output: Real-time feedback
[1828] Specific operation: The server analyzes the user's speech using speech recognition and natural language processing, and analyzes the facial expressions using an emotion engine. Based on the obtained information, it generates feedback and sends it to the device.
[1829] Step 8:
[1830] The terminal presents the feedback received from the server to the user in real time.
[1831] Input: Feedback data
[1832] Data Calculation: Feedback Display
[1833] Output: Feedback that is displayed to the user
[1834] Specific operation: The device displays the feedback received from the server to the user in voice or text.
[1835] Step 9:
[1836] The server stores the progress data in a database and generates visualization data.
[1837] Input: Training session data
[1838] Data processing: saving to database, analyzing and visualizing progress data
[1839] Output: Progress report
[1840] What it does: After the session ends, the server stores all data in a database and generates a report to visualize the progress.
[1841] Step 10:
[1842] Users can check their progress and self-assess through a dashboard.
[1843] Input: Visualized progress data
[1844] Data calculation: Check progress data and self-evaluation
[1845] Output: User self-assessment results
[1846] Specific operation: The user uses the device's dashboard function to check their training progress and perform self-evaluation.
[1847] As a specific example, if a user selects the customer service scenario "welcoming a customer," the prompt message will be displayed: "As a virtual store clerk, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." If tension is detected in the user's facial expression, the feedback will be "Relax and smile more."
[1848] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1849] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1850] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1851] [Fourth embodiment]
[1852] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1853] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1854] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1855] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1856] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1857] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1858] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1859] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1860] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1861] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1862] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1863] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1864] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1865] The present invention relates to a communication skill improvement platform, which is designed to provide a set of functions for users to effectively improve their communication skills. The system is implemented as a cross-platform application and can be used through a web browser or a mobile application.
[1866] server
[1867] The server plays a central role in this system and provides the following functions:
[1868] 1. User authentication and information management
[1869] The server verifies the user's credentials when they submit a login request, and if authentication is successful, retrieves the user's past training data from a database to determine their current skill level.
[1870] 2. Providing training scenarios
[1871] The server generates appropriate training scenarios based on the user's skill level and sends the information to the user's device. The scenarios contain various communication scenes, and the user can select a scenario that suits their level.
[1872] 3. Communication Scene Generation
[1873] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[1874] 4. Real-time analysis and feedback
[1875] The server receives the user's voice data in real time and converts it into text using natural language processing (NLP) technology. The converted text data is then analyzed, and appropriate responses and feedback are generated and sent to the device.
[1876] 5. Saving and analyzing progress data
[1877] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. Later, this data is analyzed and used to provide information for visualizing the user's progress.
[1878] Terminal
[1879] The user's device works in conjunction with the server to provide the following functions:
[1880] 1. User Interface
[1881] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[1882] 2. Audio capture and transmission
[1883] The device captures the user's speech with a microphone and transmits the audio data to the server in real time.
[1884] 3. Displaying feedback and responses
[1885] Feedback and responses received from the server are displayed as voice or text, allowing the user to understand what needs to be improved and take the next action.
[1886] User
[1887] The user performs the following actions through this system:
[1888] 1. Log in and check your profile
[1889] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[1890] 2. Selecting a training scenario
[1891] The user selects a scenario that suits their level from a list of training scenarios provided.
[1892] 3. Practice in communication situations
[1893] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[1894] 4. Receive feedback and improve
[1895] See real-time feedback and act on it to improve your skills.
[1896] 5. Track your progress over the long term
[1897] Users can view progress data on a dashboard and perform self-assessments to promote their growth.
[1898] As a concrete example, consider the case where a user selects a job interview scenario. The user experiences self-introductions and question-and-answer sessions, and receives real-time feedback from the server. For example, the server may provide feedback such as, "You speak too quickly, so you should speak more slowly," or "You need to improve your explanation, including your specific work experience." By practicing based on this feedback, the user will be able to respond confidently in an actual interview.
[1899] Thus, the present invention provides users with an environment for practical and effective communication skill development, with real-time feedback and long-term progress tracking.
[1900] The processing flow will be explained below.
[1901] Step 1:
[1902] User
[1903] The user enters their account information on the login screen and clicks the login button.
[1904] server
[1905] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[1906] Step 2:
[1907] server
[1908] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[1909] Terminal
[1910] The terminal displays the list of training scenarios received from the server on the screen.
[1911] Step 3:
[1912] User
[1913] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[1914] server
[1915] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[1916] Step 4:
[1917] Terminal
[1918] The terminal displays the scene information received from the server on the screen and prepares for audio playback.
[1919] User
[1920] The user checks the scene information and clicks the button to start training.
[1921] Step 5:
[1922] Terminal
[1923] The device captures the user's speech with a microphone and transmits the voice data to the server in real time.
[1924] server
[1925] The server receives the voice data and converts it into text data via a natural language processing (NLP) module.
[1926] Step 6:
[1927] server
[1928] The server analyzes the converted text data, generates appropriate responses and feedback, converts the generated responses and feedback into audio data, and sends it to the user's device.
[1929] Terminal
[1930] The terminal plays back the audio data received from the server and provides feedback to the user.
[1931] Step 7:
[1932] User
[1933] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[1934] Step 8:
[1935] server
[1936] After the session ends, the server stores all session data (such as speech content, feedback, and evaluation results) in a database. The stored data is analyzed and information is generated to visualize the progress.
[1937] User
[1938] After the session, users can view their progress data on a dashboard and self-evaluate.
[1939] Example 1
[1940] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1941] Current communication skill training systems lack the functionality to provide appropriate feedback in real time based on the user's skill level. They also have limited functionality to effectively track the user's progress and suggest specific ways to improve. This makes it difficult for existing systems to effectively improve skills.
[1942] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1943] In this invention, the server includes means for verifying user authentication information, means for determining the user's skill level based on past learning data, means for generating appropriate training scenarios based on the user's skill level, means for constructing dialogue scenes based on the training scenario selected by the user, means for analyzing the user's voice data in real time using natural language processing technology and generating feedback, and means for storing session data in a database and analyzing and visualizing the user's progress, thereby enabling the user to receive effective feedback in real time, thereby promoting self-evaluation and growth.
[1944] "Means for verifying user authentication information" refers to a means for checking the authentication information (user name and password) entered by the user and comparing it with a database to see if it is correct.
[1945] The "means for determining the user's skill level based on past learning data" refers to a means for analyzing the user's past training data and evaluation results to determine the user's current skill level.
[1946] The "means for generating an appropriate training scenario based on the skill level of the user" is a means for creating an optimal training scenario for the user in accordance with the determined skill level.
[1947] The "means for constructing a dialogue scene based on a training scenario selected by the user" refers to a means for constructing a scene including specific dialogue content and questions based on a scenario selected by the user.
[1948] "Means for analyzing user voice data in real time using natural language processing technology and generating feedback" refers to means for capturing a user's speech as voice data, converting it into text, analyzing it, and generating the necessary feedback in real time.
[1949] "Means for storing session data in a database and analyzing and visualizing the user's progress" refers to a means for storing data after a training session in a database and subsequently analyzing the data to visualize the user's progress in graphs or other formats.
[1950] The communication skill improvement system of the present invention can be effectively implemented by cooperation between the server, the terminal, and the user.
[1951] Server Roles
[1952] The server plays a central role in this system and provides the following functions:
[1953] 1. Verify user credentials
[1954] The server receives the authentication information entered by the user on the login screen and authenticates the user by comparing it with the database. The user can only use the system if authentication is successful.
[1955] 2. Determine the user's skill level
[1956] The server analyzes past learning data to determine the user's current skill level, including stored training session results and evaluations.
[1957] 3. Generating training scenarios
[1958] The server automatically generates an appropriate training scenario based on the user's skill level. The generated scenario is sent to the user's device. For example, a "basic self-introduction" scenario for beginners and a "job interview" scenario for intermediate users are generated.
[1959] 4. Creating a communication scene
[1960] Based on the training scenario selected by the user, the system constructs a specific dialogue scene and transmits the scene information to the device. The scene includes the actual utterances and questions the user will ask.
[1961] 5. Real-time analysis and feedback
[1962] The server receives the user's voice data in real time, converts it into text using natural language processing technology, and generates feedback, which is then immediately sent to the device and presented to the user.
[1963] 6. Session data storage and progress analysis
[1964] After each session, the server stores all relevant data in a database. By analyzing the stored data, the user's progress is visualized and displayed on a dashboard, allowing users to self-evaluate and track their progress.
[1965] Device Role
[1966] The user's device works in conjunction with the server to provide the following functions:
[1967] 1. Providing a user interface
[1968] The terminal allows the user to interact with the system through interfaces such as a login screen, scenario selection screen, and training start screen.
[1969] 2. Capture and transmit audio data
[1970] The device captures the user's speech with a microphone and transmits the audio data to a server in real time, allowing for immediate analysis and feedback.
[1971] 3. Viewing Feedback
[1972] The terminal displays the feedback or response received from the server in voice or text form and provides it to the user.
[1973] User Behavior
[1974] The user performs the following actions through this system:
[1975] 1. Log in and check your profile
[1976] Users enter their login information to access the system. After logging in, they can check their skill level and past training data on the dashboard.
[1977] 2. Selecting a training scenario
[1978] From the list of training scenarios provided, choose one that matches your skill level, for example, the intermediate "Job Interview" scenario.
[1979] 3. Practice in communication situations
[1980] The user will practice speaking in a dialogue scene generated based on the selected scenario. Once the user has finished speaking, they can move on to the next step.
[1981] 4. Receive feedback and improve
[1982] See real-time feedback and use it to improve your skills, for example, "You speak too fast; you should speak a little slower."
[1983] 5. Track your progress over the long term
[1984] Users can view progress data and self-assess on the dashboard, allowing them to understand their progress and set goals for further improvement.
[1985] Examples and prompts
[1986] Specific examples
[1987] If a user selects the "Job Interview" scenario for intermediate level learners, they will experience self-introductions and question-and-answer sessions, and receive real-time feedback from the server. For example, they may receive feedback such as, "You speak too quickly; you should speak a little more slowly," or "You need to improve your explanation, including your specific work experience." Users can then practice based on this feedback, gaining confidence in the actual interview.
[1988] Prompt Sentence Examples
[1989] The user has selected the job interview scenario. Please simulate the scene where the user introduces themselves and provide feedback on the following points:
[1990] 1. Speaking Speed
[1991] 2. Description of specific work experience
[1992] 3. Language and attitude
[1993] By including these specific details, the present invention provides users with an environment for practical and effective communication skill development, allowing for real-time feedback and long-term growth tracking.
[1994] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1995] Step 1:
[1996] Entering and validating user credentials
[1997] Input: The username and password the user enters on the login screen.
[1998] Specific behavior:
[1999] The user enters the username and password on the login screen of the device.
[2000] The terminal transmits the entered authentication information to the server.
[2001] The server checks the received authentication information against a database.
[2002] Data processing / calculation:
[2003] The server performs a calculation to match the username and password in its database.
[2004] output:
[2005] If the authentication is successful, the server sends a "user authentication successful" response to the terminal and starts a session.
[2006] If the authentication fails, the server sends a "user authentication failed" response to the terminal, prompting it to try again.
[2007] Step 2:
[2008] Obtaining user history data and determining skill level
[2009] Input: User ID and previous learning data request.
[2010] Specific behavior:
[2011] The server uses the user ID to load the user's past training data from a database.
[2012] Data processing / calculation:
[2013] The server analyzes the acquired past training data and performs calculations to evaluate the user's skill level.
[2014] output:
[2015] The user's current skill level information is transmitted to the terminal.
[2016] Step 3:
[2017] Training scenario generation and distribution
[2018] Input: User skill level information.
[2019] Specific behavior:
[2020] The server generates appropriate training scenarios based on skill level.
[2021] Data processing / calculation:
[2022] The server selects the most suitable scenario for the user's skill level from among multiple scenario candidates and constructs scenario data.
[2023] output:
[2024] The generated training scenario is sent to the terminal.
[2025] Step 4:
[2026] Building and sending communication scenes
[2027] Input: A training scenario selected by the user.
[2028] Specific behavior:
[2029] The user selects one of the scenarios provided on the terminal.
[2030] The terminal transmits the selected scenario information to the server.
[2031] The server constructs a specific dialogue scene based on the selected scenario.
[2032] output:
[2033] The constructed dialogue scene is sent to the device.
[2034] Step 5:
[2035] Capture and transmit audio data
[2036] Input: User speech.
[2037] Specific behavior:
[2038] The user speaks into the microphone of the terminal.
[2039] The device captures audio data through a microphone.
[2040] The device transmits the captured audio data to the server in real time.
[2041] output:
[2042] Send the audio data to the server.
[2043] Step 6:
[2044] Providing real-time analysis and feedback
[2045] Input: User's voice data sent from the device.
[2046] Specific behavior:
[2047] The server converts the received voice data into text using natural language processing technology.
[2048] The server analyzes the converted text data and generates appropriate feedback.
[2049] Data processing / calculation:
[2050] The system converts voice data into text and performs feedback generation calculations based on text analysis.
[2051] output:
[2052] The generated feedback is sent to the device.
[2053] Step 7:
[2054] Session data storage and progress visualization
[2055] Input: Data from each session (utterances, feedback, and evaluation results).
[2056] Specific behavior:
[2057] The server stores all relevant data in a database after the session ends.
[2058] Data processing / calculation:
[2059] The stored data is analyzed and calculations are performed to visualize the progress data.
[2060] output:
[2061] Progress data is visualized and presented to users via a dashboard.
[2062] Through the above steps, this system can effectively improve the user's communication skills.
[2063] (Application example 1)
[2064] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2065] In modern brick-and-mortar stores, improving staff communication skills is a high priority. High skills are especially required for dealing with new customers and handling complaints, as these skills are directly linked to customer satisfaction. However, in brick-and-mortar stores, staff have limited opportunities to effectively improve their skills while performing their daily tasks. Furthermore, traditional training methods make it difficult to train many staff members at once, and individual feedback is lacking. As a result, there is a problem of inefficient improvement of individual skills.
[2066] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2067] In this invention, the server includes means for authenticating a user's login information, means for determining a user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating communication scenes based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for saving the user's progress data and tracking and analyzing the growth record, means for evaluating the user's speech speed and content and suggesting specific areas for improvement, and means for creating scenarios for staff to respond to in-store situations and conducting training. This enables store staff to effectively and efficiently improve the communication skills required in-store while performing their daily work.
[2068] The means for "authenticating user login information" is a means for verifying authentication information such as a user name and password entered when a user accesses the system, and confirming that the user is a legitimate user.
[2069] The means for "determining the user's skill level based on past training data" is a means for analyzing data on training previously performed by the user and evaluating the current level of the user's skill based on the data.
[2070] The means for "providing an appropriate training scenario based on the user's skill level" is a means for automatically selecting training according to the determined skill level and presenting it to the user.
[2071] The means for "generating a communication scene based on a training scenario selected by the user" is a means for reproducing a specific communication situation or condition based on a training scenario selected by the user.
[2072] The method of "analyzing user speech in real time and generating feedback" involves collecting the user's speech with a microphone, instantly analyzing it using natural language processing technology, and presenting appropriate suggestions and areas for improvement based on the analysis results.
[2073] The means for "storing user progress data and tracking and analyzing growth records" refers to storing the user's training details and results in a database, and analyzing them over the long term to track skill improvement.
[2074] The method of "evaluating the user's speech speed and content and suggesting specific areas for improvement" involves analyzing the speed and quality of the content of the user's speech and, based on the results, providing specific instructions to the user on how to improve.
[2075] The method of "creating scenarios for staff to respond to in-store situations and conducting training" involves creating virtual scenarios for specific situations such as customer service and handling complaints in a real store, and then having staff practice using those scenarios.
[2076] This invention is a system for improving communication skills for store staff. This system mainly consists of three components: a server, a terminal, and a user. The following is a detailed description of the role of each component and its processing.
[2077] server
[2078] The server plays a central role in the system, providing the following functions:
[2079] User authentication:
[2080] When a user logs in to a system, the server authenticates them based on their username and password, ensuring that only authorized users can access the system.
[2081] Skill Level Determination:
[2082] The server analyzes past training data to assess the user's current skill level, which is stored in a database and updated in real time.
[2083] Training scenarios provided:
[2084] Based on the user's skill level, appropriate training scenarios are selected and provided to the user, including scenarios for greeting new customers and handling complaints.
[2085] Communication Scene Generation:
[2086] Based on the scenario selected by the user, specific communication scenes are generated and sent to the device, allowing the user to receive training tailored to their actual work.
[2087] Real-time analysis and feedback:
[2088] It collects user utterances in real time, analyzes them using natural language processing (NLP) technology, and generates feedback based on the analysis results and sends it to the device.
[2089] Save and analyze progress data:
[2090] Session data is stored in a database and progress is visualized for each user, allowing users to self-assess and foster long-term growth.
[2091] Terminal
[2092] The user's terminal provides the following functions in cooperation with the server.
[2093] User Interface:
[2094] The terminal provides an interface for interacting with the user, such as a login screen, a scenario selection screen, and a training scene.
[2095] Audio Capture:
[2096] The device's microphone is used to capture the user's speech in real time and send the data to the server.
[2097] Show feedback:
[2098] The analysis results and feedback received from the server are provided to the user via text or voice, allowing the user to make immediate corrections and improve their skills.
[2099] User
[2100] The users are staff members of a physical store, and by utilizing this system, they perform the following actions:
[2101] Login:
[2102] The user logs in to the terminal and checks their account information.
[2103] Select a training scenario:
[2104] Choose from the list of scenarios provided that fit your skill level.
[2105] Communication practice:
[2106] Participants will experience and speak in a series of real-life communication situations, including greeting new customers and handling complex complaints.
[2107] Receiving feedback and making corrections:
[2108] Improve your skills with real-time feedback.
[2109] Long-term progress check:
[2110] View your training history and progress data on the dashboard and self-assess.
[2111] Hardware and software used
[2112] Hardware: Smartphones, tablets
[2113] software:
[2114] Web browser (cross-platform compatible)
[2115] Python (Flask framework)
[2116] Database (MySQL or PostgreSQL)
[2117] Natural Language Processing (NLP) libraries (spaCy and nltk)
[2118] Frontend (ReactJS)
[2119] Examples of concrete examples and prompts
[2120] For example, if the user selects a scenario for handling a complaint, the following dialogue is generated:
[2121] 1. User: "Sorry, but we can't accept your return at this time."
[2122] 2. System: Performs real-time analysis and provides feedback such as "The text is too long. Make it concise" or "Intonation is important."
[2123] As described above, by using this system, store staff can efficiently improve the communication skills required in physical stores.
[2124] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2125] Step 1:
[2126] Authenticate your login details
[2127] The server receives the login information (username and password) entered by the user. It compares this information with existing data in the database, and if it matches, it completes the authentication. Based on the information entered, it determines whether the user is legitimate, and if successful, it proceeds to the next step. This process allows access to the system.
[2128] Input: Username, Password
[2129] Output: Authentication success / failure status
[2130] Step 2:
[2131] Skill level determination
[2132] The server retrieves past training data for successfully authenticated users from the database and analyzes it. It uses Python data analysis libraries (such as pandas and NumPy) to evaluate the user's current skill level. Based on this, it selects a training scenario of the appropriate level.
[2133] Input: Past training data
[2134] Output: Skill level judgment result
[2135] Step 3:
[2136] Providing training scenarios
[2137] The server selects multiple training scenarios based on the determined skill level and sends them to the user's device. The scenarios include many scenes that can be used in real stores, and the user can choose the scenario they want to train in.
[2138] Input: Skill Level
[2139] Output: List of training scenarios provided
[2140] Step 4:
[2141] Communication Scene Generation
[2142] Based on the training scenario selected by the user, the server generates specific communication scenes, including scripts and dialogue flows, to recreate real-life store situations. The generated scenes are then sent to the device for the user to practice.
[2143] Input: The selected training scenario
[2144] Output: Communication scene details
[2145] Step 5:
[2146] Audio capture and transmission
[2147] When the user presses the start button, the device uses the built-in microphone to capture the user's speech in real time. The speech data is immediately sent to the server and prepared for analysis, particularly for speech recognition and natural language processing (NLP), where it is converted into an appropriate data format (e.g., text).
[2148] Input: User utterance
[2149] Output: Sends audio data to the server
[2150] Step 6:
[2151] Real-time analysis and feedback generation
[2152] The server analyzes the received voice data using a natural language processing (NLP) engine. It analyzes the text of the speech and prosody (speech characteristics) such as speed and intonation, and generates appropriate feedback. The generated feedback is sent to the user's device in real time.
[2153] Input: Audio data
[2154] Output: Feedback
[2155] Step 7:
[2156] View Feedback
[2157] The device receives feedback from the server and displays it to the user in text or audio, allowing the user to review the feedback and learn how to improve their skills.
[2158] Input: Feedback data
[2159] Output: Display feedback
[2160] Step 8:
[2161] Save and analyze progress data
[2162] At the end of the session, the device sends the session data (such as what was said and feedback) to the server, which stores the data in a database and visualizes each user's progress. This data is analyzed to provide insights to drive long-term growth.
[2163] Input: Session data
[2164] Output: Progress data visualization and analysis results
[2165] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2166] This invention is a system that further enhances the effectiveness of training by incorporating an emotion engine into a platform for users to effectively improve their communication skills. The system has the function of recognizing the user's emotions in real time and adjusting feedback and training scenarios based on that information.
[2167] server
[2168] The server plays a central role in this system and provides the following functions:
[2169] 1. User authentication and information management
[2170] The server verifies the user's credentials when they submit a login request, and if authentication is successful, determines the user's past training data and skill level.
[2171] 2. Providing training scenarios
[2172] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[2173] 3. Communication Scene Generation
[2174] When a user selects a scenario, the server creates a specific communication scene based on that scenario and transmits that information to the terminal.
[2175] 4. Real-time analysis and feedback
[2176] The server receives and analyzes the user's voice and facial expression data in real time, and then uses a natural language processing (NLP) module to convert the voice data into text and generate appropriate responses and feedback.
[2177] 5. Emotion recognition and feedback regulation
[2178] The server uses an emotion engine to analyze the user's emotional state and adjusts the feedback accordingly. For example, if the user is nervous, it provides feedback to calm them down.
[2179] 6. Saving and analyzing progress data
[2180] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The server analyzes the stored data and provides users with information that visualizes their progress.
[2181] Terminal
[2182] The user's device works in conjunction with the server to provide the following functions:
[2183] 1. User Interface
[2184] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[2185] 2. Capture and transmit voice and facial expressions
[2186] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the data to a server in real time.
[2187] 3. Displaying feedback and responses
[2188] It displays feedback and responses received from the server in audio or text format, and also displays emotion-based feedback.
[2189] User
[2190] The user performs the following actions through this system:
[2191] 1. Log in and check your profile
[2192] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[2193] 2. Selecting a training scenario
[2194] The user selects the scenario that suits them from a list of training scenarios provided.
[2195] 3. Practice in communication situations
[2196] The user then performs actual speech and actions in a communication scene generated based on the selected scenario. Once the speech is complete, the user moves on to the next step.
[2197] 4. Receiving and improving feedback based on emotion recognition
[2198] See and improve your skills based on real-time feedback, including emotional feedback recognized by the emotion engine.
[2199] 5. Track your progress over the long term
[2200] Users can view their progress and self-assess on a dashboard, which also includes a record of their emotions, allowing them to see their mental progress.
[2201] As a concrete example, consider the case where a user selects a job interview scenario and experiences self-introduction and question-and-answer sessions. If the emotion engine recognizes that the user is nervous, the server will provide feedback such as "Take a deep breath to relieve tension." It can also provide specific advice such as "You're speaking too fast, so speak a little more slowly." By practicing based on this feedback, the user can improve their actual performance in the interview.
[2202] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[2203] The processing flow will be explained below.
[2204] Step 1:
[2205] User
[2206] The user enters their account information on the login screen and clicks the login button.
[2207] server
[2208] The server receives the user's authentication information, retrieves the corresponding user data from the database, and performs authentication. If authentication is successful, the server determines the user's past training data and skill level.
[2209] Step 2:
[2210] server
[2211] The server generates a list of appropriate training scenarios based on the determined skill level and transmits the list to the user's terminal.
[2212] Terminal
[2213] The terminal displays the list of training scenarios received from the server on the screen.
[2214] Step 3:
[2215] User
[2216] The user selects a training scenario that suits them from the displayed list and clicks the select button.
[2217] server
[2218] The server receives the user's selection, generates detailed information about the communication scene based on the selected scenario, and transmits the information to the user's terminal.
[2219] Step 4:
[2220] Terminal
[2221] The terminal displays the scene information received from the server on the screen and prepares to play audio and capture facial expressions.
[2222] User
[2223] The user checks the scene information and clicks the button to start training.
[2224] Step 5:
[2225] Terminal
[2226] The device captures the user's speech with a microphone, recognizes facial expressions with a camera, and transmits the audio and video data to a server in real time.
[2227] server
[2228] The server receives the voice data and facial expression data and converts the voice data into text via a natural language processing (NLP) module.
[2229] Step 6:
[2230] server
[2231] The server analyzes the converted text data and facial expression data and uses an emotion engine to recognize the user's emotional state. Based on the analysis results, it generates appropriate responses and feedback. The generated responses and feedback are converted into voice data and sent to the user's device.
[2232] Terminal
[2233] The terminal plays back the voice data and emotional feedback received from the server and provides them to the user.
[2234] Step 7:
[2235] User
[2236] The user reviews the feedback, improves their speech or behavior as needed, and clicks the "Next" button to move to the next stage.
[2237] Step 8:
[2238] server
[2239] After the session ends, the server stores all session data (such as speech content, feedback, evaluation results, and emotional information) in a database. The stored data is analyzed and information is generated to visualize the progress.
[2240] User
[2241] After the session, users can check their progress data on the dashboard and self-evaluate. They can also review their progress, including changes in their emotions, and use this information to improve their next training session.
[2242] Example 2
[2243] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2244] Conventional communication skill improvement systems struggle to accurately recognize a user's emotional state and provide feedback based on that, making effective training difficult. Furthermore, they often lack real-time speech analysis and feedback generation, which often results in ineffective improvement of the user's skills. Furthermore, they lack sufficient visualization of progress data and long-term growth records, leaving users with a lack of reference information for self-evaluation.
[2245] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for authenticating the user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on the training scenario selected by the user, means for analyzing the user's voice data and facial expression data in real time and generating feedback, means for analyzing the user's emotional state and adjusting the feedback based on the results, and means for saving the user's progress data and tracking and analyzing the growth record. This makes it possible to analyze the user's emotional state in real time and provide appropriate feedback based on the analysis results. Furthermore, by visualizing the progress data and providing a long-term growth record, it is possible to enhance the reference information for the user when performing self-evaluation.
[2246] "User Login Information" means the authentication data (e.g., user ID and password) provided by a User to access a System.
[2247] "Training data" refers to historical information such as the results and evaluations of past training sessions conducted by the user, as well as the content of utterances.
[2248] "Skill level" is an index that evaluates a user's communication ability and proficiency.
[2249] A "training scenario" is an exercise designed to allow a user to practice in a specific situation.
[2250] A "communication scene" is a specific situation or setting that a user simulates, which is generated based on a training scenario.
[2251] "Voice data" refers to data that digitally captures a user's speech.
[2252] "Facial Expression Data" refers to data that captures a user's facial expressions and represents them in digital form.
[2253] "Feedback" refers to the evaluation and advice provided by the system in response to the user's actions and utterances.
[2254] An "emotional state" is the emotion (e.g., tension, joy, anger, etc.) that a user is feeling at a particular moment.
[2255] "Analysis" is the process of processing collected data to extract useful information and patterns.
[2256] "Progress Data" means the performance and evaluation recorded for each of a User's training sessions.
[2257] A "growth record" is information that records the history and results of a user's training sessions over a long period of time and shows the trajectory of their growth.
[2258] "Visualization" is the process of transforming and displaying complex data into visual formats such as graphs and charts.
[2259] The present invention is a training system for improving a user's communication skills. This system recognizes the user's emotions in real time and provides feedback based on the results, thereby achieving more effective training.
[2260] System Configuration
[2261] server
[2262] The server plays a central role in the system and includes the following hardware and software components:
[2263] Authentication system: Validates user login information and performs authentication. The software used is a common authentication framework.
[2264] Database Management System (DBMS): Stores and manages users' past training data and skill levels.
[2265] Training scenario generation module: Selects an appropriate training scenario based on the user's skill level.
[2266] Communication scene generation module: Generates specific communication scenes based on the scenario selected by the user.
[2267] Real-time analysis engine: Using an NLP module, the user's voice is converted into text, and the content is then analyzed to generate an appropriate response.
[2268] Emotion Engine: Analyzes the user's facial expression data and evaluates the user's emotional state.
[2269] Feedback generation module: Generates feedback for the user based on the results of real-time analysis and sentiment analysis.
[2270] Progress Data Storage and Analysis Module: Saves and visualizes progress data after a user's training session.
[2271] Terminal
[2272] The user's device works in conjunction with the server to provide the following functions:
[2273] User Interface: Provides interfaces for users to interact with the system, such as a login screen, scenario selection screen, and training screen.
[2274] Voice capture device: Uses a microphone to capture the user's speech and sends it to the server as voice data.
[2275] Facial Expression Capture Device: Uses a camera to capture the user's facial expressions and transmits the data to a server.
[2276] Feedback display device: Displays feedback or responses received from the server in audio or text format.
[2277] User
[2278] The user performs the following actions:
[2279] 1. Log in and check your profile: Log in to the system using your device and check your skill level and past training data on the dashboard.
[2280] 2. Select a training scenario: Select an appropriate scenario from the list of training scenarios provided by the server.
[2281] 3. Practice in communication scenes: Practice your speech and actions in scenes generated based on the scenario you select.
[2282] 4. Receive real-time feedback: See feedback displayed on your device and improve your training.
[2283] 5. Track your progress: Check your progress data on the dashboard and self-assess.
[2284] Specific examples
[2285] As a concrete example, consider the case where a user selects a "job interview" scenario. The user actually introduces themselves and engages in a question-and-answer session with an AI acting as an interviewer. The device's camera and microphone are used to capture the user's facial expressions and voice. The server analyzes this data in real time and provides feedback such as "You seem nervous. Take a deep breath." It also displays specific advice such as "You speak too quickly. Please speak more slowly."
[2286] Prompt Sentence Examples
[2287] Prompt: Start a practice job interview scenario. First, introduce yourself. Then move on to questions and answers. If you're nervous, the system will provide feedback. It will also give you specific advice on how to speak.
[2288] In this way, the present invention provides users with an environment for practical and effective communication skill development, enabling real-time feedback and adjustments based on emotion recognition.
[2289] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2290] Step 1:
[2291] Enter your login information
[2292] Input: The user enters their user ID and password on the login screen.
[2293] Operation: The device sends the authentication information entered by the user to the server.
[2294] Output: The server receives the authentication information.
[2295] Step 2:
[2296] User authentication and information management
[2297] Input: The server checks the user's login information against a database.
[2298] How it works: After successful authentication, the server retrieves past training data and skill level from the database.
[2299] Output: Send the acquired training data and skill level to the device.
[2300] Step 3:
[2301] Providing training scenarios
[2302] Input: The server selects appropriate training scenarios based on the user's skill level.
[2303] Operation: The server generates a list of selected training scenarios and sends it to the terminal.
[2304] Output: A list of scenarios will be displayed on the terminal.
[2305] Step 4:
[2306] Scenario Selection
[2307] Input: The user selects one from the list of scenarios provided.
[2308] Operation: The terminal transmits the selected scenario information to the server.
[2309] Output: The server receives the scenario selection information.
[2310] Step 5:
[2311] Communication Scene Generation
[2312] Input: The server generates a scene based on the selected scenario.
[2313] Operation: The server creates detailed communication scenes based on the scenario and sends the information to the terminal.
[2314] Output: A specific training scene is displayed on the device.
[2315] Step 6:
[2316] Voice and facial expression capture
[2317] Input: The user makes utterances and expressions in the simulation.
[2318] How it works: The device uses a microphone and camera to capture the user's voice and facial expression data in real time.
[2319] Output: The captured data is sent to the server.
[2320] Step 7:
[2321] Real-time analytics
[2322] Input: The server receives the voice data and facial expression data sent from the device.
[2323] How it works: The server uses an NLP module to convert voice data into text and analyze its content, as well as analyze facial expression data to assess emotional state.
[2324] Output: Text data and emotion evaluation results are obtained as the analysis results.
[2325] Step 8:
[2326] Feedback Generation
[2327] Input: The server generates feedback based on the real-time analysis results and emotion evaluation results.
[2328] How it works: The server generates and adjusts the feedback based on the user's emotional state. For example, if the user is nervous, the server generates the feedback "Take a deep breath."
[2329] Output: The generated feedback information is sent to the terminal.
[2330] Step 9:
[2331] View Feedback
[2332] Input: The terminal obtains the feedback information received from the server.
[2333] Action: The device displays feedback in the form of audio or text.
[2334] Output: User sees feedback.
[2335] Step 10:
[2336] Save and visualize progress data
[2337] Input: The server receives all the data from the training session.
[2338] How it works: After a session ends, the server stores progress data in a database and analyzes it to generate visualizations, such as graphs and charts.
[2339] Output: Visualized progress information is sent to the device and made available to the user on a dashboard.
[2340] Through this series of processes, users receive real-time feedback to improve their communication skills, enabling effective training.
[2341] (Application example 2)
[2342] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2343] Providing effective customer service skill training is crucial in today's virtual environments. However, conventional training systems struggle to provide users with real-time feedback or feedback based on emotion recognition, resulting in delayed or insufficient improvement. They also lack a means to effectively track and visualize user progress. It is necessary to provide a system that can solve these issues and efficiently support the improvement of customer service skills.
[2344] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2345] In this invention, the server includes means for authenticating a user's login information and determining the user's skill level based on past training data, means for providing an appropriate training scenario based on the user's skill level, means for generating a communication scene based on a training scenario selected by the user, means for analyzing the user's speech in real time and generating feedback, means for capturing the user's facial expressions and recognizing their emotions, means for providing feedback based on the recognized emotions and adjusting the training scenario, means for saving the user's progress data and tracking and analyzing the growth record, means for the user to select from a plurality of customer service scenarios, and means for presenting prompt sentences for the user to train their customer service skills. This enables the user to efficiently improve their customer service skills while receiving feedback based on emotion recognition in real time.
[2346] "User Login Information" means the authentication information for a User to access the System.
[2347] "Past Training Data" means records and performance of a User's previous training sessions.
[2348] "Skill level" refers to a user's ability or proficiency in a particular skill.
[2349] A "training scenario" is a simulation or challenge that allows a user to practice a particular skill.
[2350] A "communication scene" is a virtual interaction or situation that a user experiences during training.
[2351] "User utterance" refers to the sounds and language that a user makes during training.
[2352] "Feedback" refers to evaluation of the user's actions and utterances and advice for improvement.
[2353] "Facial Expression" refers to the facial expression of the user.
[2354] "Means for recognizing emotions" refers to techniques or devices for analyzing a user's emotional state.
[2355] "Progress Data" means a record showing how much a User has improved through training.
[2356] A "customer service scenario" is a hypothetical situation that users use to train their customer service skills.
[2357] A "prompt" is an instruction or question that prompts the user to take a specific action or speak a certain word.
[2358] This invention is a system for efficiently training customer service skills in a virtual store. The system consists of three main elements: a server, a terminal, and a user.
[2359] server
[2360] The server plays a central role in this system and provides the following functions:
[2361] 1. User authentication and information management
[2362] The server verifies the user's credentials when they submit a login request, and if the login is successful, determines the user's skill level based on past training data.
[2363] 2. Providing training scenarios
[2364] The server generates a list of appropriate training scenarios based on the user's skill level and transmits the list to the user's terminal.
[2365] 3. Communication Scene Generation
[2366] When a user selects a specific training scenario, the server constructs a specific communication scene based on that scenario and transmits the information to the terminal.
[2367] 4. Real-time analysis and feedback
[2368] The server receives and analyzes the user's speech data and facial expression data in real time using OpenCV (for capturing facial expressions), Keras (for loading emotion recognition models and analyzing emotions), and NLP modules (for natural language processing).
[2369] 5. Emotion recognition and feedback regulation
[2370] The server analyzes the user's emotions using an emotion engine, provides feedback based on the results, and adjusts the training scenario. For example, if the server determines that the user is tense, it provides feedback encouraging the user to relax.
[2371] 6. Saving and analyzing progress data
[2372] After the session ends, the server stores all the user's session data in a database, analyzes the stored data, and provides the user with a visualization of their progress.
[2373] Terminal
[2374] The user's device works in conjunction with the server to provide the following functions:
[2375] 1. User Interface
[2376] The terminal provides an interface for users to interact with the system, including a login screen, a scenario selection screen, and a training start screen.
[2377] 2. Capture and transmit voice and facial expressions
[2378] The device captures the user's speech using a voice input device, recognizes facial expressions using a camera, and transmits the data to a server in real time.
[2379] 3. Displaying feedback and responses
[2380] The device displays feedback and responses received from the server in audio or text format, as well as emotion-based feedback.
[2381] User
[2382] Through this system, users can carry out the following training activities:
[2383] 1. Log in and check your profile
[2384] Users enter their account information and log in to the system. After logging in, they can check their skill level and past training data on the dashboard.
[2385] 2. Selecting a training scenario
[2386] The user selects the scenario that suits them from a list of training scenarios provided.
[2387] 3. Practice in communication situations
[2388] The user trains by actually speaking in a communication scene generated based on the selected scenario.
[2389] 4. Receiving and improving feedback based on emotion recognition
[2390] See real-time feedback and act on it to improve your skills.
[2391] 5. Track your progress over the long term
[2392] Users can view their progress and self-assess on a dashboard, which also includes tracking their emotions and showing their mental progress.
[2393] Specific examples
[2394] For example, if a user selects the customer service scenario "Welcome Customers," the prompt message will be displayed: "As a salesperson in a virtual store, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." The emotion recognition engine will detect tension in the user's facial expressions, and the server will provide feedback such as "Relax and smile more." In this way, the user can improve their customer service skills while receiving specific advice.
[2395] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2396] Step 1:
[2397] The server receives the user's login information and performs authentication.
[2398] Input: User login information (user ID and password)
[2399] Data processing: The server checks the data against the authentication database
[2400] Output: Authentication result (success or failure)
[2401] Specific operation: The server checks the login information against the authentication database, and if authentication is successful, retrieves the user's past training data.
[2402] Step 2:
[2403] The server determines the user's skill level based on past training data.
[2404] Input: Past training data
[2405] Data processing: Analysis of training data and assessment of skill level
[2406] Output: User's skill level
[2407] Specific operation: The server analyzes past training data and applies a skill assessment algorithm to calculate the user's skill level.
[2408] Step 3:
[2409] The server generates an appropriate training scenario based on the user's skill level and transmits it to the terminal.
[2410] Input: User's skill level
[2411] Data processing: Selection of appropriate scenarios from the scenario database
[2412] Output: List of training scenarios
[2413] Specific operation: The server extracts training scenarios that match the user's skill level from the scenario database and sends the list to the terminal.
[2414] Step 4:
[2415] The user selects a training scenario on the terminal.
[2416] Input: List of training scenarios
[2417] Data Calculation: Scenario Selection
[2418] Output: Selected training scenario
[2419] Specific operation: The user selects the desired scenario from the list of training scenarios displayed on the terminal.
[2420] Step 5:
[2421] The server generates a communication scene based on the selected scenario and transmits it to the terminal.
[2422] Input: Selected training scenario
[2423] Data processing: Generation of communication scenes based on scenarios
[2424] Output: Communication scene data
[2425] Specific operation: The server constructs a virtual communication scene based on the selected scenario and sends the data to the terminal.
[2426] Step 6:
[2427] The device captures the user's voice and facial expressions and transmits them to the server.
[2428] Input: User utterances and facial expressions
[2429] Data processing: Audio input device and camera capture
[2430] Output: Capture data
[2431] Specific operation: The device captures the user's voice with a microphone, recognizes facial expressions with a camera, and sends the data to a server.
[2432] Step 7:
[2433] The server analyzes the user's speech data and facial expression data in real time, generates feedback, and sends it to the device.
[2434] Input: User speech data and facial expression data
[2435] Data processing: speech recognition, natural language processing, sentiment analysis
[2436] Output: Real-time feedback
[2437] Specific operation: The server analyzes the user's speech using speech recognition and natural language processing, and analyzes the facial expressions using an emotion engine. Based on the obtained information, it generates feedback and sends it to the device.
[2438] Step 8:
[2439] The terminal presents the feedback received from the server to the user in real time.
[2440] Input: Feedback data
[2441] Data Calculation: Feedback Display
[2442] Output: Feedback that is displayed to the user
[2443] Specific operation: The device displays the feedback received from the server to the user in voice or text.
[2444] Step 9:
[2445] The server stores the progress data in a database and generates visualization data.
[2446] Input: Training session data
[2447] Data processing: saving to database, analyzing and visualizing progress data
[2448] Output: Progress report
[2449] What it does: After the session ends, the server stores all data in a database and generates a report to visualize the progress.
[2450] Step 10:
[2451] Users can check their progress and self-assess through a dashboard.
[2452] Input: Visualized progress data
[2453] Data calculation: Check progress data and self-evaluation
[2454] Output: User self-assessment results
[2455] Specific operation: The user uses the device's dashboard function to check their training progress and perform self-evaluation.
[2456] As a specific example, if a user selects the customer service scenario "welcoming a customer," the prompt message will be displayed: "As a virtual store clerk, you are greeting customers. The emotion engine is analyzing your facial expressions in real time." If tension is detected in the user's facial expression, the feedback will be "Relax and smile more."
[2457] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2458] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2459] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2460] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2461] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2462] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2463] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2464] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2465] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2466] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2467] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2468] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2469] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2470] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2471] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2472] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2473] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2474] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2475] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2476] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2477] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2478] The following is further disclosed regarding the above embodiment.
[2479] (Claim 1)
[2480] Authenticate the user's login information,
[2481] means for determining a user's skill level based on past training data;
[2482] means for providing appropriate training scenarios based on the user's skill level;
[2483] means for generating a communication scene based on a training scenario selected by a user;
[2484] A means for analyzing user utterances in real time and generating feedback;
[2485] A means to store user progress data and track and analyze growth records;
[2486] A system including:
[2487] (Claim 2)
[2488] means for analyzing the user's utterance and presenting generated feedback to the user in real time;
[2489] a means for improving the user's skills based on the provided feedback; and
[2490] 10. The system of claim 1, further comprising:
[2491] (Claim 3)
[2492] Store individual user progress data in a database,
[2493] a means for providing a user with a visualization of the saved progress data;
[2494] a means for users to self-assess and provide information to facilitate growth;
[2495] 10. The system of claim 1, further comprising:
[2496] "Example 1"
[2497] (Claim 1)
[2498] means for verifying the user...
Claims
1. Authenticate the user's login information, means for determining a user's skill level based on past training data; a means for providing appropriate training scenarios based on the user's skill level; means for generating a communication scene based on a training scenario selected by a user; A means for analyzing user utterances in real time and generating feedback; a means for storing user progress data and tracking and analyzing growth records; A system including:
2. means for analyzing the user's utterance and presenting generated feedback to the user in real time; a means for improving the user's skills based on the provided feedback; and The system of claim 1 further comprising:
3. Store individual user progress data in a database, a means for providing a user with a visualization of the saved progress data; a means for users to self-assess and provide information to facilitate growth; The system of claim 1 further comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A