system
The online platform uses natural language processing and facial recognition to provide real-time, personalized feedback on interview skills, addressing the limitations of traditional systems by offering comprehensive and immediate evaluation in simulated scenarios.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Existing systems fail to provide comprehensive and objective feedback on interview skills, particularly in simulated environments, lacking real-time evaluation and personalized advice based on specific evaluation criteria.
An online platform that captures audio and video data in a simulated test scenario, utilizing natural language processing and facial recognition algorithms to generate real-time feedback tailored to individual user needs.
Enables users to receive immediate, personalized feedback on verbal and nonverbal skills, enhancing their preparation for interviews by identifying strengths and weaknesses effectively.
Smart Images

Figure 2026070109000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005] , , ,
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Due to limited interview opportunities, there is a problem that one cannot fully grasp the strengths and weaknesses of one's interview skills. There is also a need for specific feedback based on the evaluation criteria of specific companies or universities, which is difficult to obtain. As a result, there is a problem that there is a shortage of a place where self-improvement can be carried out fairly and effectively.
Means for Solving the Problems
[0005] To address this challenge, the present invention provides an online platform that captures audio and video data based on a simulated test scenario selected by the user, and constructs a system that uses the generated data to perform evaluations in real time using natural language processing and facial recognition algorithms. This allows for the real-time presentation of automatically generated feedback based on the evaluation to the user, and further provides individualized feedback using custom evaluation criteria, thereby significantly expanding opportunities for self-improvement.
[0006] A "user" refers to an individual who uses the system to take mock exams or interviews.
[0007] "Input" refers to operations such as instructions and selections that a user performs on a system.
[0008] A "mock exam scenario" refers to a fictional interview setting tailored to a specific job or industry that the user is practicing for.
[0009] "Audio data" refers to data that is a recording of what the user says.
[0010] "Video data" refers to data that records the user's video footage.
[0011] "Capture" refers to the act of acquiring and recording audio and video data.
[0012] "Streaming" refers to a system that transmits audio and video data to a server in real time.
[0013] "Natural language processing" refers to the technology used to understand and analyze speech or text data.
[0014] A "face recognition algorithm" refers to a technology used to analyze facial expressions and movements from video data.
[0015] "Evaluation" refers to the process of analyzing and quantifying the user's performance in simulated tests.
[0016] "Feedback" refers to information indicating advice or improvement points for the user generated based on the evaluation results.
[0017] "Custom feedback" refers to feedback individually adjusted based on specific evaluation criteria.
Brief Explanation of Drawings
[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [[ID=1,7]] [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0022] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0023] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention provides an online system that allows users to effectively practice interviews using a mock exam scenario of their choice. When a user logs into the system, the terminal presents an input interface that allows the user to select a mock exam scenario based on factors such as industry, job type, and evaluation criteria of a company or university. Based on this selection, the terminal captures the user's audio and video data and streams it to the server in real time.
[0040] The server converts received audio data into text data and analyzes it using natural language processing (NLP) techniques. This evaluates the content and structure of the user's responses and quantifies the results. Simultaneously, the server processes received video data using facial recognition and emotion recognition algorithms to assess the user's nonverbal skills. This evaluates the stability of physical movements and facial expressions, as well as communication skills.
[0041] The server also automatically generates feedback based on the evaluation results. This feedback includes specific advice on areas for improvement in the answers and ways to enhance nonverbal communication. The user sees their evaluation score and advice on their device screen.
[0042] As a concrete example, consider a scenario where a user practices for an IT-related job interview. When the user selects a mock exam scenario for a "software engineer," the device records audio and video, and the server analyzes the user's technical knowledge responses and non-verbal expressions. The server provides feedback such as "make your answers to the technical questions a little more specific," and communicates this to the user through the device.
[0043] This system allows users to clearly understand their strengths and areas for improvement through mock exams, enabling them to efficiently prepare for actual interviews.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user logs into their device and selects a mock exam scenario from the interface. This includes the job type, target company, and university evaluation criteria.
[0047] Step 2:
[0048] The device activates the camera and microphone based on the user's selection, records the audio and video data of the interview session, and streams it to the server in real time.
[0049] Step 3:
[0050] The server analyzes the received audio data using a speech recognition engine and converts it into text data. The converted text is then analyzed using natural language processing technology to evaluate the logic and clarity of the content.
[0051] Step 4:
[0052] The server uses the received video data to execute face recognition and emotion recognition algorithms to evaluate the user's nonverbal expressions and emotional state.
[0053] Step 5:
[0054] The server scores the user's overall interview performance based on the analysis results. Specifically, it quantifies evaluations of the content of answers, verbal expression, and nonverbal skills.
[0055] Step 6:
[0056] The server provides automatically generated feedback based on the evaluation, including suggestions for improving responses and advice on nonverbal communication. Furthermore, if individual evaluation criteria exist, it generates custom feedback based on those criteria.
[0057] Step 7:
[0058] The device displays the score and feedback sent from the server to the user. The user reviews the feedback and creates an improvement plan for the next interview practice.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] In online interview practice, it is difficult for users to effectively and objectively evaluate their own abilities and receive feedback that helps them understand areas for improvement. In particular, feedback on nonverbal communication skills and evaluation criteria specific to particular industries and job roles is required. Traditional methods have made it difficult to provide such multifaceted evaluations in real time.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for setting up a simulated test scenario based on user input using a selection device, means for using a device for capturing and transmitting audio and image data, and means for automatically generating feedback based on evaluation results. This allows the user to be evaluated in real time from multiple perspectives and receive feedback based on specific areas for improvement.
[0064] A "selection device" is a device that has the function of setting up a mock exam scenario based on the user's input.
[0065] "Audio data and image data" refers to information that captures and records the user's voice and video in digital format.
[0066] A "conversion device" is a device that has the function of converting sound data into text data.
[0067] "Natural language processing technology" is a technology that analyzes text data to understand and evaluate its content.
[0068] A "facial feature analysis algorithm" is a technology that uses image data to analyze a user's facial expressions and nonverbal characteristics.
[0069] A "means for generating feedback" refers to a method that has the function of automatically generating areas for improvement and advice based on evaluation results.
[0070] A "device that performs real-time scored evaluations" is a device that has the function of instantly analyzing user data and quantifying the evaluation results.
[0071] This invention provides a system that allows users to select mock exam scenarios online and effectively practice for interviews. Specific embodiments of the invention are described below.
[0072] First, the user selects their desired mock exam scenario through the terminal's interface, including factors such as industry, job type, and evaluation criteria of a specific organization. Based on this selection, the terminal captures the user's voice and image data and streams this data to the server in real time.
[0073] The server converts audio data into text data using software such as Google® Cloud Speech-to-Text API. Furthermore, the server utilizes natural language processing technologies such as SpaCy and NLTK to analyze the user's responses. This analysis includes the recognition of technical terms and the understanding of sentence structure.
[0074] The server uses OpenCV and the dlib library to execute a facial feature analysis algorithm on the received image data to evaluate the user's nonverbal communication skills. This analyzes nonverbal features such as facial expressions and gaze, and provides a quantified evaluation.
[0075] Based on the evaluation results, the server automatically generates feedback through a generative AI model. This feedback includes specific areas for improvement and advice on skill enhancement. For example, it may include suggestions on how to make answers more specific or how to stabilize eye gaze.
[0076] The user's device displays their evaluation score and detailed feedback. This allows users to understand their strengths and areas for improvement, enabling them to prepare for interviews more efficiently.
[0077] A concrete example of a prompt message would be, "Based on the analysis results of the audio and video data from the mock exam scenario selected by the user, generate detailed feedback including areas for improvement." In this way, the present invention provides users with opportunities for self-improvement in a wide range of areas.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The user logs into the system, and the terminal prompts the user to select a mock exam scenario through its interface. The input consists of the user's chosen industry, job type, and evaluation criteria for a specific organization, which allows the terminal to retrieve appropriate scenario information.
[0081] Step 2:
[0082] The device captures audio and video based on user input. It uses a microphone and camera to collect data in real time. Input consists of the user's voice and video, while output consists of captured audio and image data.
[0083] Step 3:
[0084] The terminal streams captured audio and image data to the server. The input is audio and image data, and the output is the transmission of data to the server. The data is appropriately formatted to ensure accurate, real-time streaming.
[0085] Step 4:
[0086] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is streamed audio data, and the output is the corresponding text data. This process transforms the audio content into a structured text format.
[0087] Step 5:
[0088] The server analyzes the converted character data using natural language processing technologies (such as SpaCy or NLTK). The input is character data, and the output is the analysis result. This analysis identifies specialized terminology and evaluates the logical structure of sentences.
[0089] Step 6:
[0090] The server analyzes image data using the OpenCV and dlib libraries. The input is image data, and the output is face recognition and emotion analysis results. This allows for the evaluation of the user's nonverbal characteristics.
[0091] Step 7:
[0092] The server automatically generates feedback using a generative AI model based on the analysis results. The input is analysis results from natural language processing and facial recognition, and the output is a detailed feedback message. The feedback includes areas for improvement and specific advice.
[0093] Step 8:
[0094] The server sends the generated feedback to the terminal, which then displays it to the user. The input is the feedback message, and the output is the display on the terminal's screen. The user uses this information to improve their skills.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] Conventional interview practice systems struggle to simultaneously evaluate a user's voice and video in real time and provide appropriate feedback within a virtual environment. Furthermore, they have limitations in generating feedback based on individual evaluation criteria, hindering immediate and effective training.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes means for selecting a simulated test scenario within a virtual environment based on user input, means for capturing and streaming the user's voice and video data in real time, and means for performing evaluations using natural language processing and image recognition algorithms. This enables the user to receive a quantified evaluation in real time and receive feedback based on individual evaluation criteria.
[0100] "User input" refers to the information that users provide to the system, including selections and settings related to the virtual environment and simulated exam scenarios.
[0101] A "virtual environment" is an artificial environment created by a computer, serving as a platform for users to conduct tests and training.
[0102] A "mock-up test scenario" is a training scenario virtually set up based on specific situations or conditions, in which the user practices responding.
[0103] "Audio and video data" refers to audio and video information obtained from users, which is material that is analyzed by the system.
[0104] "Real-time streaming" is a process in which data such as audio and video is transmitted to a server instantly without delay.
[0105] "Natural language processing" is a technology that uses computers to analyze and understand human language, and is used to evaluate the content of user speech.
[0106] An "image recognition algorithm" is a technology that detects specific features from images and videos and analyzes their content, and is used to evaluate a user's nonverbal communication.
[0107] "Quantified evaluation" refers to an evaluation result that expresses user responses and actions as numerical values according to certain criteria.
[0108] "Feedback" refers to the system's response to user reactions and actions, indicating areas for improvement and strengths of the user.
[0109] The system for realizing this invention is designed to allow users to effectively practice simulated interviews in a virtual environment. The user begins training by first putting on a VR headset and entering a virtual interview environment. The user's voice and video data are captured in real time by a microphone and camera and streamed. The server uses Google Cloud Speech-to-Text to convert the voice data into text data. Furthermore, natural language processing technology, such as spaCy, is used to analyze the structure and content of the user's responses.
[0110] The server also uses OpenCV and TENSORFLOW® to analyze the user's facial expressions and gestures in real time, enabling the assessment of nonverbal communication skills.
[0111] The generated evaluation data is quantified and presented to the user. The device displays feedback based on the evaluation results on a visual display, providing specific advice on areas for improvement in responses and nonverbal communication.
[0112] For example, if a user chooses a mock interview simulating the role of an IT engineer, the questions will be related to technical problem-solving. After the user answers, the server generates feedback such as "Answers based on specific experience will be more persuasive" and provides it to the user.
[0113] Examples of prompts to input into a generative AI model:
[0114] "Analyze the responses to the following question and generate feedback for improvement. Question: Question content Answer: User response"
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The user puts on a VR headset, logs into the terminal, and selects an interview scenario within the virtual environment. The input here is the user's selection information, and the output is the virtual environment's configuration information. Based on the user's selection, the terminal performs actions to build an appropriate virtual interview environment.
[0118] Step 2:
[0119] The device captures the user's audio and video data in real time via the microphone and camera and streams it to the server. The input for this step is the user's real-time audio and video, and the output is streaming data. The device performs the operation of sending the data to the server without delay.
[0120] Step 3:
[0121] The server uses Google Cloud Speech-to-Text to convert received audio data into text data. The input is streamed audio data, and the output is text data. The server uses a speech recognition algorithm to convert the audio into text.
[0122] Step 4:
[0123] The server performs natural language processing based on text data to analyze the user's response. The input is converted text data, and the output is structured data of the response. The server uses spaCy for analysis, performing syntactic and semantic analysis of the response.
[0124] Step 5:
[0125] The server uses OpenCV and TensorFlow to analyze user facial expressions and nonverbal communication from video data. The input is the received video data, and the output is nonverbal evaluation data. The server uses a facial recognition algorithm to evaluate the user's facial expressions and gestures.
[0126] Step 6:
[0127] The server quantifies the user's response based on the analysis results and generates an evaluation score. The input is structured response data and non-verbal evaluation data, and the output is the numerical evaluation result. The server calculates a score based on each data point and performs an overall evaluation.
[0128] Step 7:
[0129] The server automatically generates feedback based on the evaluation results and presents it to the user. The input is the numerical evaluation result, and the output is a feedback message. The server uses a generation AI model to create feedback that includes specific areas for improvement and advice, and sends it to the terminal.
[0130] Step 8:
[0131] The terminal displays feedback sent from the server on its screen for the user to review. The input here is the feedback message, and the output is the displayed feedback. The terminal visually presents the feedback content on the screen.
[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0133] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. When a user logs into the system, the terminal displays an interface for selecting a mock exam scenario. After the user makes a selection, the terminal captures the user's audio and video data using a camera and microphone and streams this data to the server in real time.
[0134] The server converts received audio data into text data using a speech recognition engine and analyzes it using natural language processing (NLP) techniques to evaluate whether the answer is logical or not. Simultaneously, received video data is processed by a facial recognition and emotion engine to evaluate the user's nonverbal expressions and emotional state. The emotion engine analyzes the user's facial expressions in real time and reflects how the recognized emotions are affecting interview performance in the evaluation.
[0135] The server automatically generates feedback based on these evaluation results, including advice tailored to the user's emotional state. The feedback identifies moments of confidence and nervousness during the interview and provides specific ways to improve accordingly. Through this feedback, users can learn how to control their emotions and improve their overall interview performance.
[0136] As a concrete example, consider the scenario where the user selects the "management candidate" scenario. While the device streams live video and the server analyzes the audio, the emotion engine identifies the user's signs of anxiety and confidence. The results of the emotion identification are incorporated into the feedback and displayed to the user as specific advice, such as, "You showed confidence when talking about projects you manage, but nervousness was observed when asked about leadership."
[0137] This system allows users to gain a deeper understanding of how their emotional state affects the quality of their interviews and to develop better communication skills.
[0138] The following describes the processing flow.
[0139] Step 1:
[0140] Users log in to their devices to access the system and select their desired mock exam scenario from the displayed interface. Users can choose their industry, job type, and specific evaluation criteria.
[0141] Step 2:
[0142] The device activates its camera and microphone based on the selected scenario, preparing to record the interview. Once recording begins, audio and video data are streamed to the server in real time.
[0143] Step 3:
[0144] The server receives the audio data and converts it into text format using a speech recognition engine. This text data is then analyzed using natural language processing technology to evaluate the user's responses and their logical connections.
[0145] Step 4:
[0146] The server analyzes video data using facial recognition and emotion engines, tracking changes in the user's facial expressions and gaze. The analysis estimates the user's emotional state, allowing for real-time emotional shifts.
[0147] Step 5:
[0148] The server combines audio and video analysis results to comprehensively score the user's interview performance. This scoring includes factors such as accuracy of answers, confidence levels, and emotional stability.
[0149] Step 6:
[0150] The server automatically generates feedback based on an overall assessment. The feedback highlights areas for improvement in verbal and nonverbal skills, and provides specific advice, particularly based on changes in emotions.
[0151] Step 7:
[0152] The device displays the generated feedback to the user. The user understands areas for improvement based on the detailed evaluation and emotional advice, and uses this information to improve future interview practice.
[0153] (Example 2)
[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0155] In modern interview processes, there is a need to comprehensively and objectively evaluate applicants' skills and emotional state and provide appropriate feedback. However, traditional methods often rely heavily on the interviewer's subjective evaluation, making it difficult to accurately grasp the applicant's actual skills and emotional state. This leads to challenges for applicants, who may not be able to accurately recognize their own strengths and weaknesses, making improvement difficult.
[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0157] In this invention, the server includes means for presenting simulated evaluation options based on user input information, means for acquiring user voice and video information and transmitting it to a communication device, and means for converting voice information into linguistic information. This makes it possible to objectively evaluate the user's skills and emotional state and provide specific methods for improvement.
[0158] "User input information" refers to information about the choices and settings that the user provides to the system.
[0159] "Simulated evaluation options" refer to the scenarios and process options that users can choose from.
[0160] "Audio and video information" refers to data that records the user's speech and visual expressions.
[0161] "Communication equipment" refers to devices that serve as network infrastructure for sending and receiving data.
[0162] "Means of converting audio information into linguistic information" refers to the processes and technologies that convert audio data into text data.
[0163] "Language and sentiment analysis algorithms" are computational methods for analyzing natural language and emotional states.
[0164] "Methods for calculating evaluation scores" refers to the process of quantifying user performance based on data.
[0165] "Feedback information" refers to information that includes improvement suggestions and advice provided based on the evaluation results.
[0166] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. In the implementation of the system, terminals and servers play the primary roles.
[0167] When a user logs into the system on their device, they are presented with options for a practice test. The device uses a display and input device to show these options. Once the user selects their desired scenario, the device activates its camera and microphone to collect the user's audio and video information. This data is streamed to the server in real time. The camera and microphone used may be integrated into the device or external devices.
[0168] The server converts speech information into text using a speech recognition engine. Specifically, it may use a speech recognition service such as Google Speech-to-Text. In addition, natural language processing (NLP) techniques such as SpaCy or the BERT model are used to analyze the text data and evaluate whether the user's response is logical.
[0169] Simultaneously, the server utilizes facial recognition and emotion analysis algorithms to analyze video information. For example, by using OpenCV and DeepFace, it analyzes the user's facial expressions and evaluates their emotional state. This evaluation helps understand the user's nonverbal expressions and is reflected in the assessment of their performance during the interview.
[0170] The server integrates the evaluation results and automatically generates feedback. This feedback includes specific advice tailored to the user's emotional state and skills. For example, if a user undergoes a mock interview as a "management candidate," the server can provide specific examples such as "confidence when discussing projects" and suggest areas for improvement.
[0171] An example of a prompt message would be, "I want to prepare for an interview as a management candidate. Please evaluate my emotional state and responses and provide feedback." This allows users to request specific feedback from the system.
[0172] This system allows users to gain a deeper understanding of how their skills and emotional state affect interviews, and to obtain specific guidance for better performance.
[0173] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0174] Step 1:
[0175] The user logs into the system.
[0176] The user enters their username and password, and the authentication system verifies that the access is legitimate. The user's login information is used as input, and access permission is obtained as output. This prepares the terminal to display the practice test options to the user.
[0177] Step 2:
[0178] The device displays the options for the practice test scenario.
[0179] Multiple scenario options are presented on the display, and the user selects the scenario that best suits their objective. The input is scenario data stored within the system, and the output is a display that visually presents the user with choices. This allows the user to select the appropriate test conditions.
[0180] Step 3:
[0181] The user selects a scenario.
[0182] When the user presses a select button on the terminal, the selected scenario information is confirmed. The user's selection operation is registered as input, and the selected content is registered as output in the system. This prepares the system for the next data capture process.
[0183] Step 4:
[0184] The device captures audio and video data and streams it to the server.
[0185] The camera and microphone are activated, and data of the user's voice and facial expressions are collected. The input is the user's real-time audio and video, and the output is the transmission of that data to the server. This collects material for the server to perform data analysis.
[0186] Step 5:
[0187] The server analyzes the audio data.
[0188] The speech recognition engine converts the streamed audio into text and analyzes it using NLP (Neuro-Linguistic Programming) techniques. The input is the captured audio data, and the output is the text data and evaluation results obtained from that analysis. This allows for an evaluation of the consistency and logic of the user's linguistic responses.
[0189] Step 6:
[0190] The server analyzes the video data.
[0191] Facial recognition and emotion analysis algorithms process video data to evaluate facial expressions and emotional states. The input is the user's video data, and the output is an emotion evaluation result derived from nonverbal expressions. This reveals how the user's emotional state influences interview performance.
[0192] Step 7:
[0193] The server generates feedback based on the evaluation results.
[0194] The analyzed data is integrated to generate feedback that provides specific improvement suggestions and advice to the user. The input is the evaluation results of audio and video, and the output is a written feedback document for the user. This allows the user to obtain the information necessary to improve their interview skills.
[0195] Step 8:
[0196] The user receives feedback.
[0197] The feedback is displayed on the device, allowing the user to review it and implement improvements. The input is the generated feedback information, and the output is the user's understanding and skill improvement. This enables the user to acquire more effective communication techniques.
[0198] (Application Example 2)
[0199] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0200] Traditionally, a challenge has been the lack of effective systems to properly train staff in customer service and emotional control, thereby improving their skills. In particular, it has been difficult to obtain real-time feedback and specific improvement suggestions from facial expressions and voices during actual customer service situations, which has limited the efficiency of training.
[0201] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0202] In this invention, the server includes means for selecting a simulated dialogue scenario based on user input, means for acquiring and transmitting user voice and video information, means for converting voice information into text information, means for performing evaluation using natural language processing and facial recognition technology, means for automatically generating responses and improvement suggestions based on the evaluation, and means for providing responses and improvement suggestions to the user. This enables customer service staff to effectively learn skills and improve themselves in a way that is directly related to their actual work.
[0203] "Means for selecting a simulated dialogue scenario based on user input" refers to a function that allows users to access and select a specific conversation scenario via their terminal.
[0204] "Means for acquiring and transmitting user audio and video information" refers to technology that captures user audio and video data on a device and transmits it to a server in real time.
[0205] "Means for converting audio information into text information" refers to speech recognition technology that has the function of converting the user's spoken language into text data.
[0206] "Means of evaluation using natural language processing and facial recognition technology" refers to technology that processes text and video data extracted from the user's voice to identify and evaluate the logic and emotional state of the conversation.
[0207] "Means for automatically generating responses and improvement suggestions based on evaluation" refers to a function that algorithmically generates areas for improvement and specific feedback for users based on analysis results.
[0208] "Means of providing responses and improvement suggestions to users" refers to a system configuration that displays the generated feedback and advice on the user's device and conveys appropriate information.
[0209] The system used to implement this application is for training customer service staff, with the server, terminal, and user each playing a crucial role. The terminal provides an interface that allows the user to select a specific simulated conversation scenario. Once the user selects a scenario, the terminal uses its built-in camera and microphone to acquire the user's voice and video information. This information is transmitted to the server in real time. The server uses speech recognition technology to convert this voice information into text. For example, Google Cloud Speech-to-Text is used for this process.
[0210] Next, the server performs natural language processing and facial recognition technology. For natural language processing, IBM Watson® NLP is used to analyze user responses and evaluate the logic of the conversation. For facial recognition, Microsoft® Azure® emotion recognition API is used to analyze and evaluate the user's facial expressions and emotional state. Based on the analysis results, the server automatically generates responses and improvement suggestions. This feedback is generated in a way that is appropriate to the specific situation and may include suggestions such as "how to express emotions when handling customer complaints."
[0211] Finally, the server transmits this response and improvement suggestions to the terminal, providing them to the user. This allows the user to receive practical training to improve their skills in their work. This system is particularly effective for customer service staff to learn proper customer service.
[0212] As a concrete example, if a customer service staff member selects a simulated scenario of "handling a customer complaint," the server analyzes the audio and video data and provides feedback on when the user reacted in a way that reflected the customer's emotions. In this process, the following prompts are input to the generative AI model:
[0213] "In the customer service training scenario, generate feedback based on a series of customer interaction data, including staff emotional states and skill assessments."
[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0215] Step 1:
[0216] The user operates the terminal and selects a simulated dialogue scenario from the provided interface. This input process sends the user's selection information from the terminal to the server. The output at this point is the identification information of the selected scenario.
[0217] Step 2:
[0218] Upon receiving a signal to start the scenario, the device begins capturing the user's audio and video using its built-in camera and microphone. The captured data is continuously streamed to the server. The output of this step is real-time collected audio and video data.
[0219] Step 3:
[0220] The server converts the received audio data into text information using Google Cloud Speech-to-Text. The input is audio data, and the output is the corresponding text data as a result of the conversion process.
[0221] Step 4:
[0222] The server analyzes text data using IBM Watson NLP to evaluate the logical coherence and consistency of the language. The input is the converted text data, and the output is the analyzed data, including the evaluation results.
[0223] Step 5:
[0224] The server processes video data using Microsoft Azure's emotion recognition API to determine the user's facial expressions and emotional state. The input for this step is video data, and the output is an analysis result indicating the user's emotional state.
[0225] Step 6:
[0226] Based on the analyzed text data and sentiment recognition results, the server automatically generates responses and improvement suggestions using an AI model. During this process, prompt text is used as input, and specific feedback related to the user's response skills is output.
[0227] Step 7:
[0228] The server sends the generated response and improvement suggestions to the terminal, providing them to the user. The user can then use this information to improve their skills. This output consists of feedback and specific improvement suggestions.
[0229] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0230] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0231] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0232] [Second Embodiment]
[0233] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0234] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0235] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0236] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0237] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0238] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0239] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0240] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0241] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0242] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0243] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0244] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0245] This invention provides an online system that allows users to effectively practice interviews using a mock exam scenario of their choice. When a user logs into the system, the terminal presents an input interface that allows the user to select a mock exam scenario based on factors such as industry, job type, and evaluation criteria of a company or university. Based on this selection, the terminal captures the user's audio and video data and streams it to the server in real time.
[0246] The server converts received audio data into text data and analyzes it using natural language processing (NLP) techniques. This evaluates the content and structure of the user's responses and quantifies the results. Simultaneously, the server processes received video data using facial recognition and emotion recognition algorithms to assess the user's nonverbal skills. This evaluates the stability of physical movements and facial expressions, as well as communication skills.
[0247] The server also automatically generates feedback based on the evaluation results. This feedback includes specific advice on areas for improvement in the answers and ways to enhance nonverbal communication. The user sees their evaluation score and advice on their device screen.
[0248] As a concrete example, consider a scenario where a user practices for an IT-related job interview. When the user selects a mock exam scenario for a "software engineer," the device records audio and video, and the server analyzes the user's technical knowledge responses and non-verbal expressions. The server provides feedback such as "make your answers to the technical questions a little more specific," and communicates this to the user through the device.
[0249] This system allows users to clearly understand their strengths and areas for improvement through mock exams, enabling them to efficiently prepare for actual interviews.
[0250] The following describes the processing flow.
[0251] Step 1:
[0252] The user logs into their device and selects a mock exam scenario from the interface. This includes the job type, target company, and university evaluation criteria.
[0253] Step 2:
[0254] The device activates the camera and microphone based on the user's selection, records the audio and video data of the interview session, and streams it to the server in real time.
[0255] Step 3:
[0256] The server analyzes the received audio data using a speech recognition engine and converts it into text data. The converted text is then analyzed using natural language processing technology to evaluate the logic and clarity of the content.
[0257] Step 4:
[0258] The server uses the received video data to execute face recognition and emotion recognition algorithms to evaluate the user's nonverbal expressions and emotional state.
[0259] Step 5:
[0260] The server scores the user's overall interview performance based on the analysis results. Specifically, it quantifies evaluations of the content of answers, verbal expression, and nonverbal skills.
[0261] Step 6:
[0262] The server provides automatically generated feedback based on the evaluation, including suggestions for improving responses and advice on nonverbal communication. Furthermore, if individual evaluation criteria exist, it generates custom feedback based on those criteria.
[0263] Step 7:
[0264] The device displays the score and feedback sent from the server to the user. The user reviews the feedback and creates an improvement plan for the next interview practice.
[0265] (Example 1)
[0266] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0267] In online interview practice, it is difficult for users to effectively and objectively evaluate their own abilities and receive feedback that helps them understand areas for improvement. In particular, feedback on nonverbal communication skills and evaluation criteria specific to particular industries and job roles is required. Traditional methods have made it difficult to provide such multifaceted evaluations in real time.
[0268] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0269] In this invention, the server includes means for setting up a simulated test scenario based on user input using a selection device, means for using a device for capturing and transmitting audio and image data, and means for automatically generating feedback based on evaluation results. This allows the user to be evaluated in real time from multiple perspectives and receive feedback based on specific areas for improvement.
[0270] A "selection device" is a device that has the function of setting up a mock exam scenario based on the user's input.
[0271] "Audio data and image data" refers to information that captures and records the user's voice and video in digital format.
[0272] A "conversion device" is a device that has the function of converting sound data into text data.
[0273] "Natural language processing technology" is a technology that analyzes text data to understand and evaluate its content.
[0274] A "facial feature analysis algorithm" is a technology that uses image data to analyze a user's facial expressions and nonverbal characteristics.
[0275] A "means for generating feedback" refers to a method that has the function of automatically generating areas for improvement and advice based on evaluation results.
[0276] A "device that performs real-time scored evaluations" is a device that has the function of instantly analyzing user data and quantifying the evaluation results.
[0277] This invention provides a system that allows users to select mock exam scenarios online and effectively practice for interviews. Specific embodiments of the invention are described below.
[0278] First, the user selects their desired mock exam scenario through the terminal's interface, including factors such as industry, job type, and evaluation criteria of a specific organization. Based on this selection, the terminal captures the user's voice and image data and streams this data to the server in real time.
[0279] The server converts audio data into text data using software such as the Google Cloud Speech-to-Text API. Furthermore, the server utilizes natural language processing technologies such as SpaCy and NLTK to analyze the user's responses. This analysis includes the recognition of technical terms and the understanding of sentence structure.
[0280] The server uses OpenCV and the dlib library to execute a facial feature analysis algorithm on the received image data to evaluate the user's nonverbal communication skills. This analyzes nonverbal features such as facial expressions and gaze, and provides a quantified evaluation.
[0281] Based on the evaluation results, the server automatically generates feedback through a generative AI model. This feedback includes specific areas for improvement and advice on skill enhancement. For example, it may include suggestions on how to make answers more specific or how to stabilize eye gaze.
[0282] On the user's terminal, the score of the evaluation result and detailed feedback are displayed. The user can utilize this to efficiently prepare for the interview while grasping their strengths and areas for improvement.
[0283] Specific examples of the prompt sentence include "Please generate detailed feedback including areas for improvement based on the analysis results of audio and video data based on the simulated test scenario selected by the user." In this way, the present invention provides the user with opportunities for self-improvement in a wide range of aspects.
[0284] The flow of the specific process in Example 1 will be described using FIG. 11.
[0285] Step 1:
[0286] The user logs in to the system, and the terminal prompts the user to select a simulated test scenario through the interface. The input is the industry type, job type, and evaluation criteria of a specific group selected by the user, based on which the terminal obtains appropriate scenario information.
[0287] Step 2:
[0288] The terminal captures audio and video based on the user's input. This uses a microphone and a camera to collect data in real time. The input is the user's audio and video, and the output is the captured audio data and image data.
[0289] Step 3:
[0290] The terminal streams the captured audio data and image data to the server. The input is the audio data and image data, and the output is the transmission of the data to the server. The data is appropriately formatted so that it is accurately streamed in real time.
[0291] Step 4:
[0292] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is streamed audio data, and the output is the corresponding text data. This process transforms the audio content into a structured text format.
[0293] Step 5:
[0294] The server analyzes the converted character data using natural language processing technologies (such as SpaCy or NLTK). The input is character data, and the output is the analysis result. This analysis identifies specialized terminology and evaluates the logical structure of sentences.
[0295] Step 6:
[0296] The server analyzes image data using the OpenCV and dlib libraries. The input is image data, and the output is face recognition and emotion analysis results. This allows for the evaluation of the user's nonverbal characteristics.
[0297] Step 7:
[0298] The server automatically generates feedback using a generative AI model based on the analysis results. The input is analysis results from natural language processing and facial recognition, and the output is a detailed feedback message. The feedback includes areas for improvement and specific advice.
[0299] Step 8:
[0300] The server sends the generated feedback to the terminal, which then displays it to the user. The input is the feedback message, and the output is the display on the terminal's screen. The user uses this information to improve their skills.
[0301] (Application Example 1)
[0302] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0303] In a conventional interview training system, it is difficult to simultaneously evaluate a user's voice and video in real time and provide appropriate feedback in a virtual environment. In addition, there is a problem that feedback generation based on individual evaluation criteria is limited and immediate and effective training cannot be performed.
[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0305] In this invention, the server includes means for selecting a simulation test scenario in a virtual environment based on a user's input, means for capturing and streaming the user's voice and video data in real time, and means for performing an evaluation using natural language processing and image recognition algorithms. As a result, the user can receive a numerically evaluated evaluation in real time and receive feedback based on individual evaluation criteria.
[0306] The "user input" is information provided by the user to the system and refers to selections and settings related to the virtual environment and simulation test scenario.
[0307] The "virtual environment" is an artificial environment constructed by computer generation and is a stage for the user to conduct tests and training therein.
[0308] The "simulation test scenario" is a training scenario virtually set based on specific situations and conditions, and the user practices responding to it.
[0309] The "voice and video data" refers to voice information and video information acquired from the user, and these are materials to be analyzed by the system.
[0310] "Real-time streaming" is a process in which data such as audio and video is transmitted to a server instantly without delay.
[0311] "Natural language processing" is a technology that uses computers to analyze and understand human language, and is used to evaluate the content of user speech.
[0312] An "image recognition algorithm" is a technology that detects specific features from images and videos and analyzes their content, and is used to evaluate a user's nonverbal communication.
[0313] "Quantified evaluation" refers to an evaluation result that expresses user responses and actions as numerical values according to certain criteria.
[0314] "Feedback" refers to the system's response to user reactions and actions, indicating areas for improvement and strengths of the user.
[0315] The system for realizing this invention is designed to allow users to effectively practice simulated interviews in a virtual environment. The user begins training by first putting on a VR headset and entering a virtual interview environment. The user's voice and video data are captured in real time by a microphone and camera and streamed. The server uses Google Cloud Speech-to-Text to convert the voice data into text data. Furthermore, natural language processing technology, such as spaCy, is used to analyze the structure and content of the user's responses.
[0316] The server also uses OpenCV and TensorFlow to analyze the user's facial expressions and gestures in real time, enabling the assessment of nonverbal communication skills.
[0317] The generated evaluation data is quantified and presented to the user. The device displays feedback based on the evaluation results on a visual display, providing specific advice on areas for improvement in responses and nonverbal communication.
[0318] For example, if a user chooses a mock interview simulating the role of an IT engineer, the questions will be related to technical problem-solving. After the user answers, the server generates feedback such as "Answers based on specific experience will be more persuasive" and provides it to the user.
[0319] Examples of prompts to input into a generative AI model:
[0320] "Analyze the responses to the following question and generate feedback for improvement. Question: Question content Answer: User response"
[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0322] Step 1:
[0323] The user puts on a VR headset, logs into the terminal, and selects an interview scenario within the virtual environment. The input here is the user's selection information, and the output is the virtual environment's configuration information. Based on the user's selection, the terminal performs actions to build an appropriate virtual interview environment.
[0324] Step 2:
[0325] The device captures the user's audio and video data in real time via the microphone and camera and streams it to the server. The input for this step is the user's real-time audio and video, and the output is streaming data. The device performs the operation of sending the data to the server without delay.
[0326] Step 3:
[0327] The server uses Google Cloud Speech-to-Text to convert received audio data into text data. The input is streamed audio data, and the output is text data. The server uses a speech recognition algorithm to convert the audio into text.
[0328] Step 4:
[0329] The server performs natural language processing based on text data to analyze the user's response. The input is converted text data, and the output is structured data of the response. The server uses spaCy for analysis, performing syntactic and semantic analysis of the response.
[0330] Step 5:
[0331] The server uses OpenCV and TensorFlow to analyze user facial expressions and nonverbal communication from video data. The input is the received video data, and the output is nonverbal evaluation data. The server uses a facial recognition algorithm to evaluate the user's facial expressions and gestures.
[0332] Step 6:
[0333] The server quantifies the user's response based on the analysis results and generates an evaluation score. The input is structured response data and non-verbal evaluation data, and the output is the numerical evaluation result. The server calculates a score based on each data point and performs an overall evaluation.
[0334] Step 7:
[0335] The server automatically generates feedback based on the evaluation results and presents it to the user. The input is the numerical evaluation result, and the output is a feedback message. The server uses a generation AI model to create feedback that includes specific areas for improvement and advice, and sends it to the terminal.
[0336] Step 8:
[0337] The terminal displays feedback sent from the server on its screen for the user to review. The input here is the feedback message, and the output is the displayed feedback. The terminal visually presents the feedback content on the screen.
[0338] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0339] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. When a user logs into the system, the terminal displays an interface for selecting a mock exam scenario. After the user makes a selection, the terminal captures the user's audio and video data using a camera and microphone and streams this data to the server in real time.
[0340] The server converts received audio data into text data using a speech recognition engine and analyzes it using natural language processing (NLP) techniques to evaluate whether the answer is logical or not. Simultaneously, received video data is processed by a facial recognition and emotion engine to evaluate the user's nonverbal expressions and emotional state. The emotion engine analyzes the user's facial expressions in real time and reflects how the recognized emotions are affecting interview performance in the evaluation.
[0341] The server automatically generates feedback based on these evaluation results, including advice tailored to the user's emotional state. The feedback identifies moments of confidence and nervousness during the interview and provides specific ways to improve accordingly. Through this feedback, users can learn how to control their emotions and improve their overall interview performance.
[0342] As a concrete example, consider the scenario where the user selects the "management candidate" scenario. While the device streams live video and the server analyzes the audio, the emotion engine identifies the user's signs of anxiety and confidence. The results of the emotion identification are incorporated into the feedback and displayed to the user as specific advice, such as, "You showed confidence when talking about projects you manage, but nervousness was observed when asked about leadership."
[0343] This system allows users to gain a deeper understanding of how their emotional state affects the quality of their interviews and to develop better communication skills.
[0344] The following describes the processing flow.
[0345] Step 1:
[0346] Users log in to their devices to access the system and select their desired mock exam scenario from the displayed interface. Users can choose their industry, job type, and specific evaluation criteria.
[0347] Step 2:
[0348] The device activates its camera and microphone based on the selected scenario, preparing to record the interview. Once recording begins, audio and video data are streamed to the server in real time.
[0349] Step 3:
[0350] The server receives the audio data and converts it into text format using a speech recognition engine. This text data is then analyzed using natural language processing technology to evaluate the user's responses and their logical connections.
[0351] Step 4:
[0352] The server analyzes video data using facial recognition and emotion engines, tracking changes in the user's facial expressions and gaze. The analysis estimates the user's emotional state, allowing for real-time emotional shifts.
[0353] Step 5:
[0354] The server combines audio and video analysis results to comprehensively score the user's interview performance. This scoring includes factors such as accuracy of answers, confidence levels, and emotional stability.
[0355] Step 6:
[0356] The server automatically generates feedback based on an overall assessment. The feedback highlights areas for improvement in verbal and nonverbal skills, and provides specific advice, particularly based on changes in emotions.
[0357] Step 7:
[0358] The device displays the generated feedback to the user. The user understands areas for improvement based on the detailed evaluation and emotional advice, and uses this information to improve future interview practice.
[0359] (Example 2)
[0360] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0361] In modern interview processes, there is a need to comprehensively and objectively evaluate applicants' skills and emotional state and provide appropriate feedback. However, traditional methods often rely heavily on the interviewer's subjective evaluation, making it difficult to accurately grasp the applicant's actual skills and emotional state. This leads to challenges for applicants, who may not be able to accurately recognize their own strengths and weaknesses, making improvement difficult.
[0362] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0363] In this invention, the server includes means for presenting simulated evaluation options based on user input information, means for acquiring user voice and video information and transmitting it to a communication device, and means for converting voice information into linguistic information. This makes it possible to objectively evaluate the user's skills and emotional state and provide specific methods for improvement.
[0364] "User input information" refers to information about the choices and settings that the user provides to the system.
[0365] "Simulated evaluation options" refer to the scenarios and process options that users can choose from.
[0366] "Audio and video information" refers to data that records the user's speech and visual expressions.
[0367] "Communication equipment" refers to devices that serve as network infrastructure for sending and receiving data.
[0368] "Means of converting audio information into linguistic information" refers to the processes and technologies that convert audio data into text data.
[0369] "Language and sentiment analysis algorithms" are computational methods for analyzing natural language and emotional states.
[0370] "Methods for calculating evaluation scores" refers to the process of quantifying user performance based on data.
[0371] "Feedback information" refers to information that includes improvement suggestions and advice provided based on the evaluation results.
[0372] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. In the implementation of the system, terminals and servers play the primary roles.
[0373] When a user logs into the system on their device, they are presented with options for a practice test. The device uses a display and input device to show these options. Once the user selects their desired scenario, the device activates its camera and microphone to collect the user's audio and video information. This data is streamed to the server in real time. The camera and microphone used may be integrated into the device or external devices.
[0374] The server converts speech information into text using a speech recognition engine. Specifically, it may use a speech recognition service such as Google Speech-to-Text. In addition, natural language processing (NLP) techniques such as SpaCy or the BERT model are used to analyze the text data and evaluate whether the user's response is logical.
[0375] Simultaneously, the server utilizes facial recognition and emotion analysis algorithms to analyze video information. For example, by using OpenCV and DeepFace, it analyzes the user's facial expressions and evaluates their emotional state. This evaluation helps understand the user's nonverbal expressions and is reflected in the assessment of their performance during the interview.
[0376] The server integrates the evaluation results and automatically generates feedback. This feedback includes specific advice tailored to the user's emotional state and skills. For example, if a user undergoes a mock interview as a "management candidate," the server can provide specific examples such as "confidence when discussing projects" and suggest areas for improvement.
[0377] An example of a prompt message would be, "I want to prepare for an interview as a management candidate. Please evaluate my emotional state and responses and provide feedback." This allows users to request specific feedback from the system.
[0378] This system allows users to gain a deeper understanding of how their skills and emotional state affect interviews, and to obtain specific guidance for better performance.
[0379] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0380] Step 1:
[0381] The user logs into the system.
[0382] The user enters their username and password, and the authentication system verifies that the access is legitimate. The user's login information is used as input, and access permission is obtained as output. This prepares the terminal to display the practice test options to the user.
[0383] Step 2:
[0384] The device displays the options for the practice test scenario.
[0385] Multiple scenario options are presented on the display, and the user selects the scenario that best suits their objective. The input is scenario data stored within the system, and the output is a display that visually presents the user with choices. This allows the user to select the appropriate test conditions.
[0386] Step 3:
[0387] The user selects a scenario.
[0388] When the user presses a select button on the terminal, the selected scenario information is confirmed. The user's selection operation is registered as input, and the selected content is registered as output in the system. This prepares the system for the next data capture process.
[0389] Step 4:
[0390] The device captures audio and video data and streams it to the server.
[0391] The camera and microphone are activated, and data of the user's voice and facial expressions are collected. The input is the user's real-time audio and video, and the output is the transmission of that data to the server. This collects material for the server to perform data analysis.
[0392] Step 5:
[0393] The server analyzes the audio data.
[0394] The speech recognition engine converts the streamed audio into text and analyzes it using NLP (Neuro-Linguistic Programming) techniques. The input is the captured audio data, and the output is the text data and evaluation results obtained from that analysis. This allows for an evaluation of the consistency and logic of the user's linguistic responses.
[0395] Step 6:
[0396] The server analyzes the video data.
[0397] Facial recognition and emotion analysis algorithms process video data to evaluate facial expressions and emotional states. The input is the user's video data, and the output is an emotion evaluation result derived from nonverbal expressions. This reveals how the user's emotional state influences interview performance.
[0398] Step 7:
[0399] The server generates feedback based on the evaluation results.
[0400] The analyzed data is integrated to generate feedback that provides specific improvement suggestions and advice to the user. The input is the evaluation results of audio and video, and the output is a written feedback document for the user. This allows the user to obtain the information necessary to improve their interview skills.
[0401] Step 8:
[0402] The user receives feedback.
[0403] The feedback is displayed on the device, allowing the user to review it and implement improvements. The input is the generated feedback information, and the output is the user's understanding and skill improvement. This enables the user to acquire more effective communication techniques.
[0404] (Application Example 2)
[0405] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0406] Traditionally, a challenge has been the lack of effective systems to properly train staff in customer service and emotional control, thereby improving their skills. In particular, it has been difficult to obtain real-time feedback and specific improvement suggestions from facial expressions and voices during actual customer service situations, which has limited the efficiency of training.
[0407] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0408] In this invention, the server includes means for selecting a simulated dialogue scenario based on user input, means for acquiring and transmitting user voice and video information, means for converting voice information into text information, means for performing evaluation using natural language processing and facial recognition technology, means for automatically generating responses and improvement suggestions based on the evaluation, and means for providing responses and improvement suggestions to the user. This enables customer service staff to effectively learn skills and improve themselves in a way that is directly related to their actual work.
[0409] "Means for selecting a simulated dialogue scenario based on user input" refers to a function that allows users to access and select a specific conversation scenario via their terminal.
[0410] "Means for acquiring and transmitting user audio and video information" refers to technology that captures user audio and video data on a device and transmits it to a server in real time.
[0411] "Means for converting audio information into text information" refers to speech recognition technology that has the function of converting the user's spoken language into text data.
[0412] "Means of evaluation using natural language processing and facial recognition technology" refers to technology that processes text and video data extracted from the user's voice to identify and evaluate the logic and emotional state of the conversation.
[0413] "Means for automatically generating responses and improvement suggestions based on evaluation" refers to a function that algorithmically generates areas for improvement and specific feedback for users based on analysis results.
[0414] "Means of providing responses and improvement suggestions to users" refers to a system configuration that displays the generated feedback and advice on the user's device and conveys appropriate information.
[0415] The system used to implement this application is for training customer service staff, with the server, terminal, and user each playing a crucial role. The terminal provides an interface that allows the user to select a specific simulated conversation scenario. Once the user selects a scenario, the terminal uses its built-in camera and microphone to acquire the user's voice and video information. This information is transmitted to the server in real time. The server uses speech recognition technology to convert this voice information into text. For example, Google Cloud Speech-to-Text is used for this process.
[0416] Next, the server performs natural language processing and facial recognition technology. IBM Watson NLP is used for natural language processing, analyzing user responses and evaluating the logic of the conversation. For facial recognition, Microsoft Azure's emotion recognition API is used to analyze and evaluate the user's facial expressions and emotional state. Based on the analysis results, the server automatically generates responses and improvement suggestions. This feedback is generated in a way that is appropriate to the specific situation and may include suggestions such as "how to express emotions when handling customer complaints."
[0417] Finally, the server transmits this response and improvement suggestions to the terminal, providing them to the user. This allows the user to receive practical training to improve their skills in their work. This system is particularly effective for customer service staff to learn proper customer service.
[0418] As a concrete example, if a customer service staff member selects a simulated scenario of "handling a customer complaint," the server analyzes the audio and video data and provides feedback on when the user reacted in a way that reflected the customer's emotions. In this process, the following prompts are input to the generative AI model:
[0419] "In the customer service training scenario, generate feedback based on a series of customer interaction data, including staff emotional states and skill assessments."
[0420] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0421] Step 1:
[0422] The user operates the terminal and selects a simulated dialogue scenario from the provided interface. This input process sends the user's selection information from the terminal to the server. The output at this point is the identification information of the selected scenario.
[0423] Step 2:
[0424] Upon receiving a signal to start the scenario, the device begins capturing the user's audio and video using its built-in camera and microphone. The captured data is continuously streamed to the server. The output of this step is real-time collected audio and video data.
[0425] Step 3:
[0426] The server converts the received audio data into text information using Google Cloud Speech-to-Text. The input is audio data, and the output is the corresponding text data as a result of the conversion process.
[0427] Step 4:
[0428] The server analyzes text data using IBM Watson NLP to evaluate the logical coherence and consistency of the language. The input is the converted text data, and the output is the analyzed data, including the evaluation results.
[0429] Step 5:
[0430] The server processes video data using Microsoft Azure's emotion recognition API to determine the user's facial expressions and emotional state. The input for this step is video data, and the output is an analysis result indicating the user's emotional state.
[0431] Step 6:
[0432] Based on the analyzed text data and sentiment recognition results, the server automatically generates responses and improvement suggestions using an AI model. During this process, prompt text is used as input, and specific feedback related to the user's response skills is output.
[0433] Step 7:
[0434] The server sends the generated response and improvement suggestions to the terminal, providing them to the user. The user can then use this information to improve their skills. This output consists of feedback and specific improvement suggestions.
[0435] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0436] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0437] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0438] [Third Embodiment]
[0439] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0440] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0441] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0442] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0443] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0445] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0446] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0449] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0451] This invention provides an online system that allows users to effectively practice interviews using a mock exam scenario of their choice. When a user logs into the system, the terminal presents an input interface that allows the user to select a mock exam scenario based on factors such as industry, job type, and evaluation criteria of a company or university. Based on this selection, the terminal captures the user's audio and video data and streams it to the server in real time.
[0452] The server converts received audio data into text data and analyzes it using natural language processing (NLP) techniques. This evaluates the content and structure of the user's responses and quantifies the results. Simultaneously, the server processes received video data using facial recognition and emotion recognition algorithms to assess the user's nonverbal skills. This evaluates the stability of physical movements and facial expressions, as well as communication skills.
[0453] The server also automatically generates feedback based on the evaluation results. This feedback includes specific advice on areas for improvement in the answers and ways to enhance nonverbal communication. The user sees their evaluation score and advice on their device screen.
[0454] As a concrete example, consider a scenario where a user practices for an IT-related job interview. When the user selects a mock exam scenario for a "software engineer," the device records audio and video, and the server analyzes the user's technical knowledge responses and non-verbal expressions. The server provides feedback such as "make your answers to the technical questions a little more specific," and communicates this to the user through the device.
[0455] This system allows users to clearly understand their strengths and areas for improvement through mock exams, enabling them to efficiently prepare for actual interviews.
[0456] The following describes the processing flow.
[0457] Step 1:
[0458] The user logs into their device and selects a mock exam scenario from the interface. This includes the job type, target company, and university evaluation criteria.
[0459] Step 2:
[0460] The device activates the camera and microphone based on the user's selection, records the audio and video data of the interview session, and streams it to the server in real time.
[0461] Step 3:
[0462] The server analyzes the received audio data using a speech recognition engine and converts it into text data. The converted text is then analyzed using natural language processing technology to evaluate the logic and clarity of the content.
[0463] Step 4:
[0464] The server uses the received video data to execute face recognition and emotion recognition algorithms to evaluate the user's nonverbal expressions and emotional state.
[0465] Step 5:
[0466] The server scores the user's overall interview performance based on the analysis results. Specifically, it quantifies evaluations of the content of answers, verbal expression, and nonverbal skills.
[0467] Step 6:
[0468] The server provides automatically generated feedback based on the evaluation, including suggestions for improving responses and advice on nonverbal communication. Furthermore, if individual evaluation criteria exist, it generates custom feedback based on those criteria.
[0469] Step 7:
[0470] The device displays the score and feedback sent from the server to the user. The user reviews the feedback and creates an improvement plan for the next interview practice.
[0471] (Example 1)
[0472] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0473] In online interview practice, it is difficult for users to effectively and objectively evaluate their own abilities and receive feedback that helps them understand areas for improvement. In particular, feedback on nonverbal communication skills and evaluation criteria specific to particular industries and job roles is required. Traditional methods have made it difficult to provide such multifaceted evaluations in real time.
[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0475] In this invention, the server includes means for setting up a simulated test scenario based on user input using a selection device, means for using a device for capturing and transmitting audio and image data, and means for automatically generating feedback based on evaluation results. This allows the user to be evaluated in real time from multiple perspectives and receive feedback based on specific areas for improvement.
[0476] A "selection device" is a device that has the function of setting up a mock exam scenario based on the user's input.
[0477] "Audio data and image data" refers to information that captures and records the user's voice and video in digital format.
[0478] A "conversion device" is a device that has the function of converting sound data into text data.
[0479] "Natural language processing technology" is a technology that analyzes text data to understand and evaluate its content.
[0480] A "facial feature analysis algorithm" is a technology that uses image data to analyze a user's facial expressions and nonverbal characteristics.
[0481] A "means for generating feedback" refers to a method that has the function of automatically generating areas for improvement and advice based on evaluation results.
[0482] A "device that performs real-time scored evaluations" is a device that has the function of instantly analyzing user data and quantifying the evaluation results.
[0483] This invention provides a system that allows users to select mock exam scenarios online and effectively practice for interviews. Specific embodiments of the invention are described below.
[0484] First, the user selects their desired mock exam scenario through the terminal's interface, including factors such as industry, job type, and evaluation criteria of a specific organization. Based on this selection, the terminal captures the user's voice and image data and streams this data to the server in real time.
[0485] The server converts audio data into text data using software such as the Google Cloud Speech-to-Text API. Furthermore, the server utilizes natural language processing technologies such as SpaCy and NLTK to analyze the user's responses. This analysis includes the recognition of technical terms and the understanding of sentence structure.
[0486] The server uses OpenCV and the dlib library to execute a facial feature analysis algorithm on the received image data to evaluate the user's nonverbal communication skills. This analyzes nonverbal features such as facial expressions and gaze, and provides a quantified evaluation.
[0487] Based on the evaluation results, the server automatically generates feedback through a generative AI model. This feedback includes specific areas for improvement and advice on skill enhancement. For example, it may include suggestions on how to make answers more specific or how to stabilize eye gaze.
[0488] The user's device displays their evaluation score and detailed feedback. This allows users to understand their strengths and areas for improvement, enabling them to prepare for interviews more efficiently.
[0489] A concrete example of a prompt message would be, "Based on the analysis results of the audio and video data from the mock exam scenario selected by the user, generate detailed feedback including areas for improvement." In this way, the present invention provides users with opportunities for self-improvement in a wide range of areas.
[0490] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0491] Step 1:
[0492] The user logs into the system, and the terminal prompts the user to select a mock exam scenario through its interface. The input consists of the user's chosen industry, job type, and evaluation criteria for a specific organization, which allows the terminal to retrieve appropriate scenario information.
[0493] Step 2:
[0494] The device captures audio and video based on user input. It uses a microphone and camera to collect data in real time. Input consists of the user's voice and video, while output consists of captured audio and image data.
[0495] Step 3:
[0496] The terminal streams captured audio and image data to the server. The input is audio and image data, and the output is the transmission of data to the server. The data is appropriately formatted to ensure accurate, real-time streaming.
[0497] Step 4:
[0498] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is streamed audio data, and the output is the corresponding text data. This process transforms the audio content into a structured text format.
[0499] Step 5:
[0500] The server analyzes the converted character data using natural language processing technologies (such as SpaCy or NLTK). The input is character data, and the output is the analysis result. This analysis identifies specialized terminology and evaluates the logical structure of sentences.
[0501] Step 6:
[0502] The server analyzes image data using the OpenCV and dlib libraries. The input is image data, and the output is face recognition and emotion analysis results. This allows for the evaluation of the user's nonverbal characteristics.
[0503] Step 7:
[0504] The server automatically generates feedback using a generative AI model based on the analysis results. The input is analysis results from natural language processing and facial recognition, and the output is a detailed feedback message. The feedback includes areas for improvement and specific advice.
[0505] Step 8:
[0506] The server sends the generated feedback to the terminal, which then displays it to the user. The input is the feedback message, and the output is the display on the terminal's screen. The user uses this information to improve their skills.
[0507] (Application Example 1)
[0508] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0509] Conventional interview practice systems struggle to simultaneously evaluate a user's voice and video in real time and provide appropriate feedback within a virtual environment. Furthermore, they have limitations in generating feedback based on individual evaluation criteria, hindering immediate and effective training.
[0510] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0511] In this invention, the server includes means for selecting a simulated test scenario within a virtual environment based on user input, means for capturing and streaming the user's voice and video data in real time, and means for performing evaluations using natural language processing and image recognition algorithms. This enables the user to receive a quantified evaluation in real time and receive feedback based on individual evaluation criteria.
[0512] "User input" refers to the information that users provide to the system, including selections and settings related to the virtual environment and simulated exam scenarios.
[0513] A "virtual environment" is an artificial environment created by a computer, serving as a platform for users to conduct tests and training.
[0514] A "mock-up test scenario" is a training scenario virtually set up based on specific situations or conditions, in which the user practices responding.
[0515] "Audio and video data" refers to audio and video information obtained from users, which is material that is analyzed by the system.
[0516] "Real-time streaming" is a process in which data such as audio and video is transmitted to a server instantly without delay.
[0517] "Natural language processing" is a technology that uses computers to analyze and understand human language, and is used to evaluate the content of user speech.
[0518] An "image recognition algorithm" is a technology that detects specific features from images and videos and analyzes their content, and is used to evaluate a user's nonverbal communication.
[0519] "Quantified evaluation" refers to an evaluation result that expresses user responses and actions as numerical values according to certain criteria.
[0520] "Feedback" refers to the system's response to user reactions and actions, indicating areas for improvement and strengths of the user.
[0521] The system for realizing this invention is designed to allow users to effectively practice simulated interviews in a virtual environment. The user begins training by first putting on a VR headset and entering a virtual interview environment. The user's voice and video data are captured in real time by a microphone and camera and streamed. The server uses Google Cloud Speech-to-Text to convert the voice data into text data. Furthermore, natural language processing technology, such as spaCy, is used to analyze the structure and content of the user's responses.
[0522] The server also uses OpenCV and TensorFlow to analyze the user's facial expressions and gestures in real time, enabling the assessment of nonverbal communication skills.
[0523] The generated evaluation data is quantified and presented to the user. The device displays feedback based on the evaluation results on a visual display, providing specific advice on areas for improvement in responses and nonverbal communication.
[0524] For example, if a user chooses a mock interview simulating the role of an IT engineer, the questions will be related to technical problem-solving. After the user answers, the server generates feedback such as "Answers based on specific experience will be more persuasive" and provides it to the user.
[0525] Examples of prompts to input into a generative AI model:
[0526] "Analyze the responses to the following question and generate feedback for improvement. Question: Question content Answer: User response"
[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0528] Step 1:
[0529] The user puts on a VR headset, logs into the terminal, and selects an interview scenario within the virtual environment. The input here is the user's selection information, and the output is the virtual environment's configuration information. Based on the user's selection, the terminal performs actions to build an appropriate virtual interview environment.
[0530] Step 2:
[0531] The device captures the user's audio and video data in real time via the microphone and camera and streams it to the server. The input for this step is the user's real-time audio and video, and the output is streaming data. The device performs the operation of sending the data to the server without delay.
[0532] Step 3:
[0533] The server uses Google Cloud Speech-to-Text to convert received audio data into text data. The input is streamed audio data, and the output is text data. The server uses a speech recognition algorithm to convert the audio into text.
[0534] Step 4:
[0535] The server performs natural language processing based on text data to analyze the user's response. The input is converted text data, and the output is structured data of the response. The server uses spaCy for analysis, performing syntactic and semantic analysis of the response.
[0536] Step 5:
[0537] The server uses OpenCV and TensorFlow to analyze user facial expressions and nonverbal communication from video data. The input is the received video data, and the output is nonverbal evaluation data. The server uses a facial recognition algorithm to evaluate the user's facial expressions and gestures.
[0538] Step 6:
[0539] The server quantifies the user's response based on the analysis results and generates an evaluation score. The input is structured response data and non-verbal evaluation data, and the output is the numerical evaluation result. The server calculates a score based on each data point and performs an overall evaluation.
[0540] Step 7:
[0541] The server automatically generates feedback based on the evaluation results and presents it to the user. The input is the numerical evaluation result, and the output is a feedback message. The server uses a generation AI model to create feedback that includes specific areas for improvement and advice, and sends it to the terminal.
[0542] Step 8:
[0543] The terminal displays feedback sent from the server on its screen for the user to review. The input here is the feedback message, and the output is the displayed feedback. The terminal visually presents the feedback content on the screen.
[0544] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0545] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. When a user logs into the system, the terminal displays an interface for selecting a mock exam scenario. After the user makes a selection, the terminal captures the user's audio and video data using a camera and microphone and streams this data to the server in real time.
[0546] The server converts received audio data into text data using a speech recognition engine and analyzes it using natural language processing (NLP) techniques to evaluate whether the answer is logical or not. Simultaneously, received video data is processed by a facial recognition and emotion engine to evaluate the user's nonverbal expressions and emotional state. The emotion engine analyzes the user's facial expressions in real time and reflects how the recognized emotions are affecting interview performance in the evaluation.
[0547] The server automatically generates feedback based on these evaluation results, including advice tailored to the user's emotional state. The feedback identifies moments of confidence and nervousness during the interview and provides specific ways to improve accordingly. Through this feedback, users can learn how to control their emotions and improve their overall interview performance.
[0548] As a concrete example, consider the scenario where the user selects the "management candidate" scenario. While the device streams live video and the server analyzes the audio, the emotion engine identifies the user's signs of anxiety and confidence. The results of the emotion identification are incorporated into the feedback and displayed to the user as specific advice, such as, "You showed confidence when talking about projects you manage, but nervousness was observed when asked about leadership."
[0549] This system allows users to gain a deeper understanding of how their emotional state affects the quality of their interviews and to develop better communication skills.
[0550] The following describes the processing flow.
[0551] Step 1:
[0552] Users log in to their devices to access the system and select their desired mock exam scenario from the displayed interface. Users can choose their industry, job type, and specific evaluation criteria.
[0553] Step 2:
[0554] The device activates its camera and microphone based on the selected scenario, preparing to record the interview. Once recording begins, audio and video data are streamed to the server in real time.
[0555] Step 3:
[0556] The server receives the audio data and converts it into text format using a speech recognition engine. This text data is then analyzed using natural language processing technology to evaluate the user's responses and their logical connections.
[0557] Step 4:
[0558] The server analyzes video data using facial recognition and emotion engines, tracking changes in the user's facial expressions and gaze. The analysis estimates the user's emotional state, allowing for real-time emotional shifts.
[0559] Step 5:
[0560] The server combines audio and video analysis results to comprehensively score the user's interview performance. This scoring includes factors such as accuracy of answers, confidence levels, and emotional stability.
[0561] Step 6:
[0562] The server automatically generates feedback based on an overall assessment. The feedback highlights areas for improvement in verbal and nonverbal skills, and provides specific advice, particularly based on changes in emotions.
[0563] Step 7:
[0564] The device displays the generated feedback to the user. The user understands areas for improvement based on the detailed evaluation and emotional advice, and uses this information to improve future interview practice.
[0565] (Example 2)
[0566] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0567] In modern interview processes, there is a need to comprehensively and objectively evaluate applicants' skills and emotional state and provide appropriate feedback. However, traditional methods often rely heavily on the interviewer's subjective evaluation, making it difficult to accurately grasp the applicant's actual skills and emotional state. This leads to challenges for applicants, who may not be able to accurately recognize their own strengths and weaknesses, making improvement difficult.
[0568] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0569] In this invention, the server includes means for presenting simulated evaluation options based on user input information, means for acquiring user voice and video information and transmitting it to a communication device, and means for converting voice information into linguistic information. This makes it possible to objectively evaluate the user's skills and emotional state and provide specific methods for improvement.
[0570] "User input information" refers to information about the choices and settings that the user provides to the system.
[0571] "Simulated evaluation options" refer to the scenarios and process options that users can choose from.
[0572] "Audio and video information" refers to data that records the user's speech and visual expressions.
[0573] "Communication equipment" refers to devices that serve as network infrastructure for sending and receiving data.
[0574] "Means of converting audio information into linguistic information" refers to the processes and technologies that convert audio data into text data.
[0575] "Language and sentiment analysis algorithms" are computational methods for analyzing natural language and emotional states.
[0576] "Methods for calculating evaluation scores" refers to the process of quantifying user performance based on data.
[0577] "Feedback information" refers to information that includes improvement suggestions and advice provided based on the evaluation results.
[0578] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. In the implementation of the system, terminals and servers play the primary roles.
[0579] When a user logs into the system on their device, they are presented with options for a practice test. The device uses a display and input device to show these options. Once the user selects their desired scenario, the device activates its camera and microphone to collect the user's audio and video information. This data is streamed to the server in real time. The camera and microphone used may be integrated into the device or external devices.
[0580] The server converts speech information into text using a speech recognition engine. Specifically, it may use a speech recognition service such as Google Speech-to-Text. In addition, natural language processing (NLP) techniques such as SpaCy or the BERT model are used to analyze the text data and evaluate whether the user's response is logical.
[0581] Simultaneously, the server utilizes facial recognition and emotion analysis algorithms to analyze video information. For example, by using OpenCV and DeepFace, it analyzes the user's facial expressions and evaluates their emotional state. This evaluation helps understand the user's nonverbal expressions and is reflected in the assessment of their performance during the interview.
[0582] The server integrates the evaluation results and automatically generates feedback. This feedback includes specific advice tailored to the user's emotional state and skills. For example, if a user undergoes a mock interview as a "management candidate," the server can provide specific examples such as "confidence when discussing projects" and suggest areas for improvement.
[0583] An example of a prompt message would be, "I want to prepare for an interview as a management candidate. Please evaluate my emotional state and responses and provide feedback." This allows users to request specific feedback from the system.
[0584] This system allows users to gain a deeper understanding of how their skills and emotional state affect interviews, and to obtain specific guidance for better performance.
[0585] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0586] Step 1:
[0587] The user logs into the system.
[0588] The user enters their username and password, and the authentication system verifies that the access is legitimate. The user's login information is used as input, and access permission is obtained as output. This prepares the terminal to display the practice test options to the user.
[0589] Step 2:
[0590] The device displays the options for the practice test scenario.
[0591] Multiple scenario options are presented on the display, and the user selects the scenario that best suits their objective. The input is scenario data stored within the system, and the output is a display that visually presents the user with choices. This allows the user to select the appropriate test conditions.
[0592] Step 3:
[0593] The user selects a scenario.
[0594] When the user presses a select button on the terminal, the selected scenario information is confirmed. The user's selection operation is registered as input, and the selected content is registered as output in the system. This prepares the system for the next data capture process.
[0595] Step 4:
[0596] The device captures audio and video data and streams it to the server.
[0597] The camera and microphone are activated, and data of the user's voice and facial expressions are collected. The input is the user's real-time audio and video, and the output is the transmission of that data to the server. This collects material for the server to perform data analysis.
[0598] Step 5:
[0599] The server analyzes the audio data.
[0600] The speech recognition engine converts the streamed audio into text and analyzes it using NLP (Neuro-Linguistic Programming) techniques. The input is the captured audio data, and the output is the text data and evaluation results obtained from that analysis. This allows for an evaluation of the consistency and logic of the user's linguistic responses.
[0601] Step 6:
[0602] The server analyzes the video data.
[0603] Facial recognition and emotion analysis algorithms process video data to evaluate facial expressions and emotional states. The input is the user's video data, and the output is an emotion evaluation result derived from nonverbal expressions. This reveals how the user's emotional state influences interview performance.
[0604] Step 7:
[0605] The server generates feedback based on the evaluation results.
[0606] The analyzed data is integrated to generate feedback that provides specific improvement suggestions and advice to the user. The input is the evaluation results of audio and video, and the output is a written feedback document for the user. This allows the user to obtain the information necessary to improve their interview skills.
[0607] Step 8:
[0608] The user receives feedback.
[0609] The feedback is displayed on the device, allowing the user to review it and implement improvements. The input is the generated feedback information, and the output is the user's understanding and skill improvement. This enables the user to acquire more effective communication techniques.
[0610] (Application Example 2)
[0611] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0612] Traditionally, a challenge has been the lack of effective systems to properly train staff in customer service and emotional control, thereby improving their skills. In particular, it has been difficult to obtain real-time feedback and specific improvement suggestions from facial expressions and voices during actual customer service situations, which has limited the efficiency of training.
[0613] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0614] In this invention, the server includes means for selecting a simulated dialogue scenario based on user input, means for acquiring and transmitting user voice and video information, means for converting voice information into text information, means for performing evaluation using natural language processing and facial recognition technology, means for automatically generating responses and improvement suggestions based on the evaluation, and means for providing responses and improvement suggestions to the user. This enables customer service staff to effectively learn skills and improve themselves in a way that is directly related to their actual work.
[0615] "Means for selecting a simulated dialogue scenario based on user input" refers to a function that allows users to access and select a specific conversation scenario via their terminal.
[0616] "Means for acquiring and transmitting user audio and video information" refers to technology that captures user audio and video data on a device and transmits it to a server in real time.
[0617] "Means for converting audio information into text information" refers to speech recognition technology that has the function of converting the user's spoken language into text data.
[0618] "Means of evaluation using natural language processing and facial recognition technology" refers to technology that processes text and video data extracted from the user's voice to identify and evaluate the logic and emotional state of the conversation.
[0619] "Means for automatically generating responses and improvement suggestions based on evaluation" refers to a function that algorithmically generates areas for improvement and specific feedback for users based on analysis results.
[0620] "Means of providing responses and improvement suggestions to users" refers to a system configuration that displays the generated feedback and advice on the user's device and conveys appropriate information.
[0621] The system used to implement this application is for training customer service staff, with the server, terminal, and user each playing a crucial role. The terminal provides an interface that allows the user to select a specific simulated conversation scenario. Once the user selects a scenario, the terminal uses its built-in camera and microphone to acquire the user's voice and video information. This information is transmitted to the server in real time. The server uses speech recognition technology to convert this voice information into text. For example, Google Cloud Speech-to-Text is used for this process.
[0622] Next, the server performs natural language processing and facial recognition technology. IBM Watson NLP is used for natural language processing, analyzing user responses and evaluating the logic of the conversation. For facial recognition, Microsoft Azure's emotion recognition API is used to analyze and evaluate the user's facial expressions and emotional state. Based on the analysis results, the server automatically generates responses and improvement suggestions. This feedback is generated in a way that is appropriate to the specific situation and may include suggestions such as "how to express emotions when handling customer complaints."
[0623] Finally, the server transmits this response and improvement suggestions to the terminal, providing them to the user. This allows the user to receive practical training to improve their skills in their work. This system is particularly effective for customer service staff to learn proper customer service.
[0624] As a concrete example, if a customer service staff member selects a simulated scenario of "handling a customer complaint," the server analyzes the audio and video data and provides feedback on when the user reacted in a way that reflected the customer's emotions. In this process, the following prompts are input to the generative AI model:
[0625] "In the customer service training scenario, generate feedback based on a series of customer interaction data, including staff emotional states and skill assessments."
[0626] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0627] Step 1:
[0628] The user operates the terminal and selects a simulated dialogue scenario from the provided interface. This input process sends the user's selection information from the terminal to the server. The output at this point is the identification information of the selected scenario.
[0629] Step 2:
[0630] Upon receiving a signal to start the scenario, the device begins capturing the user's audio and video using its built-in camera and microphone. The captured data is continuously streamed to the server. The output of this step is real-time collected audio and video data.
[0631] Step 3:
[0632] The server converts the received audio data into text information using Google Cloud Speech-to-Text. The input is audio data, and the output is the corresponding text data as a result of the conversion process.
[0633] Step 4:
[0634] The server analyzes text data using IBM Watson NLP to evaluate the logical coherence and consistency of the language. The input is the converted text data, and the output is the analyzed data, including the evaluation results.
[0635] Step 5:
[0636] The server processes video data using Microsoft Azure's emotion recognition API to determine the user's facial expressions and emotional state. The input for this step is video data, and the output is an analysis result indicating the user's emotional state.
[0637] Step 6:
[0638] Based on the analyzed text data and sentiment recognition results, the server automatically generates responses and improvement suggestions using an AI model. During this process, prompt text is used as input, and specific feedback related to the user's response skills is output.
[0639] Step 7:
[0640] The server sends the generated response and improvement suggestions to the terminal, providing them to the user. The user can then use this information to improve their skills. This output consists of feedback and specific improvement suggestions.
[0641] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0642] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0643] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0644] [Fourth Embodiment]
[0645] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0646] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0647] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0648] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0649] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0651] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0652] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0653] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0654] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0655] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0656] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0657] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0658] This invention provides an online system that allows users to effectively practice interviews using a mock exam scenario of their choice. When a user logs into the system, the terminal presents an input interface that allows the user to select a mock exam scenario based on factors such as industry, job type, and evaluation criteria of a company or university. Based on this selection, the terminal captures the user's audio and video data and streams it to the server in real time.
[0659] The server converts received audio data into text data and analyzes it using natural language processing (NLP) techniques. This evaluates the content and structure of the user's responses and quantifies the results. Simultaneously, the server processes received video data using facial recognition and emotion recognition algorithms to assess the user's nonverbal skills. This evaluates the stability of physical movements and facial expressions, as well as communication skills.
[0660] The server also automatically generates feedback based on the evaluation results. This feedback includes specific advice on areas for improvement in the answers and ways to enhance nonverbal communication. The user sees their evaluation score and advice on their device screen.
[0661] As a concrete example, consider a scenario where a user practices for an IT-related job interview. When the user selects a mock exam scenario for a "software engineer," the device records audio and video, and the server analyzes the user's technical knowledge responses and non-verbal expressions. The server provides feedback such as "make your answers to the technical questions a little more specific," and communicates this to the user through the device.
[0662] This system allows users to clearly understand their strengths and areas for improvement through mock exams, enabling them to efficiently prepare for actual interviews.
[0663] The following describes the processing flow.
[0664] Step 1:
[0665] The user logs into their device and selects a mock exam scenario from the interface. This includes the job type, target company, and university evaluation criteria.
[0666] Step 2:
[0667] The device activates the camera and microphone based on the user's selection, records the audio and video data of the interview session, and streams it to the server in real time.
[0668] Step 3:
[0669] The server analyzes the received audio data using a speech recognition engine and converts it into text data. The converted text is then analyzed using natural language processing technology to evaluate the logic and clarity of the content.
[0670] Step 4:
[0671] The server uses the received video data to execute face recognition and emotion recognition algorithms to evaluate the user's nonverbal expressions and emotional state.
[0672] Step 5:
[0673] The server scores the user's overall interview performance based on the analysis results. Specifically, it quantifies evaluations of the content of answers, verbal expression, and nonverbal skills.
[0674] Step 6:
[0675] The server provides automatically generated feedback based on the evaluation, including suggestions for improving responses and advice on nonverbal communication. Furthermore, if individual evaluation criteria exist, it generates custom feedback based on those criteria.
[0676] Step 7:
[0677] The device displays the score and feedback sent from the server to the user. The user reviews the feedback and creates an improvement plan for the next interview practice.
[0678] (Example 1)
[0679] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0680] In online interview practice, it is difficult for users to effectively and objectively evaluate their own abilities and receive feedback that helps them understand areas for improvement. In particular, feedback on nonverbal communication skills and evaluation criteria specific to particular industries and job roles is required. Traditional methods have made it difficult to provide such multifaceted evaluations in real time.
[0681] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0682] In this invention, the server includes means for setting up a simulated test scenario based on user input using a selection device, means for using a device for capturing and transmitting audio and image data, and means for automatically generating feedback based on evaluation results. This allows the user to be evaluated in real time from multiple perspectives and receive feedback based on specific areas for improvement.
[0683] A "selection device" is a device that has the function of setting up a mock exam scenario based on the user's input.
[0684] "Audio data and image data" refers to information that captures and records the user's voice and video in digital format.
[0685] A "conversion device" is a device that has the function of converting sound data into text data.
[0686] "Natural language processing technology" is a technology that analyzes text data to understand and evaluate its content.
[0687] A "facial feature analysis algorithm" is a technology that uses image data to analyze a user's facial expressions and nonverbal characteristics.
[0688] A "means for generating feedback" refers to a method that has the function of automatically generating areas for improvement and advice based on evaluation results.
[0689] A "device that performs real-time scored evaluations" is a device that has the function of instantly analyzing user data and quantifying the evaluation results.
[0690] This invention provides a system that allows users to select mock exam scenarios online and effectively practice for interviews. Specific embodiments of the invention are described below.
[0691] First, the user selects their desired mock exam scenario through the terminal's interface, including factors such as industry, job type, and evaluation criteria of a specific organization. Based on this selection, the terminal captures the user's voice and image data and streams this data to the server in real time.
[0692] The server converts audio data into text data using software such as the Google Cloud Speech-to-Text API. Furthermore, the server utilizes natural language processing technologies such as SpaCy and NLTK to analyze the user's responses. This analysis includes the recognition of technical terms and the understanding of sentence structure.
[0693] The server uses OpenCV and the dlib library to execute a facial feature analysis algorithm on the received image data to evaluate the user's nonverbal communication skills. This analyzes nonverbal features such as facial expressions and gaze, and provides a quantified evaluation.
[0694] Based on the evaluation results, the server automatically generates feedback through a generative AI model. This feedback includes specific areas for improvement and advice on skill enhancement. For example, it may include suggestions on how to make answers more specific or how to stabilize eye gaze.
[0695] The user's device displays their evaluation score and detailed feedback. This allows users to understand their strengths and areas for improvement, enabling them to prepare for interviews more efficiently.
[0696] A concrete example of a prompt message would be, "Based on the analysis results of the audio and video data from the mock exam scenario selected by the user, generate detailed feedback including areas for improvement." In this way, the present invention provides users with opportunities for self-improvement in a wide range of areas.
[0697] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0698] Step 1:
[0699] The user logs into the system, and the terminal prompts the user to select a mock exam scenario through its interface. The input consists of the user's chosen industry, job type, and evaluation criteria for a specific organization, which allows the terminal to retrieve appropriate scenario information.
[0700] Step 2:
[0701] The device captures audio and video based on user input. It uses a microphone and camera to collect data in real time. Input consists of the user's voice and video, while output consists of captured audio and image data.
[0702] Step 3:
[0703] The terminal streams captured audio and image data to the server. The input is audio and image data, and the output is the transmission of data to the server. The data is appropriately formatted to ensure accurate, real-time streaming.
[0704] Step 4:
[0705] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is streamed audio data, and the output is the corresponding text data. This process transforms the audio content into a structured text format.
[0706] Step 5:
[0707] The server analyzes the converted character data using natural language processing technologies (such as SpaCy or NLTK). The input is character data, and the output is the analysis result. This analysis identifies specialized terminology and evaluates the logical structure of sentences.
[0708] Step 6:
[0709] The server analyzes image data using the OpenCV and dlib libraries. The input is image data, and the output is face recognition and emotion analysis results. This allows for the evaluation of the user's nonverbal characteristics.
[0710] Step 7:
[0711] The server automatically generates feedback using a generative AI model based on the analysis results. The input is analysis results from natural language processing and facial recognition, and the output is a detailed feedback message. The feedback includes areas for improvement and specific advice.
[0712] Step 8:
[0713] The server sends the generated feedback to the terminal, which then displays it to the user. The input is the feedback message, and the output is the display on the terminal's screen. The user uses this information to improve their skills.
[0714] (Application Example 1)
[0715] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] Conventional interview practice systems struggle to simultaneously evaluate a user's voice and video in real time and provide appropriate feedback within a virtual environment. Furthermore, they have limitations in generating feedback based on individual evaluation criteria, hindering immediate and effective training.
[0717] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0718] In this invention, the server includes means for selecting a simulated test scenario within a virtual environment based on user input, means for capturing and streaming the user's voice and video data in real time, and means for performing evaluations using natural language processing and image recognition algorithms. This enables the user to receive a quantified evaluation in real time and receive feedback based on individual evaluation criteria.
[0719] "User input" refers to the information that users provide to the system, including selections and settings related to the virtual environment and simulated exam scenarios.
[0720] A "virtual environment" is an artificial environment created by a computer, serving as a platform for users to conduct tests and training.
[0721] A "mock-up test scenario" is a training scenario virtually set up based on specific situations or conditions, in which the user practices responding.
[0722] "Audio and video data" refers to audio and video information obtained from users, which is material that is analyzed by the system.
[0723] "Real-time streaming" is a process in which data such as audio and video is transmitted to a server instantly without delay.
[0724] "Natural language processing" is a technology that uses computers to analyze and understand human language, and is used to evaluate the content of user speech.
[0725] An "image recognition algorithm" is a technology that detects specific features from images and videos and analyzes their content, and is used to evaluate a user's nonverbal communication.
[0726] "Quantified evaluation" refers to an evaluation result that expresses user responses and actions as numerical values according to certain criteria.
[0727] "Feedback" refers to the system's response to user reactions and actions, indicating areas for improvement and strengths of the user.
[0728] The system for realizing this invention is designed to allow users to effectively practice simulated interviews in a virtual environment. The user begins training by first putting on a VR headset and entering a virtual interview environment. The user's voice and video data are captured in real time by a microphone and camera and streamed. The server uses Google Cloud Speech-to-Text to convert the voice data into text data. Furthermore, natural language processing technology, such as spaCy, is used to analyze the structure and content of the user's responses.
[0729] The server also uses OpenCV and TensorFlow to analyze the user's facial expressions and gestures in real time, enabling the assessment of nonverbal communication skills.
[0730] The generated evaluation data is quantified and presented to the user. The device displays feedback based on the evaluation results on a visual display, providing specific advice on areas for improvement in responses and nonverbal communication.
[0731] For example, if a user chooses a mock interview simulating the role of an IT engineer, the questions will be related to technical problem-solving. After the user answers, the server generates feedback such as "Answers based on specific experience will be more persuasive" and provides it to the user.
[0732] Examples of prompts to input into a generative AI model:
[0733] "Analyze the responses to the following question and generate feedback for improvement. Question: Question content Answer: User response"
[0734] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0735] Step 1:
[0736] The user puts on a VR headset, logs into the terminal, and selects an interview scenario within the virtual environment. The input here is the user's selection information, and the output is the virtual environment's configuration information. Based on the user's selection, the terminal performs actions to build an appropriate virtual interview environment.
[0737] Step 2:
[0738] The device captures the user's audio and video data in real time via the microphone and camera and streams it to the server. The input for this step is the user's real-time audio and video, and the output is streaming data. The device performs the operation of sending the data to the server without delay.
[0739] Step 3:
[0740] The server uses Google Cloud Speech-to-Text to convert received audio data into text data. The input is streamed audio data, and the output is text data. The server uses a speech recognition algorithm to convert the audio into text.
[0741] Step 4:
[0742] The server performs natural language processing based on text data to analyze the user's response. The input is converted text data, and the output is structured data of the response. The server uses spaCy for analysis, performing syntactic and semantic analysis of the response.
[0743] Step 5:
[0744] The server uses OpenCV and TensorFlow to analyze user facial expressions and nonverbal communication from video data. The input is the received video data, and the output is nonverbal evaluation data. The server uses a facial recognition algorithm to evaluate the user's facial expressions and gestures.
[0745] Step 6:
[0746] The server quantifies the user's response based on the analysis results and generates an evaluation score. The input is structured response data and non-verbal evaluation data, and the output is the numerical evaluation result. The server calculates a score based on each data point and performs an overall evaluation.
[0747] Step 7:
[0748] The server automatically generates feedback based on the evaluation results and presents it to the user. The input is the numerical evaluation result, and the output is a feedback message. The server uses a generation AI model to create feedback that includes specific areas for improvement and advice, and sends it to the terminal.
[0749] Step 8:
[0750] The terminal displays feedback sent from the server on its screen for the user to review. The input here is the feedback message, and the output is the displayed feedback. The terminal visually presents the feedback content on the screen.
[0751] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0752] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. When a user logs into the system, the terminal displays an interface for selecting a mock exam scenario. After the user makes a selection, the terminal captures the user's audio and video data using a camera and microphone and streams this data to the server in real time.
[0753] The server converts received audio data into text data using a speech recognition engine and analyzes it using natural language processing (NLP) techniques to evaluate whether the answer is logical or not. Simultaneously, received video data is processed by a facial recognition and emotion engine to evaluate the user's nonverbal expressions and emotional state. The emotion engine analyzes the user's facial expressions in real time and reflects how the recognized emotions are affecting interview performance in the evaluation.
[0754] The server automatically generates feedback based on these evaluation results, including advice tailored to the user's emotional state. The feedback identifies moments of confidence and nervousness during the interview and provides specific ways to improve accordingly. Through this feedback, users can learn how to control their emotions and improve their overall interview performance.
[0755] As a concrete example, consider the scenario where the user selects the "management candidate" scenario. While the device streams live video and the server analyzes the audio, the emotion engine identifies the user's signs of anxiety and confidence. The results of the emotion identification are incorporated into the feedback and displayed to the user as specific advice, such as, "You showed confidence when talking about projects you manage, but nervousness was observed when asked about leadership."
[0756] This system allows users to gain a deeper understanding of how their emotional state affects the quality of their interviews and to develop better communication skills.
[0757] The following describes the processing flow.
[0758] Step 1:
[0759] Users log in to their devices to access the system and select their desired mock exam scenario from the displayed interface. Users can choose their industry, job type, and specific evaluation criteria.
[0760] Step 2:
[0761] The device activates its camera and microphone based on the selected scenario, preparing to record the interview. Once recording begins, audio and video data are streamed to the server in real time.
[0762] Step 3:
[0763] The server receives the audio data and converts it into text format using a speech recognition engine. This text data is then analyzed using natural language processing technology to evaluate the user's responses and their logical connections.
[0764] Step 4:
[0765] The server analyzes video data using facial recognition and emotion engines, tracking changes in the user's facial expressions and gaze. The analysis estimates the user's emotional state, allowing for real-time emotional shifts.
[0766] Step 5:
[0767] The server combines audio and video analysis results to comprehensively score the user's interview performance. This scoring includes factors such as accuracy of answers, confidence levels, and emotional stability.
[0768] Step 6:
[0769] The server automatically generates feedback based on an overall assessment. The feedback highlights areas for improvement in verbal and nonverbal skills, and provides specific advice, particularly based on changes in emotions.
[0770] Step 7:
[0771] The device displays the generated feedback to the user. The user understands areas for improvement based on the detailed evaluation and emotional advice, and uses this information to improve future interview practice.
[0772] (Example 2)
[0773] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0774] In modern interview processes, there is a need to comprehensively and objectively evaluate applicants' skills and emotional state and provide appropriate feedback. However, traditional methods often rely heavily on the interviewer's subjective evaluation, making it difficult to accurately grasp the applicant's actual skills and emotional state. This leads to challenges for applicants, who may not be able to accurately recognize their own strengths and weaknesses, making improvement difficult.
[0775] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0776] In this invention, the server includes means for presenting simulated evaluation options based on user input information, means for acquiring user voice and video information and transmitting it to a communication device, and means for converting voice information into linguistic information. This makes it possible to objectively evaluate the user's skills and emotional state and provide specific methods for improvement.
[0777] "User input information" refers to information about the choices and settings that the user provides to the system.
[0778] "Simulated evaluation options" refer to the scenarios and process options that users can choose from.
[0779] "Audio and video information" refers to data that records the user's speech and visual expressions.
[0780] "Communication equipment" refers to devices that serve as network infrastructure for sending and receiving data.
[0781] "Means of converting audio information into linguistic information" refers to the processes and technologies that convert audio data into text data.
[0782] "Language and sentiment analysis algorithms" are computational methods for analyzing natural language and emotional states.
[0783] "Methods for calculating evaluation scores" refers to the process of quantifying user performance based on data.
[0784] "Feedback information" refers to information that includes improvement suggestions and advice provided based on the evaluation results.
[0785] This invention provides an online interview support system that integrates user skill assessment and emotional state recognition. In the implementation of the system, terminals and servers play the primary roles.
[0786] When a user logs into the system on their device, they are presented with options for a practice test. The device uses a display and input device to show these options. Once the user selects their desired scenario, the device activates its camera and microphone to collect the user's audio and video information. This data is streamed to the server in real time. The camera and microphone used may be integrated into the device or external devices.
[0787] The server converts speech information into text using a speech recognition engine. Specifically, it may use a speech recognition service such as Google Speech-to-Text. In addition, natural language processing (NLP) techniques such as SpaCy or the BERT model are used to analyze the text data and evaluate whether the user's response is logical.
[0788] Simultaneously, the server utilizes facial recognition and emotion analysis algorithms to analyze video information. For example, by using OpenCV and DeepFace, it analyzes the user's facial expressions and evaluates their emotional state. This evaluation helps understand the user's nonverbal expressions and is reflected in the assessment of their performance during the interview.
[0789] The server integrates the evaluation results and automatically generates feedback. This feedback includes specific advice tailored to the user's emotional state and skills. For example, if a user undergoes a mock interview as a "management candidate," the server can provide specific examples such as "confidence when discussing projects" and suggest areas for improvement.
[0790] An example of a prompt message would be, "I want to prepare for an interview as a management candidate. Please evaluate my emotional state and responses and provide feedback." This allows users to request specific feedback from the system.
[0791] This system allows users to gain a deeper understanding of how their skills and emotional state affect interviews, and to obtain specific guidance for better performance.
[0792] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0793] Step 1:
[0794] The user logs into the system.
[0795] The user enters their username and password, and the authentication system verifies that the access is legitimate. The user's login information is used as input, and access permission is obtained as output. This prepares the terminal to display the practice test options to the user.
[0796] Step 2:
[0797] The device displays the options for the practice test scenario.
[0798] Multiple scenario options are presented on the display, and the user selects the scenario that best suits their objective. The input is scenario data stored within the system, and the output is a display that visually presents the user with choices. This allows the user to select the appropriate test conditions.
[0799] Step 3:
[0800] The user selects a scenario.
[0801] When the user presses a select button on the terminal, the selected scenario information is confirmed. The user's selection operation is registered as input, and the selected content is registered as output in the system. This prepares the system for the next data capture process.
[0802] Step 4:
[0803] The device captures audio and video data and streams it to the server.
[0804] The camera and microphone are activated, and data of the user's voice and facial expressions are collected. The input is the user's real-time audio and video, and the output is the transmission of that data to the server. This collects material for the server to perform data analysis.
[0805] Step 5:
[0806] The server analyzes the audio data.
[0807] The speech recognition engine converts the streamed audio into text and analyzes it using NLP (Neuro-Linguistic Programming) techniques. The input is the captured audio data, and the output is the text data and evaluation results obtained from that analysis. This allows for an evaluation of the consistency and logic of the user's linguistic responses.
[0808] Step 6:
[0809] The server analyzes the video data.
[0810] Facial recognition and emotion analysis algorithms process video data to evaluate facial expressions and emotional states. The input is the user's video data, and the output is an emotion evaluation result derived from nonverbal expressions. This reveals how the user's emotional state influences interview performance.
[0811] Step 7:
[0812] The server generates feedback based on the evaluation results.
[0813] The analyzed data is integrated to generate feedback that provides specific improvement suggestions and advice to the user. The input is the evaluation results of audio and video, and the output is a written feedback document for the user. This allows the user to obtain the information necessary to improve their interview skills.
[0814] Step 8:
[0815] The user receives feedback.
[0816] The feedback is displayed on the device, allowing the user to review it and implement improvements. The input is the generated feedback information, and the output is the user's understanding and skill improvement. This enables the user to acquire more effective communication techniques.
[0817] (Application Example 2)
[0818] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0819] Traditionally, a challenge has been the lack of effective systems to properly train staff in customer service and emotional control, thereby improving their skills. In particular, it has been difficult to obtain real-time feedback and specific improvement suggestions from facial expressions and voices during actual customer service situations, which has limited the efficiency of training.
[0820] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0821] In this invention, the server includes means for selecting a simulated dialogue scenario based on user input, means for acquiring and transmitting user voice and video information, means for converting voice information into text information, means for performing evaluation using natural language processing and facial recognition technology, means for automatically generating responses and improvement suggestions based on the evaluation, and means for providing responses and improvement suggestions to the user. This enables customer service staff to effectively learn skills and improve themselves in a way that is directly related to their actual work.
[0822] "Means for selecting a simulated dialogue scenario based on user input" refers to a function that allows users to access and select a specific conversation scenario via their terminal.
[0823] "Means for acquiring and transmitting user audio and video information" refers to technology that captures user audio and video data on a device and transmits it to a server in real time.
[0824] "Means for converting audio information into text information" refers to speech recognition technology that has the function of converting the user's spoken language into text data.
[0825] "Means of evaluation using natural language processing and facial recognition technology" refers to technology that processes text and video data extracted from the user's voice to identify and evaluate the logic and emotional state of the conversation.
[0826] "Means for automatically generating responses and improvement suggestions based on evaluation" refers to a function that algorithmically generates areas for improvement and specific feedback for users based on analysis results.
[0827] "Means of providing responses and improvement suggestions to users" refers to a system configuration that displays the generated feedback and advice on the user's device and conveys appropriate information.
[0828] The system used to implement this application is for training customer service staff, with the server, terminal, and user each playing a crucial role. The terminal provides an interface that allows the user to select a specific simulated conversation scenario. Once the user selects a scenario, the terminal uses its built-in camera and microphone to acquire the user's voice and video information. This information is transmitted to the server in real time. The server uses speech recognition technology to convert this voice information into text. For example, Google Cloud Speech-to-Text is used for this process.
[0829] Next, the server performs natural language processing and facial recognition technology. IBM Watson NLP is used for natural language processing, analyzing user responses and evaluating the logic of the conversation. For facial recognition, Microsoft Azure's emotion recognition API is used to analyze and evaluate the user's facial expressions and emotional state. Based on the analysis results, the server automatically generates responses and improvement suggestions. This feedback is generated in a way that is appropriate to the specific situation and may include suggestions such as "how to express emotions when handling customer complaints."
[0830] Finally, the server transmits this response and improvement suggestions to the terminal, providing them to the user. This allows the user to receive practical training to improve their skills in their work. This system is particularly effective for customer service staff to learn proper customer service.
[0831] As a concrete example, if a customer service staff member selects a simulated scenario of "handling a customer complaint," the server analyzes the audio and video data and provides feedback on when the user reacted in a way that reflected the customer's emotions. In this process, the following prompts are input to the generative AI model:
[0832] "In the customer service training scenario, generate feedback based on a series of customer interaction data, including staff emotional states and skill assessments."
[0833] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0834] Step 1:
[0835] The user operates the terminal and selects a simulated dialogue scenario from the provided interface. This input process sends the user's selection information from the terminal to the server. The output at this point is the identification information of the selected scenario.
[0836] Step 2:
[0837] Upon receiving a signal to start the scenario, the device begins capturing the user's audio and video using its built-in camera and microphone. The captured data is continuously streamed to the server. The output of this step is real-time collected audio and video data.
[0838] Step 3:
[0839] The server converts the received audio data into text information using Google Cloud Speech-to-Text. The input is audio data, and the output is the corresponding text data as a result of the conversion process.
[0840] Step 4:
[0841] The server analyzes text data using IBM Watson NLP to evaluate the logical coherence and consistency of the language. The input is the converted text data, and the output is the analyzed data, including the evaluation results.
[0842] Step 5:
[0843] The server processes video data using Microsoft Azure's emotion recognition API to determine the user's facial expressions and emotional state. The input for this step is video data, and the output is an analysis result indicating the user's emotional state.
[0844] Step 6:
[0845] Based on the analyzed text data and sentiment recognition results, the server automatically generates responses and improvement suggestions using an AI model. During this process, prompt text is used as input, and specific feedback related to the user's response skills is output.
[0846] Step 7:
[0847] The server sends the generated response and improvement suggestions to the terminal, providing them to the user. The user can then use this information to improve their skills. This output consists of feedback and specific improvement suggestions.
[0848] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0849] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0850] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0851] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0852] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0853] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0854] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0855] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0856] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0857] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0858] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0859] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0860] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0861] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0862] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0863] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0864] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0865] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0866] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0867] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0868] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0869] The following is further disclosed regarding the embodiments described above.
[0870] (Claim 1)
[0871] A means of selecting a mock exam scenario based on user input,
[0872] A means for capturing and streaming user audio and video data,
[0873] A means of converting audio data into text data,
[0874] A means of performing evaluation using natural language processing and facial recognition algorithms,
[0875] A means of automatically generating feedback based on evaluations,
[0876] Means of presenting feedback to users,
[0877] A system that includes this.
[0878] (Claim 2)
[0879] The system according to claim 1, comprising means for generating custom feedback based on individual evaluation criteria.
[0880] (Claim 3)
[0881] The system according to claim 1, comprising means for performing a real-time scored evaluation.
[0882] "Example 1"
[0883] (Claim 1)
[0884] A means for setting up a mock exam scenario based on user input using a selection device,
[0885] A means of using a device that captures and transmits sound data and image data,
[0886] A means of using a conversion device to convert sound data into text data,
[0887] A means for evaluating content using natural language processing technology and facial feature analysis algorithms,
[0888] A means of automatically generating feedback based on evaluation results,
[0889] A means of using a device that provides generated feedback to the user,
[0890] A system that includes this.
[0891] (Claim 2)
[0892] The system according to claim 1, which uses a device that generates custom feedback based on individual evaluation criteria.
[0893] (Claim 3)
[0894] The system according to claim 1, which uses a device that performs real-time scoring evaluations.
[0895] "Application Example 1"
[0896] (Claim 1)
[0897] A means of selecting a simulated exam scenario within a virtual environment based on user input,
[0898] A means of capturing user audio and video data and streaming it in real time,
[0899] A means of converting audio data into text data,
[0900] A means of performing evaluation using natural language processing and image recognition algorithms,
[0901] A means of automatically generating feedback based on evaluation,
[0902] A means of presenting feedback to the user through a display device,
[0903] A system that includes this.
[0904] (Claim 2)
[0905] The system according to claim 1, comprising means for generating feedback based on individual evaluation criteria.
[0906] (Claim 3)
[0907] The system according to claim 1, comprising means for performing a real-time quantified evaluation.
[0908] "Example 2 of combining an emotion engine"
[0909] (Claim 1)
[0910] A means of presenting simulated evaluation options based on user input information,
[0911] A means for acquiring user voice and video information and transmitting it to a communication device,
[0912] A means of converting audio information into linguistic information,
[0913] A means of conducting evaluation using language analysis and sentiment analysis algorithms,
[0914] A means of automatically generating feedback based on evaluations,
[0915] Means of presenting feedback information to the user,
[0916] A system that includes this.
[0917] (Claim 2)
[0918] The system according to claim 1, comprising means for generating custom feedback based on individual evaluation criteria.
[0919] (Claim 3)
[0920] The system according to claim 1, comprising means for calculating an evaluation score in real time.
[0921] "Application example 2 when combining with an emotional engine"
[0922] (Claim 1)
[0923] A means for selecting a simulated dialogue scenario based on user input,
[0924] A means for acquiring and transmitting user audio and video information,
[0925] A means of converting audio information into text information,
[0926] A means of performing evaluation using natural language processing and facial recognition technology,
[0927] A means for automatically generating responses and improvement suggestions based on evaluations,
[0928] Means for providing responses and improvement suggestions to users,
[0929] A system that includes this.
[0930] (Claim 2)
[0931] The system according to claim 1, comprising means for generating custom improvement suggestions based on individual evaluation criteria.
[0932] (Claim 3)
[0933] The system according to claim 1, comprising means for performing a real-time scored evaluation. [Explanation of symbols]
[0934] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of selecting a mock exam scenario based on user input, A means for capturing and streaming user audio and video data, A means of converting audio data into text data, A means of performing evaluation using natural language processing and facial recognition algorithms, A means of automatically generating feedback based on evaluations, Means of presenting feedback to users, A system that includes this.
2. The system according to claim 1, comprising means for generating custom feedback based on individual evaluation criteria.
3. The system according to claim 1, comprising means for performing a real-time scored evaluation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A