system

A virtual reality-based interview system addresses unconscious biases by real-time data analysis and feedback, enhancing the fairness and inclusivity of recruitment processes.

JP2026070267APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional interview processes are hindered by unconscious biases, leading to unfair evaluation of candidates with diverse backgrounds, resulting in the potential loss of talented individuals and reduced corporate competitiveness.

Method used

A simulated interview system using virtual reality that records users' voice and actions in real time, analyzes them with artificial intelligence to detect biases, and provides immediate feedback, recommending subsequent training scenarios for continuous improvement.

Benefits of technology

Enables fair and inclusive recruitment by helping users recognize and correct unconscious biases, promoting diversity and inclusivity in talent acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070267000001_ABST
    Figure 2026070267000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A computer device that provides virtual reality, comprising means for generating a simulated interview environment, A means for recording the user's voice and actions in real time, A means for analyzing generated voice and motion data and applying an artificial intelligence model to detect unconscious bias, A means of providing feedback to users based on the analysis results, A means of evaluating training progress and recommending the next training scenario, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional adoption process, many interviewers have biases in personnel selection due to unconscious biases. This bias hinders the fair evaluation of candidates with diverse backgrounds and cultures, and poses a significant obstacle to the realization of diversity and inclusivity for enterprises. As a result, there is a high possibility of overlooking talented candidates, leading to concerns about the loss of corporate competitiveness and innovation. Therefore, there is a need for training means to effectively detect and correct unconscious biases.

Means for Solving the Problems

[0005] This invention provides a simulated interview environment using virtual reality, enabling users to conduct interviews with virtual candidates. In this environment, the user's voice and actions are recorded in real time. The recorded data is analyzed by a generated artificial intelligence model to detect unconscious biases. Feedback based on the detected biases is immediately provided to the user, allowing them to recognize and correct their biases on the spot. Furthermore, the next training scenario is recommended according to the training progress, enabling continuous and effective bias correction learning. This contributes to the realization of a fair and inclusive recruitment process.

[0006] "Virtual reality" is a technology that uses computer technology to create a virtual environment that closely resembles reality, giving users a sense of immersion.

[0007] A "computer device" is an electronic device used to process, store, and transmit digital information, and to perform various functions.

[0008] A "mock interview" is a process in which a user conducts a simulated interview with a candidate in a virtual environment that replicates a real interview situation.

[0009] "Means for recording voice and actions in real time" refers to technology that instantly acquires data on a user's speech and body movements and stores it for analysis.

[0010] An "artificial intelligence model" is an algorithm that uses a computer program to mimic human-like thought patterns and perform pattern recognition and prediction on data.

[0011] "Unconscious bias" is a phenomenon in which preconceptions and prejudices that individuals hold without realizing it influence others without their awareness.

[0012] "Feedback" is the process of providing information for improvement and learning based on the user's actions and comments.

[0013] "Means for evaluating training progress" refers to technologies that measure a user's learning status and the degree of improvement in their abilities, and then analyze the results.

[0014] A "training scenario" is a simulated exercise plan designed to achieve specific learning objectives or improve skills. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a training system for recruiters to conduct fair interviews, specifically by using virtual reality (VR) to conduct mock interviews. The following details an embodiment of the system.

[0037] System Overview

[0038] The server provides various mock interview scenarios and generates virtual candidates with different cultural backgrounds and profiles. This allows users to select from a variety of scenarios according to the situation and learn practical interview techniques.

[0039] The terminal is a device that enables users to participate in a virtual environment using VR devices. The terminal records the user's voice and actions and transmits them to the server in real time.

[0040] Users select a mock interview scenario, enter a VR space, and interact with a virtual candidate. The user's goal is to recognize unconscious biases and acquire fair and neutral interviewing skills.

[0041] Program Processing Description

[0042] When the mock interview begins, the server sends a virtual candidate to the terminal based on the selected scenario. The user asks questions to the candidate in the VR space and conducts the interview through interaction.

[0043] During this time, the device records the user's voice and motion data in real time. For example, it collects information such as whether the user is using a specific gesture or if there is a change in their tone of voice. This data is immediately sent to the server.

[0044] The server inputs the received data into an artificial intelligence (AI) model to analyze whether the user's statements and actions contain unconscious biases. If a specific bias is detected as a result of the analysis, the server sends appropriate feedback to the terminal.

[0045] The device provides feedback to the user visually or audibly. For example, the user might see a notification stating, "That statement may be gender-biased."

[0046] Once a series of mock interviews is complete, the server evaluates the user's progress and recommends the next training scenario. This allows the user to continue receiving training to improve their interview skills and increase their fairness.

[0047] This system helps companies implement a talent acquisition process that promotes diversity and inclusion. One embodiment of the present invention proposes practical methods for recruiters to understand and correct unconscious biases.

[0048] The following describes the processing flow.

[0049] Step 1:

[0050] The server prepares various mock interview scenarios for the user and sends them to the terminal. The user reviews the list of scenarios on the terminal and selects the desired scenario.

[0051] Step 2:

[0052] Based on the scenario selected by the user, the device sets up the VR environment and displays virtual candidates through the user's VR device. The user then prepares to begin a mock interview based on the configured scenario.

[0053] Step 3:

[0054] When the mock interview begins, the user starts interacting with a virtual candidate in a virtual space. The user asks questions and conducts the interview based on the candidate's responses.

[0055] Step 4:

[0056] The device records the user's voice and movement data in real time. During this process, it collects various information, including gestures, speaking style, and tone, and transmits it to the server.

[0057] Step 5:

[0058] The server activates a generative AI model to analyze the collected data and check for the presence of unconscious bias. Through voice analysis and behavioral analysis, it evaluates whether the user's statements and attitudes are influenced by bias.

[0059] Step 6:

[0060] The server generates feedback based on the analysis results and sends appropriate comments and advice to the terminal. For example, it may notify users that a particular statement might be based on bias.

[0061] Step 7:

[0062] The terminal presents the user with feedback from the server, either visually or audibly. Based on this feedback, the user modifies their interview approach and seeks ways to conduct a fair evaluation.

[0063] Step 8:

[0064] Once a series of mock interviews is complete, the server evaluates the user's training progress and recommends a new scenario for the next session. This reinforces the process of continuous learning and bias correction.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] In recent years, with the increasing emphasis on diversity and inclusion in talent acquisition, conducting fair interviews free from unconscious biases has become crucial. However, traditional interview training methods have struggled to effectively analyze unconscious biases and provide feedback. As a result, obtaining concrete improvement measures to reduce interviewers' biases has been difficult.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes a configuration means for generating a virtual environment for a mock interview, a configuration means for acquiring the user's voice and actions in real time, and a configuration means for analyzing the acquired voice and action data and applying a machine learning model to detect unconscious biases. This makes it possible to accurately detect the user's unconscious biases and provide appropriate feedback. Specifically, it is possible to clarify the biases the user has and obtain concrete improvement guidelines for the next training session.

[0070] "Virtual reality" is a technology that uses computer technology to create a virtual environment that users can immerse themselves in and experience.

[0071] An "information processing device" is a device that includes hardware and software for processing data and performing specified tasks.

[0072] A "virtual environment" is an artificial environment created using digital technology that users can interact with.

[0073] "Users" refer to people who operate this system and experience mock interviews.

[0074] "Voice and motion data" refers to information about the user's speech and body movements, which are recorded in real time.

[0075] "Unconscious bias" refers to prejudices and preconceptions that individuals hold without realizing it, and which can influence decision-making.

[0076] A "machine learning model" refers to a technology that uses algorithms and statistical methods to learn patterns from data and perform predictions and classifications.

[0077] "Feedback" refers to the information and guidelines for improvement provided to users based on the analyzed results.

[0078] "Cultural background" refers to the characteristics and historical background of the culture to which an individual belongs.

[0079] "Career history" refers to an individual's record of work experience and academic achievements up to that point.

[0080] This invention is a simulated interview training system that utilizes virtual reality. Specific embodiments of the system are described below.

[0081] The server first generates a virtual environment for the mock interview. This involves creating virtual characters with different cultural backgrounds and careers based on the scenario selected by the user. Having different scenarios available allows users to engage with a variety of case studies.

[0082] The terminal is a device for acquiring user voice and actions in real time. This includes hardware such as microphones and motion sensors, and the terminal's role is to transmit this data to a server. This real-time data collection allows users to receive immediate feedback.

[0083] The server inputs collected voice and behavioral data into a machine learning model to analyze unconscious biases hidden in the user's speech and actions. This analysis uses a generative AI model, which generates feedback when specific biases are detected. This feedback is provided to the user through the terminal, allowing the user to identify specific areas for improvement.

[0084] For example, by selecting an interview scenario suitable for a technical job, users can practice how to assess technical skills. When a user asks, "Tell me about your past project experience," the AI ​​model analyzes whether the question contains any particular bias and provides feedback such as, "Try not to focus too much on technical experience."

[0085] When using a generative AI model, you can use a prompt like this: "How can I identify unconscious biases in job interviews?" This prompt helps the AI ​​model suggest bias analysis methods.

[0086] In this way, the present invention enables users to effectively recognize unconscious biases and learn fair interview techniques.

[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0088] Step 1:

[0089] The user selects a scenario according to their training objectives for the mock interview. As input, the user chooses one scenario from several options and sends relevant information to the server. Based on this information, the server generates virtual candidates with different cultural backgrounds and experience levels suitable for the specific scenario and transfers them to the terminal. The output is data on the virtual candidates based on the scenario selected by the user.

[0090] Step 2:

[0091] The terminal uses virtual candidate data received from the server to construct a virtual reality environment. The input consists of virtual candidate profile data and scenario conditions. Based on this, the terminal generates a VR scene, allowing the user to participate. The output is the virtual reality environment that the user can experience.

[0092] Step 3:

[0093] The user enters a virtual environment using a VR device and begins an interview with a virtual candidate. The user's input consists of their spoken words and actions in real time. The device records the user's voice and actions in real time. This recorded data is sent to a server, which then serves as input for the next analysis step.

[0094] Step 4:

[0095] The server inputs voice and behavioral data transmitted from the terminal into a machine learning model. The input data contains information about the user's speech and gestures. Based on this, the server uses a generative AI model to analyze unconscious biases. The output is the analysis result based on whether or not biases are present.

[0096] Step 5:

[0097] The server generates feedback based on the analysis results and sends it to the terminal. The feedback includes suggestions for improvement if the user's statements contain any particular biases. A specific example might be, "If your responses are too heavily biased towards technical experience, please try to ask more inclusive questions." The output is a feedback message for the user.

[0098] Step 6:

[0099] The terminal presents feedback received from the server to the user visually or audibly. The input is the feedback message received from the server. The user can review this and understand specific areas for improvement. The output is information the user can use for their next interview.

[0100] Step 7:

[0101] Once the series of mock interviews is complete, the server evaluates the entire training and suggests the next scenario. The server uses the user's progress data and feedback history as input. The output is a recommendation for the next training scenario aimed at improving the user's skills.

[0102] (Application Example 1)

[0103] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0104] In modern customer service, unconscious biases can influence customer interactions. However, there is a lack of effective training methods to help customer service staff recognize their own biases and provide fair and inclusive service. To address this issue, providing training methods that utilize virtual reality in physical stores is a key challenge.

[0105] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0106] In this invention, the server includes means for generating a simulated dialogue environment, means for recording the user's voice and behavior in real time, means for analyzing the generated voice and behavior data and applying a machine learning model to detect unconscious bias, and means for providing a virtual customer interaction through a visual device. This enables customer service staff to be trained to recognize and correct biases through diverse customer scenarios.

[0107] An "information processing device" is a device that handles digital information and performs various calculations and data management.

[0108] A "simulated dialogue environment" is a situation that provides users with the ability to simulate dialogues in a virtual space based on specific scenarios.

[0109] "Means for recording voice and actions in real time" refers to methods and devices for instantly acquiring data on a user's speech and body movements.

[0110] A "machine learning model" is an algorithm used to analyze collected data and learn patterns and features from it.

[0111] "Unconscious bias" refers to biased thoughts or attitudes that a person displays towards others who have certain attributes or backgrounds, even though they are not consciously aware of it.

[0112] A "visual device" is a device that presents virtual reality to the user as an image, and includes, for example, head-mounted displays and smart glasses.

[0113] A "virtual customer" is not an actual client, but rather a conversation partner generated by a program, and can have a variety of attributes and scenarios.

[0114] This invention is a system that supports customer service staff in interacting with virtual customers. The server first generates a simulated dialogue environment, and the user participates in this virtual environment using a visual device. The server delivers the generated dialogue scenario to the user's visual device and displays virtual customers with different cultural backgrounds and characteristics.

[0115] The device records the user's voice and actions in real time and sends this data to a server. The server analyzes this voice and behavior data using machine learning models to detect biased statements and actions by the user. If detected, the server generates feedback for the user based on the results and provides immediate visual and audible notifications.

[0116] For example, if a user unconsciously expresses bias towards a virtual customer with a specific cultural background, the server provides feedback such as, "That statement may reflect cultural bias. Please try to express yourself more neutrally." This allows users to learn how to respond fairly and inclusively in their daily customer service interactions.

[0117] This system could utilize commercially available cloud AI services as its information processing device and machine learning model. Furthermore, an example of a prompt message, which is part of the invention, is: "Is it possible that unconscious bias occurred in this dialogue? Please provide feedback on areas for improvement based on the user's statements and actions."

[0118] In this way, customer service staff can correct unconscious biases through systematic training and develop the ability to effectively interact with customers from diverse cultural backgrounds.

[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0120] Step 1:

[0121] The server generates a virtual dialogue environment. The input is the scenario information selected by the user. Based on this, the server generates virtual customers with different cultural backgrounds and characteristics, and sends them to the terminal as visual data. The output is the virtual customer displayed on the user's visual device.

[0122] Step 2:

[0123] The user initiates an interaction with a virtual customer via a visual device. The input is virtual customer information sent from the server. The user's speech and actions are recorded in real time and sent from the terminal to the server. The output is the recorded audio and action data.

[0124] Step 3:

[0125] The server inputs the received voice and motion data into a generating AI model. Here, the input is the user's real-time voice and motion data, which is analyzed and used in data calculations to detect unconscious biases. The output is the analysis result indicating whether or not bias is present.

[0126] Step 4:

[0127] The server generates feedback to the user, if necessary, based on the analysis results. The input consists of bias detection results and prompts for improvement. Based on this, the server generates specific feedback messages. The output is visual or audible feedback provided to the user.

[0128] Step 5:

[0129] The terminal receives feedback from the server and notifies the user. The input is the feedback message generated by the server. The terminal conveys this to the user by displaying it on a visual device or providing an audio notification. The output is the feedback information received by the user.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention combines a simulated interview system using a virtual reality environment with an emotion engine that recognizes user emotions. This system allows recruiters to receive training to correct unconscious biases and conduct fair and inclusive hiring practices.

[0132] System configuration and operation

[0133] The server generates various mock interview scenarios and sends them to the user's terminal, allowing them to select their preferred scenario. The scenarios include virtual candidates with different cultural backgrounds and profiles, realistically recreating the interview situation.

[0134] The terminal is a device that provides a virtual space to the user via a VR device, allowing the user to initiate interaction with a virtual candidate in the VR environment. It records the user's voice and actions in real time and sends that data to the server.

[0135] The emotion engine analyzes the user's emotions from input voice and motion data, recognizing indicators such as joy, surprise, anger, disgust, fear, sadness, trust, and anxiety. This reveals the user's mental state during the interview.

[0136] The server combines the output of the emotion engine with an AI model to check for any unconscious biases in the user. Based on the analyzed data, feedback is generated. The feedback is adjusted according to the user's emotional state and presented to the user through the device.

[0137] Specific example

[0138] For example, if a user is in a mock interview and the virtual candidate responds with surprise, they might ask an overly aggressive question due to nervousness. In this case, the emotion engine recognizes the user's anxiety and aggression from their tone of voice and speech patterns, and provides feedback to help the user calm down. This feedback might be conveyed through the device as a notification such as, "Your reaction to the candidate's statement may be excessive."

[0139] Once a series of sessions concludes, the server records the user's progress and guides them through the next training session based on the analysis results. This allows users to hone their interviewing skills and prepare to conduct fair and effective talent selection.

[0140] This system evolves traditional mock interviews through the use of emotional data, providing an innovative training method to correct unconscious biases.

[0141] The following describes the processing flow.

[0142] Step 1:

[0143] The server prepares a variety of mock interview scenarios and sends them to the user's terminal. The user then views the list of scenarios on their terminal and selects the scenario best suited to their training.

[0144] Step 2:

[0145] The terminal sets up the VR environment based on the selected scenario and prepares the user to enter the virtual space through the VR device. Once ready, the terminal displays the virtual candidate and signals the user to begin the interview.

[0146] Step 3:

[0147] The user begins asking questions to a virtual candidate in a VR space. The device records data such as the user's speech, tone, and actions in real time. This data is then sent to a server.

[0148] Step 4:

[0149] The server inputs the transmitted data into the emotion engine. The emotion engine analyzes the user's emotions from the intonation of their voice and body movements, and identifies emotional states such as tension, anxiety, and relief.

[0150] Step 5:

[0151] Based on the output of the emotion engine and voice and behavioral data, the server uses a generative AI model to detect the user's unconscious biases. Here, it determines whether a particular question is biased.

[0152] Step 6:

[0153] The server generates feedback based on the analyzed data. This feedback takes into account the sentiment analysis results and includes advice tailored to the user's current emotional state. The feedback is then sent to the terminal.

[0154] Step 7:

[0155] The terminal visually or audibly notifies the user of feedback from the server. The user then has the opportunity to refine their interviewing skills based on the feedback provided.

[0156] Step 8:

[0157] Once the mock interview is complete, the server evaluates the user's progress and generates data to recommend the next training scenario. The user can then use this information to plan further training.

[0158] (Example 2)

[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0160] Conventional mock interview systems have difficulty adequately detecting and correcting users' unconscious judgment biases, making it impossible to provide fair and effective feedback during interview training. Furthermore, they lacked feedback that considered the user's emotional state, resulting in insufficient improvement in the quality of training.

[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0162] In this invention, the server includes means for generating a simulated environment in an information processing system that provides virtual reality; means for sequentially recording the user's voice and actions; means for analyzing the obtained voice and action data and applying a machine learning model to detect unconscious judgment biases; and means for analyzing the user's emotional state and using the analysis results to generate feedback. This enables a detailed understanding of user behavior, including unconscious biases, and the provision of comprehensive feedback based on emotions.

[0163] "Virtual reality" is a technology that uses computer technology to create a virtual environment that is different from reality, allowing users to experience a sense of presence within that environment.

[0164] An "information processing system" is a system consisting of a series of hardware and software components that perform data input, calculation, storage, and output.

[0165] A "simulated environment" is an environment designed to allow users to gain experience by virtually recreating specific situations in the real world.

[0166] A "user" refers to a person or organization that operates or uses a system or equipment.

[0167] "Voice and behavioral data" refers to information that captures the user's voice and body movements.

[0168] "Sequential recording" refers to the process of continuously recording data in real time.

[0169] A "machine learning model" is a mathematical model that uses large amounts of data to learn patterns from algorithms and then makes predictions and judgments about new data.

[0170] "Judgment bias" refers to a systematic bias in an individual's thinking and judgment that occurs unconsciously.

[0171] "Feedback" refers to evaluations and advice provided to users by a system.

[0172] A "virtual character" is a fictional person or creature created by a computer, which enables interaction and dialogue with the user.

[0173] This invention is an information processing system that uses virtual reality technology to provide a simulated interview environment, allowing users to experience various scenarios within it. The system's hardware configuration includes a computer system, a VR device, a voice capture device, and motion sensors. The software includes an application for generating the VR environment, a machine learning algorithm for analyzing voice and behavioral data, and an interface for providing feedback.

[0174] The server first generates a virtual environment. Based on information retrieved from the database, the server creates virtual characters with different cultural backgrounds and characteristics, allowing users to dynamically select scenarios.

[0175] The terminal provides the user with a virtual environment through a VR device and initiates interaction. The terminal collects the user's voice and movement data in real time and sends that data to the server.

[0176] Users experience a simulated interview in a virtual environment. During this process, their emotional state and unconscious judgment biases are analyzed. The resulting analysis is used to improve subsequent training sessions.

[0177] For example, when a user asks an overly anxious, one-sided question to a virtual candidate, their tone of voice and speaking speed are detected. Based on this emotional state analysis, feedback such as, "Your question may be too aggressive. Try softening your tone," is provided. This feedback, based on multifaceted information, leads to a fair and emotionally balanced evaluation.

[0178] An example of a prompt sentence generated using an AI model is: "In a virtual interview, how can you remain calm and respond flexibly even when you receive an unexpected answer?"

[0179] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0180] Step 1:

[0181] The server generates mock interview scenarios. The server retrieves profile information with different cultural backgrounds and characteristics from a database and creates multiple virtual characters based on this information. It receives user preferences and training objectives as input, generates a list of selectable scenarios as output, and sends it to the terminal.

[0182] Step 2:

[0183] The terminal provides the user with a virtual interview environment. Using scenario data received from the server, the terminal displays a virtual scenario to the user through a VR device. The user then begins interacting with a virtual character within the VR environment. It receives scenario data from the server as input and provides the virtual environment experienced by the user as output.

[0184] Step 3:

[0185] The user interacts with a virtual character. During this time, the user's voice and actions are continuously recorded by the device. The user's voice and actions are captured in real time as input, and this data is processed and sent to the server. The raw data is sent to the server as output.

[0186] Step 4:

[0187] The server analyzes the received audio and behavioral data. Using an emotion analysis engine, it estimates the user's emotions and stress levels. It analyzes the user data received as input and generates emotional states and potential judgment bias characteristics as output. These results are then fed into a generative AI model.

[0188] Step 5:

[0189] The server uses a generative AI model to create feedback for the user. Based on sentiment data and bias detection results, it identifies specific actions and reactions the user performed unconsciously and suggests improvements. It receives analysis results as input and sends the feedback content to the terminal as output.

[0190] Step 6:

[0191] The device presents the generated feedback to the user. The feedback is delivered to the user visually or aurally. It receives feedback information from the server as input and provides feedback to the user as output.

[0192] Step 7:

[0193] Based on the feedback, users review their interviewing skills and apply them to their next training session. The server further accumulates progress data and suggests the next training scenario. User responses and progress are recorded as input, and guidelines for the next step are generated as output.

[0194] (Application Example 2)

[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0196] In modern security operations, security personnel often face problems where unconscious biases or excessive stress prevent them from making accurate judgments when faced with emergencies. This can lead to inappropriate responses or unnecessary escalations. This invention aims to cultivate appropriate judgment skills in security personnel by simulating these situations in a virtual reality environment and receiving real-time feedback on their emotions and psychological state.

[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0198] In this invention, the server includes means for generating a simulated interview environment, means for recording the user's voice and actions in real time, means for providing a simulated situation to enhance the security officer's judgment in emergency or intimidating situations, and means for analyzing the user's psychological state in real time and providing feedback to help maintain composure. This enables security officers to learn appropriate responses in a virtual reality environment and eliminate unconscious biases, thereby enabling effective security responses.

[0199] "Virtual reality" is a technology that uses computer technology to create virtual environments for users, enabling them to have experiences that are as close to reality as possible.

[0200] A "mock interview" is a method that uses virtual reality technology to recreate a real interview situation, allowing users to practice and learn within that environment.

[0201] "Unconscious bias" is a phenomenon in which prejudices and preconceptions held by individuals without their awareness influence their judgments and actions.

[0202] "Real-time recording" is the process of recording data such as user voice and actions the moment they occur.

[0203] An "artificial intelligence model" is a mathematical or computational model that learns from data analysis and user behavior and automatically performs specific tasks.

[0204] "Feedback" is a means of providing users with information and advice based on analysis results to promote behavioral improvement and learning.

[0205] An "emergency situation" is a situation in which an unexpected event or dangerous situation occurs and requires a swift response.

[0206] "Psychological state analysis" is a process of analyzing emotions and mental responses to assess a user's current emotions and mental health.

[0207] The system implementing this invention provides a simulated environment using virtual reality, enabling security personnel to undergo training to enhance their decision-making abilities in emergency situations. The system operates using a VR headset and emotion recognition sensors. Specifically, it integrates a VR device such as Oculus Quest with emotion recognition software such as Affectiva.

[0208] The server delivers data modeling mock interviews and emergency situations to VR devices. Users wear VR headsets and immerse themselves in the virtual environment to experience realistic emergency situations. Sensors collect audio and motion data in real time, which is then transmitted to the server.

[0209] The server analyzes the collected data using a generative AI model. This analysis assesses the user's psychological state and unconscious biases, and provides feedback based on the results. The feedback is presented to the user visually or audibly, allowing the user to understand their own reactions and improve their ability to maintain composure.

[0210] For example, if a user experiences panic or anxiety when a suspicious person approaches them in VR, they will be given feedback such as, "Stay calm and continue to observe the person's movements." In this way, users can train their reactions to potential risks.

[0211] An example of a prompt sentence to be input into the generating AI model might be, "Please provide feedback on how security personnel can remain calm when encountering a suspicious person." This prompt sentence is used to provide guidance on how the user should respond.

[0212] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0213] Step 1:

[0214] The device presents a virtual environment to the user through a VR headset. The user becomes immersed in the VR environment, and an emergency scenario is displayed on the screen. During this process, the device receives input from the user to initiate interaction and loads the data that constitutes the VR environment.

[0215] Step 2:

[0216] While the user operates within the virtual environment, the terminal uses emotion recognition sensors to record the user's voice and behavioral data in real time. This collects biometric data, which is then immediately transmitted to the server. In this step, the collected biometric data becomes the input, and an output is generated that packages and transfers the data to the server.

[0217] Step 3:

[0218] The server inputs the received data into a generating AI model to analyze the user's psychological state. The analysis process infers stress and emotional fluctuations from biometric data and verifies for any unconscious biases. The determined psychological state is generated as output, which is then used to generate feedback.

[0219] Step 4:

[0220] The server generates appropriate feedback for the user based on the analysis results. The generating AI model uses prompts to create advice and warnings regarding the user's actions, and compiles this content into a feedback message. This feedback message is output and sent to the terminal in either a visual or auditory form.

[0221] Step 5:

[0222] The device presents the transmitted feedback to the user through the VR headset. The user receives the feedback and gains guidance for improving their actions. In this step, a feedback message is input, and the process of conveying the feedback to the user as a visual or audio output is performed.

[0223] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0224] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0225] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0226] [Second Embodiment]

[0227] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0228] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0229] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0230] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0231] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0232] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0233] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0234] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0235] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0236] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0237] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0238] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0239] This invention is a training system for recruiters to conduct fair interviews, specifically by using virtual reality (VR) to conduct mock interviews. The following details an embodiment of the system.

[0240] System Overview

[0241] The server provides various mock interview scenarios and generates virtual candidates with different cultural backgrounds and profiles. This allows users to select from a variety of scenarios according to the situation and learn practical interview techniques.

[0242] The terminal is a device that enables users to participate in a virtual environment using VR devices. The terminal records the user's voice and actions and transmits them to the server in real time.

[0243] Users select a mock interview scenario, enter a VR space, and interact with a virtual candidate. The user's goal is to recognize unconscious biases and acquire fair and neutral interviewing skills.

[0244] Program Processing Description

[0245] When the mock interview begins, the server sends a virtual candidate to the terminal based on the selected scenario. The user asks questions to the candidate in the VR space and conducts the interview through interaction.

[0246] During this time, the device records the user's voice and motion data in real time. For example, it collects information such as whether the user is using a specific gesture or if there is a change in their tone of voice. This data is immediately sent to the server.

[0247] The server inputs the received data into an artificial intelligence (AI) model to analyze whether the user's statements and actions contain unconscious biases. If a specific bias is detected as a result of the analysis, the server sends appropriate feedback to the terminal.

[0248] The device provides feedback to the user visually or audibly. For example, the user might see a notification stating, "That statement may be gender-biased."

[0249] Once a series of mock interviews is complete, the server evaluates the user's progress and recommends the next training scenario. This allows the user to continue receiving training to improve their interview skills and increase their fairness.

[0250] This system helps companies implement a talent acquisition process that promotes diversity and inclusion. One embodiment of the present invention proposes practical methods for recruiters to understand and correct unconscious biases.

[0251] The following describes the processing flow.

[0252] Step 1:

[0253] The server prepares various mock interview scenarios for the user and sends them to the terminal. The user reviews the list of scenarios on the terminal and selects the desired scenario.

[0254] Step 2:

[0255] Based on the scenario selected by the user, the device sets up the VR environment and displays virtual candidates through the user's VR device. The user then prepares to begin a mock interview based on the configured scenario.

[0256] Step 3:

[0257] When the mock interview begins, the user starts interacting with a virtual candidate in a virtual space. The user asks questions and conducts the interview based on the candidate's responses.

[0258] Step 4:

[0259] The device records the user's voice and movement data in real time. During this process, it collects various information, including gestures, speaking style, and tone, and transmits it to the server.

[0260] Step 5:

[0261] The server activates a generative AI model to analyze the collected data and check for the presence of unconscious bias. Through voice analysis and behavioral analysis, it evaluates whether the user's statements and attitudes are influenced by bias.

[0262] Step 6:

[0263] The server generates feedback based on the analysis results and sends appropriate comments and advice to the terminal. For example, it may notify users that a particular statement might be based on bias.

[0264] Step 7:

[0265] The terminal presents the user with feedback from the server, either visually or audibly. Based on this feedback, the user modifies their interview approach and seeks ways to conduct a fair evaluation.

[0266] Step 8:

[0267] Once a series of mock interviews is complete, the server evaluates the user's training progress and recommends a new scenario for the next session. This reinforces the process of continuous learning and bias correction.

[0268] (Example 1)

[0269] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0270] In recent years, with the increasing emphasis on diversity and inclusion in talent acquisition, conducting fair interviews free from unconscious biases has become crucial. However, traditional interview training methods have struggled to effectively analyze unconscious biases and provide feedback. As a result, obtaining concrete improvement measures to reduce interviewers' biases has been difficult.

[0271] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0272] In this invention, the server includes a configuration means for generating a virtual environment for a mock interview, a configuration means for acquiring the user's voice and actions in real time, and a configuration means for analyzing the acquired voice and action data and applying a machine learning model to detect unconscious biases. This makes it possible to accurately detect the user's unconscious biases and provide appropriate feedback. Specifically, it is possible to clarify the biases the user has and obtain concrete improvement guidelines for the next training session.

[0273] "Virtual reality" is a technology that uses computer technology to create a virtual environment that users can immerse themselves in and experience.

[0274] An "information processing device" is a device that includes hardware and software for processing data and performing specified tasks.

[0275] A "virtual environment" is an artificial environment created using digital technology that users can interact with.

[0276] "Users" refer to people who operate this system and experience mock interviews.

[0277] "Voice and motion data" refers to information about the user's speech and body movements, which are recorded in real time.

[0278] "Unconscious bias" refers to prejudices and preconceptions that individuals hold without realizing it, and which can influence decision-making.

[0279] A "machine learning model" refers to a technology that uses algorithms and statistical methods to learn patterns from data and perform predictions and classifications.

[0280] "Feedback" refers to the information and guidelines for improvement provided to users based on the analyzed results.

[0281] "Cultural background" refers to the characteristics and historical background of the culture to which an individual belongs.

[0282] "Experience" refers to the work and academic history that an individual has experienced so far.

[0283] The present invention is a simulated interview training system that utilizes virtual reality. The specific embodiments of the system will be described below.

[0284] The server first generates a virtual environment for the simulated interview. This includes the process of preparing virtual characters with different cultural backgrounds and experiences based on the scenario selected by the user. By preparing different scenarios, the user can handle various case studies.

[0285] The terminal is a device for acquiring the user's voice and actions in real time. This includes hardware such as microphones and motion sensors, and the terminal plays the role of sending this data to the server. Through this real-time data collection, the user can receive feedback on the spot.

[0286] The server inputs the collected voice and motion data into a machine learning model and analyzes the unconscious biases hidden in the user's speech and actions. A generative AI model is used for this analysis, and feedback is generated when a specific bias is detected. This feedback is provided to the user through the terminal, and the user can confirm specific points for improvement.

[0287] For example, by selecting an interview scenario suitable for a technical-related position, the user can practice the method of evaluating technical skills. When the user asks "Please tell me about your past project experience," the AI model analyzes whether this question contains a specific bias and provides feedback such as "Please avoid being too biased towards technical experience."

[0288] When using a generative AI model, you can use a prompt like this: "How can I identify unconscious biases in job interviews?" This prompt helps the AI ​​model suggest bias analysis methods.

[0289] In this way, the present invention enables users to effectively recognize unconscious biases and learn fair interview techniques.

[0290] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0291] Step 1:

[0292] The user selects a scenario according to their training objectives for the mock interview. As input, the user chooses one scenario from several options and sends relevant information to the server. Based on this information, the server generates virtual candidates with different cultural backgrounds and experience levels suitable for the specific scenario and transfers them to the terminal. The output is data on the virtual candidates based on the scenario selected by the user.

[0293] Step 2:

[0294] The terminal uses virtual candidate data received from the server to construct a virtual reality environment. The input consists of virtual candidate profile data and scenario conditions. Based on this, the terminal generates a VR scene, allowing the user to participate. The output is the virtual reality environment that the user can experience.

[0295] Step 3:

[0296] The user enters a virtual environment using a VR device and begins an interview with a virtual candidate. The user's input consists of their spoken words and actions in real time. The device records the user's voice and actions in real time. This recorded data is sent to a server, which then serves as input for the next analysis step.

[0297] Step 4:

[0298] The server inputs voice and behavioral data transmitted from the terminal into a machine learning model. The input data contains information about the user's speech and gestures. Based on this, the server uses a generative AI model to analyze unconscious biases. The output is the analysis result based on whether or not biases are present.

[0299] Step 5:

[0300] The server generates feedback based on the analysis results and sends it to the terminal. The feedback includes suggestions for improvement if the user's statements contain any particular biases. A specific example might be, "If your responses are too heavily biased towards technical experience, please try to ask more inclusive questions." The output is a feedback message for the user.

[0301] Step 6:

[0302] The terminal presents feedback received from the server to the user visually or audibly. The input is the feedback message received from the server. The user can review this and understand specific areas for improvement. The output is information the user can use for their next interview.

[0303] Step 7:

[0304] Once the series of mock interviews is complete, the server evaluates the entire training and suggests the next scenario. The server uses the user's progress data and feedback history as input. The output is a recommendation for the next training scenario aimed at improving the user's skills.

[0305] (Application Example 1)

[0306] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0307] In modern customer service operations, unconscious bias can sometimes affect customer interactions. However, there is a lack of effective training means for customer service staff to recognize their own biases and provide fair and inclusive responses. To address this issue, it is a challenge to provide a training method that utilizes virtual reality in physical stores.

[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0309] In this invention, the server includes means for generating a simulated dialogue environment, means for real-time recording of the user's voice and actions, means for analyzing the generated voice and action data and applying a machine learning model for detecting unconscious bias, and means for providing a dialogue with a virtual customer through a visual device. This enables customer service staff to undergo training to recognize and correct biases through various customer scenarios.

[0310] An "information processing device" is a device that handles digital information and performs various calculations and data management.

[0311] A "simulated dialogue environment" is to provide a situation where a user can imitate a dialogue based on a specific scenario in a virtual space.

[0312] "Means for real-time recording of voice and actions" refers to methods and devices for immediately acquiring the user's speech content and body movements as data.

[0313] A "machine learning model" is an algorithm used to analyze the collected data and learn patterns and features therefrom.

[0314] "Unconscious bias" refers to biased thoughts and attitudes shown towards a person with specific attributes or backgrounds, despite the person being unaware of them.

[0315] A "visual device" is a device that presents virtual reality to the user as an image, and includes, for example, head-mounted displays and smart glasses.

[0316] A "virtual customer" is not an actual client, but rather a conversation partner generated by a program, and can have a variety of attributes and scenarios.

[0317] This invention is a system that supports customer service staff in interacting with virtual customers. The server first generates a simulated dialogue environment, and the user participates in this virtual environment using a visual device. The server delivers the generated dialogue scenario to the user's visual device and displays virtual customers with different cultural backgrounds and characteristics.

[0318] The device records the user's voice and actions in real time and sends this data to a server. The server analyzes this voice and behavior data using machine learning models to detect biased statements and actions by the user. If detected, the server generates feedback for the user based on the results and provides immediate visual and audible notifications.

[0319] For example, if a user unconsciously expresses bias towards a virtual customer with a specific cultural background, the server provides feedback such as, "That statement may reflect cultural bias. Please try to express yourself more neutrally." This allows users to learn how to respond fairly and inclusively in their daily customer service interactions.

[0320] This system could utilize commercially available cloud AI services as its information processing device and machine learning model. Furthermore, an example of a prompt message, which is part of the invention, is: "Is it possible that unconscious bias occurred in this dialogue? Please provide feedback on areas for improvement based on the user's statements and actions."

[0321] In this way, customer service staff can correct unconscious biases through systematic training and develop the ability to effectively interact with customers from diverse cultural backgrounds.

[0322] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0323] Step 1:

[0324] The server generates a virtual dialogue environment. The input is the scenario information selected by the user. Based on this, the server generates virtual customers with different cultural backgrounds and characteristics, and sends them to the terminal as visual data. The output is the virtual customer displayed on the user's visual device.

[0325] Step 2:

[0326] The user initiates an interaction with a virtual customer via a visual device. The input is virtual customer information sent from the server. The user's speech and actions are recorded in real time and sent from the terminal to the server. The output is the recorded audio and action data.

[0327] Step 3:

[0328] The server inputs the received voice and motion data into a generating AI model. Here, the input is the user's real-time voice and motion data, which is analyzed and used in data calculations to detect unconscious biases. The output is the analysis result indicating whether or not bias is present.

[0329] Step 4:

[0330] The server generates feedback to the user, if necessary, based on the analysis results. The input consists of bias detection results and prompts for improvement. Based on this, the server generates specific feedback messages. The output is visual or audible feedback provided to the user.

[0331] Step 5:

[0332] The terminal receives feedback from the server and notifies the user. The input is the feedback message generated by the server. The terminal conveys this to the user by displaying it on a visual device or providing an audio notification. The output is the feedback information received by the user.

[0333] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0334] This invention combines a simulated interview system using a virtual reality environment with an emotion engine that recognizes user emotions. This system allows recruiters to receive training to correct unconscious biases and conduct fair and inclusive hiring practices.

[0335] System configuration and operation

[0336] The server generates various mock interview scenarios and sends them to the user's terminal, allowing them to select their preferred scenario. The scenarios include virtual candidates with different cultural backgrounds and profiles, realistically recreating the interview situation.

[0337] The terminal is a device that provides a virtual space to the user via a VR device, allowing the user to initiate interaction with a virtual candidate in the VR environment. It records the user's voice and actions in real time and sends that data to the server.

[0338] The emotion engine analyzes the user's emotions from input voice and motion data, recognizing indicators such as joy, surprise, anger, disgust, fear, sadness, trust, and anxiety. This reveals the user's mental state during the interview.

[0339] The server combines the output of the emotion engine with an AI model to check for any unconscious biases in the user. Based on the analyzed data, feedback is generated. The feedback is adjusted according to the user's emotional state and presented to the user through the device.

[0340] Specific example

[0341] For example, if a user is in a mock interview and the virtual candidate responds with surprise, they might ask an overly aggressive question due to nervousness. In this case, the emotion engine recognizes the user's anxiety and aggression from their tone of voice and speech patterns, and provides feedback to help the user calm down. This feedback might be conveyed through the device as a notification such as, "Your reaction to the candidate's statement may be excessive."

[0342] Once a series of sessions concludes, the server records the user's progress and guides them through the next training session based on the analysis results. This allows users to hone their interviewing skills and prepare to conduct fair and effective talent selection.

[0343] This system evolves traditional mock interviews through the use of emotional data, providing an innovative training method to correct unconscious biases.

[0344] The following describes the processing flow.

[0345] Step 1:

[0346] The server prepares a variety of mock interview scenarios and sends them to the user's terminal. The user then views the list of scenarios on their terminal and selects the scenario best suited to their training.

[0347] Step 2:

[0348] The terminal sets up the VR environment based on the selected scenario and prepares the user to enter the virtual space through the VR device. Once ready, the terminal displays the virtual candidate and signals the user to begin the interview.

[0349] Step 3:

[0350] The user begins asking questions to a virtual candidate in a VR space. The device records data such as the user's speech, tone, and actions in real time. This data is then sent to a server.

[0351] Step 4:

[0352] The server inputs the transmitted data into the emotion engine. The emotion engine analyzes the user's emotions from the intonation of their voice and body movements, and identifies emotional states such as tension, anxiety, and relief.

[0353] Step 5:

[0354] Based on the output of the emotion engine and voice and behavioral data, the server uses a generative AI model to detect the user's unconscious biases. Here, it determines whether a particular question is biased.

[0355] Step 6:

[0356] The server generates feedback based on the analyzed data. This feedback takes into account the sentiment analysis results and includes advice tailored to the user's current emotional state. The feedback is then sent to the terminal.

[0357] Step 7:

[0358] The terminal visually or audibly notifies the user of feedback from the server. The user then has the opportunity to refine their interviewing skills based on the feedback provided.

[0359] Step 8:

[0360] Once the mock interview is complete, the server evaluates the user's progress and generates data to recommend the next training scenario. The user can then use this information to plan further training.

[0361] (Example 2)

[0362] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0363] Conventional mock interview systems have difficulty adequately detecting and correcting users' unconscious judgment biases, making it impossible to provide fair and effective feedback during interview training. Furthermore, they lacked feedback that considered the user's emotional state, resulting in insufficient improvement in the quality of training.

[0364] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0365] In this invention, the server includes means for generating a simulated environment in an information processing system that provides virtual reality; means for sequentially recording the user's voice and actions; means for analyzing the obtained voice and action data and applying a machine learning model to detect unconscious judgment biases; and means for analyzing the user's emotional state and using the analysis results to generate feedback. This enables a detailed understanding of user behavior, including unconscious biases, and the provision of comprehensive feedback based on emotions.

[0366] "Virtual reality" is a technology that uses computer technology to create a virtual environment that is different from reality, allowing users to experience a sense of presence within that environment.

[0367] An "information processing system" is a system consisting of a series of hardware and software components that perform data input, calculation, storage, and output.

[0368] A "simulated environment" is an environment designed to allow users to gain experience by virtually recreating specific situations in the real world.

[0369] A "user" refers to a person or organization that operates or uses a system or equipment.

[0370] "Voice and behavioral data" refers to information that captures the user's voice and body movements.

[0371] "Sequential recording" refers to the process of continuously recording data in real time.

[0372] A "machine learning model" is a mathematical model that uses large amounts of data to learn patterns from algorithms and then makes predictions and judgments about new data.

[0373] "Judgment bias" refers to a systematic bias in an individual's thinking and judgment that occurs unconsciously.

[0374] "Feedback" refers to evaluations and advice provided to users by a system.

[0375] A "virtual character" is a fictional person or creature created by a computer, which enables interaction and dialogue with the user.

[0376] This invention is an information processing system that uses virtual reality technology to provide a simulated interview environment, allowing users to experience various scenarios within it. The system's hardware configuration includes a computer system, a VR device, a voice capture device, and motion sensors. The software includes an application for generating the VR environment, a machine learning algorithm for analyzing voice and behavioral data, and an interface for providing feedback.

[0377] The server first generates a virtual environment. Based on information retrieved from the database, the server creates virtual characters with different cultural backgrounds and characteristics, allowing users to dynamically select scenarios.

[0378] The terminal provides the user with a virtual environment through a VR device and initiates interaction. The terminal collects the user's voice and movement data in real time and sends that data to the server.

[0379] Users experience a simulated interview in a virtual environment. During this process, their emotional state and unconscious judgment biases are analyzed. The resulting analysis is used to improve subsequent training sessions.

[0380] For example, when a user asks an overly anxious, one-sided question to a virtual candidate, their tone of voice and speaking speed are detected. Based on this emotional state analysis, feedback such as, "Your question may be too aggressive. Try softening your tone," is provided. This feedback, based on multifaceted information, leads to a fair and emotionally balanced evaluation.

[0381] An example of a prompt sentence generated using an AI model is: "In a virtual interview, how can you remain calm and respond flexibly even when you receive an unexpected answer?"

[0382] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0383] Step 1:

[0384] The server generates mock interview scenarios. The server retrieves profile information with different cultural backgrounds and characteristics from a database and creates multiple virtual characters based on this information. It receives user preferences and training objectives as input, generates a list of selectable scenarios as output, and sends it to the terminal.

[0385] Step 2:

[0386] The terminal provides the user with a virtual interview environment. Using scenario data received from the server, the terminal displays a virtual scenario to the user through a VR device. The user then begins interacting with a virtual character within the VR environment. It receives scenario data from the server as input and provides the virtual environment experienced by the user as output.

[0387] Step 3:

[0388] The user interacts with a virtual character. During this time, the user's voice and actions are continuously recorded by the device. The user's voice and actions are captured in real time as input, and this data is processed and sent to the server. The raw data is sent to the server as output.

[0389] Step 4:

[0390] The server analyzes the received audio and behavioral data. Using an emotion analysis engine, it estimates the user's emotions and stress levels. It analyzes the user data received as input and generates emotional states and potential judgment bias characteristics as output. These results are then fed into a generative AI model.

[0391] Step 5:

[0392] The server uses a generative AI model to create feedback for the user. Based on sentiment data and bias detection results, it identifies specific actions and reactions the user performed unconsciously and suggests improvements. It receives analysis results as input and sends the feedback content to the terminal as output.

[0393] Step 6:

[0394] The device presents the generated feedback to the user. The feedback is delivered to the user visually or aurally. It receives feedback information from the server as input and provides feedback to the user as output.

[0395] Step 7:

[0396] Based on the feedback, users review their interviewing skills and apply them to their next training session. The server further accumulates progress data and suggests the next training scenario. User responses and progress are recorded as input, and guidelines for the next step are generated as output.

[0397] (Application Example 2)

[0398] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0399] In modern security operations, security personnel often face problems where unconscious biases or excessive stress prevent them from making accurate judgments when faced with emergencies. This can lead to inappropriate responses or unnecessary escalations. This invention aims to cultivate appropriate judgment skills in security personnel by simulating these situations in a virtual reality environment and receiving real-time feedback on their emotions and psychological state.

[0400] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0401] In this invention, the server includes means for generating a simulated interview environment, means for recording the user's voice and actions in real time, means for providing a simulated situation to enhance the security officer's judgment in emergency or intimidating situations, and means for analyzing the user's psychological state in real time and providing feedback to help maintain composure. This enables security officers to learn appropriate responses in a virtual reality environment and eliminate unconscious biases, thereby enabling effective security responses.

[0402] "Virtual reality" is a technology that uses computer technology to create virtual environments for users, enabling them to have experiences that are as close to reality as possible.

[0403] A "mock interview" is a method that uses virtual reality technology to recreate a real interview situation, allowing users to practice and learn within that environment.

[0404] "Unconscious bias" is a phenomenon in which prejudices and preconceptions held by individuals without their awareness influence their judgments and actions.

[0405] "Real-time recording" is the process of recording data such as user voice and actions the moment they occur.

[0406] An "artificial intelligence model" is a mathematical or computational model that learns from data analysis and user behavior and automatically performs specific tasks.

[0407] "Feedback" is a means of providing users with information and advice based on analysis results to promote behavioral improvement and learning.

[0408] An "emergency situation" is a situation in which an unexpected event or dangerous situation occurs and requires a swift response.

[0409] "Psychological state analysis" is a process of analyzing emotions and mental responses to assess a user's current emotions and mental health.

[0410] The system implementing this invention provides a simulated environment using virtual reality, enabling security personnel to undergo training to enhance their decision-making abilities in emergency situations. The system operates using a VR headset and emotion recognition sensors. Specifically, it integrates a VR device such as Oculus Quest with emotion recognition software such as Affectiva.

[0411] The server delivers data modeling mock interviews and emergency situations to VR devices. Users wear VR headsets and immerse themselves in the virtual environment to experience realistic emergency situations. Sensors collect audio and motion data in real time, which is then transmitted to the server.

[0412] The server analyzes the collected data using a generative AI model. This analysis assesses the user's psychological state and unconscious biases, and provides feedback based on the results. The feedback is presented to the user visually or audibly, allowing the user to understand their own reactions and improve their ability to maintain composure.

[0413] For example, if a user experiences panic or anxiety when a suspicious person approaches them in VR, they will be given feedback such as, "Stay calm and continue to observe the person's movements." In this way, users can train their reactions to potential risks.

[0414] An example of a prompt sentence to be input into the generating AI model might be, "Please provide feedback on how security personnel can remain calm when encountering a suspicious person." This prompt sentence is used to provide guidance on how the user should respond.

[0415] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0416] Step 1:

[0417] The device presents a virtual environment to the user through a VR headset. The user becomes immersed in the VR environment, and an emergency scenario is displayed on the screen. During this process, the device receives input from the user to initiate interaction and loads the data that constitutes the VR environment.

[0418] Step 2:

[0419] While the user operates within the virtual environment, the terminal uses emotion recognition sensors to record the user's voice and behavioral data in real time. This collects biometric data, which is then immediately transmitted to the server. In this step, the collected biometric data becomes the input, and an output is generated that packages and transfers the data to the server.

[0420] Step 3:

[0421] The server inputs the received data into a generating AI model to analyze the user's psychological state. The analysis process infers stress and emotional fluctuations from biometric data and verifies for any unconscious biases. The determined psychological state is generated as output, which is then used to generate feedback.

[0422] Step 4:

[0423] The server generates appropriate feedback for the user based on the analysis results. The generating AI model uses prompts to create advice and warnings regarding the user's actions, and compiles this content into a feedback message. This feedback message is output and sent to the terminal in either a visual or auditory form.

[0424] Step 5:

[0425] The device presents the transmitted feedback to the user through the VR headset. The user receives the feedback and gains guidance for improving their actions. In this step, a feedback message is input, and the process of conveying the feedback to the user as a visual or audio output is performed.

[0426] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0427] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0428] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0429] [Third Embodiment]

[0430] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0431] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0432] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0433] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0434] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0436] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0437] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0438] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0439] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0440] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0441] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0442] This invention is a training system for recruiters to conduct fair interviews, specifically by using virtual reality (VR) to conduct mock interviews. The following details an embodiment of the system.

[0443] System Overview

[0444] The server provides various mock interview scenarios and generates virtual candidates with different cultural backgrounds and profiles. This allows users to select from a variety of scenarios according to the situation and learn practical interview techniques.

[0445] The terminal is a device that enables users to participate in a virtual environment using VR devices. The terminal records the user's voice and actions and transmits them to the server in real time.

[0446] Users select a mock interview scenario, enter a VR space, and interact with a virtual candidate. The user's goal is to recognize unconscious biases and acquire fair and neutral interviewing skills.

[0447] Program Processing Description

[0448] When the mock interview begins, the server sends a virtual candidate to the terminal based on the selected scenario. The user asks questions to the candidate in the VR space and conducts the interview through interaction.

[0449] During this time, the device records the user's voice and motion data in real time. For example, it collects information such as whether the user is using a specific gesture or if there is a change in their tone of voice. This data is immediately sent to the server.

[0450] The server inputs the received data into an artificial intelligence (AI) model to analyze whether the user's statements and actions contain unconscious biases. If a specific bias is detected as a result of the analysis, the server sends appropriate feedback to the terminal.

[0451] The device provides feedback to the user visually or audibly. For example, the user might see a notification stating, "That statement may be gender-biased."

[0452] Once a series of mock interviews is complete, the server evaluates the user's progress and recommends the next training scenario. This allows the user to continue receiving training to improve their interview skills and increase their fairness.

[0453] This system helps companies implement a talent acquisition process that promotes diversity and inclusion. One embodiment of the present invention proposes practical methods for recruiters to understand and correct unconscious biases.

[0454] The following describes the processing flow.

[0455] Step 1:

[0456] The server prepares various mock interview scenarios for the user and sends them to the terminal. The user reviews the list of scenarios on the terminal and selects the desired scenario.

[0457] Step 2:

[0458] Based on the scenario selected by the user, the device sets up the VR environment and displays virtual candidates through the user's VR device. The user then prepares to begin a mock interview based on the configured scenario.

[0459] Step 3:

[0460] When the mock interview begins, the user starts interacting with a virtual candidate in a virtual space. The user asks questions and conducts the interview based on the candidate's responses.

[0461] Step 4:

[0462] The device records the user's voice and movement data in real time. During this process, it collects various information, including gestures, speaking style, and tone, and transmits it to the server.

[0463] Step 5:

[0464] The server activates a generative AI model to analyze the collected data and check for the presence of unconscious bias. Through voice analysis and behavioral analysis, it evaluates whether the user's statements and attitudes are influenced by bias.

[0465] Step 6:

[0466] The server generates feedback based on the analysis results and sends appropriate comments and advice to the terminal. For example, it may notify users that a particular statement might be based on bias.

[0467] Step 7:

[0468] The terminal presents the user with feedback from the server, either visually or audibly. Based on this feedback, the user modifies their interview approach and seeks ways to conduct a fair evaluation.

[0469] Step 8:

[0470] Once a series of mock interviews is complete, the server evaluates the user's training progress and recommends a new scenario for the next session. This reinforces the process of continuous learning and bias correction.

[0471] (Example 1)

[0472] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0473] In recent years, with the increasing emphasis on diversity and inclusion in talent acquisition, conducting fair interviews free from unconscious biases has become crucial. However, traditional interview training methods have struggled to effectively analyze unconscious biases and provide feedback. As a result, obtaining concrete improvement measures to reduce interviewers' biases has been difficult.

[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0475] In this invention, the server includes a configuration means for generating a virtual environment for a mock interview, a configuration means for acquiring the user's voice and actions in real time, and a configuration means for analyzing the acquired voice and action data and applying a machine learning model to detect unconscious biases. This makes it possible to accurately detect the user's unconscious biases and provide appropriate feedback. Specifically, it is possible to clarify the biases the user has and obtain concrete improvement guidelines for the next training session.

[0476] "Virtual reality" is a technology that uses computer technology to create a virtual environment that users can immerse themselves in and experience.

[0477] An "information processing device" is a device that includes hardware and software for processing data and performing specified tasks.

[0478] A "virtual environment" is an artificial environment created using digital technology that users can interact with.

[0479] "Users" refer to people who operate this system and experience mock interviews.

[0480] "Voice and motion data" refers to information about the user's speech and body movements, which are recorded in real time.

[0481] "Unconscious bias" refers to prejudices and preconceptions that individuals hold without realizing it, and which can influence decision-making.

[0482] A "machine learning model" refers to a technology that uses algorithms and statistical methods to learn patterns from data and perform predictions and classifications.

[0483] "Feedback" refers to the information and guidelines for improvement provided to users based on the analyzed results.

[0484] "Cultural background" refers to the characteristics and historical background of the culture to which an individual belongs.

[0485] "Career history" refers to an individual's record of work experience and academic achievements up to that point.

[0486] This invention is a simulated interview training system that utilizes virtual reality. Specific embodiments of the system are described below.

[0487] The server first generates a virtual environment for the mock interview. This involves creating virtual characters with different cultural backgrounds and careers based on the scenario selected by the user. Having different scenarios available allows users to engage with a variety of case studies.

[0488] The terminal is a device for acquiring user voice and actions in real time. This includes hardware such as microphones and motion sensors, and the terminal's role is to transmit this data to a server. This real-time data collection allows users to receive immediate feedback.

[0489] The server inputs collected voice and behavioral data into a machine learning model to analyze unconscious biases hidden in the user's speech and actions. This analysis uses a generative AI model, which generates feedback when specific biases are detected. This feedback is provided to the user through the terminal, allowing the user to identify specific areas for improvement.

[0490] For example, by selecting an interview scenario suitable for a technical job, users can practice how to assess technical skills. When a user asks, "Tell me about your past project experience," the AI ​​model analyzes whether the question contains any particular bias and provides feedback such as, "Try not to focus too much on technical experience."

[0491] When using a generative AI model, you can use a prompt like this: "How can I identify unconscious biases in job interviews?" This prompt helps the AI ​​model suggest bias analysis methods.

[0492] In this way, the present invention enables users to effectively recognize unconscious biases and learn fair interview techniques.

[0493] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0494] Step 1:

[0495] The user selects a scenario according to their training objectives for the mock interview. As input, the user chooses one scenario from several options and sends relevant information to the server. Based on this information, the server generates virtual candidates with different cultural backgrounds and experience levels suitable for the specific scenario and transfers them to the terminal. The output is data on the virtual candidates based on the scenario selected by the user.

[0496] Step 2:

[0497] The terminal uses virtual candidate data received from the server to construct a virtual reality environment. The input consists of virtual candidate profile data and scenario conditions. Based on this, the terminal generates a VR scene, allowing the user to participate. The output is the virtual reality environment that the user can experience.

[0498] Step 3:

[0499] The user enters a virtual environment using a VR device and begins an interview with a virtual candidate. The user's input consists of their spoken words and actions in real time. The device records the user's voice and actions in real time. This recorded data is sent to a server, which then serves as input for the next analysis step.

[0500] Step 4:

[0501] The server inputs voice and behavioral data transmitted from the terminal into a machine learning model. The input data contains information about the user's speech and gestures. Based on this, the server uses a generative AI model to analyze unconscious biases. The output is the analysis result based on whether or not biases are present.

[0502] Step 5:

[0503] The server generates feedback based on the analysis results and sends it to the terminal. The feedback includes suggestions for improvement if the user's statements contain any particular biases. A specific example might be, "If your responses are too heavily biased towards technical experience, please try to ask more inclusive questions." The output is a feedback message for the user.

[0504] Step 6:

[0505] The terminal presents feedback received from the server to the user visually or audibly. The input is the feedback message received from the server. The user can review this and understand specific areas for improvement. The output is information the user can use for their next interview.

[0506] Step 7:

[0507] Once the series of mock interviews is complete, the server evaluates the entire training and suggests the next scenario. The server uses the user's progress data and feedback history as input. The output is a recommendation for the next training scenario aimed at improving the user's skills.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] In modern customer service, unconscious biases can influence customer interactions. However, there is a lack of effective training methods to help customer service staff recognize their own biases and provide fair and inclusive service. To address this issue, providing training methods that utilize virtual reality in physical stores is a key challenge.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes means for generating a simulated dialogue environment, means for recording the user's voice and behavior in real time, means for analyzing the generated voice and behavior data and applying a machine learning model to detect unconscious bias, and means for providing a virtual customer interaction through a visual device. This enables customer service staff to be trained to recognize and correct biases through diverse customer scenarios.

[0513] An "information processing device" is a device that handles digital information and performs various calculations and data management.

[0514] A "simulated dialogue environment" is a situation that provides users with the ability to simulate dialogues in a virtual space based on specific scenarios.

[0515] "Means for recording voice and actions in real time" refers to methods and devices for instantly acquiring data on a user's speech and body movements.

[0516] A "machine learning model" is an algorithm used to analyze collected data and learn patterns and features from it.

[0517] "Unconscious bias" refers to biased thoughts or attitudes that a person displays towards others who have certain attributes or backgrounds, even though they are not consciously aware of it.

[0518] A "visual device" is a device that presents virtual reality to the user as an image, and includes, for example, head-mounted displays and smart glasses.

[0519] A "virtual customer" is not an actual client, but rather a conversation partner generated by a program, and can have a variety of attributes and scenarios.

[0520] This invention is a system that supports customer service staff in interacting with virtual customers. The server first generates a simulated dialogue environment, and the user participates in this virtual environment using a visual device. The server delivers the generated dialogue scenario to the user's visual device and displays virtual customers with different cultural backgrounds and characteristics.

[0521] The device records the user's voice and actions in real time and sends this data to a server. The server analyzes this voice and behavior data using machine learning models to detect biased statements and actions by the user. If detected, the server generates feedback for the user based on the results and provides immediate visual and audible notifications.

[0522] For example, if a user unconsciously expresses bias towards a virtual customer with a specific cultural background, the server provides feedback such as, "That statement may reflect cultural bias. Please try to express yourself more neutrally." This allows users to learn how to respond fairly and inclusively in their daily customer service interactions.

[0523] This system could utilize commercially available cloud AI services as its information processing device and machine learning model. Furthermore, an example of a prompt message, which is part of the invention, is: "Is it possible that unconscious bias occurred in this dialogue? Please provide feedback on areas for improvement based on the user's statements and actions."

[0524] In this way, customer service staff can correct unconscious biases through systematic training and develop the ability to effectively interact with customers from diverse cultural backgrounds.

[0525] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0526] Step 1:

[0527] The server generates a virtual dialogue environment. The input is the scenario information selected by the user. Based on this, the server generates virtual customers with different cultural backgrounds and characteristics, and sends them to the terminal as visual data. The output is the virtual customer displayed on the user's visual device.

[0528] Step 2:

[0529] The user initiates an interaction with a virtual customer via a visual device. The input is virtual customer information sent from the server. The user's speech and actions are recorded in real time and sent from the terminal to the server. The output is the recorded audio and action data.

[0530] Step 3:

[0531] The server inputs the received voice and motion data into a generating AI model. Here, the input is the user's real-time voice and motion data, which is analyzed and used in data calculations to detect unconscious biases. The output is the analysis result indicating whether or not bias is present.

[0532] Step 4:

[0533] The server generates feedback to the user, if necessary, based on the analysis results. The input consists of bias detection results and prompts for improvement. Based on this, the server generates specific feedback messages. The output is visual or audible feedback provided to the user.

[0534] Step 5:

[0535] The terminal receives feedback from the server and notifies the user. The input is the feedback message generated by the server. The terminal conveys this to the user by displaying it on a visual device or providing an audio notification. The output is the feedback information received by the user.

[0536] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0537] This invention combines a simulated interview system using a virtual reality environment with an emotion engine that recognizes user emotions. This system allows recruiters to receive training to correct unconscious biases and conduct fair and inclusive hiring practices.

[0538] System configuration and operation

[0539] The server generates various mock interview scenarios and sends them to the user's terminal, allowing them to select their preferred scenario. The scenarios include virtual candidates with different cultural backgrounds and profiles, realistically recreating the interview situation.

[0540] The terminal is a device that provides a virtual space to the user via a VR device, allowing the user to initiate interaction with a virtual candidate in the VR environment. It records the user's voice and actions in real time and sends that data to the server.

[0541] The emotion engine analyzes the user's emotions from input voice and motion data, recognizing indicators such as joy, surprise, anger, disgust, fear, sadness, trust, and anxiety. This reveals the user's mental state during the interview.

[0542] The server combines the output of the emotion engine with an AI model to check for any unconscious biases in the user. Based on the analyzed data, feedback is generated. The feedback is adjusted according to the user's emotional state and presented to the user through the device.

[0543] Specific example

[0544] For example, if a user is in a mock interview and the virtual candidate responds with surprise, they might ask an overly aggressive question due to nervousness. In this case, the emotion engine recognizes the user's anxiety and aggression from their tone of voice and speech patterns, and provides feedback to help the user calm down. This feedback might be conveyed through the device as a notification such as, "Your reaction to the candidate's statement may be excessive."

[0545] Once a series of sessions concludes, the server records the user's progress and guides them through the next training session based on the analysis results. This allows users to hone their interviewing skills and prepare to conduct fair and effective talent selection.

[0546] This system evolves traditional mock interviews through the use of emotional data, providing an innovative training method to correct unconscious biases.

[0547] The following describes the processing flow.

[0548] Step 1:

[0549] The server prepares a variety of mock interview scenarios and sends them to the user's terminal. The user then views the list of scenarios on their terminal and selects the scenario best suited to their training.

[0550] Step 2:

[0551] The terminal sets up the VR environment based on the selected scenario and prepares the user to enter the virtual space through the VR device. Once ready, the terminal displays the virtual candidate and signals the user to begin the interview.

[0552] Step 3:

[0553] The user begins asking questions to a virtual candidate in a VR space. The device records data such as the user's speech, tone, and actions in real time. This data is then sent to a server.

[0554] Step 4:

[0555] The server inputs the transmitted data into the emotion engine. The emotion engine analyzes the user's emotions from the intonation of their voice and body movements, and identifies emotional states such as tension, anxiety, and relief.

[0556] Step 5:

[0557] Based on the output of the emotion engine and voice and behavioral data, the server uses a generative AI model to detect the user's unconscious biases. Here, it determines whether a particular question is biased.

[0558] Step 6:

[0559] The server generates feedback based on the analyzed data. This feedback takes into account the sentiment analysis results and includes advice tailored to the user's current emotional state. The feedback is then sent to the terminal.

[0560] Step 7:

[0561] The terminal visually or audibly notifies the user of feedback from the server. The user then has the opportunity to refine their interviewing skills based on the feedback provided.

[0562] Step 8:

[0563] Once the mock interview is complete, the server evaluates the user's progress and generates data to recommend the next training scenario. The user can then use this information to plan further training.

[0564] (Example 2)

[0565] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0566] Conventional mock interview systems have difficulty adequately detecting and correcting users' unconscious judgment biases, making it impossible to provide fair and effective feedback during interview training. Furthermore, they lacked feedback that considered the user's emotional state, resulting in insufficient improvement in the quality of training.

[0567] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0568] In this invention, the server includes means for generating a simulated environment in an information processing system that provides virtual reality; means for sequentially recording the user's voice and actions; means for analyzing the obtained voice and action data and applying a machine learning model to detect unconscious judgment biases; and means for analyzing the user's emotional state and using the analysis results to generate feedback. This enables a detailed understanding of user behavior, including unconscious biases, and the provision of comprehensive feedback based on emotions.

[0569] "Virtual reality" is a technology that uses computer technology to create a virtual environment that is different from reality, allowing users to experience a sense of presence within that environment.

[0570] An "information processing system" is a system consisting of a series of hardware and software components that perform data input, calculation, storage, and output.

[0571] A "simulated environment" is an environment designed to allow users to gain experience by virtually recreating specific situations in the real world.

[0572] A "user" refers to a person or organization that operates or uses a system or equipment.

[0573] "Voice and behavioral data" refers to information that captures the user's voice and body movements.

[0574] "Sequential recording" refers to the process of continuously recording data in real time.

[0575] A "machine learning model" is a mathematical model that uses large amounts of data to learn patterns from algorithms and then makes predictions and judgments about new data.

[0576] "Judgment bias" refers to a systematic bias in an individual's thinking and judgment that occurs unconsciously.

[0577] "Feedback" refers to evaluations and advice provided to users by a system.

[0578] A "virtual character" is a fictional person or creature created by a computer, which enables interaction and dialogue with the user.

[0579] This invention is an information processing system that uses virtual reality technology to provide a simulated interview environment, allowing users to experience various scenarios within it. The system's hardware configuration includes a computer system, a VR device, a voice capture device, and motion sensors. The software includes an application for generating the VR environment, a machine learning algorithm for analyzing voice and behavioral data, and an interface for providing feedback.

[0580] The server first generates a virtual environment. Based on information retrieved from the database, the server creates virtual characters with different cultural backgrounds and characteristics, allowing users to dynamically select scenarios.

[0581] The terminal provides the user with a virtual environment through a VR device and initiates interaction. The terminal collects the user's voice and movement data in real time and sends that data to the server.

[0582] Users experience a simulated interview in a virtual environment. During this process, their emotional state and unconscious judgment biases are analyzed. The resulting analysis is used to improve subsequent training sessions.

[0583] For example, when a user asks an overly anxious, one-sided question to a virtual candidate, their tone of voice and speaking speed are detected. Based on this emotional state analysis, feedback such as, "Your question may be too aggressive. Try softening your tone," is provided. This feedback, based on multifaceted information, leads to a fair and emotionally balanced evaluation.

[0584] An example of a prompt sentence generated using an AI model is: "In a virtual interview, how can you remain calm and respond flexibly even when you receive an unexpected answer?"

[0585] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0586] Step 1:

[0587] The server generates mock interview scenarios. The server retrieves profile information with different cultural backgrounds and characteristics from a database and creates multiple virtual characters based on this information. It receives user preferences and training objectives as input, generates a list of selectable scenarios as output, and sends it to the terminal.

[0588] Step 2:

[0589] The terminal provides the user with a virtual interview environment. Using scenario data received from the server, the terminal displays a virtual scenario to the user through a VR device. The user then begins interacting with a virtual character within the VR environment. It receives scenario data from the server as input and provides the virtual environment experienced by the user as output.

[0590] Step 3:

[0591] The user interacts with a virtual character. During this time, the user's voice and actions are continuously recorded by the device. The user's voice and actions are captured in real time as input, and this data is processed and sent to the server. The raw data is sent to the server as output.

[0592] Step 4:

[0593] The server analyzes the received audio and behavioral data. Using an emotion analysis engine, it estimates the user's emotions and stress levels. It analyzes the user data received as input and generates emotional states and potential judgment bias characteristics as output. These results are then fed into a generative AI model.

[0594] Step 5:

[0595] The server uses a generative AI model to create feedback for the user. Based on sentiment data and bias detection results, it identifies specific actions and reactions the user performed unconsciously and suggests improvements. It receives analysis results as input and sends the feedback content to the terminal as output.

[0596] Step 6:

[0597] The device presents the generated feedback to the user. The feedback is delivered to the user visually or aurally. It receives feedback information from the server as input and provides feedback to the user as output.

[0598] Step 7:

[0599] Based on the feedback, users review their interviewing skills and apply them to their next training session. The server further accumulates progress data and suggests the next training scenario. User responses and progress are recorded as input, and guidelines for the next step are generated as output.

[0600] (Application Example 2)

[0601] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0602] In modern security operations, security personnel often face problems where unconscious biases or excessive stress prevent them from making accurate judgments when faced with emergencies. This can lead to inappropriate responses or unnecessary escalations. This invention aims to cultivate appropriate judgment skills in security personnel by simulating these situations in a virtual reality environment and receiving real-time feedback on their emotions and psychological state.

[0603] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0604] In this invention, the server includes means for generating a simulated interview environment, means for recording the user's voice and actions in real time, means for providing a simulated situation to enhance the security officer's judgment in emergency or intimidating situations, and means for analyzing the user's psychological state in real time and providing feedback to help maintain composure. This enables security officers to learn appropriate responses in a virtual reality environment and eliminate unconscious biases, thereby enabling effective security responses.

[0605] "Virtual reality" is a technology that uses computer technology to create virtual environments for users, enabling them to have experiences that are as close to reality as possible.

[0606] A "mock interview" is a method that uses virtual reality technology to recreate a real interview situation, allowing users to practice and learn within that environment.

[0607] "Unconscious bias" is a phenomenon in which prejudices and preconceptions held by individuals without their awareness influence their judgments and actions.

[0608] "Real-time recording" is the process of recording data such as user voice and actions the moment they occur.

[0609] An "artificial intelligence model" is a mathematical or computational model that learns from data analysis and user behavior and automatically performs specific tasks.

[0610] "Feedback" is a means of providing users with information and advice based on analysis results to promote behavioral improvement and learning.

[0611] An "emergency situation" is a situation in which an unexpected event or dangerous situation occurs and requires a swift response.

[0612] "Psychological state analysis" is a process of analyzing emotions and mental responses to assess a user's current emotions and mental health.

[0613] The system implementing this invention provides a simulated environment using virtual reality, enabling security personnel to undergo training to enhance their decision-making abilities in emergency situations. The system operates using a VR headset and emotion recognition sensors. Specifically, it integrates a VR device such as Oculus Quest with emotion recognition software such as Affectiva.

[0614] The server delivers data modeling mock interviews and emergency situations to VR devices. Users wear VR headsets and immerse themselves in the virtual environment to experience realistic emergency situations. Sensors collect audio and motion data in real time, which is then transmitted to the server.

[0615] The server analyzes the collected data using a generative AI model. This analysis assesses the user's psychological state and unconscious biases, and provides feedback based on the results. The feedback is presented to the user visually or audibly, allowing the user to understand their own reactions and improve their ability to maintain composure.

[0616] For example, if a user experiences panic or anxiety when a suspicious person approaches them in VR, they will be given feedback such as, "Stay calm and continue to observe the person's movements." In this way, users can train their reactions to potential risks.

[0617] An example of a prompt sentence to be input into the generating AI model might be, "Please provide feedback on how security personnel can remain calm when encountering a suspicious person." This prompt sentence is used to provide guidance on how the user should respond.

[0618] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0619] Step 1:

[0620] The device presents a virtual environment to the user through a VR headset. The user becomes immersed in the VR environment, and an emergency scenario is displayed on the screen. During this process, the device receives input from the user to initiate interaction and loads the data that constitutes the VR environment.

[0621] Step 2:

[0622] While the user operates within the virtual environment, the terminal uses emotion recognition sensors to record the user's voice and behavioral data in real time. This collects biometric data, which is then immediately transmitted to the server. In this step, the collected biometric data becomes the input, and an output is generated that packages and transfers the data to the server.

[0623] Step 3:

[0624] The server inputs the received data into a generating AI model to analyze the user's psychological state. The analysis process infers stress and emotional fluctuations from biometric data and verifies for any unconscious biases. The determined psychological state is generated as output, which is then used to generate feedback.

[0625] Step 4:

[0626] The server generates appropriate feedback for the user based on the analysis results. The generating AI model uses prompts to create advice and warnings regarding the user's actions, and compiles this content into a feedback message. This feedback message is output and sent to the terminal in either a visual or auditory form.

[0627] Step 5:

[0628] The device presents the transmitted feedback to the user through the VR headset. The user receives the feedback and gains guidance for improving their actions. In this step, a feedback message is input, and the process of conveying the feedback to the user as a visual or audio output is performed.

[0629] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0630] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0631] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0632] [Fourth Embodiment]

[0633] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0634] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0635] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0636] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0637] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0638] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0639] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0640] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0641] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0642] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0643] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0644] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0645] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0646] This invention is a training system for recruiters to conduct fair interviews, specifically by using virtual reality (VR) to conduct mock interviews. The following details an embodiment of the system.

[0647] System Overview

[0648] The server provides various mock interview scenarios and generates virtual candidates with different cultural backgrounds and profiles. This allows users to select from a variety of scenarios according to the situation and learn practical interview techniques.

[0649] The terminal is a device that enables users to participate in a virtual environment using VR devices. The terminal records the user's voice and actions and transmits them to the server in real time.

[0650] Users select a mock interview scenario, enter a VR space, and interact with a virtual candidate. The user's goal is to recognize unconscious biases and acquire fair and neutral interviewing skills.

[0651] Program Processing Description

[0652] When the mock interview begins, the server sends a virtual candidate to the terminal based on the selected scenario. The user asks questions to the candidate in the VR space and conducts the interview through interaction.

[0653] During this time, the device records the user's voice and motion data in real time. For example, it collects information such as whether the user is using a specific gesture or if there is a change in their tone of voice. This data is immediately sent to the server.

[0654] The server inputs the received data into an artificial intelligence (AI) model to analyze whether the user's statements and actions contain unconscious biases. If a specific bias is detected as a result of the analysis, the server sends appropriate feedback to the terminal.

[0655] The device provides feedback to the user visually or audibly. For example, the user might see a notification stating, "That statement may be gender-biased."

[0656] Once a series of mock interviews is complete, the server evaluates the user's progress and recommends the next training scenario. This allows the user to continue receiving training to improve their interview skills and increase their fairness.

[0657] This system helps companies implement a talent acquisition process that promotes diversity and inclusion. One embodiment of the present invention proposes practical methods for recruiters to understand and correct unconscious biases.

[0658] The following describes the processing flow.

[0659] Step 1:

[0660] The server prepares various mock interview scenarios for the user and sends them to the terminal. The user reviews the list of scenarios on the terminal and selects the desired scenario.

[0661] Step 2:

[0662] Based on the scenario selected by the user, the device sets up the VR environment and displays virtual candidates through the user's VR device. The user then prepares to begin a mock interview based on the configured scenario.

[0663] Step 3:

[0664] When the mock interview begins, the user starts interacting with a virtual candidate in a virtual space. The user asks questions and conducts the interview based on the candidate's responses.

[0665] Step 4:

[0666] The device records the user's voice and movement data in real time. During this process, it collects various information, including gestures, speaking style, and tone, and transmits it to the server.

[0667] Step 5:

[0668] The server activates a generative AI model to analyze the collected data and check for the presence of unconscious bias. Through voice analysis and behavioral analysis, it evaluates whether the user's statements and attitudes are influenced by bias.

[0669] Step 6:

[0670] The server generates feedback based on the analysis results and sends appropriate comments and advice to the terminal. For example, it may notify users that a particular statement might be based on bias.

[0671] Step 7:

[0672] The terminal presents the user with feedback from the server, either visually or audibly. Based on this feedback, the user modifies their interview approach and seeks ways to conduct a fair evaluation.

[0673] Step 8:

[0674] Once a series of mock interviews is complete, the server evaluates the user's training progress and recommends a new scenario for the next session. This reinforces the process of continuous learning and bias correction.

[0675] (Example 1)

[0676] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] In recent years, with the increasing emphasis on diversity and inclusion in talent acquisition, conducting fair interviews free from unconscious biases has become crucial. However, traditional interview training methods have struggled to effectively analyze unconscious biases and provide feedback. As a result, obtaining concrete improvement measures to reduce interviewers' biases has been difficult.

[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0679] In this invention, the server includes a configuration means for generating a virtual environment for a mock interview, a configuration means for acquiring the user's voice and actions in real time, and a configuration means for analyzing the acquired voice and action data and applying a machine learning model to detect unconscious biases. This makes it possible to accurately detect the user's unconscious biases and provide appropriate feedback. Specifically, it is possible to clarify the biases the user has and obtain concrete improvement guidelines for the next training session.

[0680] "Virtual reality" is a technology that uses computer technology to create a virtual environment that users can immerse themselves in and experience.

[0681] An "information processing device" is a device that includes hardware and software for processing data and performing specified tasks.

[0682] A "virtual environment" is an artificial environment created using digital technology that users can interact with.

[0683] "Users" refer to people who operate this system and experience mock interviews.

[0684] "Voice and motion data" refers to information about the user's speech and body movements, which are recorded in real time.

[0685] "Unconscious bias" refers to prejudices and preconceptions that individuals hold without realizing it, and which can influence decision-making.

[0686] A "machine learning model" refers to a technology that uses algorithms and statistical methods to learn patterns from data and perform predictions and classifications.

[0687] "Feedback" refers to the information and guidelines for improvement provided to users based on the analyzed results.

[0688] "Cultural background" refers to the characteristics and historical background of the culture to which an individual belongs.

[0689] "Career history" refers to an individual's record of work experience and academic achievements up to that point.

[0690] This invention is a simulated interview training system that utilizes virtual reality. Specific embodiments of the system are described below.

[0691] The server first generates a virtual environment for the mock interview. This involves creating virtual characters with different cultural backgrounds and careers based on the scenario selected by the user. Having different scenarios available allows users to engage with a variety of case studies.

[0692] The terminal is a device for acquiring user voice and actions in real time. This includes hardware such as microphones and motion sensors, and the terminal's role is to transmit this data to a server. This real-time data collection allows users to receive immediate feedback.

[0693] The server inputs collected voice and behavioral data into a machine learning model to analyze unconscious biases hidden in the user's speech and actions. This analysis uses a generative AI model, which generates feedback when specific biases are detected. This feedback is provided to the user through the terminal, allowing the user to identify specific areas for improvement.

[0694] For example, by selecting an interview scenario suitable for a technical job, users can practice how to assess technical skills. When a user asks, "Tell me about your past project experience," the AI ​​model analyzes whether the question contains any particular bias and provides feedback such as, "Try not to focus too much on technical experience."

[0695] When using a generative AI model, you can use a prompt like this: "How can I identify unconscious biases in job interviews?" This prompt helps the AI ​​model suggest bias analysis methods.

[0696] In this way, the present invention enables users to effectively recognize unconscious biases and learn fair interview techniques.

[0697] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0698] Step 1:

[0699] The user selects a scenario according to their training objectives for the mock interview. As input, the user chooses one scenario from several options and sends relevant information to the server. Based on this information, the server generates virtual candidates with different cultural backgrounds and experience levels suitable for the specific scenario and transfers them to the terminal. The output is data on the virtual candidates based on the scenario selected by the user.

[0700] Step 2:

[0701] The terminal uses virtual candidate data received from the server to construct a virtual reality environment. The input consists of virtual candidate profile data and scenario conditions. Based on this, the terminal generates a VR scene, allowing the user to participate. The output is the virtual reality environment that the user can experience.

[0702] Step 3:

[0703] The user enters a virtual environment using a VR device and begins an interview with a virtual candidate. The user's input consists of their spoken words and actions in real time. The device records the user's voice and actions in real time. This recorded data is sent to a server, which then serves as input for the next analysis step.

[0704] Step 4:

[0705] The server inputs voice and behavioral data transmitted from the terminal into a machine learning model. The input data contains information about the user's speech and gestures. Based on this, the server uses a generative AI model to analyze unconscious biases. The output is the analysis result based on whether or not biases are present.

[0706] Step 5:

[0707] The server generates feedback based on the analysis results and sends it to the terminal. The feedback includes suggestions for improvement if the user's statements contain any particular biases. A specific example might be, "If your responses are too heavily biased towards technical experience, please try to ask more inclusive questions." The output is a feedback message for the user.

[0708] Step 6:

[0709] The terminal presents feedback received from the server to the user visually or audibly. The input is the feedback message received from the server. The user can review this and understand specific areas for improvement. The output is information the user can use for their next interview.

[0710] Step 7:

[0711] Once the series of mock interviews is complete, the server evaluates the entire training and suggests the next scenario. The server uses the user's progress data and feedback history as input. The output is a recommendation for the next training scenario aimed at improving the user's skills.

[0712] (Application Example 1)

[0713] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0714] In modern customer service, unconscious biases can influence customer interactions. However, there is a lack of effective training methods to help customer service staff recognize their own biases and provide fair and inclusive service. To address this issue, providing training methods that utilize virtual reality in physical stores is a key challenge.

[0715] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0716] In this invention, the server includes means for generating a simulated dialogue environment, means for recording the user's voice and behavior in real time, means for analyzing the generated voice and behavior data and applying a machine learning model to detect unconscious bias, and means for providing a virtual customer interaction through a visual device. This enables customer service staff to be trained to recognize and correct biases through diverse customer scenarios.

[0717] An "information processing device" is a device that handles digital information and performs various calculations and data management.

[0718] A "simulated dialogue environment" is a situation that provides users with the ability to simulate dialogues in a virtual space based on specific scenarios.

[0719] "Means for recording voice and actions in real time" refers to methods and devices for instantly acquiring data on a user's speech and body movements.

[0720] A "machine learning model" is an algorithm used to analyze collected data and learn patterns and features from it.

[0721] "Unconscious bias" refers to biased thoughts or attitudes that a person displays towards others who have certain attributes or backgrounds, even though they are not consciously aware of it.

[0722] A "visual device" is a device that presents virtual reality to the user as an image, and includes, for example, head-mounted displays and smart glasses.

[0723] A "virtual customer" is not an actual client, but rather a conversation partner generated by a program, and can have a variety of attributes and scenarios.

[0724] This invention is a system that supports customer service staff in interacting with virtual customers. The server first generates a simulated dialogue environment, and the user participates in this virtual environment using a visual device. The server delivers the generated dialogue scenario to the user's visual device and displays virtual customers with different cultural backgrounds and characteristics.

[0725] The device records the user's voice and actions in real time and sends this data to a server. The server analyzes this voice and behavior data using machine learning models to detect biased statements and actions by the user. If detected, the server generates feedback for the user based on the results and provides immediate visual and audible notifications.

[0726] For example, if a user unconsciously expresses bias towards a virtual customer with a specific cultural background, the server provides feedback such as, "That statement may reflect cultural bias. Please try to express yourself more neutrally." This allows users to learn how to respond fairly and inclusively in their daily customer service interactions.

[0727] This system could utilize commercially available cloud AI services as its information processing device and machine learning model. Furthermore, an example of a prompt message, which is part of the invention, is: "Is it possible that unconscious bias occurred in this dialogue? Please provide feedback on areas for improvement based on the user's statements and actions."

[0728] In this way, customer service staff can correct unconscious biases through systematic training and develop the ability to effectively interact with customers from diverse cultural backgrounds.

[0729] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0730] Step 1:

[0731] The server generates a virtual dialogue environment. The input is the scenario information selected by the user. Based on this, the server generates virtual customers with different cultural backgrounds and characteristics, and sends them to the terminal as visual data. The output is the virtual customer displayed on the user's visual device.

[0732] Step 2:

[0733] The user initiates an interaction with a virtual customer via a visual device. The input is virtual customer information sent from the server. The user's speech and actions are recorded in real time and sent from the terminal to the server. The output is the recorded audio and action data.

[0734] Step 3:

[0735] The server inputs the received voice and motion data into a generating AI model. Here, the input is the user's real-time voice and motion data, which is analyzed and used in data calculations to detect unconscious biases. The output is the analysis result indicating whether or not bias is present.

[0736] Step 4:

[0737] The server generates feedback to the user, if necessary, based on the analysis results. The input consists of bias detection results and prompts for improvement. Based on this, the server generates specific feedback messages. The output is visual or audible feedback provided to the user.

[0738] Step 5:

[0739] The terminal receives feedback from the server and notifies the user. The input is the feedback message generated by the server. The terminal conveys this to the user by displaying it on a visual device or providing an audio notification. The output is the feedback information received by the user.

[0740] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0741] This invention combines a simulated interview system using a virtual reality environment with an emotion engine that recognizes user emotions. This system allows recruiters to receive training to correct unconscious biases and conduct fair and inclusive hiring practices.

[0742] System configuration and operation

[0743] The server generates various mock interview scenarios and sends them to the user's terminal, allowing them to select their preferred scenario. The scenarios include virtual candidates with different cultural backgrounds and profiles, realistically recreating the interview situation.

[0744] The terminal is a device that provides a virtual space to the user via a VR device, allowing the user to initiate interaction with a virtual candidate in the VR environment. It records the user's voice and actions in real time and sends that data to the server.

[0745] The emotion engine analyzes the user's emotions from input voice and motion data, recognizing indicators such as joy, surprise, anger, disgust, fear, sadness, trust, and anxiety. This reveals the user's mental state during the interview.

[0746] The server combines the output of the emotion engine with an AI model to check for any unconscious biases in the user. Based on the analyzed data, feedback is generated. The feedback is adjusted according to the user's emotional state and presented to the user through the device.

[0747] Specific example

[0748] For example, if a user is in a mock interview and the virtual candidate responds with surprise, they might ask an overly aggressive question due to nervousness. In this case, the emotion engine recognizes the user's anxiety and aggression from their tone of voice and speech patterns, and provides feedback to help the user calm down. This feedback might be conveyed through the device as a notification such as, "Your reaction to the candidate's statement may be excessive."

[0749] Once a series of sessions concludes, the server records the user's progress and guides them through the next training session based on the analysis results. This allows users to hone their interviewing skills and prepare to conduct fair and effective talent selection.

[0750] This system evolves traditional mock interviews through the use of emotional data, providing an innovative training method to correct unconscious biases.

[0751] The following describes the processing flow.

[0752] Step 1:

[0753] The server prepares a variety of mock interview scenarios and sends them to the user's terminal. The user then views the list of scenarios on their terminal and selects the scenario best suited to their training.

[0754] Step 2:

[0755] The terminal sets up the VR environment based on the selected scenario and prepares the user to enter the virtual space through the VR device. Once ready, the terminal displays the virtual candidate and signals the user to begin the interview.

[0756] Step 3:

[0757] The user begins asking questions to a virtual candidate in a VR space. The device records data such as the user's speech, tone, and actions in real time. This data is then sent to a server.

[0758] Step 4:

[0759] The server inputs the transmitted data into the emotion engine. The emotion engine analyzes the user's emotions from the intonation of their voice and body movements, and identifies emotional states such as tension, anxiety, and relief.

[0760] Step 5:

[0761] Based on the output of the emotion engine and voice and behavioral data, the server uses a generative AI model to detect the user's unconscious biases. Here, it determines whether a particular question is biased.

[0762] Step 6:

[0763] The server generates feedback based on the analyzed data. This feedback takes into account the sentiment analysis results and includes advice tailored to the user's current emotional state. The feedback is then sent to the terminal.

[0764] Step 7:

[0765] The terminal visually or audibly notifies the user of feedback from the server. The user then has the opportunity to refine their interviewing skills based on the feedback provided.

[0766] Step 8:

[0767] Once the mock interview is complete, the server evaluates the user's progress and generates data to recommend the next training scenario. The user can then use this information to plan further training.

[0768] (Example 2)

[0769] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0770] Conventional mock interview systems have difficulty adequately detecting and correcting users' unconscious judgment biases, making it impossible to provide fair and effective feedback during interview training. Furthermore, they lacked feedback that considered the user's emotional state, resulting in insufficient improvement in the quality of training.

[0771] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0772] In this invention, the server includes means for generating a simulated environment in an information processing system that provides virtual reality; means for sequentially recording the user's voice and actions; means for analyzing the obtained voice and action data and applying a machine learning model to detect unconscious judgment biases; and means for analyzing the user's emotional state and using the analysis results to generate feedback. This enables a detailed understanding of user behavior, including unconscious biases, and the provision of comprehensive feedback based on emotions.

[0773] "Virtual reality" is a technology that uses computer technology to create a virtual environment that is different from reality, allowing users to experience a sense of presence within that environment.

[0774] An "information processing system" is a system consisting of a series of hardware and software components that perform data input, calculation, storage, and output.

[0775] A "simulated environment" is an environment designed to allow users to gain experience by virtually recreating specific situations in the real world.

[0776] A "user" refers to a person or organization that operates or uses a system or equipment.

[0777] "Voice and behavioral data" refers to information that captures the user's voice and body movements.

[0778] "Sequential recording" refers to the process of continuously recording data in real time.

[0779] A "machine learning model" is a mathematical model that uses large amounts of data to learn patterns from algorithms and then makes predictions and judgments about new data.

[0780] "Judgment bias" refers to a systematic bias in an individual's thinking and judgment that occurs unconsciously.

[0781] "Feedback" refers to evaluations and advice provided to users by a system.

[0782] A "virtual character" is a fictional person or creature created by a computer, which enables interaction and dialogue with the user.

[0783] This invention is an information processing system that uses virtual reality technology to provide a simulated interview environment, allowing users to experience various scenarios within it. The system's hardware configuration includes a computer system, a VR device, a voice capture device, and motion sensors. The software includes an application for generating the VR environment, a machine learning algorithm for analyzing voice and behavioral data, and an interface for providing feedback.

[0784] The server first generates a virtual environment. Based on information retrieved from the database, the server creates virtual characters with different cultural backgrounds and characteristics, allowing users to dynamically select scenarios.

[0785] The terminal provides the user with a virtual environment through a VR device and initiates interaction. The terminal collects the user's voice and movement data in real time and sends that data to the server.

[0786] Users experience a simulated interview in a virtual environment. During this process, their emotional state and unconscious judgment biases are analyzed. The resulting analysis is used to improve subsequent training sessions.

[0787] For example, when a user asks an overly anxious, one-sided question to a virtual candidate, their tone of voice and speaking speed are detected. Based on this emotional state analysis, feedback such as, "Your question may be too aggressive. Try softening your tone," is provided. This feedback, based on multifaceted information, leads to a fair and emotionally balanced evaluation.

[0788] An example of a prompt sentence generated using an AI model is: "In a virtual interview, how can you remain calm and respond flexibly even when you receive an unexpected answer?"

[0789] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0790] Step 1:

[0791] The server generates mock interview scenarios. The server retrieves profile information with different cultural backgrounds and characteristics from a database and creates multiple virtual characters based on this information. It receives user preferences and training objectives as input, generates a list of selectable scenarios as output, and sends it to the terminal.

[0792] Step 2:

[0793] The terminal provides the user with a virtual interview environment. Using scenario data received from the server, the terminal displays a virtual scenario to the user through a VR device. The user then begins interacting with a virtual character within the VR environment. It receives scenario data from the server as input and provides the virtual environment experienced by the user as output.

[0794] Step 3:

[0795] The user interacts with a virtual character. During this time, the user's voice and actions are continuously recorded by the device. The user's voice and actions are captured in real time as input, and this data is processed and sent to the server. The raw data is sent to the server as output.

[0796] Step 4:

[0797] The server analyzes the received audio and behavioral data. Using an emotion analysis engine, it estimates the user's emotions and stress levels. It analyzes the user data received as input and generates emotional states and potential judgment bias characteristics as output. These results are then fed into a generative AI model.

[0798] Step 5:

[0799] The server uses a generative AI model to create feedback for the user. Based on sentiment data and bias detection results, it identifies specific actions and reactions the user performed unconsciously and suggests improvements. It receives analysis results as input and sends the feedback content to the terminal as output.

[0800] Step 6:

[0801] The device presents the generated feedback to the user. The feedback is delivered to the user visually or aurally. It receives feedback information from the server as input and provides feedback to the user as output.

[0802] Step 7:

[0803] Based on the feedback, users review their interviewing skills and apply them to their next training session. The server further accumulates progress data and suggests the next training scenario. User responses and progress are recorded as input, and guidelines for the next step are generated as output.

[0804] (Application Example 2)

[0805] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0806] In modern security operations, security personnel often face problems where unconscious biases or excessive stress prevent them from making accurate judgments when faced with emergencies. This can lead to inappropriate responses or unnecessary escalations. This invention aims to cultivate appropriate judgment skills in security personnel by simulating these situations in a virtual reality environment and receiving real-time feedback on their emotions and psychological state.

[0807] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0808] In this invention, the server includes means for generating a simulated interview environment, means for recording the user's voice and actions in real time, means for providing a simulated situation to enhance the security officer's judgment in emergency or intimidating situations, and means for analyzing the user's psychological state in real time and providing feedback to help maintain composure. This enables security officers to learn appropriate responses in a virtual reality environment and eliminate unconscious biases, thereby enabling effective security responses.

[0809] "Virtual reality" is a technology that uses computer technology to create virtual environments for users, enabling them to have experiences that are as close to reality as possible.

[0810] A "mock interview" is a method that uses virtual reality technology to recreate a real interview situation, allowing users to practice and learn within that environment.

[0811] "Unconscious bias" is a phenomenon in which prejudices and preconceptions held by individuals without their awareness influence their judgments and actions.

[0812] "Real-time recording" is the process of recording data such as user voice and actions the moment they occur.

[0813] An "artificial intelligence model" is a mathematical or computational model that learns from data analysis and user behavior and automatically performs specific tasks.

[0814] "Feedback" is a means of providing users with information and advice based on analysis results to promote behavioral improvement and learning.

[0815] An "emergency situation" is a situation in which an unexpected event or dangerous situation occurs and requires a swift response.

[0816] "Psychological state analysis" is a process of analyzing emotions and mental responses to assess a user's current emotions and mental health.

[0817] The system implementing this invention provides a simulated environment using virtual reality, enabling security personnel to undergo training to enhance their decision-making abilities in emergency situations. The system operates using a VR headset and emotion recognition sensors. Specifically, it integrates a VR device such as Oculus Quest with emotion recognition software such as Affectiva.

[0818] The server delivers data modeling mock interviews and emergency situations to VR devices. Users wear VR headsets and immerse themselves in the virtual environment to experience realistic emergency situations. Sensors collect audio and motion data in real time, which is then transmitted to the server.

[0819] The server analyzes the collected data using a generative AI model. This analysis assesses the user's psychological state and unconscious biases, and provides feedback based on the results. The feedback is presented to the user visually or audibly, allowing the user to understand their own reactions and improve their ability to maintain composure.

[0820] For example, if a user experiences panic or anxiety when a suspicious person approaches them in VR, they will be given feedback such as, "Stay calm and continue to observe the person's movements." In this way, users can train their reactions to potential risks.

[0821] An example of a prompt sentence to be input into the generating AI model might be, "Please provide feedback on how security personnel can remain calm when encountering a suspicious person." This prompt sentence is used to provide guidance on how the user should respond.

[0822] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0823] Step 1:

[0824] The device presents a virtual environment to the user through a VR headset. The user becomes immersed in the VR environment, and an emergency scenario is displayed on the screen. During this process, the device receives input from the user to initiate interaction and loads the data that constitutes the VR environment.

[0825] Step 2:

[0826] While the user operates within the virtual environment, the terminal uses emotion recognition sensors to record the user's voice and behavioral data in real time. This collects biometric data, which is then immediately transmitted to the server. In this step, the collected biometric data becomes the input, and an output is generated that packages and transfers the data to the server.

[0827] Step 3:

[0828] The server inputs the received data into a generating AI model to analyze the user's psychological state. The analysis process infers stress and emotional fluctuations from biometric data and verifies for any unconscious biases. The determined psychological state is generated as output, which is then used to generate feedback.

[0829] Step 4:

[0830] The server generates appropriate feedback for the user based on the analysis results. The generating AI model uses prompts to create advice and warnings regarding the user's actions, and compiles this content into a feedback message. This feedback message is output and sent to the terminal in either a visual or auditory form.

[0831] Step 5:

[0832] The device presents the transmitted feedback to the user through the VR headset. The user receives the feedback and gains guidance for improving their actions. In this step, a feedback message is input, and the process of conveying the feedback to the user as a visual or audio output is performed.

[0833] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0834] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0835] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0836] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0837] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0838] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0839] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0840] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0841] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0842] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0843] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0844] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0845] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0846] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0847] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0848] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0849] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0850] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0851] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0852] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0853] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0854] The following is further disclosed regarding the embodiments described above.

[0855] (Claim 1)

[0856] A computer device that provides virtual reality, comprising means for generating a simulated interview environment,

[0857] A means for recording the user's voice and actions in real time,

[0858] A means for analyzing generated voice and motion data and applying an artificial intelligence model to detect unconscious bias,

[0859] A means of providing feedback to users based on the analysis results,

[0860] A means of evaluating training progress and recommending the next training scenario,

[0861] A system that includes this.

[0862] (Claim 2)

[0863] The system according to claim 1, further comprising means for generating virtual candidates with different cultural backgrounds and profiles based on a mock interview scenario selected by the user.

[0864] (Claim 3)

[0865] The system according to claim 1, comprising means for providing the generated feedback to the user through visual or auditory notification.

[0866] "Example 1"

[0867] (Claim 1)

[0868] An information processing device that provides virtual reality includes a configuration means for generating a virtual environment for a mock interview,

[0869] A configuration means for acquiring the user's voice and actions in real time,

[0870] A configuration means for analyzing acquired voice and motion data and applying a machine learning model to detect unconscious biases,

[0871] A configuration means for providing notifications to users based on analysis results,

[0872] A configuration means for evaluating the progress of training and proposing the next training scenario,

[0873] Information technology systems including

[0874] (Claim 2)

[0875] The information technology system according to claim 1, further comprising a means for generating virtual individuals with diverse cultural backgrounds and careers based on a mock interview scenario selected by the user.

[0876] (Claim 3)

[0877] The information technology system according to claim 1, comprising means for presenting the provided notice to the user in a visual or auditory manner.

[0878] "Application Example 1"

[0879] (Claim 1)

[0880] An information processing device that provides a virtual environment includes means for generating a simulated dialogue environment,

[0881] A means of recording user voice and behavior in real time,

[0882] A means for analyzing generated audio and behavioral data and applying machine learning models to detect unconscious bias,

[0883] A means of providing feedback to users based on the analysis results,

[0884] A means to evaluate training progress and recommend the next learning scenario,

[0885] A means of providing virtual customer interaction through visual devices,

[0886] A system that includes this.

[0887] (Claim 2)

[0888] The system according to claim 1, further comprising means for generating virtual characters with different cultural backgrounds and characteristics based on a simulated dialogue scenario selected by the user.

[0889] (Claim 3)

[0890] The system according to claim 1, comprising means for providing the generated feedback to the user through visual or audible notification.

[0891] "Example 2 of combining an emotion engine"

[0892] (Claim 1)

[0893] In an information processing system that provides virtual reality, means for generating a simulated environment,

[0894] A means for sequentially recording the user's voice and actions,

[0895] A means for analyzing the obtained voice and behavioral data and applying a machine learning model to detect unconscious judgment biases,

[0896] A means of providing feedback to users based on the analysis results,

[0897] A means of evaluating training progress and recommending the content of the next training session,

[0898] A means of analyzing the emotional state of users and using the analysis results to generate feedback,

[0899] A device that includes this.

[0900] (Claim 2)

[0901] The apparatus according to claim 1, further comprising means for generating virtual characters having different cultural backgrounds and characteristics based on a simulated environment selected by the user.

[0902] (Claim 3)

[0903] The apparatus according to claim 1, comprising means for providing the generated feedback to the user through visual or auditory notification.

[0904] "Application example 2 when combining with an emotional engine"

[0905] (Claim 1)

[0906] A computer device that provides virtual reality, comprising means for generating a simulated interview environment,

[0907] A means for recording the user's voice and actions in real time,

[0908] A means for analyzing generated voice and motion data and applying an artificial intelligence model to detect unconscious bias,

[0909] A means of providing feedback to users based on the analysis results,

[0910] A means of evaluating training progress and recommending the next training scenario,

[0911] A means of providing simulated situations to enhance the judgment of security personnel in emergency and intimidating situations,

[0912] A means of analyzing the user's psychological state in real time and providing feedback to help them maintain composure,

[0913] A system that includes this.

[0914] (Claim 2)

[0915] The system according to claim 1, comprising means for generating virtual candidates with different cultural backgrounds and profiles based on a mock interview scenario selected by the user, and further providing a simulated emergency scenario.

[0916] (Claim 3)

[0917] The system according to claim 1, further comprising means for providing the generated feedback to the user through notifications that assist in visual or auditory and real-time situational judgment. [Explanation of Symbols]

[0918] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A computer device that provides virtual reality, comprising means for generating a simulated interview environment, A means for recording the user's voice and actions in real time, A means for analyzing generated voice and motion data and applying an artificial intelligence model to detect unconscious bias, A means of providing feedback to users based on the analysis results, A means of evaluating training progress and recommending the next training scenario, A system that includes this.

2. The system according to claim 1, further comprising means for generating virtual candidates with different cultural backgrounds and profiles based on a mock interview scenario selected by the user.

3. The system according to claim 1, comprising means for providing the generated feedback to the user through visual or auditory notification.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A