system

The system provides a realistic interview simulation by using AI-generated virtual interviewers and detailed feedback to enhance user performance.

JP2026037460APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Conventional interview training methods fail to provide a realistic simulation environment, making it difficult to evaluate user responses and nervousness, and feedback is limited, leading to unexpected issues during actual interviews.

Method used

A system that includes authentication, selection of a virtual interviewer and question set, generation of a virtual interviewer's image and voice using AI, recording and analyzing user answers, and providing detailed feedback visually and audibly to simulate a realistic interview environment.

Benefits of technology

Enables users to train in a realistic setting, receiving detailed feedback to improve their interview skills effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026037460000001_ABST
    Figure 2026037460000001_ABST
Patent Text Reader

Abstract

Provide a system. [Solution] an authentication means for users to log in; means for selecting an appropriate virtual interviewer and question set from the received user information; A means for generating an image and voice of a virtual interviewer using generative AI; a means for displaying the generated virtual interviewer and playing the questions aloud; a means for registering a user's answer by voice and analyzing the voice; A means for generating feedback based on the analysis results and providing it visually and audibly; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional interview training methods have had the problem of being unable to provide effective training due to the difficulty of simulating a realistic interview environment. Furthermore, it is difficult to evaluate in detail the user's responses and level of nervousness, and feedback is limited. As a result, users often encounter unexpected problems during the actual interview. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides the following means. It includes an authentication means for a user to log in, and a means for selecting an appropriate virtual interviewer and question set from received user information. It also includes a means for generating an image and voice of a virtual interviewer using a generation AI, and a means for displaying the generated virtual interviewer and playing questions aloud. It provides a means for recording the user's answers by voice and analyzing the voice, and generates feedback based on the analysis results, which is provided both visually and aloud. This allows the user to train in a realistic interview environment and receive detailed feedback.

[0006] "User" refers to an individual who uses the system to receive interview training.

[0007] "Authentication means" refers to a means that has the function of verifying a user's login information and determining whether or not they can access the system.

[0008] "Virtual Interviewer" refers to a character with the image and voice of a virtual interviewer created by generative AI.

[0009] A "question set" refers to a series of interview questions generated based on the user's information and past data.

[0010] "Generative AI" refers to artificial intelligence technology that uses a computer to generate the image and voice of a virtual interviewer.

[0011] "Display means" refers to a means having a function for displaying an image of a virtual interviewer on a user's terminal.

[0012] "Audio playback means" refers to a means having a function for playing back the voice of the virtual interviewer on the user's terminal.

[0013] The "voice registration means" refers to a means having a function for recording the user's response as voice data.

[0014] "Analysis means" refers to means having a function for analyzing the voice data of a registered user and evaluating the content, voice quality, volume, speaking speed, and level of tension.

[0015] The "feedback generation means" refers to a means having a function for generating feedback to be provided to the user based on the analysis results.

[0016] "Display and audio providing means" refers to means having the function of providing the generated feedback to the user visually and audibly. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The system of the present invention is designed to allow users to receive training in a realistic interview environment. The system analyzes the user's spoken responses and provides detailed feedback to improve the user's interview skills. The following describes the specific processing of the program as an embodiment of the present invention.

[0039] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the most appropriate virtual interviewer and question set based on the user's past training history and the interview format desired by the user.

[0040] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0041] Once the user's response is recorded, the device sends the audio data to a server. The server then activates a voice analysis system to analyze the recorded audio data. The audio data is converted into text, and its content, voice quality, volume, speaking speed, and level of tension are evaluated. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Corporation," the appropriateness of the content, clarity of the voice, and speed of speech are evaluated.

[0042] Based on the analysis results, the server generates feedback, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0043] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0044] The processing flow will be explained below.

[0045] Step 1:

[0046] The user enters their ID and password on the login screen and clicks the login button.

[0047] The terminal sends the entered login information to the server.

[0048] Step 2:

[0049] The server authenticates the received login information, and if the user authentication is successful, obtains the user's profile information.

[0050] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0051] Step 3:

[0052] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0053] The device displays the generated virtual interviewer and prepares for audio output.

[0054] Step 4:

[0055] The device will play the first question asked by the virtual interviewer.

[0056] The user answers the questions verbally.

[0057] The device records the user's response and sends the audio data to the server.

[0058] Step 5:

[0059] The server starts the voice analysis system and analyzes the recorded voice data.

[0060] The server converts the voice data into text and analyzes the content.

[0061] Step 6:

[0062] The server evaluates the content of the user's answers, voice quality, volume, speaking speed, and level of tension based on the analysis results.

[0063] Examples: Evaluate whether the content is appropriate, whether the speaker speaks quickly, and whether the speaker speaks quietly.

[0064] Step 7:

[0065] The server generates feedback based on the analysis results and sends it to the device.

[0066] Example: Feedback such as, "The content is appropriate, but you speak a little too quickly. You should try to speak more slowly."

[0067] Step 8:

[0068] The device displays the feedback to the user and also plays it back as audio.

[0069] The user reviews the feedback and retrains if necessary.

[0070] Example 1

[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0072] In conventional interview training systems, it was difficult for users to receive training in a realistic interview environment. It was also difficult to analyze users' voice responses and provide specific feedback. As a result, it was not possible to effectively improve users' interview skills.

[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0074] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generation AI, means for displaying the generated virtual interviewer and playing questions aloud, means for registering the user's answers aloud and analyzing the voice data, means for converting the voice data into text and evaluating the content, voice quality, volume, speaking speed, and level of tension, and means for generating feedback based on the analysis results and providing it as a display and voice. This allows users to train in an environment similar to a real interview and receive specific feedback, thereby effectively improving their interview skills.

[0075] "Authentication means" refers to a method or device used to verify the identity of a user when accessing a system, such as entering an ID and password.

[0076] A "virtual interviewer" is a digital character or system that mimics a real interviewer when users are undergoing interview training. Images and sounds are generated by generative AI.

[0077] A "question set" is a series of questions that the virtual interviewer asks the user.

[0078] "Generative AI" refers to systems or algorithms that use artificial intelligence techniques to generate the image and voice of a virtual interviewer.

[0079] The "audio playback means" refers to a method or device that allows the virtual interviewer to audibly communicate questions to the user, such as a speaker or an audio player.

[0080] "Voice data" refers to information that is a digital recording of a user's speech.

[0081] "Means of analysis" refers to methods or techniques for converting the content of voice data into text and analyzing voice quality, volume, speaking speed, level of tension, etc., in order to evaluate the data.

[0082] "Convert to text" is the process of converting audio data into written information, for example using voice recognition technology.

[0083] "Feedback" refers to specific advice or comments to the user that are generated based on the analysis results.

[0084] The "display means" refers to a method or device for visually presenting the generated feedback to the user, such as a display or monitor.

[0085] The "means for providing by voice" refers to a method or device for transmitting the generated feedback to the user by voice, such as a speaker or a headset.

[0086] The system of the present invention is designed to allow users to receive training in a realistic interview environment, and the system analyzes users' spoken responses and provides detailed feedback to help improve their interview skills.

[0087] First, the user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server compares the received login information with its database and performs authentication. If authentication is successful, the server obtains the user's profile information.

[0088] Next, the server selects the optimal virtual interviewer and question set based on the user's past training history and the interview format desired by the user. To do this, the server analyzes the user's information and uses generative AI to generate an image and voice of the virtual interviewer. The image and voice data of this generated virtual interviewer are sent to the device. The device displays this virtual interviewer and notifies the user that the interview is ready.

[0089] When the interview begins, the device plays a voice message asking the virtual interviewer an initial question. For example, the question might be, "Please tell us about yourself." The user answers the question aloud, and the answer is recorded by the device.

[0090] Once the user's response is recorded, the device sends the voice data to the server. The server then activates the voice analysis system and analyzes the recorded voice data. The voice data is first converted into text. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Co., Ltd.", this voice will be converted into text format.

[0091] The server uses this text data to evaluate the appropriateness of the content, voice quality, volume, speaking speed, and level of tension. Based on the evaluation, the server generates analysis results and provides appropriate feedback. For example, it may generate specific advice such as, "The content is appropriate, but you speak a little too quickly. Please try to speak a little more slowly."

[0092] The generated feedback is sent to the device and provided to the user visually and audibly, allowing the user to review the feedback and retrain if necessary.

[0093] Below is an example of a prompt sentence to input to the generative AI model.

[0094] "I'm Tanaka from XYZ Co., Ltd. I'm the manager of the sales department.

[0095] Analyze this audio data and rate it on the following:

[0096] 1. Content Appropriateness

[0097] 2. Voice quality

[0098] 3. Volume

[0099] 4. Speaking Speed

[0100] 5. Tension

[0101] To implement the system of the present invention, hardware such as a general computer, microphone, speaker, display, etc. is required, and software such as a generative AI model (e.g., GPT-3 (registered trademark)) and a voice analysis system (e.g., Google (registered trademark) Cloud Speech-to-Text API, IBM Watson (registered trademark)) is used.

[0102] The system of the present invention allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0103] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0104] Step 1:

[0105] The user enters their ID and password on the login screen and clicks the login button. The entered ID and password are sent from the terminal to the server. The server receives the input and compares it with the authentication information stored in the database. If the authentication is successful, the server retrieves the profile information and returns permission to the terminal to access the training system.

[0106] Input: User ID and password

[0107] Output: Authentication success / failure result, profile information

[0108] Specific operation: When the user clicks the login button, the device sends a request to the server via the browser or app, and the server performs a database search.

[0109] Step 2:

[0110] The server analyzes the user's profile information, past training history, and the user's desired interview format, and selects the optimal virtual interviewer and question set. This selection is based on the user's past training data and user requests.

[0111] Input: Profile information, past training data, user requests

[0112] Output: Selected virtual interviewer and question set

[0113] Specific operation: The relevant information is obtained from a database provided by the server, and the optimal settings are selected using an algorithm.

[0114] Step 3:

[0115] The server uses a generative AI to generate images and voices of virtual interviewers, using a pre-trained model, and sends the generated images and voice data to the device.

[0116] Input: A selected virtual interviewer and a set of questions

[0117] Output: Images and audio data of the generated virtual interviewer

[0118] Specific operation: A generative AI model (e.g., GPT-3 or an image generation model) is used and data is delivered to the device.

[0119] Step 4:

[0120] The terminal displays the image and voice of the virtual interviewer and notifies the user that the interview is ready. After the notification, the terminal plays the first question in the voice of the virtual interviewer.

[0121] Input: Image and audio data of a virtual interviewer

[0122] Output: Questions played back and notification that the interview is ready

[0123] Specific operation: The device displays an image of a virtual interviewer on the screen and plays questions through the speaker.

[0124] Step 5:

[0125] The user answers the questions posed by the virtual interviewer by voice, and this voice is recorded by the device.

[0126] Input: User's spoken response

[0127] Output: Recorded audio data

[0128] Specific operation: Records the user's speech through the device's microphone and saves it as an audio file.

[0129] Step 6:

[0130] The device sends the recorded voice data to a server, which then activates a voice analysis system to convert the recorded voice data into text, and analyzes the content, voice quality, volume, speaking speed, and level of tension.

[0131] Input: Recorded audio data

[0132] Output: Analysis results (text data, evaluation items)

[0133] What it does: The server uses a speech analysis service such as the Google Cloud Speech-to-Text API to convert the speech into text and run the evaluation algorithm.

[0134] Step 7:

[0135] The server generates feedback based on the analysis results, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device.

[0136] Input: Analysis results

[0137] Output: Generated feedback

[0138] Specific operation: Feedback is created using a template based on the analysis results and sent to the device.

[0139] Step 8:

[0140] The device provides the generated feedback to the user visually and audibly, allowing the user to review it and retrain if necessary.

[0141] Input: Generated feedback

[0142] Output: Visual and audio feedback

[0143] Specific behavior: Display feedback on the device display and play it audibly through the speaker.

[0144] These are the specific processing steps of this system. At each step, the user, terminal, and server work together to provide the user with the optimal interview training environment.

[0145] (Application example 1)

[0146] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0147] Conventional driver training systems lack advanced technology to simulate real-world driving environments, making it difficult for users to receive training that is in line with actual driving. In addition, because voice instructions and feedback are not provided in real time, it is difficult to make immediate improvements, and effective improvement of driving skills cannot be expected.

[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0149] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual instructor and instruction set from received user information, a means for generating an image and audio of the virtual instructor using a generation AI, a means for displaying the generated virtual instructor and playing questions by audio, a means for recording the user's answers by audio and analyzing the audio, a means for generating feedback based on the analysis results and providing it by display and audio, a means for generating a virtual driving environment and issuing driving instructions to the user, and a means for analyzing the user's driving situation in real time. This allows the user to receive training in a situation similar to a real driving environment, and is expected to improve their skills through real-time feedback.

[0150] "Authentication means for user login" refers to the process of sending authentication information, such as an ID and password used by the user to identify themselves, to the server and performing authentication.

[0151] The "means for selecting an appropriate virtual instructor and set of instructions from the received user information" is an algorithm for selecting the most appropriate virtual instructor and driving instructions based on the user's profile information and past training history.

[0152] "Generative AI" is a system that uses artificial intelligence technology to generate images and sounds of virtual characters.

[0153] A "virtual instructor" is a virtual character generated by the generation AI, whose role is to give driving instructions to the user.

[0154] The "means for playing back questions by voice" is a voice synthesis technology that allows the virtual instructor to communicate instructions and questions to the user by voice.

[0155] The "means for recording a user's voice response and analyzing the voice" is a system for recording a user's voice response and analyzing the voice data.

[0156] "Means for generating feedback based on the analysis results and providing it visually and audibly" refers to a system that generates appropriate feedback based on the speech analysis results and provides it to the user visually and audibly.

[0157] The "means for generating a virtual driving environment and issuing driving instructions to a user" is a system for generating a simulation environment for a user to undergo driving training in a virtual environment and for a virtual instructor to issue driving instructions.

[0158] The "means for analyzing the user's driving status in real time" is a system that monitors the driving behavior of the user in a virtual environment in real time and analyzes the data.

[0159] The system of the present invention allows a user to receive training in a driving environment similar to a real driving environment, thereby effectively improving driving skills. This system analyzes the user's voice responses and driving situation and provides detailed feedback, thereby improving the user's driving skills. Specific processing of the system will be described below as an embodiment of the present invention.

[0160] First, the user logs into the system using their smartphone or head-mounted display (HMD). They enter their ID and password on the login screen and click the login button. The device then sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the optimal virtual instructor and instruction set based on the user's past training history and the user's desired driving environment.

[0161] The server then uses the generative AI model to generate an image and voice of the selected virtual instructor. This generated virtual instructor is displayed on the device, informing the user that they are ready to drive. When the driving simulation begins, the device plays the initial instructions from the virtual instructor by voice. The user performs driving maneuvers according to these instructions, and the driving situation is monitored and recorded in real time by the device.

[0162] Once the user's driving behavior is recorded, the device sends the data to a server. The server then activates a voice analysis system and driving data analysis system to analyze the recorded data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, level of tension, and driving skill are evaluated. For example, if a user says, "I missed the red light," the system evaluates whether the content was appropriate and whether the driving maneuver was safe.

[0163] Based on the analysis results, the server generates feedback, which may include specific advice such as "Pay more attention to traffic lights next time." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0164] These processes are achieved through speech recognition using the speech_recognition library, speech synthesis technology using the pyttsx3 library, and a generative AI model using the ai_model.

[0165] As a specific example, the prompt sentence for "if the user ignores the traffic light" is as follows:

[0166] "Generate appropriate feedback when a user ignores a signal."

[0167] "Generate feedback when the user brakes suddenly while driving."

[0168] In this way, the system of the present invention allows users to receive high-quality driving training in real time, and is expected to improve their driving skills safely and effectively in actual driving.

[0169] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0170] Step 1:

[0171] An authentication method is used for users to log in. The user accesses the login screen via a smartphone or HMD, enters their ID and password, and clicks the login button. The login information is sent from the device to the server, and the server authenticates the received login information. If authentication is successful, the server obtains the user's profile information. The input data is the user's ID and password, and the output data is the user's profile information.

[0172] Step 2:

[0173] The server selects an appropriate virtual instructor and instruction set from the received user information. The server selects the optimal virtual instructor and instruction set based on the user's past training history and desired driving environment. The input data is the user's profile information and past training history, and the output data is the virtual instructor and instruction set.

[0174] Step 3:

[0175] The generative AI model is used to generate the image and voice of the virtual instructor. The server inputs the image and voice of the selected virtual instructor into the generative AI model to generate the image and voice of the virtual instructor. The input data is the virtual instructor's profile, and the output data is the generated image and voice.

[0176] Step 4:

[0177] The generated virtual instructor is displayed and driving instructions are played back by voice.The terminal displays an image of the generated virtual instructor and plays back driving instructions by voice.The input data are the generated image and voice, and the output data are the display and voice playback of the virtual instructor.

[0178] Step 5:

[0179] The user's driving status is recorded by voice, and the voice is analyzed. The user performs driving operations according to the instructions of the virtual instructor. The device monitors and records the driving status in real time. The recorded voice data is sent from the device to a server, which analyzes the data using a voice analysis system. The input data is the user's driving status and voice data, and the output data is the analysis results.

[0180] Step 6:

[0181] Feedback is generated based on the analysis results and provided visually and audibly. The server generates feedback for the user based on the analysis results. The feedback includes specific instructions for improving driving. The generated feedback is sent to the terminal, which then provides it to the user visually and audibly. The input data is the analysis results, and the output data is the feedback.

[0182] These processing steps allow users to receive effective training in a situation that closely resembles a real driving environment, and real-time feedback is expected to improve users' driving skills.

[0183] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0184] The system of the present invention is designed to enable users to receive training in a realistic interview environment. In particular, by recognizing emotions from the user's voice and facial expressions and incorporating them into the analysis results for feedback, the system provides more realistic and effective interview training. Below, the processing of the program will be specifically described as an embodiment of the present invention.

[0185] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information. The server analyzes the user's past training history and desired interview format, and selects an appropriate virtual interviewer and question set.

[0186] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0187] The recorded voice data is simultaneously sent to the server along with the user's facial expression data. The server then activates a voice analysis system and emotion engine to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data to recognize the user's emotional state. For example, if the user's facial expression is tense, that information is also added to the analysis results.

[0188] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend taking a deep breath and relaxing."

[0189] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0190] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0191] The processing flow will be explained below.

[0192] Step 1:

[0193] The user enters their ID and password on the login screen and clicks the login button.

[0194] The terminal sends the entered login information to the server.

[0195] Step 2:

[0196] The server authenticates the received login information, and if the authentication is successful, obtains the user's profile information.

[0197] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0198] Step 3:

[0199] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0200] The device displays the generated virtual interviewer and prepares for audio output.

[0201] Step 4:

[0202] The device will play the first question asked by the virtual interviewer.

[0203] The user answers the questions verbally.

[0204] The device records the user's response and sends the audio data to the server.

[0205] The terminal also transmits the user's facial expression data (e.g., video captured by a camera) to the server.

[0206] Step 5:

[0207] The server activates the voice analysis system and emotion engine to analyze the received voice data and facial expression data.

[0208] Step 6:

[0209] The server converts the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[0210] The server uses an emotion engine to analyze the facial expression data and recognize the user's emotional state (joy, anger, sadness, surprise, fear, etc.).

[0211] Step 7:

[0212] The server generates comprehensive feedback based on the analysis results and emotion recognition results.

[0213] Example: Generate feedback such as, "Your content is good, but you speak too quickly and should slow down. You also seem nervous, so I would recommend you relax."

[0214] Step 8:

[0215] The server generates feedback and sends it to the terminal.

[0216] Step 9:

[0217] The device displays the feedback to the user and also plays it back as audio.

[0218] The user reviews the feedback and retrains if necessary.

[0219] Example: The user undergoes relaxation training and then repeats the interview simulation to reduce tension.

[0220] The above is the specific processing flow of the system of the present invention. By using this system, users can have a more realistic interview experience and receive detailed feedback, thereby improving their interview skills.

[0221] Example 2

[0222] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0223] Conventional interview training systems have difficulty providing users with an experience similar to that of a real interview environment, and in particular lack the functionality to recognize and provide feedback on emotions from the user's voice and facial expressions, making it difficult to provide effective interview training. This has prevented them from providing sufficient support for improving users' interview performance.

[0224] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0225] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generative AI model, means for transmitting voice data and facial expression data to the server for analysis, and means for generating feedback based on the analysis results and providing it as display and voice. This allows the user to receive training in a similar environment to a real interview, and provides multifaceted feedback based on the analysis results of their voice and facial expressions, thereby improving their interview performance.

[0226] "Authentication means" refers to the means for verifying the ID and password required for a user to log in to a system and confirming that the user is a legitimate user.

[0227] "Received user information" is data relating to the user who has logged into the system, and includes profile information, past training history, and the like.

[0228] A "virtual interviewer" is a virtual interviewer generated within the system for the purpose of conducting interview training, and is represented by an image and voice.

[0229] A "question set" is a collection of questions that the virtual interviewer asks the user and is used to conduct interview training.

[0230] A "generative AI model" is an artificial intelligence algorithm used to generate generated data (such as images or audio), and is a system that generates optimal data based on various parameters.

[0231] A "voice analysis system" is a system that analyzes voice data provided by a user and converts it into text, and evaluates voice quality, volume, speaking speed, etc.

[0232] An "emotion engine" is a system that has the function of analyzing a user's facial expression data and recognizing the user's emotional state (for example, tension or joy).

[0233] "Feedback" refers to evaluations and advice provided to the user based on the results of analysis of voice data and facial expression data, and includes areas for improvement and points to note in training.

[0234] "Means for providing by display and sound" refers to means for conveying feedback generated based on the analysis results to the user visually (such as a display) and audibly (such as a speaker).

[0235] The present invention is a system that allows users to receive training in a similar environment to a real interview, and in particular, realizes more realistic and effective interview training by recognizing emotions from the user's voice and facial expressions, incorporating these into the analysis results, and providing feedback. Specific processing of the system will be described below as an embodiment of the present invention.

[0236] First, the user logs into the system using a terminal. The user enters their ID and password on the login screen and clicks the login button. The terminal sends this entered information to the server. The server authenticates the received login information, and if authentication is successful, obtains the user's profile information. This information includes past training history and preferred interview format.

[0237] Next, the server analyzes the user's profile information and selects an appropriate virtual interviewer and question set. This selection is performed using a generative AI model (e.g., a general natural language generation model). For example, a prompt such as "Please generate a virtual interviewer and question set that are optimal for this interview scenario" is sent to the generative AI model. The image and audio data of the virtual interviewer returned by the generative AI model are then acquired and sent to the device.

[0238] The device displays the generated virtual interviewer on the screen and informs the user that the interview is ready. When the interview begins, the virtual interviewer plays the first question aloud. For example, it might ask, "Please introduce yourself." The user answers this question aloud, and the answer is recorded by the device.

[0239] Along with the recorded voice data, the user's facial expression data is also sent from the device to the server. The server then activates a voice analysis system (e.g., a general voice recognition system) and an emotion engine (e.g., a general emotion analysis system) to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data and recognizes the user's emotional state.

[0240] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the user's emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend you take a deep breath and relax."

[0241] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0242] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0243] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0244] Step 1:

[0245] The user enters their ID and password on the login screen and clicks the login button.

[0246] Input: The ID and password entered by the user.

[0247] Processing: The device sends the entered information to the server via SSL / TLS. The server then checks the database to see if the entered ID and password match.

[0248] Output: If authentication is successful, the user's profile information is retrieved from the server. If authentication fails, an error message is returned.

[0249] Step 2:

[0250] The server acquires the user's profile information and analyzes the user's past training history and desired interview format.

[0251] Input: User profile information and past training history.

[0252] Processing: The server analyzes this data and selects an appropriate virtual interviewer and question set, using a generative AI model.

[0253] Output: Selection of virtual interviewers and question sets.

[0254] Step 3:

[0255] The server uses the generative AI model to generate an image and voice of the virtual interviewer and transmits them to the device.

[0256] Input: The virtual interviewer and the selected question set.

[0257] Processing: The server sends a prompt to the generation AI model to generate images and audio data of a virtual interviewer. Example prompt: "Please generate a virtual interviewer and question set that are optimal for this interview scenario."

[0258] Output: Image and audio data of the generated virtual interviewer.

[0259] Step 4:

[0260] The terminal displays the generated virtual interviewer on the screen and notifies the user that the interview is ready.

[0261] Input: Image and audio data of the virtual interviewer.

[0262] Processing: The terminal displays these data and notifies the user visually and audibly that the interview is ready to begin.

[0263] Output: The user sees the virtual interviewer and is ready to start the interview.

[0264] Step 5:

[0265] The virtual interviewer plays the initial question aloud and asks it to the user, who then answers aloud, and the answers are recorded by the device.

[0266] Input: Audio data of virtual interviewer questions.

[0267] Processing: The device plays the question aloud, the user answers aloud, and the device records the answer in real time.

[0268] Output: User's answer audio data.

[0269] Step 6:

[0270] The device sends the recorded voice data and the user's facial expression data collected by the camera to the server.

[0271] Input: Recorded voice data and user facial expression data.

[0272] Processing: The device sends this data to the server using a secure protocol (e.g. HTTPS).

[0273] Output: Voice and facial expression data sent to the server.

[0274] Step 7:

[0275] The server activates a voice analysis system and an emotion engine to analyze the voice data and facial expression data.

[0276] Input: Voice and facial expression data sent to the server.

[0277] Processing: Converts voice data into text and evaluates its content, voice quality, volume, and speaking rate. At the same time, an emotion engine analyzes facial expression data to recognize the user's emotional state.

[0278] Output: User's voice content assessment and emotional state assessment.

[0279] Step 8:

[0280] The server generates feedback based on the analysis results and sends it to the device.

[0281] Input: User's voice content rating and emotional state rating.

[0282] Processing: The server generates feedback based on these evaluation results, such as specific advice like "The content is appropriate, but you should speak more slowly."

[0283] Output: The generated feedback data.

[0284] Step 9:

[0285] The terminal provides the generated feedback to the user visually and audibly.

[0286] Input: Feedback data.

[0287] Action: The device displays the feedback as text and plays it as audio.

[0288] Output: User can review the feedback and retrain based on it.

[0289] (Application example 2)

[0290] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0291] The present invention aims to address the problem of providing a comfortable environment for passengers in autonomous vehicles. Conventional in-vehicle entertainment systems and passenger service systems have had difficulty recognizing passengers' voices and facial expressions in real time and providing services tailored to their individual needs. In particular, it has been difficult to properly grasp passengers' emotional states during long drives and provide appropriate feedback, such as relaxation and stress reduction.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0293] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual service provider and question set from received user information, a means for generating an image and sound of the virtual service provider using a generation AI, a means for displaying the generated virtual service provider and playing questions by voice, a means for registering a user's answer by voice and analyzing the voice, a means for acquiring video information of the user and recognizing emotions from the video information, and a means for generating feedback based on the analysis results and providing it by display and voice, thereby making it possible to improve the comfort and satisfaction of passengers in self-driving vehicles.

[0294] "Authentication means for users to log in" is a function that verifies the ID and password required when a user accesses the system and confirms legitimate access.

[0295] "Means for selecting an appropriate virtual interviewer and question set from received user information" is a function that automatically selects the most appropriate virtual interviewer and question set based on the user's past history and current objectives.

[0296] "Means for generating images and voices of virtual interviewers using generative AI" refers to a function that uses AI technology to generate the visual and audio of virtual interviewers in real time.

[0297] The "means for displaying the generated virtual interviewer and playing back questions by voice" is a function for displaying an image of the generated virtual interviewer on the user's screen and playing back interview questions by voice.

[0298] The "means for recording the user's response as voice and analyzing the voice" is a function for recording the user's verbal response, analyzing the voice data, and evaluating the content and speaking style.

[0299] The "means for acquiring video information of the user and recognizing emotions from that video information" is a function for acquiring video data such as the user's facial expressions, analyzing it, and recognizing the current emotional state.

[0300] "Means for generating feedback based on the analysis results and providing it by display or voice" refers to a function that creates appropriate advice or feedback from the results of voice analysis and emotion recognition, and provides it to the user by screen display or voice.

[0301] The system of the present invention aims to provide a comfortable environment for passengers in autonomous vehicles. In particular, it aims to reduce stress and improve passenger satisfaction by recognizing passengers' voices and facial expressions in real time and providing feedback based on that.

[0302] 1. System Configuration

[0303] This includes servers, terminals, and various devices such as users' smart glasses and head-mounted displays.

[0304] 2. Hardware and Software

[0305] The system uses the following hardware and software:

[0306] Hardware:

[0307] Camera (video information acquisition)

[0308] Microphone (audio input)

[0309] Smart glasses, head-mounted displays

[0310] software:

[0311] TENSORFLOW (registered trademark) (emotion recognition model)

[0312] OpenCV (video analysis)

[0313] Google Speech API (voice recognition)

[0314] Hugging Face Transformers (generative AI model, QA pipeline)

[0315] 3. System Operation

[0316] 1. User logs in: The user logs in to the system using authentication methods. An ID and password are required to log in, which confirms that the user is a legitimate user.

[0317] 2. Information Selection: The server selects an appropriate virtual service provider and question set based on the received user information. This information is processed based on the user's past history and current goals.

[0318] 3. Use of generative AI: The server uses generative AI models to generate images and sounds of virtual service providers, which are then displayed on devices, smart glasses, or head-mounted displays.

[0319] 4. Initiating an interaction: When a passenger asks a question or makes a request, the device registers this via voice, converts it to text using the Google Speech API, and then uses a generative AI model to generate an appropriate response.

[0320] 5. Emotion Recognition: The video information acquired by the camera is analyzed through OpenCV, and the emotions of passengers are recognized from their facial expressions using TensorFlow's emotion recognition model.

[0321] 6. Providing feedback: Based on the analysis results, appropriate feedback is generated, such as "Relax and enjoy yourself" or "We will arrive in about 30 minutes."

[0322] Examples of concrete examples and prompts

[0323] Example: If a passenger asks, "How long until the car arrives?", you can use a Q&A pipeline to answer, "The journey time to your destination will be about 30 minutes, depending on..."

[0324] Example prompt sentence:

[0325] Question: "How soon will this car arrive?"

[0326] Context: "The estimated travel time to your destination will be approximately 30 minutes, based on current traffic conditions."

[0327] This system will enable increased passenger comfort and satisfaction in autonomous vehicles.

[0328] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0329] Step 1:

[0330] A user logs in.

[0331] The user accesses the login screen and enters their ID and password. The device sends this information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information.

[0332] Input: User ID and password

[0333] Output: User authentication status and profile information

[0334] Step 2:

[0335] The server selects a virtual service provider and a set of questions based on the user information.

[0336] The server analyzes the user's past training history and current goals, and selects the most suitable virtual service provider and question set.

[0337] Input: User profile information and historical data

[0338] Output: Selected virtual service providers and question sets

[0339] Step 3:

[0340] The server uses a generative AI model to generate images and sounds of virtual service providers.

[0341] Based on the selected virtual service provider and question set, a generative AI model generates images and voices of the virtual service provider in real time.

[0342] Input: A hypothetical service provider and a set of questions

[0343] Output: Images and audio of the generated virtual service provider

[0344] Step 4:

[0345] The terminal displays the generated virtual service provider, and the virtual service provider plays the question aloud.

[0346] An image of the virtual service provider is displayed on the device's display, smart glasses, or head-mounted display, and the question is played back aloud.

[0347] Input: Image and voice of the generated virtual service provider

[0348] Output: Display and audio playback

[0349] Step 5:

[0350] The user answers questions by voice, and the device records the answers by voice.

[0351] The user answers questions from the virtual service provider by voice, and the terminal records the voice.

[0352] Input: User's spoken response

[0353] Output: Recorded audio data

[0354] Step 6:

[0355] The server analyzes the recorded audio data.

[0356] The server uses the Google Speech API to convert the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[0357] Input: Recorded audio data

[0358] Output: Text-converted speech data and its evaluation results

[0359] Step 7:

[0360] Emotions are recognized from the user's video data captured by the camera.

[0361] The user's video data acquired by the camera is analyzed using OpenCV, and the emotional state is determined using a TensorFlow model.

[0362] Input: User's video data

[0363] Output: Perceived emotional state

[0364] Step 8:

[0365] The server generates feedback based on the analysis results and sends it to the device.

[0366] The server generates feedback content for the user based on the results of voice and facial expression analysis and sends it to the terminal.

[0367] Input: Results of speech analysis and emotion recognition

[0368] Output: Generated feedback

[0369] Step 9:

[0370] The device displays the feedback to the user and also plays it audibly.

[0371] The feedback content is displayed on the terminal display and also played back as audio to provide to the user.

[0372] Input: Generated feedback

[0373] Output: Display and audio playback

[0374] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0375] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0376] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0377] [Second embodiment]

[0378] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0379] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0380] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0381] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0382] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0383] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0384] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0385] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0386] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0387] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0388] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0389] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0390] The system of the present invention is designed to allow users to receive training in a realistic interview environment. The system analyzes the user's spoken responses and provides detailed feedback to improve the user's interview skills. The following describes the specific processing of the program as an embodiment of the present invention.

[0391] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the most appropriate virtual interviewer and question set based on the user's past training history and the interview format desired by the user.

[0392] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0393] Once the user's response is recorded, the device sends the audio data to a server. The server then activates a voice analysis system to analyze the recorded audio data. The audio data is converted into text, and its content, voice quality, volume, speaking speed, and level of tension are evaluated. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Corporation," the appropriateness of the content, clarity of the voice, and speed of speech are evaluated.

[0394] Based on the analysis results, the server generates feedback, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0395] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0396] The processing flow will be explained below.

[0397] Step 1:

[0398] The user enters their ID and password on the login screen and clicks the login button.

[0399] The terminal sends the entered login information to the server.

[0400] Step 2:

[0401] The server authenticates the received login information, and if the user authentication is successful, obtains the user's profile information.

[0402] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0403] Step 3:

[0404] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0405] The device displays the generated virtual interviewer and prepares for audio output.

[0406] Step 4:

[0407] The device will play the first question asked by the virtual interviewer.

[0408] The user answers the questions verbally.

[0409] The device records the user's response and sends the audio data to the server.

[0410] Step 5:

[0411] The server starts the voice analysis system and analyzes the recorded voice data.

[0412] The server converts the voice data into text and analyzes the content.

[0413] Step 6:

[0414] The server evaluates the content of the user's answers, voice quality, volume, speaking speed, and level of tension based on the analysis results.

[0415] Examples: Evaluate whether the content is appropriate, whether the speaker speaks quickly, and whether the speaker speaks quietly.

[0416] Step 7:

[0417] The server generates feedback based on the analysis results and sends it to the device.

[0418] Example: Feedback such as, "The content is appropriate, but you speak a little too quickly. You should try to speak more slowly."

[0419] Step 8:

[0420] The device displays the feedback to the user and also plays it back as audio.

[0421] The user reviews the feedback and retrains if necessary.

[0422] Example 1

[0423] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0424] In conventional interview training systems, it was difficult for users to receive training in a realistic interview environment. It was also difficult to analyze users' voice responses and provide specific feedback. As a result, it was not possible to effectively improve users' interview skills.

[0425] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0426] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generation AI, means for displaying the generated virtual interviewer and playing questions aloud, means for registering the user's answers aloud and analyzing the voice data, means for converting the voice data into text and evaluating the content, voice quality, volume, speaking speed, and level of tension, and means for generating feedback based on the analysis results and providing it as a display and voice. This allows users to train in an environment similar to a real interview and receive specific feedback, thereby effectively improving their interview skills.

[0427] "Authentication means" refers to a method or device used to verify the identity of a user when accessing a system, such as entering an ID and password.

[0428] A "virtual interviewer" is a digital character or system that mimics a real interviewer when users are undergoing interview training. Images and sounds are generated by generative AI.

[0429] A "question set" is a series of questions that the virtual interviewer asks the user.

[0430] "Generative AI" refers to systems or algorithms that use artificial intelligence techniques to generate the image and voice of a virtual interviewer.

[0431] The "audio playback means" refers to a method or device that allows the virtual interviewer to audibly communicate questions to the user, such as a speaker or an audio player.

[0432] "Voice data" refers to information that is a digital recording of a user's speech.

[0433] "Means of analysis" refers to methods or techniques for converting the content of voice data into text and analyzing voice quality, volume, speaking speed, level of tension, etc., in order to evaluate the data.

[0434] "Convert to text" is the process of converting audio data into written information, for example using voice recognition technology.

[0435] "Feedback" refers to specific advice or comments to the user that are generated based on the analysis results.

[0436] The "display means" refers to a method or device for visually presenting the generated feedback to the user, such as a display or monitor.

[0437] The "means for providing by voice" refers to a method or device for transmitting the generated feedback to the user by voice, such as a speaker or a headset.

[0438] The system of the present invention is designed to allow users to receive training in a realistic interview environment, and the system analyzes users' spoken responses and provides detailed feedback to help improve their interview skills.

[0439] First, the user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server compares the received login information with its database and performs authentication. If authentication is successful, the server obtains the user's profile information.

[0440] Next, the server selects the optimal virtual interviewer and question set based on the user's past training history and the interview format desired by the user. To do this, the server analyzes the user's information and uses generative AI to generate an image and voice of the virtual interviewer. The image and voice data of this generated virtual interviewer are sent to the device. The device displays this virtual interviewer and notifies the user that the interview is ready.

[0441] When the interview begins, the device plays a voice message asking the virtual interviewer an initial question. For example, the question might be, "Please tell us about yourself." The user answers the question aloud, and the answer is recorded by the device.

[0442] Once the user's response is recorded, the device sends the voice data to the server. The server then activates the voice analysis system and analyzes the recorded voice data. The voice data is first converted into text. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Co., Ltd.", this voice will be converted into text format.

[0443] The server uses this text data to evaluate the appropriateness of the content, voice quality, volume, speaking speed, and level of tension. Based on the evaluation, the server generates analysis results and provides appropriate feedback. For example, it may generate specific advice such as, "The content is appropriate, but you speak a little too quickly. Please try to speak a little more slowly."

[0444] The generated feedback is sent to the device and provided to the user visually and audibly, allowing the user to review the feedback and retrain if necessary.

[0445] Below is an example of a prompt sentence to input to the generative AI model.

[0446] "I'm Tanaka from XYZ Co., Ltd. I'm the manager of the sales department.

[0447] Analyze this audio data and rate it on the following:

[0448] 1. Content Appropriateness

[0449] 2. Voice quality

[0450] 3. Volume

[0451] 4. Speaking Speed

[0452] 5. Tension

[0453] To implement the system of the present invention, hardware such as a general computer, microphone, speaker, display, etc. is required, and software such as a generative AI model (e.g., GPT-3) and a voice analysis system (e.g., Google Cloud Speech-to-Text API, IBM Watson, etc.) is used.

[0454] The system of the present invention allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0455] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0456] Step 1:

[0457] The user enters their ID and password on the login screen and clicks the login button. The entered ID and password are sent from the terminal to the server. The server receives the input and compares it with the authentication information stored in the database. If the authentication is successful, the server retrieves the profile information and returns permission to the terminal to access the training system.

[0458] Input: User ID and password

[0459] Output: Authentication success / failure result, profile information

[0460] Specific operation: When the user clicks the login button, the device sends a request to the server via the browser or app, and the server performs a database search.

[0461] Step 2:

[0462] The server analyzes the user's profile information, past training history, and the user's desired interview format, and selects the optimal virtual interviewer and question set. This selection is based on the user's past training data and user requests.

[0463] Input: Profile information, past training data, user requests

[0464] Output: Selected virtual interviewer and question set

[0465] Specific operation: The relevant information is obtained from a database provided by the server, and the optimal settings are selected using an algorithm.

[0466] Step 3:

[0467] The server uses a generative AI to generate images and voices of virtual interviewers, using a pre-trained model, and sends the generated images and voice data to the device.

[0468] Input: A selected virtual interviewer and a set of questions

[0469] Output: Images and audio data of the generated virtual interviewer

[0470] Specific operation: A generative AI model (e.g., GPT-3 or an image generation model) is used and data is delivered to the device.

[0471] Step 4:

[0472] The terminal displays the image and voice of the virtual interviewer and notifies the user that the interview is ready. After the notification, the terminal plays the first question in the voice of the virtual interviewer.

[0473] Input: Image and audio data of a virtual interviewer

[0474] Output: Questions played back and notification that the interview is ready

[0475] Specific operation: The device displays an image of a virtual interviewer on the screen and plays questions through the speaker.

[0476] Step 5:

[0477] The user answers the questions posed by the virtual interviewer by voice, and this voice is recorded by the device.

[0478] Input: User's spoken response

[0479] Output: Recorded audio data

[0480] Specific operation: Records the user's speech through the device's microphone and saves it as an audio file.

[0481] Step 6:

[0482] The device sends the recorded voice data to a server, which then activates a voice analysis system to convert the recorded voice data into text, and analyzes the content, voice quality, volume, speaking speed, and level of tension.

[0483] Input: Recorded audio data

[0484] Output: Analysis results (text data, evaluation items)

[0485] What it does: The server uses a speech analysis service such as the Google Cloud Speech-to-Text API to convert the speech into text and run the evaluation algorithm.

[0486] Step 7:

[0487] The server generates feedback based on the analysis results, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device.

[0488] Input: Analysis results

[0489] Output: Generated feedback

[0490] Specific operation: Feedback is created using a template based on the analysis results and sent to the device.

[0491] Step 8:

[0492] The device provides the generated feedback to the user visually and audibly, allowing the user to review it and retrain if necessary.

[0493] Input: Generated feedback

[0494] Output: Visual and audio feedback

[0495] Specific behavior: Display feedback on the device display and play it audibly through the speaker.

[0496] These are the specific processing steps of this system. At each step, the user, terminal, and server work together to provide the user with the optimal interview training environment.

[0497] (Application example 1)

[0498] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0499] Conventional driver training systems lack advanced technology to simulate real-world driving environments, making it difficult for users to receive training that is in line with actual driving. In addition, because voice instructions and feedback are not provided in real time, it is difficult to make immediate improvements, and effective improvement of driving skills cannot be expected.

[0500] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0501] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual instructor and instruction set from received user information, a means for generating an image and audio of the virtual instructor using a generation AI, a means for displaying the generated virtual instructor and playing questions by audio, a means for recording the user's answers by audio and analyzing the audio, a means for generating feedback based on the analysis results and providing it by display and audio, a means for generating a virtual driving environment and issuing driving instructions to the user, and a means for analyzing the user's driving situation in real time. This allows the user to receive training in a situation similar to a real driving environment, and is expected to improve their skills through real-time feedback.

[0502] "Authentication means for user login" refers to the process of sending authentication information, such as an ID and password used by the user to identify themselves, to the server and performing authentication.

[0503] The "means for selecting an appropriate virtual instructor and set of instructions from the received user information" is an algorithm for selecting the most appropriate virtual instructor and driving instructions based on the user's profile information and past training history.

[0504] "Generative AI" is a system that uses artificial intelligence technology to generate images and sounds of virtual characters.

[0505] A "virtual instructor" is a virtual character generated by the generation AI, whose role is to give driving instructions to the user.

[0506] The "means for playing back questions by voice" is a voice synthesis technology that allows the virtual instructor to communicate instructions and questions to the user by voice.

[0507] The "means for recording a user's voice response and analyzing the voice" is a system for recording a user's voice response and analyzing the voice data.

[0508] "Means for generating feedback based on the analysis results and providing it visually and audibly" refers to a system that generates appropriate feedback based on the speech analysis results and provides it to the user visually and audibly.

[0509] The "means for generating a virtual driving environment and issuing driving instructions to a user" is a system for generating a simulation environment for a user to undergo driving training in a virtual environment and for a virtual instructor to issue driving instructions.

[0510] The "means for analyzing the user's driving status in real time" is a system that monitors the driving behavior of the user in a virtual environment in real time and analyzes the data.

[0511] The system of the present invention allows a user to receive training in a driving environment similar to a real driving environment, thereby effectively improving driving skills. This system analyzes the user's voice responses and driving situation and provides detailed feedback, thereby improving the user's driving skills. Specific processing of the system will be described below as an embodiment of the present invention.

[0512] First, the user logs into the system using their smartphone or head-mounted display (HMD). They enter their ID and password on the login screen and click the login button. The device then sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the optimal virtual instructor and instruction set based on the user's past training history and the user's desired driving environment.

[0513] The server then uses the generative AI model to generate an image and voice of the selected virtual instructor. This generated virtual instructor is displayed on the device, informing the user that they are ready to drive. When the driving simulation begins, the device plays the initial instructions from the virtual instructor by voice. The user performs driving maneuvers according to these instructions, and the driving situation is monitored and recorded in real time by the device.

[0514] Once the user's driving behavior is recorded, the device sends the data to a server. The server then activates a voice analysis system and driving data analysis system to analyze the recorded data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, level of tension, and driving skill are evaluated. For example, if a user says, "I missed the red light," the system evaluates whether the content was appropriate and whether the driving maneuver was safe.

[0515] Based on the analysis results, the server generates feedback, which may include specific advice such as "Pay more attention to traffic lights next time." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0516] These processes are achieved through speech recognition using the speech_recognition library, speech synthesis technology using the pyttsx3 library, and a generative AI model using the ai_model.

[0517] As a specific example, the prompt sentence for "if the user ignores the traffic light" is as follows:

[0518] "Generate appropriate feedback when a user ignores a signal."

[0519] "Generate feedback when the user brakes suddenly while driving."

[0520] In this way, the system of the present invention allows users to receive high-quality driving training in real time, and is expected to improve their driving skills safely and effectively in actual driving.

[0521] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0522] Step 1:

[0523] An authentication method is used for users to log in. The user accesses the login screen via a smartphone or HMD, enters their ID and password, and clicks the login button. The login information is sent from the device to the server, and the server authenticates the received login information. If authentication is successful, the server obtains the user's profile information. The input data is the user's ID and password, and the output data is the user's profile information.

[0524] Step 2:

[0525] The server selects an appropriate virtual instructor and instruction set from the received user information. The server selects the optimal virtual instructor and instruction set based on the user's past training history and desired driving environment. The input data is the user's profile information and past training history, and the output data is the virtual instructor and instruction set.

[0526] Step 3:

[0527] The generative AI model is used to generate the image and voice of the virtual instructor. The server inputs the image and voice of the selected virtual instructor into the generative AI model to generate the image and voice of the virtual instructor. The input data is the virtual instructor's profile, and the output data is the generated image and voice.

[0528] Step 4:

[0529] The generated virtual instructor is displayed and driving instructions are played back by voice.The terminal displays an image of the generated virtual instructor and plays back driving instructions by voice.The input data are the generated image and voice, and the output data are the display and voice playback of the virtual instructor.

[0530] Step 5:

[0531] The user's driving status is recorded by voice, and the voice is analyzed. The user performs driving operations according to the instructions of the virtual instructor. The device monitors and records the driving status in real time. The recorded voice data is sent from the device to a server, which analyzes the data using a voice analysis system. The input data is the user's driving status and voice data, and the output data is the analysis results.

[0532] Step 6:

[0533] Feedback is generated based on the analysis results and provided visually and audibly. The server generates feedback for the user based on the analysis results. The feedback includes specific instructions for improving driving. The generated feedback is sent to the terminal, which then provides it to the user visually and audibly. The input data is the analysis results, and the output data is the feedback.

[0534] These processing steps allow users to receive effective training in a situation that closely resembles a real driving environment, and real-time feedback is expected to improve users' driving skills.

[0535] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0536] The system of the present invention is designed to enable users to receive training in a realistic interview environment. In particular, by recognizing emotions from the user's voice and facial expressions and incorporating them into the analysis results for feedback, the system provides more realistic and effective interview training. Below, the processing of the program will be specifically described as an embodiment of the present invention.

[0537] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information. The server analyzes the user's past training history and desired interview format, and selects an appropriate virtual interviewer and question set.

[0538] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0539] The recorded voice data is simultaneously sent to the server along with the user's facial expression data. The server then activates a voice analysis system and emotion engine to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data to recognize the user's emotional state. For example, if the user's facial expression is tense, that information is also added to the analysis results.

[0540] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend taking a deep breath and relaxing."

[0541] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0542] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0543] The processing flow will be explained below.

[0544] Step 1:

[0545] The user enters their ID and password on the login screen and clicks the login button.

[0546] The terminal sends the entered login information to the server.

[0547] Step 2:

[0548] The server authenticates the received login information, and if the authentication is successful, obtains the user's profile information.

[0549] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0550] Step 3:

[0551] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0552] The device displays the generated virtual interviewer and prepares for audio output.

[0553] Step 4:

[0554] The device will play the first question asked by the virtual interviewer.

[0555] The user answers the questions verbally.

[0556] The device records the user's response and sends the audio data to the server.

[0557] The terminal also transmits the user's facial expression data (e.g., video captured by a camera) to the server.

[0558] Step 5:

[0559] The server activates the voice analysis system and emotion engine to analyze the received voice data and facial expression data.

[0560] Step 6:

[0561] The server converts the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[0562] The server uses an emotion engine to analyze the facial expression data and recognize the user's emotional state (joy, anger, sadness, surprise, fear, etc.).

[0563] Step 7:

[0564] The server generates comprehensive feedback based on the analysis results and emotion recognition results.

[0565] Example: Generate feedback such as, "Your content is good, but you speak too quickly and should slow down. You also seem nervous, so I would recommend you relax."

[0566] Step 8:

[0567] The server generates feedback and sends it to the terminal.

[0568] Step 9:

[0569] The device displays the feedback to the user and also plays it back as audio.

[0570] The user reviews the feedback and retrains if necessary.

[0571] Example: The user undergoes relaxation training and then repeats the interview simulation to reduce tension.

[0572] The above is the specific processing flow of the system of the present invention. By using this system, users can have a more realistic interview experience and receive detailed feedback, thereby improving their interview skills.

[0573] Example 2

[0574] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0575] Conventional interview training systems have difficulty providing users with an experience similar to that of a real interview environment, and in particular lack the functionality to recognize and provide feedback on emotions from the user's voice and facial expressions, making it difficult to provide effective interview training. This has prevented them from providing sufficient support for improving users' interview performance.

[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0577] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generative AI model, means for transmitting voice data and facial expression data to the server for analysis, and means for generating feedback based on the analysis results and providing it as display and voice. This allows the user to receive training in a similar environment to a real interview, and provides multifaceted feedback based on the analysis results of their voice and facial expressions, thereby improving their interview performance.

[0578] "Authentication means" refers to the means for verifying the ID and password required for a user to log in to a system and confirming that the user is a legitimate user.

[0579] "Received user information" is data relating to the user who has logged into the system, and includes profile information, past training history, and the like.

[0580] A "virtual interviewer" is a virtual interviewer generated within the system for the purpose of conducting interview training, and is represented by an image and voice.

[0581] A "question set" is a collection of questions that the virtual interviewer asks the user and is used to conduct interview training.

[0582] A "generative AI model" is an artificial intelligence algorithm used to generate generated data (such as images or audio), and is a system that generates optimal data based on various parameters.

[0583] A "voice analysis system" is a system that analyzes voice data provided by a user and converts it into text, and evaluates voice quality, volume, speaking speed, etc.

[0584] An "emotion engine" is a system that has the function of analyzing a user's facial expression data and recognizing the user's emotional state (for example, tension or joy).

[0585] "Feedback" refers to evaluations and advice provided to the user based on the results of analysis of voice data and facial expression data, and includes areas for improvement and points to note in training.

[0586] "Means for providing by display and sound" refers to means for conveying feedback generated based on the analysis results to the user visually (such as a display) and audibly (such as a speaker).

[0587] The present invention is a system that allows users to receive training in a similar environment to a real interview, and in particular, realizes more realistic and effective interview training by recognizing emotions from the user's voice and facial expressions, incorporating these into the analysis results, and providing feedback. Specific processing of the system will be described below as an embodiment of the present invention.

[0588] First, the user logs into the system using a terminal. The user enters their ID and password on the login screen and clicks the login button. The terminal sends this entered information to the server. The server authenticates the received login information, and if authentication is successful, obtains the user's profile information. This information includes past training history and preferred interview format.

[0589] Next, the server analyzes the user's profile information and selects an appropriate virtual interviewer and question set. This selection is performed using a generative AI model (e.g., a general natural language generation model). For example, a prompt such as "Please generate a virtual interviewer and question set that are optimal for this interview scenario" is sent to the generative AI model. The image and audio data of the virtual interviewer returned by the generative AI model are then acquired and sent to the device.

[0590] The device displays the generated virtual interviewer on the screen and informs the user that the interview is ready. When the interview begins, the virtual interviewer plays the first question aloud. For example, it might ask, "Please introduce yourself." The user answers this question aloud, and the answer is recorded by the device.

[0591] Along with the recorded voice data, the user's facial expression data is also sent from the device to the server. The server then activates a voice analysis system (e.g., a general voice recognition system) and an emotion engine (e.g., a general emotion analysis system) to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data and recognizes the user's emotional state.

[0592] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the user's emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend you take a deep breath and relax."

[0593] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0594] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0595] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0596] Step 1:

[0597] The user enters their ID and password on the login screen and clicks the login button.

[0598] Input: The ID and password entered by the user.

[0599] Processing: The device sends the entered information to the server via SSL / TLS. The server then checks the database to see if the entered ID and password match.

[0600] Output: If authentication is successful, the user's profile information is retrieved from the server. If authentication fails, an error message is returned.

[0601] Step 2:

[0602] The server acquires the user's profile information and analyzes the user's past training history and desired interview format.

[0603] Input: User profile information and past training history.

[0604] Processing: The server analyzes this data and selects an appropriate virtual interviewer and question set, using a generative AI model.

[0605] Output: Selection of virtual interviewers and question sets.

[0606] Step 3:

[0607] The server uses the generative AI model to generate an image and voice of the virtual interviewer and transmits them to the device.

[0608] Input: The virtual interviewer and the selected question set.

[0609] Processing: The server sends a prompt to the generation AI model to generate images and audio data of a virtual interviewer. Example prompt: "Please generate a virtual interviewer and question set that are optimal for this interview scenario."

[0610] Output: Image and audio data of the generated virtual interviewer.

[0611] Step 4:

[0612] The terminal displays the generated virtual interviewer on the screen and notifies the user that the interview is ready.

[0613] Input: Image and audio data of the virtual interviewer.

[0614] Processing: The terminal displays these data and notifies the user visually and audibly that the interview is ready to begin.

[0615] Output: The user sees the virtual interviewer and is ready to start the interview.

[0616] Step 5:

[0617] The virtual interviewer plays the initial question aloud and asks it to the user, who then answers aloud, and the answers are recorded by the device.

[0618] Input: Audio data of virtual interviewer questions.

[0619] Processing: The device plays the question aloud, the user answers aloud, and the device records the answer in real time.

[0620] Output: User's answer audio data.

[0621] Step 6:

[0622] The device sends the recorded voice data and the user's facial expression data collected by the camera to the server.

[0623] Input: Recorded voice data and user facial expression data.

[0624] Processing: The device sends this data to the server using a secure protocol (e.g. HTTPS).

[0625] Output: Voice and facial expression data sent to the server.

[0626] Step 7:

[0627] The server activates a voice analysis system and an emotion engine to analyze the voice data and facial expression data.

[0628] Input: Voice and facial expression data sent to the server.

[0629] Processing: Converts voice data into text and evaluates its content, voice quality, volume, and speaking rate. At the same time, an emotion engine analyzes facial expression data to recognize the user's emotional state.

[0630] Output: User's voice content assessment and emotional state assessment.

[0631] Step 8:

[0632] The server generates feedback based on the analysis results and sends it to the device.

[0633] Input: User's voice content rating and emotional state rating.

[0634] Processing: The server generates feedback based on these evaluation results, such as specific advice like "The content is appropriate, but you should speak more slowly."

[0635] Output: The generated feedback data.

[0636] Step 9:

[0637] The terminal provides the generated feedback to the user visually and audibly.

[0638] Input: Feedback data.

[0639] Action: The device displays the feedback as text and plays it as audio.

[0640] Output: User can review the feedback and retrain based on it.

[0641] (Application example 2)

[0642] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0643] The present invention aims to address the problem of providing a comfortable environment for passengers in autonomous vehicles. Conventional in-vehicle entertainment systems and passenger service systems have had difficulty recognizing passengers' voices and facial expressions in real time and providing services tailored to their individual needs. In particular, it has been difficult to properly grasp passengers' emotional states during long drives and provide appropriate feedback, such as relaxation and stress reduction.

[0644] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0645] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual service provider and question set from received user information, a means for generating an image and sound of the virtual service provider using a generation AI, a means for displaying the generated virtual service provider and playing questions by voice, a means for registering a user's answer by voice and analyzing the voice, a means for acquiring video information of the user and recognizing emotions from the video information, and a means for generating feedback based on the analysis results and providing it by display and voice, thereby making it possible to improve the comfort and satisfaction of passengers in self-driving vehicles.

[0646] "Authentication means for users to log in" is a function that verifies the ID and password required when a user accesses the system and confirms legitimate access.

[0647] "Means for selecting an appropriate virtual interviewer and question set from received user information" is a function that automatically selects the most appropriate virtual interviewer and question set based on the user's past history and current objectives.

[0648] "Means for generating images and voices of virtual interviewers using generative AI" refers to a function that uses AI technology to generate the visual and audio of virtual interviewers in real time.

[0649] The "means for displaying the generated virtual interviewer and playing back questions by voice" is a function for displaying an image of the generated virtual interviewer on the user's screen and playing back interview questions by voice.

[0650] The "means for recording the user's response as voice and analyzing the voice" is a function for recording the user's verbal response, analyzing the voice data, and evaluating the content and speaking style.

[0651] The "means for acquiring video information of the user and recognizing emotions from that video information" is a function for acquiring video data such as the user's facial expressions, analyzing it, and recognizing the current emotional state.

[0652] "Means for generating feedback based on the analysis results and providing it by display or voice" refers to a function that creates appropriate advice or feedback from the results of voice analysis and emotion recognition, and provides it to the user by screen display or voice.

[0653] The system of the present invention aims to provide a comfortable environment for passengers in autonomous vehicles. In particular, it aims to reduce stress and improve passenger satisfaction by recognizing passengers' voices and facial expressions in real time and providing feedback based on that.

[0654] 1. System Configuration

[0655] This includes servers, terminals, and various devices such as users' smart glasses and head-mounted displays.

[0656] 2. Hardware and Software

[0657] The system uses the following hardware and software:

[0658] Hardware:

[0659] Camera (video information acquisition)

[0660] Microphone (audio input)

[0661] Smart glasses, head-mounted displays

[0662] software:

[0663] TensorFlow (emotion recognition model)

[0664] OpenCV (video analysis)

[0665] Google Speech API (voice recognition)

[0666] Hugging Face Transformers (generative AI model, QA pipeline)

[0667] 3. System Operation

[0668] 1. User logs in: The user logs in to the system using authentication methods. An ID and password are required to log in, which confirms that the user is a legitimate user.

[0669] 2. Information Selection: The server selects an appropriate virtual service provider and question set based on the received user information. This information is processed based on the user's past history and current goals.

[0670] 3. Use of generative AI: The server uses generative AI models to generate images and sounds of virtual service providers, which are then displayed on devices, smart glasses, or head-mounted displays.

[0671] 4. Initiating an interaction: When a passenger asks a question or makes a request, the device registers this via voice, converts it to text using the Google Speech API, and then uses a generative AI model to generate an appropriate response.

[0672] 5. Emotion Recognition: The video information acquired by the camera is analyzed through OpenCV, and the emotions of passengers are recognized from their facial expressions using TensorFlow's emotion recognition model.

[0673] 6. Providing feedback: Based on the analysis results, appropriate feedback is generated, such as "Relax and enjoy yourself" or "We will arrive in about 30 minutes."

[0674] Examples of concrete examples and prompts

[0675] Example: If a passenger asks, "How long until the car arrives?", you can use a Q&A pipeline to answer, "The journey time to your destination will be about 30 minutes, depending on..."

[0676] Example prompt sentence:

[0677] Question: "How soon will this car arrive?"

[0678] Context: "The estimated travel time to your destination will be approximately 30 minutes, based on current traffic conditions."

[0679] This system will enable increased passenger comfort and satisfaction in autonomous vehicles.

[0680] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0681] Step 1:

[0682] A user logs in.

[0683] The user accesses the login screen and enters their ID and password. The device sends this information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information.

[0684] Input: User ID and password

[0685] Output: User authentication status and profile information

[0686] Step 2:

[0687] The server selects a virtual service provider and a set of questions based on the user information.

[0688] The server analyzes the user's past training history and current goals, and selects the most suitable virtual service provider and question set.

[0689] Input: User profile information and historical data

[0690] Output: Selected virtual service providers and question sets

[0691] Step 3:

[0692] The server uses a generative AI model to generate images and sounds of virtual service providers.

[0693] Based on the selected virtual service provider and question set, a generative AI model generates images and voices of the virtual service provider in real time.

[0694] Input: A hypothetical service provider and a set of questions

[0695] Output: Images and audio of the generated virtual service provider

[0696] Step 4:

[0697] The terminal displays the generated virtual service provider, and the virtual service provider plays the question aloud.

[0698] An image of the virtual service provider is displayed on the device's display, smart glasses, or head-mounted display, and the question is played back aloud.

[0699] Input: Image and voice of the generated virtual service provider

[0700] Output: Display and audio playback

[0701] Step 5:

[0702] The user answers questions by voice, and the device records the answers by voice.

[0703] The user answers questions from the virtual service provider by voice, and the terminal records the voice.

[0704] Input: User's spoken response

[0705] Output: Recorded audio data

[0706] Step 6:

[0707] The server analyzes the recorded audio data.

[0708] The server uses the Google Speech API to convert the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[0709] Input: Recorded audio data

[0710] Output: Text-converted speech data and its evaluation results

[0711] Step 7:

[0712] Emotions are recognized from the user's video data captured by the camera.

[0713] The user's video data acquired by the camera is analyzed using OpenCV, and the emotional state is determined using a TensorFlow model.

[0714] Input: User's video data

[0715] Output: Perceived emotional state

[0716] Step 8:

[0717] The server generates feedback based on the analysis results and sends it to the device.

[0718] The server generates feedback content for the user based on the results of voice and facial expression analysis and sends it to the terminal.

[0719] Input: Results of speech analysis and emotion recognition

[0720] Output: Generated feedback

[0721] Step 9:

[0722] The device displays the feedback to the user and also plays it audibly.

[0723] The feedback content is displayed on the terminal display and also played back as audio to provide to the user.

[0724] Input: Generated feedback

[0725] Output: Display and audio playback

[0726] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0727] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0728] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0729] [Third embodiment]

[0730] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0731] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0732] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0733] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0734] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0735] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0736] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0737] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0738] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0739] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0740] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0741] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0742] The system of the present invention is designed to allow users to receive training in a realistic interview environment. The system analyzes the user's spoken responses and provides detailed feedback to improve the user's interview skills. The following describes the specific processing of the program as an embodiment of the present invention.

[0743] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the most appropriate virtual interviewer and question set based on the user's past training history and the interview format desired by the user.

[0744] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0745] Once the user's response is recorded, the device sends the audio data to a server. The server then activates a voice analysis system to analyze the recorded audio data. The audio data is converted into text, and its content, voice quality, volume, speaking speed, and level of tension are evaluated. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Corporation," the appropriateness of the content, clarity of the voice, and speed of speech are evaluated.

[0746] Based on the analysis results, the server generates feedback, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0747] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0748] The processing flow will be explained below.

[0749] Step 1:

[0750] The user enters their ID and password on the login screen and clicks the login button.

[0751] The terminal sends the entered login information to the server.

[0752] Step 2:

[0753] The server authenticates the received login information, and if the user authentication is successful, obtains the user's profile information.

[0754] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0755] Step 3:

[0756] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0757] The device displays the generated virtual interviewer and prepares for audio output.

[0758] Step 4:

[0759] The device will play the first question asked by the virtual interviewer.

[0760] The user answers the questions verbally.

[0761] The device records the user's response and sends the audio data to the server.

[0762] Step 5:

[0763] The server starts the voice analysis system and analyzes the recorded voice data.

[0764] The server converts the voice data into text and analyzes the content.

[0765] Step 6:

[0766] The server evaluates the content of the user's answers, voice quality, volume, speaking speed, and level of tension based on the analysis results.

[0767] Examples: Evaluate whether the content is appropriate, whether the speaker speaks quickly, and whether the speaker speaks quietly.

[0768] Step 7:

[0769] The server generates feedback based on the analysis results and sends it to the device.

[0770] Example: Feedback such as, "The content is appropriate, but you speak a little too quickly. You should try to speak more slowly."

[0771] Step 8:

[0772] The device displays the feedback to the user and also plays it back as audio.

[0773] The user reviews the feedback and retrains if necessary.

[0774] Example 1

[0775] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0776] In conventional interview training systems, it was difficult for users to receive training in a realistic interview environment. It was also difficult to analyze users' voice responses and provide specific feedback. As a result, it was not possible to effectively improve users' interview skills.

[0777] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0778] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generation AI, means for displaying the generated virtual interviewer and playing questions aloud, means for registering the user's answers aloud and analyzing the voice data, means for converting the voice data into text and evaluating the content, voice quality, volume, speaking speed, and level of tension, and means for generating feedback based on the analysis results and providing it as a display and voice. This allows users to train in an environment similar to a real interview and receive specific feedback, thereby effectively improving their interview skills.

[0779] "Authentication means" refers to a method or device used to verify the identity of a user when accessing a system, such as entering an ID and password.

[0780] A "virtual interviewer" is a digital character or system that mimics a real interviewer when users are undergoing interview training. Images and sounds are generated by generative AI.

[0781] A "question set" is a series of questions that the virtual interviewer asks the user.

[0782] "Generative AI" refers to systems or algorithms that use artificial intelligence techniques to generate the image and voice of a virtual interviewer.

[0783] The "audio playback means" refers to a method or device that allows the virtual interviewer to audibly communicate questions to the user, such as a speaker or an audio player.

[0784] "Voice data" refers to information that is a digital recording of a user's speech.

[0785] "Means of analysis" refers to methods or techniques for converting the content of voice data into text and analyzing voice quality, volume, speaking speed, level of tension, etc., in order to evaluate the data.

[0786] "Convert to text" is the process of converting audio data into written information, for example using voice recognition technology.

[0787] "Feedback" refers to specific advice or comments to the user that are generated based on the analysis results.

[0788] The "display means" refers to a method or device for visually presenting the generated feedback to the user, such as a display or monitor.

[0789] The "means for providing by voice" refers to a method or device for transmitting the generated feedback to the user by voice, such as a speaker or a headset.

[0790] The system of the present invention is designed to allow users to receive training in a realistic interview environment, and the system analyzes users' spoken responses and provides detailed feedback to help improve their interview skills.

[0791] First, the user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server compares the received login information with its database and performs authentication. If authentication is successful, the server obtains the user's profile information.

[0792] Next, the server selects the optimal virtual interviewer and question set based on the user's past training history and the interview format desired by the user. To do this, the server analyzes the user's information and uses generative AI to generate an image and voice of the virtual interviewer. The image and voice data of this generated virtual interviewer are sent to the device. The device displays this virtual interviewer and notifies the user that the interview is ready.

[0793] When the interview begins, the device plays a voice message asking the virtual interviewer an initial question. For example, the question might be, "Please tell us about yourself." The user answers the question aloud, and the answer is recorded by the device.

[0794] Once the user's response is recorded, the device sends the voice data to the server. The server then activates the voice analysis system and analyzes the recorded voice data. The voice data is first converted into text. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Co., Ltd.", this voice will be converted into text format.

[0795] The server uses this text data to evaluate the appropriateness of the content, voice quality, volume, speaking speed, and level of tension. Based on the evaluation, the server generates analysis results and provides appropriate feedback. For example, it may generate specific advice such as, "The content is appropriate, but you speak a little too quickly. Please try to speak a little more slowly."

[0796] The generated feedback is sent to the device and provided to the user visually and audibly, allowing the user to review the feedback and retrain if necessary.

[0797] Below is an example of a prompt sentence to input to the generative AI model.

[0798] "I'm Tanaka from XYZ Co., Ltd. I'm the manager of the sales department.

[0799] Analyze this audio data and rate it on the following:

[0800] 1. Content Appropriateness

[0801] 2. Voice quality

[0802] 3. Volume

[0803] 4. Speaking Speed

[0804] 5. Tension

[0805] To implement the system of the present invention, hardware such as a general computer, microphone, speaker, display, etc. is required, and software such as a generative AI model (e.g., GPT-3) and a voice analysis system (e.g., Google Cloud Speech-to-Text API, IBM Watson, etc.) is used.

[0806] The system of the present invention allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[0807] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0808] Step 1:

[0809] The user enters their ID and password on the login screen and clicks the login button. The entered ID and password are sent from the terminal to the server. The server receives the input and compares it with the authentication information stored in the database. If the authentication is successful, the server retrieves the profile information and returns permission to the terminal to access the training system.

[0810] Input: User ID and password

[0811] Output: Authentication success / failure result, profile information

[0812] Specific operation: When the user clicks the login button, the device sends a request to the server via the browser or app, and the server performs a database search.

[0813] Step 2:

[0814] The server analyzes the user's profile information, past training history, and the user's desired interview format, and selects the optimal virtual interviewer and question set. This selection is based on the user's past training data and user requests.

[0815] Input: Profile information, past training data, user requests

[0816] Output: Selected virtual interviewer and question set

[0817] Specific operation: The relevant information is obtained from a database provided by the server, and the optimal settings are selected using an algorithm.

[0818] Step 3:

[0819] The server uses a generative AI to generate images and voices of virtual interviewers, using a pre-trained model, and sends the generated images and voice data to the device.

[0820] Input: A selected virtual interviewer and a set of questions

[0821] Output: Images and audio data of the generated virtual interviewer

[0822] Specific operation: A generative AI model (e.g., GPT-3 or an image generation model) is used and data is delivered to the device.

[0823] Step 4:

[0824] The terminal displays the image and voice of the virtual interviewer and notifies the user that the interview is ready. After the notification, the terminal plays the first question in the voice of the virtual interviewer.

[0825] Input: Image and audio data of a virtual interviewer

[0826] Output: Questions played back and notification that the interview is ready

[0827] Specific operation: The device displays an image of a virtual interviewer on the screen and plays questions through the speaker.

[0828] Step 5:

[0829] The user answers the questions posed by the virtual interviewer by voice, and this voice is recorded by the device.

[0830] Input: User's spoken response

[0831] Output: Recorded audio data

[0832] Specific operation: Records the user's speech through the device's microphone and saves it as an audio file.

[0833] Step 6:

[0834] The device sends the recorded voice data to a server, which then activates a voice analysis system to convert the recorded voice data into text, and analyzes the content, voice quality, volume, speaking speed, and level of tension.

[0835] Input: Recorded audio data

[0836] Output: Analysis results (text data, evaluation items)

[0837] What it does: The server uses a speech analysis service such as the Google Cloud Speech-to-Text API to convert the speech into text and run the evaluation algorithm.

[0838] Step 7:

[0839] The server generates feedback based on the analysis results, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device.

[0840] Input: Analysis results

[0841] Output: Generated feedback

[0842] Specific operation: Feedback is created using a template based on the analysis results and sent to the device.

[0843] Step 8:

[0844] The device provides the generated feedback to the user visually and audibly, allowing the user to review it and retrain if necessary.

[0845] Input: Generated feedback

[0846] Output: Visual and audio feedback

[0847] Specific behavior: Display feedback on the device display and play it audibly through the speaker.

[0848] These are the specific processing steps of this system. At each step, the user, terminal, and server work together to provide the user with the optimal interview training environment.

[0849] (Application example 1)

[0850] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0851] Conventional driver training systems lack advanced technology to simulate real-world driving environments, making it difficult for users to receive training that is in line with actual driving. In addition, because voice instructions and feedback are not provided in real time, it is difficult to make immediate improvements, and effective improvement of driving skills cannot be expected.

[0852] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0853] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual instructor and instruction set from received user information, a means for generating an image and audio of the virtual instructor using a generation AI, a means for displaying the generated virtual instructor and playing questions by audio, a means for recording the user's answers by audio and analyzing the audio, a means for generating feedback based on the analysis results and providing it by display and audio, a means for generating a virtual driving environment and issuing driving instructions to the user, and a means for analyzing the user's driving situation in real time. This allows the user to receive training in a situation similar to a real driving environment, and is expected to improve their skills through real-time feedback.

[0854] "Authentication means for user login" refers to the process of sending authentication information, such as an ID and password used by the user to identify themselves, to the server and performing authentication.

[0855] The "means for selecting an appropriate virtual instructor and set of instructions from the received user information" is an algorithm for selecting the most appropriate virtual instructor and driving instructions based on the user's profile information and past training history.

[0856] "Generative AI" is a system that uses artificial intelligence technology to generate images and sounds of virtual characters.

[0857] A "virtual instructor" is a virtual character generated by the generation AI, whose role is to give driving instructions to the user.

[0858] The "means for playing back questions by voice" is a voice synthesis technology that allows the virtual instructor to communicate instructions and questions to the user by voice.

[0859] The "means for recording a user's voice response and analyzing the voice" is a system for recording a user's voice response and analyzing the voice data.

[0860] "Means for generating feedback based on the analysis results and providing it visually and audibly" refers to a system that generates appropriate feedback based on the speech analysis results and provides it to the user visually and audibly.

[0861] The "means for generating a virtual driving environment and issuing driving instructions to a user" is a system for generating a simulation environment for a user to undergo driving training in a virtual environment and for a virtual instructor to issue driving instructions.

[0862] The "means for analyzing the user's driving status in real time" is a system that monitors the driving behavior of the user in a virtual environment in real time and analyzes the data.

[0863] The system of the present invention allows a user to receive training in a driving environment similar to a real driving environment, thereby effectively improving driving skills. This system analyzes the user's voice responses and driving situation and provides detailed feedback, thereby improving the user's driving skills. Specific processing of the system will be described below as an embodiment of the present invention.

[0864] First, the user logs into the system using their smartphone or head-mounted display (HMD). They enter their ID and password on the login screen and click the login button. The device then sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the optimal virtual instructor and instruction set based on the user's past training history and the user's desired driving environment.

[0865] The server then uses the generative AI model to generate an image and voice of the selected virtual instructor. This generated virtual instructor is displayed on the device, informing the user that they are ready to drive. When the driving simulation begins, the device plays the initial instructions from the virtual instructor by voice. The user performs driving maneuvers according to these instructions, and the driving situation is monitored and recorded in real time by the device.

[0866] Once the user's driving behavior is recorded, the device sends the data to a server. The server then activates a voice analysis system and driving data analysis system to analyze the recorded data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, level of tension, and driving skill are evaluated. For example, if a user says, "I missed the red light," the system evaluates whether the content was appropriate and whether the driving maneuver was safe.

[0867] Based on the analysis results, the server generates feedback, which may include specific advice such as "Pay more attention to traffic lights next time." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[0868] These processes are achieved through speech recognition using the speech_recognition library, speech synthesis technology using the pyttsx3 library, and a generative AI model using the ai_model.

[0869] As a specific example, the prompt sentence for "if the user ignores the traffic light" is as follows:

[0870] "Generate appropriate feedback when a user ignores a signal."

[0871] "Generate feedback when the user brakes suddenly while driving."

[0872] In this way, the system of the present invention allows users to receive high-quality driving training in real time, and is expected to improve their driving skills safely and effectively in actual driving.

[0873] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0874] Step 1:

[0875] An authentication method is used for users to log in. The user accesses the login screen via a smartphone or HMD, enters their ID and password, and clicks the login button. The login information is sent from the device to the server, and the server authenticates the received login information. If authentication is successful, the server obtains the user's profile information. The input data is the user's ID and password, and the output data is the user's profile information.

[0876] Step 2:

[0877] The server selects an appropriate virtual instructor and instruction set from the received user information. The server selects the optimal virtual instructor and instruction set based on the user's past training history and desired driving environment. The input data is the user's profile information and past training history, and the output data is the virtual instructor and instruction set.

[0878] Step 3:

[0879] The generative AI model is used to generate the image and voice of the virtual instructor. The server inputs the image and voice of the selected virtual instructor into the generative AI model to generate the image and voice of the virtual instructor. The input data is the virtual instructor's profile, and the output data is the generated image and voice.

[0880] Step 4:

[0881] The generated virtual instructor is displayed and driving instructions are played back by voice.The terminal displays an image of the generated virtual instructor and plays back driving instructions by voice.The input data are the generated image and voice, and the output data are the display and voice playback of the virtual instructor.

[0882] Step 5:

[0883] The user's driving status is recorded by voice, and the voice is analyzed. The user performs driving operations according to the instructions of the virtual instructor. The device monitors and records the driving status in real time. The recorded voice data is sent from the device to a server, which analyzes the data using a voice analysis system. The input data is the user's driving status and voice data, and the output data is the analysis results.

[0884] Step 6:

[0885] Feedback is generated based on the analysis results and provided visually and audibly. The server generates feedback for the user based on the analysis results. The feedback includes specific instructions for improving driving. The generated feedback is sent to the terminal, which then provides it to the user visually and audibly. The input data is the analysis results, and the output data is the feedback.

[0886] These processing steps allow users to receive effective training in a situation that closely resembles a real driving environment, and real-time feedback is expected to improve users' driving skills.

[0887] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0888] The system of the present invention is designed to enable users to receive training in a realistic interview environment. In particular, by recognizing emotions from the user's voice and facial expressions and incorporating them into the analysis results for feedback, the system provides more realistic and effective interview training. Below, the processing of the program will be specifically described as an embodiment of the present invention.

[0889] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information. The server analyzes the user's past training history and desired interview format, and selects an appropriate virtual interviewer and question set.

[0890] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[0891] The recorded voice data is simultaneously sent to the server along with the user's facial expression data. The server then activates a voice analysis system and emotion engine to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data to recognize the user's emotional state. For example, if the user's facial expression is tense, that information is also added to the analysis results.

[0892] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend taking a deep breath and relaxing."

[0893] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0894] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0895] The processing flow will be explained below.

[0896] Step 1:

[0897] The user enters their ID and password on the login screen and clicks the login button.

[0898] The terminal sends the entered login information to the server.

[0899] Step 2:

[0900] The server authenticates the received login information, and if the authentication is successful, obtains the user's profile information.

[0901] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[0902] Step 3:

[0903] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[0904] The device displays the generated virtual interviewer and prepares for audio output.

[0905] Step 4:

[0906] The device will play the first question asked by the virtual interviewer.

[0907] The user answers the questions verbally.

[0908] The device records the user's response and sends the audio data to the server.

[0909] The terminal also transmits the user's facial expression data (e.g., video captured by a camera) to the server.

[0910] Step 5:

[0911] The server activates the voice analysis system and emotion engine to analyze the received voice data and facial expression data.

[0912] Step 6:

[0913] The server converts the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[0914] The server uses an emotion engine to analyze the facial expression data and recognize the user's emotional state (joy, anger, sadness, surprise, fear, etc.).

[0915] Step 7:

[0916] The server generates comprehensive feedback based on the analysis results and emotion recognition results.

[0917] Example: Generate feedback such as, "Your content is good, but you speak too quickly and should slow down. You also seem nervous, so I would recommend you relax."

[0918] Step 8:

[0919] The server generates feedback and sends it to the terminal.

[0920] Step 9:

[0921] The device displays the feedback to the user and also plays it back as audio.

[0922] The user reviews the feedback and retrains if necessary.

[0923] Example: The user undergoes relaxation training and then repeats the interview simulation to reduce tension.

[0924] The above is the specific processing flow of the system of the present invention. By using this system, users can have a more realistic interview experience and receive detailed feedback, thereby improving their interview skills.

[0925] Example 2

[0926] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0927] Conventional interview training systems have difficulty providing users with an experience similar to that of a real interview environment, and in particular lack the functionality to recognize and provide feedback on emotions from the user's voice and facial expressions, making it difficult to provide effective interview training. This has prevented them from providing sufficient support for improving users' interview performance.

[0928] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0929] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generative AI model, means for transmitting voice data and facial expression data to the server for analysis, and means for generating feedback based on the analysis results and providing it as display and voice. This allows the user to receive training in a similar environment to a real interview, and provides multifaceted feedback based on the analysis results of their voice and facial expressions, thereby improving their interview performance.

[0930] "Authentication means" refers to the means for verifying the ID and password required for a user to log in to a system and confirming that the user is a legitimate user.

[0931] "Received user information" is data relating to the user who has logged into the system, and includes profile information, past training history, and the like.

[0932] A "virtual interviewer" is a virtual interviewer generated within the system for the purpose of conducting interview training, and is represented by an image and voice.

[0933] A "question set" is a collection of questions that the virtual interviewer asks the user and is used to conduct interview training.

[0934] A "generative AI model" is an artificial intelligence algorithm used to generate generated data (such as images or audio), and is a system that generates optimal data based on various parameters.

[0935] A "voice analysis system" is a system that analyzes voice data provided by a user and converts it into text, and evaluates voice quality, volume, speaking speed, etc.

[0936] An "emotion engine" is a system that has the function of analyzing a user's facial expression data and recognizing the user's emotional state (for example, tension or joy).

[0937] "Feedback" refers to evaluations and advice provided to the user based on the results of analysis of voice data and facial expression data, and includes areas for improvement and points to note in training.

[0938] "Means for providing by display and sound" refers to means for conveying feedback generated based on the analysis results to the user visually (such as a display) and audibly (such as a speaker).

[0939] The present invention is a system that allows users to receive training in a similar environment to a real interview, and in particular, realizes more realistic and effective interview training by recognizing emotions from the user's voice and facial expressions, incorporating these into the analysis results, and providing feedback. Specific processing of the system will be described below as an embodiment of the present invention.

[0940] First, the user logs into the system using a terminal. The user enters their ID and password on the login screen and clicks the login button. The terminal sends this entered information to the server. The server authenticates the received login information, and if authentication is successful, obtains the user's profile information. This information includes past training history and preferred interview format.

[0941] Next, the server analyzes the user's profile information and selects an appropriate virtual interviewer and question set. This selection is performed using a generative AI model (e.g., a general natural language generation model). For example, a prompt such as "Please generate a virtual interviewer and question set that are optimal for this interview scenario" is sent to the generative AI model. The image and audio data of the virtual interviewer returned by the generative AI model are then acquired and sent to the device.

[0942] The device displays the generated virtual interviewer on the screen and informs the user that the interview is ready. When the interview begins, the virtual interviewer plays the first question aloud. For example, it might ask, "Please introduce yourself." The user answers this question aloud, and the answer is recorded by the device.

[0943] Along with the recorded voice data, the user's facial expression data is also sent from the device to the server. The server then activates a voice analysis system (e.g., a general voice recognition system) and an emotion engine (e.g., a general emotion analysis system) to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data and recognizes the user's emotional state.

[0944] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the user's emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend you take a deep breath and relax."

[0945] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[0946] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[0947] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0948] Step 1:

[0949] The user enters their ID and password on the login screen and clicks the login button.

[0950] Input: The ID and password entered by the user.

[0951] Processing: The device sends the entered information to the server via SSL / TLS. The server then checks the database to see if the entered ID and password match.

[0952] Output: If authentication is successful, the user's profile information is retrieved from the server. If authentication fails, an error message is returned.

[0953] Step 2:

[0954] The server acquires the user's profile information and analyzes the user's past training history and desired interview format.

[0955] Input: User profile information and past training history.

[0956] Processing: The server analyzes this data and selects an appropriate virtual interviewer and question set, using a generative AI model.

[0957] Output: Selection of virtual interviewers and question sets.

[0958] Step 3:

[0959] The server uses the generative AI model to generate an image and voice of the virtual interviewer and transmits them to the device.

[0960] Input: The virtual interviewer and the selected question set.

[0961] Processing: The server sends a prompt to the generation AI model to generate images and audio data of a virtual interviewer. Example prompt: "Please generate a virtual interviewer and question set that are optimal for this interview scenario."

[0962] Output: Image and audio data of the generated virtual interviewer.

[0963] Step 4:

[0964] The terminal displays the generated virtual interviewer on the screen and notifies the user that the interview is ready.

[0965] Input: Image and audio data of the virtual interviewer.

[0966] Processing: The terminal displays these data and notifies the user visually and audibly that the interview is ready to begin.

[0967] Output: The user sees the virtual interviewer and is ready to start the interview.

[0968] Step 5:

[0969] The virtual interviewer plays the initial question aloud and asks it to the user, who then answers aloud, and the answers are recorded by the device.

[0970] Input: Audio data of virtual interviewer questions.

[0971] Processing: The device plays the question aloud, the user answers aloud, and the device records the answer in real time.

[0972] Output: User's answer audio data.

[0973] Step 6:

[0974] The device sends the recorded voice data and the user's facial expression data collected by the camera to the server.

[0975] Input: Recorded voice data and user facial expression data.

[0976] Processing: The device sends this data to the server using a secure protocol (e.g. HTTPS).

[0977] Output: Voice and facial expression data sent to the server.

[0978] Step 7:

[0979] The server activates a voice analysis system and an emotion engine to analyze the voice data and facial expression data.

[0980] Input: Voice and facial expression data sent to the server.

[0981] Processing: Converts voice data into text and evaluates its content, voice quality, volume, and speaking rate. At the same time, an emotion engine analyzes facial expression data to recognize the user's emotional state.

[0982] Output: User's voice content assessment and emotional state assessment.

[0983] Step 8:

[0984] The server generates feedback based on the analysis results and sends it to the device.

[0985] Input: User's voice content rating and emotional state rating.

[0986] Processing: The server generates feedback based on these evaluation results, such as specific advice like "The content is appropriate, but you should speak more slowly."

[0987] Output: The generated feedback data.

[0988] Step 9:

[0989] The terminal provides the generated feedback to the user visually and audibly.

[0990] Input: Feedback data.

[0991] Action: The device displays the feedback as text and plays it as audio.

[0992] Output: User can review the feedback and retrain based on it.

[0993] (Application example 2)

[0994] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0995] The present invention aims to address the problem of providing a comfortable environment for passengers in autonomous vehicles. Conventional in-vehicle entertainment systems and passenger service systems have had difficulty recognizing passengers' voices and facial expressions in real time and providing services tailored to their individual needs. In particular, it has been difficult to properly grasp passengers' emotional states during long drives and provide appropriate feedback, such as relaxation and stress reduction.

[0996] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0997] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual service provider and question set from received user information, a means for generating an image and sound of the virtual service provider using a generation AI, a means for displaying the generated virtual service provider and playing questions by voice, a means for registering a user's answer by voice and analyzing the voice, a means for acquiring video information of the user and recognizing emotions from the video information, and a means for generating feedback based on the analysis results and providing it by display and voice, thereby making it possible to improve the comfort and satisfaction of passengers in self-driving vehicles.

[0998] "Authentication means for users to log in" is a function that verifies the ID and password required when a user accesses the system and confirms legitimate access.

[0999] "Means for selecting an appropriate virtual interviewer and question set from received user information" is a function that automatically selects the most appropriate virtual interviewer and question set based on the user's past history and current objectives.

[1000] "Means for generating images and voices of virtual interviewers using generative AI" refers to a function that uses AI technology to generate the visual and audio of virtual interviewers in real time.

[1001] The "means for displaying the generated virtual interviewer and playing back questions by voice" is a function for displaying an image of the generated virtual interviewer on the user's screen and playing back interview questions by voice.

[1002] The "means for recording the user's response as voice and analyzing the voice" is a function for recording the user's verbal response, analyzing the voice data, and evaluating the content and speaking style.

[1003] The "means for acquiring video information of the user and recognizing emotions from that video information" is a function for acquiring video data such as the user's facial expressions, analyzing it, and recognizing the current emotional state.

[1004] "Means for generating feedback based on the analysis results and providing it by display or voice" refers to a function that creates appropriate advice or feedback from the results of voice analysis and emotion recognition, and provides it to the user by screen display or voice.

[1005] The system of the present invention aims to provide a comfortable environment for passengers in autonomous vehicles. In particular, it aims to reduce stress and improve passenger satisfaction by recognizing passengers' voices and facial expressions in real time and providing feedback based on that.

[1006] 1. System Configuration

[1007] This includes servers, terminals, and various devices such as users' smart glasses and head-mounted displays.

[1008] 2. Hardware and Software

[1009] The system uses the following hardware and software:

[1010] Hardware:

[1011] Camera (video information acquisition)

[1012] Microphone (audio input)

[1013] Smart glasses, head-mounted displays

[1014] software:

[1015] TensorFlow (emotion recognition model)

[1016] OpenCV (video analysis)

[1017] Google Speech API (voice recognition)

[1018] Hugging Face Transformers (generative AI model, QA pipeline)

[1019] 3. System Operation

[1020] 1. User logs in: The user logs in to the system using authentication methods. An ID and password are required to log in, which confirms that the user is a legitimate user.

[1021] 2. Information Selection: The server selects an appropriate virtual service provider and question set based on the received user information. This information is processed based on the user's past history and current goals.

[1022] 3. Use of generative AI: The server uses generative AI models to generate images and sounds of virtual service providers, which are then displayed on devices, smart glasses, or head-mounted displays.

[1023] 4. Initiating an interaction: When a passenger asks a question or makes a request, the device registers this via voice, converts it to text using the Google Speech API, and then uses a generative AI model to generate an appropriate response.

[1024] 5. Emotion Recognition: The video information acquired by the camera is analyzed through OpenCV, and the emotions of passengers are recognized from their facial expressions using TensorFlow's emotion recognition model.

[1025] 6. Providing feedback: Based on the analysis results, appropriate feedback is generated, such as "Relax and enjoy yourself" or "We will arrive in about 30 minutes."

[1026] Examples of concrete examples and prompts

[1027] Example: If a passenger asks, "How long until the car arrives?", you can use a Q&A pipeline to answer, "The journey time to your destination will be about 30 minutes, depending on..."

[1028] Example prompt sentence:

[1029] Question: "How soon will this car arrive?"

[1030] Context: "The estimated travel time to your destination will be approximately 30 minutes, based on current traffic conditions."

[1031] This system will enable increased passenger comfort and satisfaction in autonomous vehicles.

[1032] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1033] Step 1:

[1034] A user logs in.

[1035] The user accesses the login screen and enters their ID and password. The device sends this information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information.

[1036] Input: User ID and password

[1037] Output: User authentication status and profile information

[1038] Step 2:

[1039] The server selects a virtual service provider and a set of questions based on the user information.

[1040] The server analyzes the user's past training history and current goals, and selects the most suitable virtual service provider and question set.

[1041] Input: User profile information and historical data

[1042] Output: Selected virtual service providers and question sets

[1043] Step 3:

[1044] The server uses a generative AI model to generate images and sounds of virtual service providers.

[1045] Based on the selected virtual service provider and question set, a generative AI model generates images and voices of the virtual service provider in real time.

[1046] Input: A hypothetical service provider and a set of questions

[1047] Output: Images and audio of the generated virtual service provider

[1048] Step 4:

[1049] The terminal displays the generated virtual service provider, and the virtual service provider plays the question aloud.

[1050] An image of the virtual service provider is displayed on the device's display, smart glasses, or head-mounted display, and the question is played back aloud.

[1051] Input: Image and voice of the generated virtual service provider

[1052] Output: Display and audio playback

[1053] Step 5:

[1054] The user answers questions by voice, and the device records the answers by voice.

[1055] The user answers questions from the virtual service provider by voice, and the terminal records the voice.

[1056] Input: User's spoken response

[1057] Output: Recorded audio data

[1058] Step 6:

[1059] The server analyzes the recorded audio data.

[1060] The server uses the Google Speech API to convert the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[1061] Input: Recorded audio data

[1062] Output: Text-converted speech data and its evaluation results

[1063] Step 7:

[1064] Emotions are recognized from the user's video data captured by the camera.

[1065] The user's video data acquired by the camera is analyzed using OpenCV, and the emotional state is determined using a TensorFlow model.

[1066] Input: User's video data

[1067] Output: Perceived emotional state

[1068] Step 8:

[1069] The server generates feedback based on the analysis results and sends it to the device.

[1070] The server generates feedback content for the user based on the results of voice and facial expression analysis and sends it to the terminal.

[1071] Input: Results of speech analysis and emotion recognition

[1072] Output: Generated feedback

[1073] Step 9:

[1074] The device displays the feedback to the user and also plays it audibly.

[1075] The feedback content is displayed on the terminal display and also played back as audio to provide to the user.

[1076] Input: Generated feedback

[1077] Output: Display and audio playback

[1078] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1079] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1080] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1081] [Fourth embodiment]

[1082] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1083] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1084] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1085] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1086] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1087] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1088] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1089] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1090] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1091] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1092] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1093] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1094] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1095] The system of the present invention is designed to allow users to receive training in a realistic interview environment. The system analyzes the user's spoken responses and provides detailed feedback to improve the user's interview skills. The following describes the specific processing of the program as an embodiment of the present invention.

[1096] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the most appropriate virtual interviewer and question set based on the user's past training history and the interview format desired by the user.

[1097] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[1098] Once the user's response is recorded, the device sends the audio data to a server. The server then activates a voice analysis system to analyze the recorded audio data. The audio data is converted into text, and its content, voice quality, volume, speaking speed, and level of tension are evaluated. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Corporation," the appropriateness of the content, clarity of the voice, and speed of speech are evaluated.

[1099] Based on the analysis results, the server generates feedback, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[1100] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[1101] The processing flow will be explained below.

[1102] Step 1:

[1103] The user enters their ID and password on the login screen and clicks the login button.

[1104] The terminal sends the entered login information to the server.

[1105] Step 2:

[1106] The server authenticates the received login information, and if the user authentication is successful, obtains the user's profile information.

[1107] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[1108] Step 3:

[1109] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[1110] The device displays the generated virtual interviewer and prepares for audio output.

[1111] Step 4:

[1112] The device will play the first question asked by the virtual interviewer.

[1113] The user answers the questions verbally.

[1114] The device records the user's response and sends the audio data to the server.

[1115] Step 5:

[1116] The server starts the voice analysis system and analyzes the recorded voice data.

[1117] The server converts the voice data into text and analyzes the content.

[1118] Step 6:

[1119] The server evaluates the content of the user's answers, voice quality, volume, speaking speed, and level of tension based on the analysis results.

[1120] Examples: Evaluate whether the content is appropriate, whether the speaker speaks quickly, and whether the speaker speaks quietly.

[1121] Step 7:

[1122] The server generates feedback based on the analysis results and sends it to the device.

[1123] Example: Feedback such as, "The content is appropriate, but you speak a little too quickly. You should try to speak more slowly."

[1124] Step 8:

[1125] The device displays the feedback to the user and also plays it back as audio.

[1126] The user reviews the feedback and retrains if necessary.

[1127] Example 1

[1128] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1129] In conventional interview training systems, it was difficult for users to receive training in a realistic interview environment. It was also difficult to analyze users' voice responses and provide specific feedback. As a result, it was not possible to effectively improve users' interview skills.

[1130] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1131] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generation AI, means for displaying the generated virtual interviewer and playing questions aloud, means for registering the user's answers aloud and analyzing the voice data, means for converting the voice data into text and evaluating the content, voice quality, volume, speaking speed, and level of tension, and means for generating feedback based on the analysis results and providing it as a display and voice. This allows users to train in an environment similar to a real interview and receive specific feedback, thereby effectively improving their interview skills.

[1132] "Authentication means" refers to a method or device used to verify the identity of a user when accessing a system, such as entering an ID and password.

[1133] A "virtual interviewer" is a digital character or system that mimics a real interviewer when users are undergoing interview training. Images and sounds are generated by generative AI.

[1134] A "question set" is a series of questions that the virtual interviewer asks the user.

[1135] "Generative AI" refers to systems or algorithms that use artificial intelligence techniques to generate the image and voice of a virtual interviewer.

[1136] The "audio playback means" refers to a method or device that allows the virtual interviewer to audibly communicate questions to the user, such as a speaker or an audio player.

[1137] "Voice data" refers to information that is a digital recording of a user's speech.

[1138] "Means of analysis" refers to methods or techniques for converting the content of voice data into text and analyzing voice quality, volume, speaking speed, level of tension, etc., in order to evaluate the data.

[1139] "Convert to text" is the process of converting audio data into written information, for example using voice recognition technology.

[1140] "Feedback" refers to specific advice or comments to the user that are generated based on the analysis results.

[1141] The "display means" refers to a method or device for visually presenting the generated feedback to the user, such as a display or monitor.

[1142] The "means for providing by voice" refers to a method or device for transmitting the generated feedback to the user by voice, such as a speaker or a headset.

[1143] The system of the present invention is designed to allow users to receive training in a realistic interview environment, and the system analyzes users' spoken responses and provides detailed feedback to help improve their interview skills.

[1144] First, the user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server compares the received login information with its database and performs authentication. If authentication is successful, the server obtains the user's profile information.

[1145] Next, the server selects the optimal virtual interviewer and question set based on the user's past training history and the interview format desired by the user. To do this, the server analyzes the user's information and uses generative AI to generate an image and voice of the virtual interviewer. The image and voice data of this generated virtual interviewer are sent to the device. The device displays this virtual interviewer and notifies the user that the interview is ready.

[1146] When the interview begins, the device plays a voice message asking the virtual interviewer an initial question. For example, the question might be, "Please tell us about yourself." The user answers the question aloud, and the answer is recorded by the device.

[1147] Once the user's response is recorded, the device sends the voice data to the server. The server then activates the voice analysis system and analyzes the recorded voice data. The voice data is first converted into text. For example, if a user introduces themselves by saying, "Hello, I'm Tanaka from XYZ Co., Ltd.", this voice will be converted into text format.

[1148] The server uses this text data to evaluate the appropriateness of the content, voice quality, volume, speaking speed, and level of tension. Based on the evaluation, the server generates analysis results and provides appropriate feedback. For example, it may generate specific advice such as, "The content is appropriate, but you speak a little too quickly. Please try to speak a little more slowly."

[1149] The generated feedback is sent to the device and provided to the user visually and audibly, allowing the user to review the feedback and retrain if necessary.

[1150] Below is an example of a prompt sentence to input to the generative AI model.

[1151] "I'm Tanaka from XYZ Co., Ltd. I'm the manager of the sales department.

[1152] Analyze this audio data and rate it on the following:

[1153] 1. Content Appropriateness

[1154] 2. Voice quality

[1155] 3. Volume

[1156] 4. Speaking Speed

[1157] 5. Tension

[1158] To implement the system of the present invention, hardware such as a general computer, microphone, speaker, display, etc. is required, and software such as a generative AI model (e.g., GPT-3) and a voice analysis system (e.g., Google Cloud Speech-to-Text API, IBM Watson, etc.) is used.

[1159] The system of the present invention allows users to receive effective interview training, and is expected to improve their performance in actual interviews.

[1160] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1161] Step 1:

[1162] The user enters their ID and password on the login screen and clicks the login button. The entered ID and password are sent from the terminal to the server. The server receives the input and compares it with the authentication information stored in the database. If the authentication is successful, the server retrieves the profile information and returns permission to the terminal to access the training system.

[1163] Input: User ID and password

[1164] Output: Authentication success / failure result, profile information

[1165] Specific operation: When the user clicks the login button, the device sends a request to the server via the browser or app, and the server performs a database search.

[1166] Step 2:

[1167] The server analyzes the user's profile information, past training history, and the user's desired interview format, and selects the optimal virtual interviewer and question set. This selection is based on the user's past training data and user requests.

[1168] Input: Profile information, past training data, user requests

[1169] Output: Selected virtual interviewer and question set

[1170] Specific operation: The relevant information is obtained from a database provided by the server, and the optimal settings are selected using an algorithm.

[1171] Step 3:

[1172] The server uses a generative AI to generate images and voices of virtual interviewers, using a pre-trained model, and sends the generated images and voice data to the device.

[1173] Input: A selected virtual interviewer and a set of questions

[1174] Output: Images and audio data of the generated virtual interviewer

[1175] Specific operation: A generative AI model (e.g., GPT-3 or an image generation model) is used and data is delivered to the device.

[1176] Step 4:

[1177] The terminal displays the image and voice of the virtual interviewer and notifies the user that the interview is ready. After the notification, the terminal plays the first question in the voice of the virtual interviewer.

[1178] Input: Image and audio data of a virtual interviewer

[1179] Output: Questions played back and notification that the interview is ready

[1180] Specific operation: The device displays an image of a virtual interviewer on the screen and plays questions through the speaker.

[1181] Step 5:

[1182] The user answers the questions posed by the virtual interviewer by voice, and this voice is recorded by the device.

[1183] Input: User's spoken response

[1184] Output: Recorded audio data

[1185] Specific operation: Records the user's speech through the device's microphone and saves it as an audio file.

[1186] Step 6:

[1187] The device sends the recorded voice data to a server, which then activates a voice analysis system to convert the recorded voice data into text, and analyzes the content, voice quality, volume, speaking speed, and level of tension.

[1188] Input: Recorded audio data

[1189] Output: Analysis results (text data, evaluation items)

[1190] What it does: The server uses a speech analysis service such as the Google Cloud Speech-to-Text API to convert the speech into text and run the evaluation algorithm.

[1191] Step 7:

[1192] The server generates feedback based on the analysis results, including specific advice such as, "Your content is appropriate, but you speak a little too quickly. Please try to speak more slowly." The generated feedback is sent to the device.

[1193] Input: Analysis results

[1194] Output: Generated feedback

[1195] Specific operation: Feedback is created using a template based on the analysis results and sent to the device.

[1196] Step 8:

[1197] The device provides the generated feedback to the user visually and audibly, allowing the user to review it and retrain if necessary.

[1198] Input: Generated feedback

[1199] Output: Visual and audio feedback

[1200] Specific behavior: Display feedback on the device display and play it audibly through the speaker.

[1201] These are the specific processing steps of this system. At each step, the user, terminal, and server work together to provide the user with the optimal interview training environment.

[1202] (Application example 1)

[1203] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1204] Conventional driver training systems lack advanced technology to simulate real-world driving environments, making it difficult for users to receive training that is in line with actual driving. In addition, because voice instructions and feedback are not provided in real time, it is difficult to make immediate improvements, and effective improvement of driving skills cannot be expected.

[1205] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1206] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual instructor and instruction set from received user information, a means for generating an image and audio of the virtual instructor using a generation AI, a means for displaying the generated virtual instructor and playing questions by audio, a means for recording the user's answers by audio and analyzing the audio, a means for generating feedback based on the analysis results and providing it by display and audio, a means for generating a virtual driving environment and issuing driving instructions to the user, and a means for analyzing the user's driving situation in real time. This allows the user to receive training in a situation similar to a real driving environment, and is expected to improve their skills through real-time feedback.

[1207] "Authentication means for user login" refers to the process of sending authentication information, such as an ID and password used by the user to identify themselves, to the server and performing authentication.

[1208] The "means for selecting an appropriate virtual instructor and set of instructions from the received user information" is an algorithm for selecting the most appropriate virtual instructor and driving instructions based on the user's profile information and past training history.

[1209] "Generative AI" is a system that uses artificial intelligence technology to generate images and sounds of virtual characters.

[1210] A "virtual instructor" is a virtual character generated by the generation AI, whose role is to give driving instructions to the user.

[1211] The "means for playing back questions by voice" is a voice synthesis technology that allows the virtual instructor to communicate instructions and questions to the user by voice.

[1212] The "means for recording a user's voice response and analyzing the voice" is a system for recording a user's voice response and analyzing the voice data.

[1213] "Means for generating feedback based on the analysis results and providing it visually and audibly" refers to a system that generates appropriate feedback based on the speech analysis results and provides it to the user visually and audibly.

[1214] The "means for generating a virtual driving environment and issuing driving instructions to a user" is a system for generating a simulation environment for a user to undergo driving training in a virtual environment and for a virtual instructor to issue driving instructions.

[1215] The "means for analyzing the user's driving status in real time" is a system that monitors the driving behavior of the user in a virtual environment in real time and analyzes the data.

[1216] The system of the present invention allows a user to receive training in a driving environment similar to a real driving environment, thereby effectively improving driving skills. This system analyzes the user's voice responses and driving situation and provides detailed feedback, thereby improving the user's driving skills. Specific processing of the system will be described below as an embodiment of the present invention.

[1217] First, the user logs into the system using their smartphone or head-mounted display (HMD). They enter their ID and password on the login screen and click the login button. The device then sends the entered information to the server. The server authenticates the received login information and, if successful, obtains the user's profile information. The server then selects the optimal virtual instructor and instruction set based on the user's past training history and the user's desired driving environment.

[1218] The server then uses the generative AI model to generate an image and voice of the selected virtual instructor. This generated virtual instructor is displayed on the device, informing the user that they are ready to drive. When the driving simulation begins, the device plays the initial instructions from the virtual instructor by voice. The user performs driving maneuvers according to these instructions, and the driving situation is monitored and recorded in real time by the device.

[1219] Once the user's driving behavior is recorded, the device sends the data to a server. The server then activates a voice analysis system and driving data analysis system to analyze the recorded data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, level of tension, and driving skill are evaluated. For example, if a user says, "I missed the red light," the system evaluates whether the content was appropriate and whether the driving maneuver was safe.

[1220] Based on the analysis results, the server generates feedback, which may include specific advice such as "Pay more attention to traffic lights next time." The generated feedback is sent to the device and provided to the user visually and audibly. The user can review this feedback and retrain if necessary.

[1221] These processes are achieved through speech recognition using the speech_recognition library, speech synthesis technology using the pyttsx3 library, and a generative AI model using the ai_model.

[1222] As a specific example, the prompt sentence for "if the user ignores the traffic light" is as follows:

[1223] "Generate appropriate feedback when a user ignores a signal."

[1224] "Generate feedback when the user brakes suddenly while driving."

[1225] In this way, the system of the present invention allows users to receive high-quality driving training in real time, and is expected to improve their driving skills safely and effectively in actual driving.

[1226] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1227] Step 1:

[1228] An authentication method is used for users to log in. The user accesses the login screen via a smartphone or HMD, enters their ID and password, and clicks the login button. The login information is sent from the device to the server, and the server authenticates the received login information. If authentication is successful, the server obtains the user's profile information. The input data is the user's ID and password, and the output data is the user's profile information.

[1229] Step 2:

[1230] The server selects an appropriate virtual instructor and instruction set from the received user information. The server selects the optimal virtual instructor and instruction set based on the user's past training history and desired driving environment. The input data is the user's profile information and past training history, and the output data is the virtual instructor and instruction set.

[1231] Step 3:

[1232] The generative AI model is used to generate the image and voice of the virtual instructor. The server inputs the image and voice of the selected virtual instructor into the generative AI model to generate the image and voice of the virtual instructor. The input data is the virtual instructor's profile, and the output data is the generated image and voice.

[1233] Step 4:

[1234] The generated virtual instructor is displayed and driving instructions are played back by voice.The terminal displays an image of the generated virtual instructor and plays back driving instructions by voice.The input data are the generated image and voice, and the output data are the display and voice playback of the virtual instructor.

[1235] Step 5:

[1236] The user's driving status is recorded by voice, and the voice is analyzed. The user performs driving operations according to the instructions of the virtual instructor. The device monitors and records the driving status in real time. The recorded voice data is sent from the device to a server, which analyzes the data using a voice analysis system. The input data is the user's driving status and voice data, and the output data is the analysis results.

[1237] Step 6:

[1238] Feedback is generated based on the analysis results and provided visually and audibly. The server generates feedback for the user based on the analysis results. The feedback includes specific instructions for improving driving. The generated feedback is sent to the terminal, which then provides it to the user visually and audibly. The input data is the analysis results, and the output data is the feedback.

[1239] These processing steps allow users to receive effective training in a situation that closely resembles a real driving environment, and real-time feedback is expected to improve users' driving skills.

[1240] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1241] The system of the present invention is designed to enable users to receive training in a realistic interview environment. In particular, by recognizing emotions from the user's voice and facial expressions and incorporating them into the analysis results for feedback, the system provides more realistic and effective interview training. Below, the processing of the program will be specifically described as an embodiment of the present invention.

[1242] First, the user logs in to the system. The user enters their ID and password on the login screen and clicks the login button. The device sends this entered information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information. The server analyzes the user's past training history and desired interview format, and selects an appropriate virtual interviewer and question set.

[1243] Next, the server uses a generation AI to generate an image and voice of the selected virtual interviewer. This generated virtual interviewer is displayed on the device, informing the user that the interview is ready. When the interview begins, the device plays the first question posed by the virtual interviewer aloud. The user answers this question aloud, and the answer is recorded by the device.

[1244] The recorded voice data is simultaneously sent to the server along with the user's facial expression data. The server then activates a voice analysis system and emotion engine to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data to recognize the user's emotional state. For example, if the user's facial expression is tense, that information is also added to the analysis results.

[1245] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend taking a deep breath and relaxing."

[1246] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[1247] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[1248] The processing flow will be explained below.

[1249] Step 1:

[1250] The user enters their ID and password on the login screen and clicks the login button.

[1251] The terminal sends the entered login information to the server.

[1252] Step 2:

[1253] The server authenticates the received login information, and if the authentication is successful, obtains the user's profile information.

[1254] The server analyzes the user's past training history and desired interview format and selects an appropriate virtual interviewer and question set.

[1255] Step 3:

[1256] The server uses generation AI to generate images and voices of the selected virtual interviewer.

[1257] The device displays the generated virtual interviewer and prepares for audio output.

[1258] Step 4:

[1259] The device will play the first question asked by the virtual interviewer.

[1260] The user answers the questions verbally.

[1261] The device records the user's response and sends the audio data to the server.

[1262] The terminal also transmits the user's facial expression data (e.g., video captured by a camera) to the server.

[1263] Step 5:

[1264] The server activates the voice analysis system and emotion engine to analyze the received voice data and facial expression data.

[1265] Step 6:

[1266] The server converts the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[1267] The server uses an emotion engine to analyze the facial expression data and recognize the user's emotional state (joy, anger, sadness, surprise, fear, etc.).

[1268] Step 7:

[1269] The server generates comprehensive feedback based on the analysis results and emotion recognition results.

[1270] Example: Generate feedback such as, "Your content is good, but you speak too quickly and should slow down. You also seem nervous, so I would recommend you relax."

[1271] Step 8:

[1272] The server generates feedback and sends it to the terminal.

[1273] Step 9:

[1274] The device displays the feedback to the user and also plays it back as audio.

[1275] The user reviews the feedback and retrains if necessary.

[1276] Example: The user undergoes relaxation training and then repeats the interview simulation to reduce tension.

[1277] The above is the specific processing flow of the system of the present invention. By using this system, users can have a more realistic interview experience and receive detailed feedback, thereby improving their interview skills.

[1278] Example 2

[1279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1280] Conventional interview training systems have difficulty providing users with an experience similar to that of a real interview environment, and in particular lack the functionality to recognize and provide feedback on emotions from the user's voice and facial expressions, making it difficult to provide effective interview training. This has prevented them from providing sufficient support for improving users' interview performance.

[1281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1282] In this invention, the server includes authentication means for allowing a user to log in, means for selecting an appropriate virtual interviewer and question set from received user information, means for generating an image and voice of the virtual interviewer using a generative AI model, means for transmitting voice data and facial expression data to the server for analysis, and means for generating feedback based on the analysis results and providing it as display and voice. This allows the user to receive training in a similar environment to a real interview, and provides multifaceted feedback based on the analysis results of their voice and facial expressions, thereby improving their interview performance.

[1283] "Authentication means" refers to the means for verifying the ID and password required for a user to log in to a system and confirming that the user is a legitimate user.

[1284] "Received user information" is data relating to the user who has logged into the system, and includes profile information, past training history, and the like.

[1285] A "virtual interviewer" is a virtual interviewer generated within the system for the purpose of conducting interview training, and is represented by an image and voice.

[1286] A "question set" is a collection of questions that the virtual interviewer asks the user and is used to conduct interview training.

[1287] A "generative AI model" is an artificial intelligence algorithm used to generate generated data (such as images or audio), and is a system that generates optimal data based on various parameters.

[1288] A "voice analysis system" is a system that analyzes voice data provided by a user and converts it into text, and evaluates voice quality, volume, speaking speed, etc.

[1289] An "emotion engine" is a system that has the function of analyzing a user's facial expression data and recognizing the user's emotional state (for example, tension or joy).

[1290] "Feedback" refers to evaluations and advice provided to the user based on the results of analysis of voice data and facial expression data, and includes areas for improvement and points to note in training.

[1291] "Means for providing by display and sound" refers to means for conveying feedback generated based on the analysis results to the user visually (such as a display) and audibly (such as a speaker).

[1292] The present invention is a system that allows users to receive training in a similar environment to a real interview, and in particular, realizes more realistic and effective interview training by recognizing emotions from the user's voice and facial expressions, incorporating these into the analysis results, and providing feedback. Specific processing of the system will be described below as an embodiment of the present invention.

[1293] First, the user logs into the system using a terminal. The user enters their ID and password on the login screen and clicks the login button. The terminal sends this entered information to the server. The server authenticates the received login information, and if authentication is successful, obtains the user's profile information. This information includes past training history and preferred interview format.

[1294] Next, the server analyzes the user's profile information and selects an appropriate virtual interviewer and question set. This selection is performed using a generative AI model (e.g., a general natural language generation model). For example, a prompt such as "Please generate a virtual interviewer and question set that are optimal for this interview scenario" is sent to the generative AI model. The image and audio data of the virtual interviewer returned by the generative AI model are then acquired and sent to the device.

[1295] The device displays the generated virtual interviewer on the screen and informs the user that the interview is ready. When the interview begins, the virtual interviewer plays the first question aloud. For example, it might ask, "Please introduce yourself." The user answers this question aloud, and the answer is recorded by the device.

[1296] Along with the recorded voice data, the user's facial expression data is also sent from the device to the server. The server then activates a voice analysis system (e.g., a general voice recognition system) and an emotion engine (e.g., a general emotion analysis system) to analyze the recorded voice data and facial expression data. The voice data is converted into text, and its content, voice quality, volume, speaking speed, etc. are evaluated. At the same time, the emotion engine analyzes the facial expression data and recognizes the user's emotional state.

[1297] Based on the analysis results, the server generates feedback. This feedback includes not only an evaluation of the audio data and content, but also advice on the user's emotional state. For example, the server might generate feedback such as, "The content is appropriate, but you speak a little too quickly, so you should try to speak more slowly. Also, you seem nervous, so I recommend you take a deep breath and relax."

[1298] The generated feedback is sent to the terminal and provided to the user visually and audibly. The user can check this feedback and retrain if necessary. For example, the user can conduct relaxation training based on the feedback and simulate an interview again to improve their nervous state.

[1299] The above is the program processing of the system of the present invention and its specific embodiment. This system allows users to receive more realistic and effective interview training, and is expected to improve their performance in actual interviews.

[1300] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1301] Step 1:

[1302] The user enters their ID and password on the login screen and clicks the login button.

[1303] Input: The ID and password entered by the user.

[1304] Processing: The device sends the entered information to the server via SSL / TLS. The server then checks the database to see if the entered ID and password match.

[1305] Output: If authentication is successful, the user's profile information is retrieved from the server. If authentication fails, an error message is returned.

[1306] Step 2:

[1307] The server acquires the user's profile information and analyzes the user's past training history and desired interview format.

[1308] Input: User profile information and past training history.

[1309] Processing: The server analyzes this data and selects an appropriate virtual interviewer and question set, using a generative AI model.

[1310] Output: Selection of virtual interviewers and question sets.

[1311] Step 3:

[1312] The server uses the generative AI model to generate an image and voice of the virtual interviewer and transmits them to the device.

[1313] Input: The virtual interviewer and the selected question set.

[1314] Processing: The server sends a prompt to the generation AI model to generate images and audio data of a virtual interviewer. Example prompt: "Please generate a virtual interviewer and question set that are optimal for this interview scenario."

[1315] Output: Image and audio data of the generated virtual interviewer.

[1316] Step 4:

[1317] The terminal displays the generated virtual interviewer on the screen and notifies the user that the interview is ready.

[1318] Input: Image and audio data of the virtual interviewer.

[1319] Processing: The terminal displays these data and notifies the user visually and audibly that the interview is ready to begin.

[1320] Output: The user sees the virtual interviewer and is ready to start the interview.

[1321] Step 5:

[1322] The virtual interviewer plays the initial question aloud and asks it to the user, who then answers aloud, and the answers are recorded by the device.

[1323] Input: Audio data of virtual interviewer questions.

[1324] Processing: The device plays the question aloud, the user answers aloud, and the device records the answer in real time.

[1325] Output: User's answer audio data.

[1326] Step 6:

[1327] The device sends the recorded voice data and the user's facial expression data collected by the camera to the server.

[1328] Input: Recorded voice data and user facial expression data.

[1329] Processing: The device sends this data to the server using a secure protocol (e.g. HTTPS).

[1330] Output: Voice and facial expression data sent to the server.

[1331] Step 7:

[1332] The server activates a voice analysis system and an emotion engine to analyze the voice data and facial expression data.

[1333] Input: Voice and facial expression data sent to the server.

[1334] Processing: Converts voice data into text and evaluates its content, voice quality, volume, and speaking rate. At the same time, an emotion engine analyzes facial expression data to recognize the user's emotional state.

[1335] Output: User's voice content assessment and emotional state assessment.

[1336] Step 8:

[1337] The server generates feedback based on the analysis results and sends it to the device.

[1338] Input: User's voice content rating and emotional state rating.

[1339] Processing: The server generates feedback based on these evaluation results, such as specific advice like "The content is appropriate, but you should speak more slowly."

[1340] Output: The generated feedback data.

[1341] Step 9:

[1342] The terminal provides the generated feedback to the user visually and audibly.

[1343] Input: Feedback data.

[1344] Action: The device displays the feedback as text and plays it as audio.

[1345] Output: User can review the feedback and retrain based on it.

[1346] (Application example 2)

[1347] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1348] The present invention aims to address the problem of providing a comfortable environment for passengers in autonomous vehicles. Conventional in-vehicle entertainment systems and passenger service systems have had difficulty recognizing passengers' voices and facial expressions in real time and providing services tailored to their individual needs. In particular, it has been difficult to properly grasp passengers' emotional states during long drives and provide appropriate feedback, such as relaxation and stress reduction.

[1349] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1350] In this invention, the server includes an authentication means for a user to log in, a means for selecting an appropriate virtual service provider and question set from received user information, a means for generating an image and sound of the virtual service provider using a generation AI, a means for displaying the generated virtual service provider and playing questions by voice, a means for registering a user's answer by voice and analyzing the voice, a means for acquiring video information of the user and recognizing emotions from the video information, and a means for generating feedback based on the analysis results and providing it by display and voice, thereby making it possible to improve the comfort and satisfaction of passengers in self-driving vehicles.

[1351] "Authentication means for users to log in" is a function that verifies the ID and password required when a user accesses the system and confirms legitimate access.

[1352] "Means for selecting an appropriate virtual interviewer and question set from received user information" is a function that automatically selects the most appropriate virtual interviewer and question set based on the user's past history and current objectives.

[1353] "Means for generating images and voices of virtual interviewers using generative AI" refers to a function that uses AI technology to generate the visual and audio of virtual interviewers in real time.

[1354] The "means for displaying the generated virtual interviewer and playing back questions by voice" is a function for displaying an image of the generated virtual interviewer on the user's screen and playing back interview questions by voice.

[1355] The "means for recording the user's response as voice and analyzing the voice" is a function for recording the user's verbal response, analyzing the voice data, and evaluating the content and speaking style.

[1356] The "means for acquiring video information of the user and recognizing emotions from that video information" is a function for acquiring video data such as the user's facial expressions, analyzing it, and recognizing the current emotional state.

[1357] "Means for generating feedback based on the analysis results and providing it by display or voice" refers to a function that creates appropriate advice or feedback from the results of voice analysis and emotion recognition, and provides it to the user by screen display or voice.

[1358] The system of the present invention aims to provide a comfortable environment for passengers in autonomous vehicles. In particular, it aims to reduce stress and improve passenger satisfaction by recognizing passengers' voices and facial expressions in real time and providing feedback based on that.

[1359] 1. System Configuration

[1360] This includes servers, terminals, and various devices such as users' smart glasses and head-mounted displays.

[1361] 2. Hardware and Software

[1362] The system uses the following hardware and software:

[1363] Hardware:

[1364] Camera (video information acquisition)

[1365] Microphone (audio input)

[1366] Smart glasses, head-mounted displays

[1367] software:

[1368] TensorFlow (emotion recognition model)

[1369] OpenCV (video analysis)

[1370] Google Speech API (voice recognition)

[1371] Hugging Face Transformers (generative AI model, QA pipeline)

[1372] 3. System Operation

[1373] 1. User logs in: The user logs in to the system using authentication methods. An ID and password are required to log in, which confirms that the user is a legitimate user.

[1374] 2. Information Selection: The server selects an appropriate virtual service provider and question set based on the received user information. This information is processed based on the user's past history and current goals.

[1375] 3. Use of generative AI: The server uses generative AI models to generate images and sounds of virtual service providers, which are then displayed on devices, smart glasses, or head-mounted displays.

[1376] 4. Initiating an interaction: When a passenger asks a question or makes a request, the device registers this via voice, converts it to text using the Google Speech API, and then uses a generative AI model to generate an appropriate response.

[1377] 5. Emotion Recognition: The video information acquired by the camera is analyzed through OpenCV, and the emotions of passengers are recognized from their facial expressions using TensorFlow's emotion recognition model.

[1378] 6. Providing feedback: Based on the analysis results, appropriate feedback is generated, such as "Relax and enjoy yourself" or "We will arrive in about 30 minutes."

[1379] Examples of concrete examples and prompts

[1380] Example: If a passenger asks, "How long until the car arrives?", you can use a Q&A pipeline to answer, "The journey time to your destination will be about 30 minutes, depending on..."

[1381] Example prompt sentence:

[1382] Question: "How soon will this car arrive?"

[1383] Context: "The estimated travel time to your destination will be approximately 30 minutes, based on current traffic conditions."

[1384] This system will enable increased passenger comfort and satisfaction in autonomous vehicles.

[1385] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1386] Step 1:

[1387] A user logs in.

[1388] The user accesses the login screen and enters their ID and password. The device sends this information to the server. The server authenticates the received login information and, if authentication is successful, obtains the user's profile information.

[1389] Input: User ID and password

[1390] Output: User authentication status and profile information

[1391] Step 2:

[1392] The server selects a virtual service provider and a set of questions based on the user information.

[1393] The server analyzes the user's past training history and current goals, and selects the most suitable virtual service provider and question set.

[1394] Input: User profile information and historical data

[1395] Output: Selected virtual service providers and question sets

[1396] Step 3:

[1397] The server uses a generative AI model to generate images and sounds of virtual service providers.

[1398] Based on the selected virtual service provider and question set, a generative AI model generates images and voices of the virtual service provider in real time.

[1399] Input: A hypothetical service provider and a set of questions

[1400] Output: Images and audio of the generated virtual service provider

[1401] Step 4:

[1402] The terminal displays the generated virtual service provider, and the virtual service provider plays the question aloud.

[1403] An image of the virtual service provider is displayed on the device's display, smart glasses, or head-mounted display, and the question is played back aloud.

[1404] Input: Image and voice of the generated virtual service provider

[1405] Output: Display and audio playback

[1406] Step 5:

[1407] The user answers questions by voice, and the device records the answers by voice.

[1408] The user answers questions from the virtual service provider by voice, and the terminal records the voice.

[1409] Input: User's spoken response

[1410] Output: Recorded audio data

[1411] Step 6:

[1412] The server analyzes the recorded audio data.

[1413] The server uses the Google Speech API to convert the audio data into text and evaluates the content, voice quality, volume, and speaking speed.

[1414] Input: Recorded audio data

[1415] Output: Text-converted speech data and its evaluation results

[1416] Step 7:

[1417] Emotions are recognized from the user's video data captured by the camera.

[1418] The user's video data acquired by the camera is analyzed using OpenCV, and the emotional state is determined using a TensorFlow model.

[1419] Input: User's video data

[1420] Output: Perceived emotional state

[1421] Step 8:

[1422] The server generates feedback based on the analysis results and sends it to the device.

[1423] The server generates feedback content for the user based on the results of voice and facial expression analysis and sends it to the terminal.

[1424] Input: Results of speech analysis and emotion recognition

[1425] Output: Generated feedback

[1426] Step 9:

[1427] The device displays the feedback to the user and also plays it audibly.

[1428] The feedback content is displayed on the terminal display and also played back as audio to provide to the user.

[1429] Input: Generated feedback

[1430] Output: Display and audio playback

[1431] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1432] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1433] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1434] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1435] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1436] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1437] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1438] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1439] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1440] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1441] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1442] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1443] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1444] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1445] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1446] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1447] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1448] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1449] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1450] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1451] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1452] The following is further disclosed regarding the above embodiment.

[1453] (Claim 1)

[1454] an authentication means for users to log in;

[1455] means for selecting an appropriate virtual interviewer and question set from the received user information;

[1456] A means for generating an image and voice of a virtual interviewer using generative AI;

[1457] a means for displaying the generated virtual interviewer and playing the questions aloud;

[1458] a means for registering a user's answer by voice and analyzing the voice;

[1459] A means for generating feedback based on the analysis results and providing it visually and audibly;

[1460] A system including:

[1461] (Claim 2)

[1462] 10. The system of claim 1, further comprising means for analyzing the user's voice data to assess its content, voice quality, volume, speaking rate, and tension.

[1463] (Claim 3)

[1464] 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and also playing it aloud.

[1465] "Example 1"

[1466] (Claim 1)

[1467] an authentication means for users to log in;

[1468] means for selecting an appropriate virtual interviewer and question set from the received user information;

[1469] A means for generating an image and voice of a virtual interviewer using generative AI;

[1470] a means for displaying the generated virtual interviewer and playing the questions aloud;

[1471] a means for registering a user's answer by voice and analyzing the voice data;

[1472] means for converting the speech data into text and assessing the content, voice quality, volume, speaking rate, and stress level;

[1473] means for generating and providing visual and audio feedback based on the analysis results;

[1474] A system including:

[1475] (Claim 2)

[1476] 10. The system of claim 1, further comprising means for analyzing the user's voice data to assess its content, voice quality, volume, speaking rate, and tension.

[1477] (Claim 3)

[1478] 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and also playing it aloud.

[1479] "Application Example 1"

[1480] (Claim 1)

[1481] an authentication means for users to log in;

[1482] means for selecting an appropriate virtual instructor and instruction set from the received user information;

[1483] a means for generating an image and audio of a virtual instructor by a generative AI;

[1484] a means for displaying the generated virtual instructor and playing questions aloud;

[1485] a means for registering a user's answer by voice and analyzing the voice;

[1486] A means for generating feedback based on the analysis results and providing it visually and audibly;

[1487] means for generating a virtual driving environment and providing driving instructions to a user;

[1488] A means for analyzing the user's driving status in real time;

[1489] A system including:

[1490] (Claim 2)

[1491] 10. The system of claim 1, further comprising means for analyzing the user's voice data and driving data to evaluate the content, voice quality, volume, speaking rate, and driving skills.

[1492] (Claim 3)

[1493] 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and also playing it aloud.

[1494] "Example 2: Combining Emotion Engines"

[1495] (Claim 1)

[1496] an authentication means for users to log in;

[1497] means for selecting an appropriate virtual interviewer and question set from the received user information;

[1498] a means for generating an image and voice of a virtual interviewer using a generative AI model;

[1499] a means for displaying the generated virtual interviewer and playing the questions aloud;

[1500] a means for registering a user's answer by voice and analyzing the voice;

[1501] means for transmitting the voice data and facial expression data to a server for analysis;

[1502] means for generating and providing visual and audio feedback based on the analysis results;

[1503] A system including:

[1504] (Claim 2)

[1505] 10. The system of claim 1, further comprising means for evaluating the speech data and facial expression data for analysis.

[1506] (Claim 3)

[1507] 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and also playing it aloud.

[1508] "Application example 2 when combining emotion engines"

[1509] (Claim 1)

[1510] an authentication means for users to log in;

[1511] means for selecting an appropriate virtual interviewer and question set from the received user information;

[1512] A means for generating an image and voice of a virtual interviewer using generative AI;

[1513] a means for displaying the generated virtual interviewer and playing the questions aloud;

[1514] a means for registering a user's answer by voice and analyzing the voice;

[1515] A means for acquiring video information of a user and recognizing emotions from the video information;

[1516] A means for generating feedback based on the analysis results and providing it visually and audibly;

[1517] A system including:

[1518] (Claim 2)

[1519] 10. The system of claim 1, further comprising means for analyzing the user's voice data to assess its content, voice quality, volume, speaking rate, and tension.

[1520] (Claim 3)

[1521] 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and also playing it aloud. [Explanation of symbols]

[1522] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. an authentication means for users to log in; means for selecting an appropriate virtual interviewer and question set from the received user information; A means for generating an image and voice of a virtual interviewer using generative AI; a means for displaying the generated virtual interviewer and playing the questions aloud; a means for registering a user's answer by voice and analyzing the voice; A means for generating feedback based on the analysis results and providing it visually and audibly; A system including:

2. 10. The system of claim 1, further comprising means for analyzing the user's voice data to assess its content, voice quality, volume, speaking rate, and stress level.

3. 2. The system according to claim 1, further comprising means for displaying to the user feedback generated based on the analysis results and reproducing the feedback as audio.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A