Interview assistance system
The interview support system uses generative AI to simulate interviews and analyze biological reactions, addressing the limitations of conventional preparation methods by providing realistic feedback for improved interview performance.
Patent Information
- Application Number
- PCT/JP2024/016105
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional interview preparation methods fail to account for the interviewer's actual reactions and questions, leading to insufficient preparation, and lack a system for objectively evaluating a user's biological reactions during an interview.
An interview support system using generative AI technology to simulate an interview with an avatar based on employer documents, analyze biological reactions from video footage, and provide feedback on confident and unconfident responses.
Enables effective interview preparation by simulating real-world scenarios and objectively identifying areas for improvement through biological response analysis.
Smart Images

Figure JP2024016105_30102025_PF_FP_ABST
Abstract
Description
Interview support system
[0001] The present invention relates to an interview support system using AI technology, and in particular to an interview support system that combines generative AI technology and biological response analysis technology.
[0002] There is known a technique for analyzing the emotions felt by others in response to a speaker's comments (see, for example, Patent Document 1). There is also known a technique for analyzing changes in a subject's facial expression over a long period of time and estimating the emotions felt during that time (see, for example, Patent Document 2). There is also known a technique for identifying the factors that most influenced changes in emotions (see, for example, Patent Documents 3 to 5). There is also known a technique for comparing a subject's usual facial expression with their current facial expression and issuing an alert if the facial expression is gloomy (see, for example, Patent Document 6). There is also known a technique for comparing a subject's normal (expressionless) facial expression with their current facial expression to determine the subject's level of emotion (see, for example, Patent Documents 7 to 9). There are also known techniques for analyzing organizational emotions and the atmosphere felt by individuals within a group (see, for example, Patent Documents 10 and 11).
[0003] JP 2019-58625 A JP 2016-149063 A JP 2020-86559 A JP 2000-76421 A JP 2017-201499 A JP 2018-112831 A JP 2011-154665 A JP 2012-8949 A JP 2013-300 A JP 2011-186521 A WO15 / 174426
[0004] Attempts are being made to apply this type of analysis method to various fields. For example, there is a need for practice in preparation for an actual interview, such as through mock interviews, and for feedback based on self-analysis during the process.
[0005] Specifically, in conventional interview preparation, users typically think up and practice answers to anticipated questions. However, this method has the problem of making it difficult to anticipate the interviewer's actual reactions and questions, resulting in insufficient preparation. Furthermore, there has been no system that objectively evaluates the user's biological reactions during an interview and suggests areas for improvement.
[0006] Therefore, an object of the present invention is to provide a technology for analyzing a user's state based on communication, such as a mock interview, and supporting the interview by informing the user of the analysis results.
[0007] According to the present invention, an interview support system is obtained, comprising: a document acquisition unit that acquires documents related to the employer of choice from the user, a document assignment unit that assigns the documents and prompts for conducting the interview based on the documents to a generative AI, an avatar generation unit that generates an avatar that will conduct the interview, an interview execution unit that causes the avatar to conduct the interview, a video acquisition unit that acquires video footage of the interview, an analysis unit that analyzes changes in the user's biological reactions based on the video footage, and an output unit that outputs the results of the analysis.
[0008] According to the present invention, an interviewer avatar can be generated using generative AI technology, and an interview can be simulated based on documents related to the user's desired company. This allows the user to practice in an environment that is similar to a real interview, enabling more effective interview preparation.
[0009] Furthermore, by analyzing changes in the user's biological responses from video footage of the interview, it is possible to identify scenes in which the user answered confidently and scenes in which they answered unconfidently, and output video footage containing those scenes. This allows the user to objectively understand areas for improvement in their attitude and answers during the interview, enabling them to effectively prepare for their next interview.
[0010] In particular, according to the present invention, in a situation where online communication is the norm, it is possible to objectively evaluate the communication that has been exchanged in order to carry out more efficient communication.
[0011] FIG. 1 is a diagram showing an overall system diagram according to an embodiment of the present invention. FIG. 2 is another diagram showing an overall system diagram according to an embodiment of the present invention. FIG. 3 is a diagram showing functional configuration example 1 of an evaluation terminal according to an embodiment of the present invention. FIG. 4 is a diagram showing functional configuration example 2 of an evaluation terminal according to an embodiment of the present invention. FIG. 5 is a diagram showing functional configuration example 3 of an evaluation terminal according to an embodiment of the present invention. FIG. 6 is a diagram showing another configuration of functional configuration example 3 of an evaluation terminal according to an embodiment of the present invention. FIG. 7 is a functional block diagram of an interview support system according to an embodiment of the present invention. FIG. 8 is a sequence diagram showing the processing flow of the interview support system in an embodiment of the present invention. FIG. 9 is an image diagram of how the system is used in an embodiment of the present invention. FIG. 10 is an image diagram of how the system is used in an embodiment of the present invention.
[0012] The contents of embodiments of the present disclosure will be described below. The present disclosure has the following configuration. [Item 1] An interview support system comprising: a document acquisition unit that acquires a document related to a desired employer from the user; a document assignment unit that assigns the document and a prompt for conducting an interview based on the document to a generative AI; an avatar generation unit that generates an avatar that will conduct the interview; an interview execution unit that causes the avatar to conduct the interview; a video acquisition unit that acquires video footage of the interview; an analysis unit that analyzes changes in the user's biological reactions based on the video footage; and an output unit that outputs the results of the analysis. [Item 2] The interview support system according to claim 1, wherein the analysis unit analyzes at least one of scenes in which the user answered confidently or scenes in which the user answered unconfidently, and the output unit outputs the video footage including the scene.
[0013] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0014] <Summary> This invention relates to a system that supports interviews by simulating an interview using generative AI technology based on documents related to the user's desired employer, and analyzing changes in the user's biological responses from video footage of the interview.
[0015] <Hardware Configuration Example> Each functional block, functional unit, and functional module described below can be configured using, for example, hardware, a DSP (Digital Signal Processor), or software provided in a computer. For example, when configured using software, the system is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by running a program stored in a recording medium such as RAM, ROM, a hard disk, or a semiconductor memory. The series of processes performed by the system and terminal described herein can be realized using software, hardware, or a combination of software and hardware. A computer program for implementing each function of the information sharing support device 10 according to this embodiment can be created and installed on a PC or the like. A computer-readable recording medium storing such a computer program can also be provided. Examples of the recording medium include a magnetic disk, an optical disk, a magneto-optical disk, and a flash memory. The computer program may also be distributed, for example, via a network, without using a recording medium.
[0016] The evaluation terminal according to this embodiment acquires moving images from a video session service terminal, identifies at least facial images contained in the moving images for each predetermined frame, and calculates an evaluation value for the facial images (details will be described later).
[0017] <Video Acquisition Method> As shown in Figure 2, the video session service (hereinafter simply referred to as "this service") provided by the video session service terminal enables two-way image and audio communication with user terminals 10 and 20. This service displays video captured by the camera of the other user terminal on the display of the user terminal, and can output audio captured by the microphone of the other user terminal from the speaker. This service is also configured to enable both or either user terminal to record video and audio (collectively referred to as "video, etc.") in the memory of at least one of the user terminals. The recorded video information Vs (hereinafter referred to as "recorded information") is cached on the user terminal that initiated the recording and is recorded only locally on one of the user terminals. If necessary, users can view the recorded information themselves or share it with others within the scope of their use of this service.
[0018] <Functional Configuration Example 1> Fig. 3 is a block diagram showing an example configuration according to this embodiment. As shown in Fig. 3, the video session evaluation system of this embodiment is realized as a functional configuration possessed by a user terminal 10. That is, the user terminal 10 has, as its functions, a video image acquisition unit 11, a biological response analysis unit 12, a peculiar determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16.
[0019] The video acquisition unit 11 acquires video from each terminal, which is obtained by capturing images of multiple people (multiple users) using a camera provided in each terminal during an online session. The video acquired from each terminal may or may not be set to be displayed on the screen of each terminal. In other words, the video acquisition unit 11 acquires video from each terminal, including video currently being displayed and video currently not being displayed on each terminal.
[0020] The biological response analysis unit 12 analyzes changes in biological responses for each of multiple people based on the video images acquired by the video image acquisition unit 11 (regardless of whether they are currently being displayed on the screen). In this embodiment, the biological response analysis unit 12 separates the video images acquired by the video image acquisition unit 11 into an image set (a collection of frame images) and audio, and analyzes changes in biological responses from each of them. For example, the biological response analysis unit 12 analyzes changes in biological responses related to at least one of facial expression, eye movement, pulse rate, and facial movement by analyzing the user's facial image using the frame images separated from the video images acquired by the video image acquisition unit 11. Furthermore, the biological response analysis unit 12 analyzes changes in biological responses related to at least one of the user's speech content and voice quality by analyzing the audio separated from the video images acquired by the video image acquisition unit 11.
[0021] When a person's emotions change, this is reflected in changes in biological reactions such as facial expressions, eye movements, pulse rate, facial movements, speech content, and voice quality. In this embodiment, changes in the user's emotions are analyzed by analyzing changes in the user's biological reactions. One example of the emotion analyzed in this embodiment is the degree of comfort / discomfort. In this embodiment, the biological reaction analysis unit 12 quantifies changes in biological reactions according to a predetermined standard, thereby calculating a biological reaction index value that reflects the details of the changes in biological reactions.
[0022] The analysis of facial expression changes is performed, for example, as follows: For each frame image, a facial region is identified within the frame image, and the identified facial expressions are classified into multiple categories according to an image analysis model that has been trained in advance by machine learning. Based on the classification results, the system analyzes whether a positive or negative facial expression change has occurred between consecutive frame images, and the magnitude of the change, and outputs a facial expression change index value according to the analysis results.
[0023] The analysis of changes in gaze is performed, for example, as follows. That is, for each frame image, the eye area is identified within the frame image, and the direction of both eyes is analyzed to analyze where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker currently being displayed, at the shared document currently being displayed, or looking off-screen. It may also be possible to analyze whether the gaze movements are large or small, and whether the movements are frequent or infrequent. The gaze changes are also related to the user's concentration level. The biological response analysis unit 12 outputs a gaze change index value according to the analysis results of the gaze changes.
[0024] The analysis of pulse rate changes is performed, for example, as follows. That is, for each frame image, the facial area is identified within the frame image. Then, using a trained image analysis model that captures the numerical value of facial color information (G in RGB), changes in the G color of the facial surface are analyzed. The results are arranged along the time axis to form a waveform representing changes in color information, and the pulse is identified from this waveform. When a person is nervous, their pulse rate increases, and when they feel calm, their pulse rate decreases. The biological response analysis unit 12 outputs a pulse rate change index value according to the analysis results of the pulse rate changes.
[0025] The analysis of changes in facial movement is performed, for example, as follows. That is, for each frame image, a facial area is identified within the frame image, and the facial direction is analyzed to analyze where the user is looking. For example, it is analyzed whether the user is looking at the face of the currently displayed speaker, the currently displayed shared material, or looking off-screen. It may also be analyzed whether the facial movement is large or small, or whether the movement is frequent or infrequent. It may also be analyzed by combining facial movement and eye movement. For example, it may be analyzed whether the user is looking directly at the currently displayed speaker's face, looking up or down, or looking at an angle. The biological response analysis unit 12 outputs a facial direction change index value according to the analysis result of the change in facial direction.
[0026] The analysis of speech content is performed, for example, as follows. That is, the biological response analysis unit 12 converts speech for a specified period of time (for example, approximately 30 to 150 seconds) into a string of characters by performing known speech recognition processing, and then performs morphological analysis on the string of characters to remove words unnecessary for expressing the conversation, such as particles and articles. The remaining words are then vectorized, and an analysis is performed to determine whether a positive or negative emotional change has occurred, and the extent of the emotional change, and a speech content index value corresponding to the analysis result is output.
[0027] Voice quality analysis is performed, for example, as follows: The biological response analysis unit 12 identifies the acoustic features of the voice by performing known voice analysis processing on the voice for a specified period of time (for example, approximately 30 to 150 seconds). Based on the acoustic features, the unit then analyzes whether a positive or negative voice quality change has occurred and the magnitude of the voice quality change, and outputs a voice quality change index value according to the analysis results.
[0028] The biological response analysis unit 12 calculates a biological response index value using at least one of the facial expression change index value, eye direction change index value, pulse rate change index value, facial direction change index value, speech content index value, and voice quality change index value calculated as described above. For example, the biological response index value is calculated by weighting the facial expression change index value, eye direction change index value, pulse rate change index value, facial direction change index value, speech content index value, and voice quality change index value.
[0029] The peculiar determination unit 13 determines whether or not the change in biological reaction analyzed for the subject of analysis is peculiar compared to the change in biological reaction analyzed for other people other than the subject of analysis. In this embodiment, the peculiar determination unit 13 determines whether or not the change in biological reaction analyzed for the subject of analysis is peculiar compared to other people based on the biological reaction index values calculated for each of the multiple users by the biological reaction analysis unit 12.
[0030] For example, the unique determination unit 13 calculates the variance of the biological reaction index values calculated for each of multiple people by the biological reaction analysis unit 12, and by comparing the biological reaction index value calculated for the person being analyzed with the variance, determines whether the changes in the biological reactions analyzed for the person being analyzed are unique compared to others.
[0031] The following three patterns can be considered when changes in the analyzed biological reactions of the subject are unique compared to others. The first is when no particularly large changes in biological reactions occur in others, but a relatively large change in biological reactions occurs in the subject. The second is when no particularly large changes in biological reactions occur in the subject, but a relatively large change in biological reactions occurs in others. The third is when relatively large changes in biological reactions occur in both the subject and others, but the content of the change differs between the subject and others.
[0032] The associated event identification unit 14 identifies an event occurring with respect to at least one of the subject, other people, and the environment when a change in a biological reaction determined to be unique by the unique determination unit 13 occurs. For example, the associated event identification unit 14 identifies, from video images, the words and actions of the subject when a unique change in a biological reaction occurs in the subject. The associated event identification unit 14 also identifies, from video images, the words and actions of other people when a unique change in a biological reaction occurs in the subject. The associated event identification unit 14 also identifies, from video images, the environment when a unique change in a biological reaction occurs in the subject. The environment may be, for example, shared materials displayed on the screen or something that appears in the background of the subject.
[0033] The clustering unit 15 analyzes the degree of correlation between a change in biological reaction determined to be unique by the unique determination unit 13 (for example, one or more combinations of eye contact, pulse rate, facial movement, speech content, and voice quality) and an event occurring when the unique change in biological reaction occurs (an event identified by the related event identification unit 14), and if it is determined that the correlation is at a certain level or above, it clusters the person or event being analyzed based on the analysis results of the correlation.
[0034] For example, if a specific change in a biological reaction corresponds to a negative emotional change and the event occurring when the specific change in the biological reaction occurs is also a negative event, a correlation of a certain level or higher is detected. The clustering unit 15 clusters the analysis subject or event into one of a plurality of pre-segmented classifications according to the content of the event, the degree of negativity, the magnitude of correlation, etc.
[0035] Similarly, if a specific change in biological reaction corresponds to a positive emotional change and the event occurring when the specific change in biological reaction occurs is also a positive event, a correlation of a certain level or higher is detected. The clustering unit 15 clusters the analysis subject or event into one of a plurality of pre-segmented classifications according to the content, degree of positivity, magnitude of correlation, etc. of the event.
[0036] The analysis result notification unit 16 notifies the person designating the subject of analysis (the subject of analysis or the organizer of the online session) of at least one of the changes in biological reactions determined to be specific by the specific determination unit 13, the events identified by the related event identification unit 14, and the classifications clustered by the clustering unit 15.
[0037] For example, the analysis result notification unit 16 notifies the analysis subject of the analysis subject's own words and actions as an event occurring when a unique change in biological reaction occurs in the analysis subject that is different from that of others (one of the three patterns described above; the same applies below). This allows the analysis subject to understand that when he or she behaves in a certain way, he or she has different emotions than others. At this time, the analysis subject may also be notified of the unique changes in biological reaction identified for the analysis subject. Furthermore, the analysis subject may also be notified of changes in the biological reaction of others to be compared.
[0038] For example, if the emotions felt by others in response to words or actions made by the subject without any particular awareness and with normal emotions, or words or actions made by the subject with a particular awareness and with a certain emotion differ from the emotions felt by the subject himself at the time of the words or actions, the subject will be notified of his or her own words or actions at that time. This makes it possible to discover words or actions that are well-received by others or that are not well-received by others, despite the subject's own awareness.
[0039] Furthermore, the analysis result notification unit 16 notifies the organizer of the online session of events occurring when a unique change in biological reaction occurs in the analysis subject that is different from that of others, along with the unique change in biological reaction. This allows the organizer of the online session to know what events are influencing what emotional changes as phenomena unique to the designated analysis subject. Then, it becomes possible to take appropriate measures for the analysis subject based on the information obtained.
[0040] Furthermore, the analysis result notification unit 16 notifies the organizer of the online session of events occurring when a unique change in the biological reaction of the analysis subject occurs that is different from that of others, or of the clustering results of the analysis subject. This allows the organizer of the online session to understand the behavioral tendencies unique to the analysis subject and predict possible future behaviors and conditions, etc., depending on which category the specified analysis subject is clustered into. This then makes it possible to take appropriate measures for the analysis subject.
[0041] In the above embodiment, an example has been described in which a biological reaction index value is calculated by quantifying changes in biological reactions according to a predetermined standard, and whether or not the changes in biological reactions analyzed for the subject of analysis are unique compared to others is determined based on the biological reaction index values calculated for each of a plurality of people, but the present invention is not limited to this example. For example, the following may be used.
[0042] That is, the biological reaction analysis unit 12 analyzes the eye movement of each of the multiple people and generates a heat map showing the eye direction. The peculiar determination unit 13 compares the heat map generated by the biological reaction analysis unit 12 for the analysis subject with the heat map generated for other people, and determines whether the change in the biological reaction analyzed for the analysis subject is more peculiar than the change in the biological reaction analyzed for other people.
[0043] In this manner, in this embodiment, the video of the video session is stored in the local storage of the user terminal 10, and the above-described analysis is performed on the user terminal 10. Although it may depend on the machine specifications of the user terminal 10, it is possible to analyze the video information without providing it to an external party.
[0044] <Functional Configuration Example 2> As shown in FIG. 4 , the video session evaluation system of this embodiment may include, as its functional configuration, a video image acquisition unit 11, a biological response analysis unit 12, and a reaction information presentation unit 13a. The reaction information presentation unit 13a presents information indicating changes in biological responses analyzed by the biological response analysis unit 12a, including participants not displayed on the screen. For example, the reaction information presentation unit 13a presents information indicating changes in biological responses to a leader, facilitator, or manager of the online session (hereinafter collectively referred to as the organizer). The organizer of the online session may be, for example, a lecturer of an online class, a chairperson or facilitator of an online conference, or a coach of a session for coaching purposes. The organizer of the online session is typically one of multiple users participating in the online session, but may also be a different person who does not participate in the online session.
[0045] In this way, the host of an online session can grasp the status of participants who are not displayed on the screen in an environment where an online session is being held with multiple people.
[0046] <Functional Configuration Example 3> Fig. 5 is a block diagram showing a configuration example according to this embodiment. As shown in Fig. 5, the video session evaluation system of this embodiment has a functional configuration in which functions similar to those of the first embodiment described above are assigned the same reference numerals, and descriptions thereof may be omitted. The system according to this embodiment includes a camera unit that acquires video of the video session, a microphone unit that acquires audio, an analysis unit that analyzes and evaluates the video, an object generation unit that generates a display object (described later) based on information obtained by evaluating the acquired video, and a display unit that displays both the video of the video session and the display object during execution of the video session.
[0047] As explained above, the analysis unit includes a video image acquisition unit 11, a biological reaction analysis unit 12, a peculiar determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16. The functions of each element are as described above.
[0048] Based on the analysis results of the video acquired from the video session by the analysis unit, the object generation unit displays an object representing the recognized face and information representing the analyzed and evaluated content on the video, as necessary. When multiple faces appear in the video, the object generation unit may identify and display all of the faces. Furthermore, even if the camera function of the other party's device is disabled (i.e., the camera is disabled by software within the video session application, rather than by physically covering it), the object generation unit may display the object in the area where the other party's face is located if the other party's face is recognized by the camera. This allows both parties to confirm that the other party is in front of the device even if the camera function is disabled. In this case, for example, the video session application may hide information acquired from the camera while displaying an object corresponding to the face recognized by the analysis unit. Furthermore, the video information acquired from the video session and the information recognized by the analysis unit may be displayed on different display layers, and the layer related to the former information may be hidden. When there are areas for displaying multiple videos, the object may be displayed in all areas or only in some areas. For example, the object may be displayed only in the video on the guest side.
[0049] The embodiments of the invention described in Basic Configuration Examples 1 to 3 above may be implemented as a single device, or as multiple devices (e.g., cloud servers) connected in part or in whole via a network. For example, the control unit 110 and storage 130 of each terminal 10 may be implemented as different servers connected to each other via a network. That is, the present system includes user terminals 10 and 20, a video session service terminal 30 that provides two-way video sessions to the user terminals 10 and 20, and an evaluation terminal 40 that evaluates the video sessions. The following variations of the configuration are possible: (1) Processing entirely on the user terminal: As shown in FIG. 6, by performing processing by the analysis unit on the terminal performing the video session, analysis and evaluation results can be obtained simultaneously (in real time) with the video session (although a certain amount of processing power is required). (2) Processing on the user terminal and evaluation terminal: As shown in FIG. 7, the analysis unit may be included in an evaluation terminal connected via a network or the like. In this case, the video images acquired on the user terminal are shared with the evaluation terminal simultaneously with the video session or afterwards, and after being analyzed and evaluated by the analysis unit in the evaluation terminal, information on objects 50 and 100 is shared with the user terminal together with the video image data or separately (i.e., information including at least the analysis data) and displayed on the display unit.
[0050] The following interview support system is realized using each of the configurations of Functional Configuration Example 1 to Functional Configuration Example 3 described above or a combination thereof.
[0051] <Interview support system> As shown in Figure 8, the interview support system according to this embodiment includes a document acquisition unit, a document assignment unit, an avatar generation unit, an interview execution unit, a video acquisition unit, an analysis unit, and an output unit.
[0052] <Document Acquisition Unit> In the interview support system, the document acquisition unit plays an important role in acquiring documents related to the employer of choice from the user.
[0053] The Document Acquisition Unit collects various documents required when a user prepares for an interview. These documents include resumes, curriculum vitae, motivation letters, personal statements, portfolios, etc. This information is essential for assessing a user's background, skills, goals, and compatibility with potential employers.
[0054] The document acquisition section provides an interface that allows users to directly upload documents to the system. Users can select files stored on their devices (PC, smartphone, tablet, etc.) and upload them to the system. Common file formats are supported, making it easy for users to provide documents.
[0055] The document acquisition unit may also have the ability to automatically collect documents from users' devices or cloud storage. By granting users access to the system, it can periodically scan designated folders and files and automatically retrieve relevant documents. This saves users the trouble of manually uploading files, allowing them to always prepare for interviews with the most up-to-date documents.
[0056] The retrieved documents are properly classified and stored within the system. They are categorized according to document type (resume, curriculum vitae, etc.) and indexed for quick search and reference as needed. The document content is also analyzed using natural language processing technology to extract keywords and important information. This allows the document content to be effectively utilized in subsequent processing (prompt generation, interviews with avatars, etc.).
[0057] The document acquisition unit also takes user privacy and security into consideration. Acquired documents are stored securely under appropriate access control. Users can also delete documents or disconnect them from the system.
[0058] <Document Assignment Unit> In the interview support system, the document assignment unit assigns documents related to the user's desired workplace, collected by the document acquisition unit, to the generative AI, and plays an important role in generating prompts for conducting interviews.
[0059] The document annotation unit analyzes the documents received from the document acquisition unit and provides the information obtained from the analysis to the generative AI, which uses natural language processing and machine learning techniques to understand the content of the given documents and generate prompts for acting as an interviewer based on that understanding.
[0060] Prompts play a key role in controlling the flow of the interview and guiding the interaction with the user. Prompts include questions based on the content of the document and probing questions about the user's experience and motivation. These prompts are essential for interacting with the user to assess their capabilities and suitability.
[0061] The document annotation unit follows the steps below to have the generative AI generate a prompt.
[0062] Document preprocessing: The retrieved documents are converted into a format that is easy to analyze. Processing such as text extraction, format conversion, and noise removal is performed. Document analysis: Important information is extracted from the preprocessed documents. Natural language processing technology is used to understand keywords and context, and to grasp the user's background, skills, motivation for applying, etc.
[0063] Prompt generation: Based on the extracted information, a generative AI generates prompts. The generative AI is a model trained on a large number of interview cases and question patterns, and is able to generate appropriate questions based on the user's information.
[0064] Prompt optimization: Generated prompts are adapted to the interview context, optimizing the order and wording of questions to ensure a natural interview flow.
[0065] Generative AI is implemented using large-scale language models such as GPT (Generative Pre-trained Transformer). GPT can generate natural-sounding sentences based on context by learning from large amounts of text data. Generative AI can generate appropriate prompts by fine-tuning these models specifically for interviews.
[0066] The document assignment unit passes the prompts generated by the generative AI to the avatar generation unit, which then generates an avatar to conduct the interview based on the prompts and begins interacting with the user.
[0067] <Avatar Generation Unit> The avatar generation unit plays an important role in the interview support system in generating an avatar (a virtual interviewer) who will conduct the interview. The avatar plays a role in simulating the interview through interaction with the user.
[0068] The avatar generator generates an avatar that acts as an interviewer based on the prompts received from the document generator. The avatar can be expressed in various formats, such as a 3D model, a 2D image, or a video. The avatar is often generated using computer graphics (CG) technology or machine learning.
[0069] The avatar generation unit generates an avatar in the following procedure.
[0070] 1. Selecting an avatar base model: Select an avatar base model with an appearance and voice appropriate for an interviewer. The most suitable model is selected from several pre-prepared base models based on the user's desired company and the type of interview. 2. Customizing the avatar's appearance: Customize the appearance of the selected base model. Adjust hairstyle, clothing, accessories, etc. to create a more unique and friendly avatar. 3. Generating avatar behavior: Generate avatar behavior. Using natural language processing technology, appropriate behavior (gestures and facial expressions) for an interviewer is inferred from the prompt and reflected in the avatar. This allows the avatar to behave naturally in accordance with the context of the interview. 4. Generating avatar voice: Generate avatar voice. Using text-to-speech synthesis (TTS) technology, the content of the prompt is converted into the avatar's voice. Synchronizing the avatar's mouth movements with the voice achieves more natural interaction. 5. Generating avatar facial expressions: Generate avatar facial expressions. Using emotion recognition technology, the emotion the avatar should express is inferred from the content of the prompt and reflected in the facial expressions. This allows the avatar to respond appropriately to what the user says.
[0071] The avatar generation unit delivers the generated avatar to the interview execution unit, which uses the avatar to conduct an interactive interview with the user.
[0072] The avatar generator has the flexibility to customize the avatar to suit the user's preferences and needs. By adjusting the avatar's appearance, voice, speaking style, etc., the user can simulate an interview with an interviewer that suits them best.
[0073] The avatar generation unit can also generate avatars with specialized knowledge depending on the type of interview or the job the candidate is applying for. For example, in an interview for a technical job, an avatar with expertise in that field is generated, and it can ask specialized questions and provide feedback.
[0074] <Interview Execution Unit> The interview execution unit in the interview support system has the important function of simulating an actual interview between the interviewer avatar generated by the avatar generation unit and the user.
[0075] The interview execution unit receives the prompt generated by the document assignment unit and the avatar generated by the avatar generation unit, and uses them to conduct an interactive interview with the user. The interview can be conducted in various ways, such as text chat, voice call, or video call.
[0076] The interview execution unit executes an interview in the following steps: 1. Starting the interview: When the user instructs the start of the interview, the interview execution unit begins interaction with the avatar. The avatar introduces itself and explains the purpose of the interview, introducing the user to the interview atmosphere. 2. Presenting questions: The avatar asks the user questions based on prompts generated by the document annotation unit. The questions cover a wide range of interview-related topics, such as the user's background, skills, and motivation. 3. Processing the user's answers: When the user answers the avatar's questions, the interview execution unit analyzes the answers using natural language processing technology. Based on the answers, the interview execution unit infers the user's abilities, aptitude, and personality, and uses them to generate the next questions. 4. Generating additional questions: Based on the user's answers, the interview execution unit generates additional questions. These questions may include in-depth questions based on the user's answers or expansions to other related topics. These questions make the interaction with the user more natural and in-depth. 5. Interview Progress: The interview progresses through a question-and-answer session between the avatar and the user. The interview execution unit controls the flow of the interview and manages the time. It takes breaks and ends the interview as necessary. 6. Providing Feedback: After the interview is over, the interview execution unit provides feedback on the user's answers and behavior. It points out good points and areas for improvement and supports the user in preparing for the next interview. The interview execution unit uses advanced natural language processing and dialogue management technologies to achieve natural interaction between the user and the avatar. The avatar can understand what the user says and generate appropriate responses based on the context. It can also dynamically adjust the difficulty of the interview and change the pace of the interview to suit the user's condition by analyzing the user's reactions. The interview execution unit passes the interview footage to the video capture unit. The video capture unit records the interview footage for later analysis.
[0077] <Processing Flow> The processing flow of the system according to this embodiment will be described with reference to Figure 9. First, the user uploads documents related to the interview (such as a resume or curriculum vitae) to the document acquisition unit. The document acquisition unit passes these documents to the document annotation unit. Next, the document annotation unit provides the received documents to the generative AI. The generative AI analyzes the documents and generates prompts (questions or topics) for the interview. The generated prompts are returned to the document annotation unit.
[0078] The document annotation module passes prompts received from the generative AI to the avatar generation module. The avatar generation module generates an interviewer avatar and passes it to the interview execution module. The interview execution module starts an interactive interview with the user. As the interview progresses, the interview execution module presents questions to the user based on the prompts, and the user answers them. The interview execution module processes the user's answers and generates additional questions as needed. This process is repeated until the interview is completed.
[0079] After the interview is completed, the interview execution unit transfers the recorded video of the interview to the video acquisition unit, which then transfers the video to the analysis unit.
[0080] The analysis unit analyzes the video of the interview (see above) and evaluates the user's biological reactions and performance. The analysis results are passed to the output unit.
[0081] Finally, the output unit presents the analysis results from the analysis unit to the user, allowing the user to reflect on their own interview performance.
[0082] <Screen Examples> Examples of screen displays to the user in this embodiment will be described with reference to FIGS. 10 to 12. FIG.
[0083] As shown in FIG. 10, this screen shows a situation in which a practice interview is in progress. A label indicating that a practice interview is in progress is displayed at the top of the screen. An image of an avatar playing the role of interviewer is displayed in the center of the screen. The avatar image is thought to have been generated by the avatar generation unit. A small image of the user being interviewed is displayed at the bottom left of the screen. Questions from the interviewer to the user and the user's answers to those questions are displayed at the bottom of the screen. An "End Interview" button is displayed at the bottom right of the screen, indicating that the user can end the interview by pressing this button.
[0084] As shown in Figure 11, the left side of the screen displays an overview of the interview. This includes information such as the interview date and time, practice time, overall evaluation, and end time. In the "Interviewer's Comments" section, an avatar playing the role of interviewer provides feedback on the interview and comments on areas for improvement. In the upper right corner of the screen, there is a "View Statement of Purpose" button, which allows the user to review the statement of purpose submitted before the interview. In the center of the screen, a video of the interview is displayed. This is presumably a recording of the interview conducted by the interview execution unit. Below the video, a transcript of the interview content is displayed, allowing the user to review the exchanges that took place during the interview. In the lower right corner of the screen, "Good Points" and "Check Points" are displayed. It is presumed that "Good Points" indicate positive aspects of the interview, while "Check Points" indicate areas for improvement. These points are based on the results of an analysis of the interview video by the analysis unit.
[0085] As shown in Figure 12, this screen shows the settings screen before starting the interview practice. The title "Interview Practice" is displayed at the top of the screen. The interview setting options are displayed in the center of the screen. The user can configure the following settings: 1. Difficulty: The user can select the interview difficulty level from "Easy" or "Hard." 2. Preferred School Selection: The user can select the preferred school for the interview practice. 3. Camera Settings: The user can select whether or not to use the camera during the interview. 4. Environment Test: The user can perform an environment test for the interview. 5. Transcription Test: The user can test the voice transcription function. 6. Audio Test: The user can perform an audio test to check the microphone operation. The bottom of the screen displays a note explaining the purpose of the interview practice: "Interview practice will begin. This is not a re-creation of an actual situation before the interview, but rather an objective evaluation of interview performance." The bottom right of the screen displays a "Start Interview Practice" button. The user can press this button to begin the interview practice based on the set conditions.
[0086] The above-described embodiments may be combined as appropriate. Furthermore, the effects described in this specification are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that are apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0087] 10, 20 User terminal 30 Video session service terminal 40 Evaluation terminal
Claims
1. An interview support system comprising: a document acquisition unit that acquires documents related to the user's desired employer from the user; a document assignment unit that assigns the documents and prompts to a generative AI for conducting an interview based on the documents; an avatar generation unit that generates an avatar that will conduct the interview; an interview execution unit that causes the avatar to conduct the interview; a video acquisition unit that acquires video footage of the interview; an analysis unit that analyzes changes in the user's biological reactions based on the video footage; and an output unit that outputs the results of the analysis.
2. An interview support system as described in claim 1, wherein the analysis unit analyzes at least one of a scene in which the user answered with confidence or a scene in which the user answered without confidence, and the output unit outputs the video including the scene.
Citation Information
Patent Citations
Training system
JP2022014188A
Diagnostic apparatus, diagnostic method, diagnostic system, and program
JP2022114789A
Pseudo-interview system, pseudo-interview method, pseudo-interview apparatus, and program
JP2023000937A