Interview system, method, program and device.
The interview system addresses the challenge of acclimating interviewees to AI avatars by incorporating an ice-breaking phase, ensuring a comfortable interaction with the AI avatar from the outset.
Patent Information
- Application Number
- JP2025114172
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-07-05
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-05
AI Technical Summary
Conventional interview systems using AI avatars often fail to acclimate interviewees to interacting with them, leading to suboptimal interview experiences.
An interview system that includes an acquisition unit, prompt generation unit, and output unit to facilitate interaction with an AI avatar through a machine learning model, enabling an ice-breaking phase to familiarize the interviewee with the AI avatar before the main interview.
The interviewee can start the interview in a state of familiarity with the AI avatar, enhancing the interview experience.
Smart Images

Figure 0007763553000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an interview system, a method, a program, and an apparatus. [Background technology]
[0002] In recruitment activities, companies review documents submitted by job seekers or applicants, and then decide whether to hire them after one or more interviews. However, in traditional interviews, companies and interviewers had to coordinate their schedules. This meant that companies or interviewers sometimes missed opportunities for interviews.
[0003] In this regard, for example, a system is known in which a company assists in the selection of personnel by conducting interviews with interviewees using AI (Artificial Intelligence) (for example, see Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7665123 Summary of the Invention [Problem to be solved by the invention]
[0005] Although the conventional system described above can efficiently select personnel, when conducting interviews using AI, there are cases where the interviewee is not used to interacting with an AI avatar.
[0006] The present invention has been made in consideration of the above-mentioned problems, and aims to enable the interviewee to start the interview in a state where he or she is accustomed to interacting with an AI avatar. [Means for solving the problem]
[0007] According to one embodiment, the interview system includes an acquisition unit that acquires an interviewee's utterance from the interviewee's terminal device, a prompt generation unit that generates a prompt that instructs the creation of a response to the utterance based on the utterance and topic information indicating the topic before the interview with the interviewee, an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device, and an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition for the conversation is met. [Effects of the Invention]
[0008] According to one embodiment, the interviewee can start the interview in a state where he or she is familiar with interacting with an AI avatar. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an interview system 1000. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of an information processing device 100. FIG. [Figure 3] FIG. 2 is a diagram illustrating an example of the functional configuration of the interview device 1. [Figure 4] FIG. 10 is a diagram illustrating an example of template information. [Figure 5] FIG. 2 is a diagram illustrating an example of a functional configuration of a terminal device 2. [Figure 6] FIG. 2 is a diagram illustrating an example of a functional configuration of a language model device 3. [Figure 7] 10 is a flowchart showing the first half of an example of an interview method. [Figure 8] 10 is a flowchart showing the second half of an example of an interview method. [Figure 9] FIG. 10 is a diagram showing an example of a display screen of a terminal device during an interview with an AI avatar. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the description of the embodiment and the drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0011] <System configuration> First, an overview of the interview system 1000 according to this embodiment will be described. The interview system 1000 is an information processing system in which an interviewer and an AI avatar conduct an interview with the interviewer.
[0012] An interview is conducted between an interviewer and an AI avatar. The interview is conducted by the interviewer conversing with the AI avatar via a predetermined terminal device 2 connected to the network N. The interview is conducted, for example, by the interviewer repeatedly responding to statements made by the AI avatar. The interview may be, for example, an interview to hire a job seeker or applicant, or a one-on-one interview conducted between company employees. For example, if the interview is a job interview, the interviewer is a job seeker or applicant. For example, if the interview is one-on-one, the interviewer may be a company employee. Note that the type of interview that can be applied in the interview system 1000 is not limited to a job interview or one-on-one.
[0013] The AI avatar is a character that acts as a proxy for the person conducting the interview with the interviewer. During the interview, the AI avatar may perform any action, such as nodding, clapping, or shaking its head. The AI avatar is controlled by known techniques such as voice recognition, natural language processing, or emotion estimation. The AI avatar is displayed on the display unit of the terminal device 2. The appearance of the AI avatar, such as clothing, gender, and facial expression, can be changed according to the system settings. In this embodiment, the AI avatar is an image that resembles a human, but is not limited to a human.
[0014] The interview system 1000 can be used by a person who has conducted interviews with interviewees such as job seekers or employees, who acts as a user of the interview system 1000 and has an AI avatar conduct the interview on his or her behalf. The interview system 1000 can be used, for example, to conduct interviews in a company or the like, but the use of the interview system 1000 is not limited to this.
[0015] Figure 1 is a diagram showing an example of the configuration of an interview system 1000. As shown in Figure 1, the interview system 1000 includes an interview device 1, a terminal device 2, and a language model device 3, which are communicably connected to each other via a network N. The network N is, for example, a wired LAN (Local Area Network), a wireless LAN, the Internet, a public line network, a mobile data communication network, or a combination of these. In the example of Figure 1, the interview system 1000 includes one each of the interview device 1, the terminal device 2, and the language model device 3, but may include multiple of each.
[0016] The interview device 1 is an information processing device that conducts an interview with an interviewee using an AI avatar. The interview device 1 is, for example, but not limited to, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer. In the example of FIG. 1, the interview device 1 is a single information processing device, but may also be realized as a system consisting of multiple information processing devices connected via a network N.
[0017] The terminal device 2 is an information processing device used by a user to conduct an interview with an AI avatar. The terminal device 2 is, for example, a PC, a smartphone, a tablet terminal, a headset, or a head-mounted display, but is not limited to these.
[0018] The language model device 3 is an information processing device that receives instructions (prompts) expressed in natural language from an external device and generates response information to the prompts. The language model device 3 is an information processing device that transmits the generated response information to the external device. In this embodiment, the language model device 3 receives instructions (prompts) expressed in natural language from the interview device 1 and generates conversation content to be conveyed by the AI avatar to the interviewer. For example, the language model device 3 inputs the prompts received from the interview device 1 through an API (Application Programming Interface) into a machine learning model. The language model device 3 transmits information output from the machine learning model to the interview device 1 through the API. The language model device 3 is, for example, a PC, a server device, or a microcomputer, but is not limited to these. In the example of FIG. 1, the language model device 3 is a single information processing device, but may also be realized as a system consisting of multiple information processing devices connected via a network N.
[0019] <Hardware configuration of information processing device 100> Next, a description will be given of the hardware configuration of the information processing device 100. Fig. 2 is a diagram showing an example of the hardware configuration of the information processing device 100. As shown in Fig. 2, the information processing device 100 includes a processor 101, a memory 102, a storage 103, a communication I / F 104, an input device 105, an output device 106, and a drive device 107, which are connected to each other via a bus B.
[0020] The processor 101 controls each component of the information processing device 100 and realizes the functions of the information processing device 100 by loading various programs including an OS (Operating System) stored in the storage 103 into the memory 102 and executing the programs. The processor 101 is, for example, but not limited to, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), or a DSP (Digital Signal Processor).
[0021] The memory 102 is, for example, a read-only memory (ROM), a random access memory (RAM), or a combination thereof. The ROM is, for example, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a combination thereof. The RAM is, for example, but not limited to, a dynamic random access memory (DRAM) or a static random access memory (SRAM).
[0022] The storage 103 stores various programs including an OS and data. The storage 103 is, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or storage class memories (SCM), but is not limited to these.
[0023] The communication I / F 104 is an interface for connecting the information processing device 100 to an external device via the network N and controlling communication. The communication I / F 104 is, for example, Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), or Ethernet (registered trademark), but is not limited to these.
[0024] The input device 105 is a device for inputting information to the information processing device 100. The input device 105 is, for example, a mouse, a keyboard, a microphone, a scanner, a photographing device (camera), various sensors, or an operation button, but is not limited to these.
[0025] The output device 106 is a device for outputting information from the information processing device 100. The output device 106 is, for example, a display device such as a display, a projector, a printer, a speaker, or a vibrator, but is not limited to these. Note that in this embodiment, the input device 105 and the output device 106 may be configured as a touch panel integrated into the display device.
[0026] The drive device 107 is a device that reads and writes data from and to the recording medium 108. The drive device 107 is, for example, but not limited to, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or an SD card reader. The recording medium 108 is, for example, but not limited to, a CD (Compact Disc), a DVD (Digital Versatile Disc), an FD (Floppy Disk), an MO (Magneto-Optical disk), a BD (Blu-ray (registered trademark) Disc), a USB (registered trademark) memory, or an SD card.
[0027] In this embodiment, the program may be written to the memory 102 or the storage 103 during the manufacturing stage of the information processing device 100, or may be provided to the information processing device 100 via the network N, or may be provided to the information processing device 100 via a non-transitory computer-readable recording medium such as the recording medium 108.
[0028] <Functional configuration of interview device 1> Next, we will explain the functional configuration of the interview device 1. Figure 3 is a diagram showing an example of the functional configuration of the interview device 1. As shown in Figure 3, the interview device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0029] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information to and from the terminal device 2 or the language model device 3 via the network N.
[0030] The storage unit 12 is realized by the memory 102 and the storage 103. The storage unit 12 stores topic information 121, template information 122, prompts 123, conversation information 124, utterance information 125, and termination conditions 126.
[0031] The topic information 121 indicates a topic of conversation before the interview with the interviewer. The topic information 121 includes one or more of the following: weather, calendar, news, information about the interviewer, and information about past interviews conducted with the interviewer. The topic information 121 may be pre-recorded in the storage unit 12. The topic information 121 may be acquired from an external source, such as the Internet, via the network N when an interview start request is acquired from the terminal device 2. The information about the interviewer may be, for example, information written on a resume sent by the interviewer (e.g., address, name, date of birth, hobbies, qualifications, etc.), information registered in the interviewer's account, or the interviewer's facial expression. The interviewer's facial expression may be, for example, an image of the interviewer captured by an imaging device such as a camera built into or connected to the terminal device 2. The information about past interviews may be any information related to past interviews conducted with the interviewer. The information about past interviews may be, for example, text indicating the content of the conversation between the AI avatar and the interviewer in the past interview, or an evaluation given to the interview.
[0032] The template information 122 is information indicating a template of an instruction for a predetermined machine learning model. The template information 122 is composed of a predetermined natural language and an input area. The template information 122 is a template of a prompt to be input to the machine learning model. Information recorded in the storage unit 12 or information acquired by the acquisition unit 131 is input to the input area. The information recorded in the storage unit 12 may be, for example, topic information 121, prompt 123, conversation information 124, or utterance information 125. The natural language used for the template information 122 may be a natural language other than Japanese. For example, the natural language may be any language that can be processed by a machine learning model, such as English, German, or Chinese. The number of input areas may be zero, one, or multiple. The storage unit 12 may store one or multiple types of template information 122 in advance.
[0033] FIG. 4 is a diagram illustrating an example of template information 122. Template information 122 is an example of a template written in Japanese as natural language. The template information 122 illustrated in FIG. 4 is a template for generating dialogue content to help an interviewee, such as an interview applicant, become accustomed to interacting with an AI avatar before starting the interview with the AI avatar during a job interview. Specifically, template information 122 generates words that an AI avatar conducting an interview with an interviewee, such as an interview applicant, will convey to the interviewee, such as an interview applicant, before starting the interview. Template information 122 may include words to avoid discussing inappropriate topics with the interviewee. Inappropriate topics may be, for example, topics that are legally problematic, such as topics that constitute harassment, or topics that make the interviewee feel uncomfortable.
[0034] In the template information 122 shown in Fig. 4, the input area is described in an area surrounded by {#}. The template information 122 shown in Fig. 4 includes {#utterance information} and {#topic information} as input areas. For example, utterance information 125 stored in the storage unit 12 is input into {#utterance information}. Topic information 121 stored in the storage unit 12 is input into {#topic information}. Note that the template information 122 shown in Fig. 4 has one each of {#utterance information} and {#topic information}, but may have multiple.
[0035] The template information 122 may have {#conversation information} or {#prompt} in addition to {#utterance information} and {#topic information}. A prompt 123 stored in the storage unit 12 is input to {#prompt}. Conversation information 124 stored in the storage unit 12 is input to {#conversation information}. Furthermore, a plurality of pieces of topic information 121 may be input to one input field in the template information 122. For example, two pieces of topic information 121, namely, weather and calendar, may be input to {#topic information} in the template information 122. Note that in FIG. 4, the input field is indicated by an area surrounded by {#}, but is not limited thereto. The input field may also be indicated by an area surrounded by another symbol. Note that the content of the template information 122 is not limited thereto, and a plurality of pieces of template information 122 having different content may be stored in the storage unit 12.
[0036] 4 also includes the sentence, "Please be careful to avoid topics that are inappropriate for the interview applicant." Such a sentence can prevent the AI avatar from engaging in inappropriate conversation with the interviewee, such as the interview applicant.
[0037] Returning to FIG. 3 , the description of the interview device 1 will be continued. The prompt 123 is information expressed in a predetermined natural language for creating a conversation to be conveyed from the AI avatar to the interviewer. The prompt 123 is transmitted to the language model device 3 and input to a predetermined machine learning model. The prompt 123 is generated by inputting information stored in the storage unit 12 into an input field of the template information 122. The information input to the input field is, for example, topic information 121, a previously generated prompt 123, conversation information 124, or utterance information 125. The generated prompt 123 may be, for example, a prompt that instructs the AI avatar to create a response to the content of an utterance by the interviewer. The generated prompt 123 may be a prompt that instructs the AI avatar to create the content of an initial conversation to be conveyed to the interviewer based on the topic information 121. One or more prompts 123 may be stored in the storage unit 12. The multiple prompts 123 are, for example, prompts that have been input to a machine learning model in the past.
[0038] The conversation information 124 is information that represents the content of the conversation conveyed from the AI avatar to the interviewer. The conversation information 124 is information that is generated by inputting the prompt 123 into a predetermined machine learning model. The conversation information 124 is information expressed in a predetermined natural language to be conveyed to the interviewer. The conversation information 124 is stored in the memory unit 12 each time it is generated by the machine learning model. One or more pieces of conversation information 124 are stored, including information that has been generated by the machine learning model in the past.
[0039] The speech information 125 is information that represents the content of the interviewee's utterance. The speech information 125 may be acquired from the terminal device 2, or may be information that has been converted from a voice signal acquired from the terminal device 2 into text. The speech information 125 is expressed in a predetermined natural language. The speech information 125 is stored in the storage unit 12 every time it is acquired from the terminal device 2. Therefore, one or more pieces of speech information 125 are stored, including information acquired in the past.
[0040] The termination condition 126 is a condition for terminating the dialogue before the interview. The termination condition 126 may be, for example, the number of dialogues between the AI avatar and the interviewee. In this case, the termination condition is satisfied when the number of dialogues between the AI avatar and the interviewee exceeds a predetermined number. The number of dialogues may be the number of conversations conveyed by the AI avatar to the interviewee, the number of utterances by the interviewee acquired from the terminal device 2, or the sum of these numbers.
[0041] The termination condition 126 may be, for example, the duration of the dialogue between the AI avatar and the interviewee. In this case, the termination condition is satisfied when the duration of the dialogue between the AI avatar and the interviewee exceeds a predetermined time. Note that the duration of the dialogue may be the cumulative duration of the conversation conveyed from the AI avatar to the interviewee, the cumulative duration of the utterances by the interviewee acquired from the terminal device 2, or the sum of these cumulative durations.
[0042] The termination condition 126 may be, for example, the number of characters processed in the dialogue between the AI avatar and the interviewer. In this case, the termination condition is met when the number of characters processed in the dialogue between the AI avatar and the interviewer exceeds a predetermined number of characters. Note that the number of characters processed in the dialogue may be the cumulative number of characters in the conversation conveyed from the AI avatar to the interviewer, the cumulative number of characters in the utterances by the interviewer acquired from the terminal device 2, or the sum of these cumulative numbers of characters.
[0043] The termination condition 126 may be, for example, that the emotional state of the interviewee is determined to be in a state where the interview is possible during a dialogue between the AI avatar and the interviewee. The emotional state of the interviewee is estimated using a known method. A case will be described in which the emotional state of the interviewee is expressed numerically using parameters such as joy, sadness, anger, and happiness. In this case, the termination condition 126 may be satisfied when a predetermined parameter such as joy or happiness exceeds a predetermined threshold, or when the sum of all parameters exceeds a threshold. The emotional state of the interviewee may be estimated based on the content of the interviewee's utterance acquired from the terminal device 2. The emotional state of the interviewee may be estimated based on a facial image of the interviewee acquired from the terminal device 2. The facial image may be, for example, facial expression or the amount of change in facial expression.
[0044] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware configurations. The control unit 13 controls the overall operation of the interview device 1. The control unit 13 includes an acquisition unit 131, an interview unit 132, a prompt unit 133, an output unit 134, and a feeling estimation unit 135.
[0045] The acquisition unit 131 acquires information from various devices such as the terminal device 2 or the language model device 3 connected to the network N. For example, the acquisition unit 131 acquires an interview start request from the terminal device 2. The interview start request is a process of requesting the interview device 1 to start an interview. The acquisition unit 131 acquires utterance information indicating the content of the utterance made by the interviewer from the interview device 2 and records it as utterance information 125 in the memory unit 12. The acquisition unit 131 acquires conversation information 323 from the language model device 3 and records it as conversation information 124 in the memory unit 12. The acquisition unit 131 acquires topic information 121 from the memory unit 12 or a device capable of communicating via the network N. The device capable of communicating via the network N is, for example, an arbitrary web server connected to the Internet.
[0046] The interview unit 132 conducts an interview with the interviewee. Specifically, in response to receiving an interview start request, the interview unit 132 transmits an interview program to the terminal device 2. The interview program is a program for conducting an interview between the terminal device 2 and the interview device 1. The interview program includes a function for displaying an AI avatar on the terminal device 2 and a function for controlling the AI avatar displayed on the terminal device 2. The interview unit 132 conducts an interview with the interviewee by controlling the AI avatar displayed on the terminal device 2. The interview unit 132 may be configured to convert an audio signal acquired from the terminal device 2 into text. The audio signal may be, for example, speech information indicating the content of the interviewee's utterance collected from the input unit 26 of the terminal device 2. In this case, the interview unit 132 records the information obtained by converting the audio signal into text as speech information 125 in the memory unit 12. The interview is conducted between the AI avatar and the interviewee. The interview is conducted in two phases: an ice-breaking phase and an interview phase. The interview phase is an interview between the AI avatar and the interviewee. For example, the interview may be a job interview if the interviewee is a job seeker or applicant, or a one-on-one interview if the interviewee is a company employee. The interview may be any type of interview.
[0047] The ice-breaking phase is a phase before transitioning to the interview phase when conducting an interview with an interviewee. The ice-breaking phase is a phase in which a dialogue is conducted to allow the interviewee to become accustomed to dialogue with an AI avatar. In the ice-breaking phase, a dialogue between the interviewee and the AI avatar is conducted based on topic information 121. By conducting the ice-breaking phase, the interviewee can begin the interview in a state where they are accustomed to dialogue with an AI avatar. When the interview program is executed on the terminal device 2, the interview unit 132 first starts the ice-breaking phase. When a predetermined termination condition 126 is satisfied in the ice-breaking phase, the interview unit 132 transitions from the ice-breaking phase to the interview phase. The interview unit 132 conducts a dialogue between the interviewee and the AI avatar in the ice-breaking phase until the termination condition 126 is satisfied.
[0048] The prompt generation unit 133 generates a prompt 123 based on various information recorded in the memory unit 12. The content of the prompt 123 is an instruction to generate a conversation to be conveyed by the AI avatar to the interviewee. The prompt 123 is displayed in a predetermined natural language. Specifically, the prompt generation unit 133 acquires template information 122 from the memory unit 12. The prompt generation unit 133 inputs various pieces of information stored in the memory unit 12 into the input area of the acquired template information 122. In this way, the prompt generation unit 133 generates the prompt 123 by inputting information into the input area of the template information 122.
[0049] Next, a specific example of a prompt generated by the prompt generation unit 123 will be described. The prompt generation unit 123 generates the prompt 123 in the ice-breaking phase. The prompt generation unit 133 generates the prompt 123 that instructs the interviewer to create a response to the utterance by inputting the utterance information 125 and the topic information 121 into the template information 122. The prompt generation unit 133 generates the prompt 123 that instructs the interviewer to create an initial conversation to be conveyed by inputting the topic information 121 into the template information 122. The prompt generation unit 133 further inputs the interviewer's past utterances included in the utterance information 125 and the past responses included in the conversation information 124 into the template information 122, thereby generating a prompt that instructs the interviewer to create a response to the utterance. Note that the combination of information used by the prompt generation unit 133 to generate a prompt is not limited to these. The prompt generation unit 133 may generate the prompt 123 by inputting any information stored in the storage unit 12 into the input field of the template information 122. The prompt generation unit 133 records the generated prompt 123 in the storage unit 12. The prompt generation unit 133 transmits the generated prompt 123 to the language model device 3.
[0050] The output unit 134 inputs the prompt 123 into a predetermined machine learning model, and outputs the content of the dialogue that the AI avatar conveys to the interviewer to the terminal device 2. For example, the output unit 134 transmits the prompt 123 to the language model device 3 via an API. The output unit 134 records the conversation information 323 acquired from the language model device 3 by the acquisition unit 131 via the API as conversation information 124 in the storage unit 12. The output unit 134 transmits the conversation information 124 to the terminal device 2.
[0051] The emotion estimation unit 135 estimates the emotion of the interviewee during the ice-breaking phase. The emotion estimation unit 135 estimates the emotion of the interviewee based on the natural language indicated by the utterance information 125 or a facial image of the interviewee acquired from the terminal device 2. The emotion estimation unit 135 estimates the emotion using a known method. For example, when the emotion estimation unit 135 estimates the emotion based on natural language, the emotion may be estimated by performing natural language processing such as grammar analysis or key phrase extraction. Alternatively, the emotion estimation unit 135 may estimate the emotion using a convolutional neural network for text (TextCNN) or a long short-term memory (LSTM). For example, when the emotion estimation unit 135 estimates the emotion based on a facial image of the interviewee, the emotion may be estimated based on features obtained by image processing such as facial feature extraction. The emotion estimation unit 135 may indicate multiple emotions, such as joy, anger, sadness, fear, and enjoyment, using an evaluable index such as a score.
[0052] In this embodiment, the output unit 134 outputs the conversation information 124 by transmitting the prompt 123 to the language model device 3, but the means for outputting the conversation information 124 is not limited to this. For example, the storage unit 12 may be configured to store a machine learning model such as a large language model (LLM), and the output unit 134 may input the prompt 123 to the machine learning model stored in the storage unit 12.
[0053] <Functional configuration of terminal device 2> Next, the functional configuration of the terminal device 2 will be described. Fig. 5 is a diagram showing an example of the functional configuration of the terminal device 2. Explanations that overlap with those of the interview device 1 shown in Fig. 3 will be omitted as appropriate. As shown in Fig. 5, the terminal device 2 includes a communication unit 21, a storage unit 22, a control unit 23, a display unit 24, an operation unit 25, an input unit 26, and an audio output unit 27.
[0054] The communication unit 21 is realized by the communication I / F 104. The communication unit 21 transmits and receives information to and from the interview device 1 via the network N.
[0055] The storage unit 22 is realized by the memory 102 and the storage 103. The storage unit 22 stores conversation information 221.
[0056] The conversation information 221 is information that represents the content of the conversation conveyed from the AI avatar to the interviewer. The conversation information 221 is information transmitted from the interview device 1. The conversation information 221 is information expressed in a predetermined natural language (for example, the interviewer's native language) so that it can be conveyed to the interviewer. The conversation information 221 is stored in the memory unit 22 each time it is transmitted from the interview device 1. One or more pieces of conversation information 221 are stored, including information transmitted from the terminal device 2 in the past.
[0057] The control unit 23 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware components. The control unit 23 controls the overall operation of the terminal device 2. The control unit 23 includes an acquisition unit 231 and an application unit 232.
[0058] The acquisition unit 231 acquires information from various devices connected to the network N. For example, the acquisition unit 231 acquires an interview program from the interview device 1. The acquisition unit 231 acquires conversation information 124 from the interview device 1 and records it in the memory unit 22 as conversation information 221.
[0059] The application unit 232 is communicably connected to the terminal device 2 via the network N, and thereby conducts an interview between the interviewer and an AI avatar. Specifically, the application unit 232 sends an interview start request to the interview device 1 via operation of the operation unit 25. Thereafter, the interview program acquired by the acquisition unit 231 is loaded into the control unit 23 of the terminal device 2. The application unit 232 executes the loaded interview program. When the interview program is executed, the application unit 232 synchronizes with the interview unit 132. When the application unit 232 and the interview unit 132 are synchronized, the interview begins. The application unit 232 first generates an AI avatar. The application unit 232 controls the AI avatar based on control instructions from the interview unit 132. The application unit 232 generates a screen for conducting the interview. The screen includes the AI avatar. The application unit 232 outputs the generated screen to the display unit 24.
[0060] For example, the application unit 232 converts the conversation information 221 into an audio signal. The application unit 232 outputs the converted audio signal to the audio output unit 27. The audio output unit 27 outputs the content of the conversation included in the conversation information 124 as audio. At this time, the application unit 232 can control the AI avatar by moving its mouth or body so that the AI avatar appears to be speaking the content of the conversation indicated by the conversation information 124 to the interviewer. The application unit 232 also acquires the content of the interviewer's utterance from the input unit 26 as speech information. The application unit 232 may transmit the acquired speech information as an audio signal to the interview device 1. The application unit 232 may also transcribe the acquired speech information using a known method and transmit the transcribed information to the interview device 1.
[0061] The display unit 24 is, for example, a display, and is a device that outputs information as a screen of a graphical user interface that can be operated by the user. The display unit 24 may be included in the housing of the terminal device 2, or may be externally attached. Specifically, the display unit 24 may be implemented as a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, an organic EL display, or a plasma display. It is preferable that these display devices are implemented by selectively using them depending on the type of the terminal device 2. The display unit 24 displays various screens generated by the application unit 232.
[0062] The operation unit 25 is a device that inputs information to the terminal device 2 in response to operations by the interviewer. The operation unit 25 accepts operation inputs made by the interviewer. The operation inputs are transferred to the control unit 23 as command signals. The operation unit 25 may be implemented as a touch panel integrated with the display unit 24. When the operation unit 25 is implemented as a touch panel, the interviewer can input tap operations, swipe operations, etc. to the operation unit 25. Instead of a touch panel, switch buttons, a mouse, a trackpad, a QWERTY keyboard, etc. can be used as the operation unit 25. It is preferable that these input devices are used appropriately depending on the type of terminal device 2.
[0063] The input unit 26 is, for example, an imaging device such as a microphone or a camera. If the input unit 26 is a microphone, the input unit 26 is a device that inputs the speech of the interviewee or surrounding sounds as an audio signal to the terminal device 2. If the input unit 26 is an imaging device such as a camera, the input unit 26 is a device that inputs the interviewee or the background as a video signal to the terminal device 2. The input unit 26 may be included in the housing of the terminal device 2 or may be externally attached.
[0064] The audio output unit 27 is, for example, a speaker. The audio output unit 26 is a device that outputs a predetermined audio signal. The predetermined audio signal may be, for example, an audio signal converted from the conversation information 221. The audio output unit 27 may be included in the housing of the terminal device 2 or may be externally attached.
[0065] The functional configuration of the terminal device 2 is not limited to the above example. For example, the terminal device 2 may have some of the above functional configurations, with the interview device 1 having the rest. The terminal device 2 may also have functional configurations other than those described above. Each functional configuration of the terminal device 2 may be realized by software, as described above, or by hardware such as an IC chip, SoC, LSI, or microcomputer.
[0066] <Functional Configuration of Language Model Device 3> Next, a description will be given of the functional configuration of the language model device 3. Fig. 6 is a diagram showing an example of the functional configuration of the language model device 3. As shown in Fig. 6, the language model device 3 includes a communication unit 31, a storage unit 32, and a control unit 33.
[0067] The communication unit 31 is realized by the communication I / F 104. The communication unit 31 transmits and receives information to and from the interview device 1 via the network N.
[0068] The storage unit 32 is realized by the memory 102 and the storage 103. The storage unit 32 stores a prompt 321, a large-scale language model 322, and conversation information 323.
[0069] The prompt 321 is information expressed in a predetermined natural language to create a conversation to be conveyed from the AI avatar to the interviewer. The prompt 321 is transmitted from the interview device 1. The prompt 321 is input to the large-scale language model 322. One or more prompts 321 are stored in the memory unit 32. The multiple prompts 321 are, for example, prompts that have been input to the large-scale language model 322 in the past.
[0070] The large-scale language model 322 is a machine learning model that can perform machine learning on large amounts of text data and generate natural-looking words and images. The large-scale language model 322 generates information in cooperation with each functional unit of the control unit 33. For example, when information is input from each functional unit of the control unit 33, the large-scale language model 322 generates information based on the input information. The generated information is, for example, conversation information 323 that indicates the content of the dialogue conveyed from the AI avatar to the interviewer. The large-scale language model 322 outputs the generated information to each functional unit. The memory unit 32 may store one large-scale language model 322, or multiple large-scale language models 322 may be stored depending on the application. The type of large-scale language model 322 and the machine learning method thereof are not limited. In this embodiment, a prompt 123 is input to the large-scale language model 322.
[0071] Conversation information 323 is information representing the content of the conversation conveyed from the AI avatar to the interviewer. Conversation information 323 is information generated by large-scale language model 322 to which prompt 321 is input. Conversation information 323 is expressed in a predetermined natural language to be conveyed to the interviewer. Each time conversation information 323 is generated by large-scale language model 322, it is stored in memory unit 12. Therefore, one or more pieces of conversation information 323 are stored, including information generated in the past.
[0072] The control unit 33 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware components. The control unit 33 controls the overall operation of the language model device 3. The control unit 33 includes an acquisition unit 331 and a conversation generation unit 332.
[0073] The acquisition unit 331 acquires information from various devices such as the interview device 1 connected to the network N. For example, the acquisition unit 331 acquires the prompt 123 from the interview device 1 and records it as the prompt 321 in the memory unit 32.
[0074] The conversation generation unit 332 generates conversation information 323 based on the prompt 321. Specifically, the conversation generation unit 332 acquires the prompt 321 from the storage unit 12. Next, the conversation generation unit 332 inputs the prompt 321 into the large-scale language model 322. The conversation generation unit 332 acquires the conversation information 323 generated by the large-scale language model 322. The conversation generation unit 332 records the generated conversation information 323 in the storage unit 32. The conversation generation unit 332 transmits the generated conversation information 323 to the interview device 1.
[0075] The functional configuration of the language model device 3 is not limited to the above example. For example, the language model device 3 may have some of the above functional configuration, with the interview device 1 or the terminal device 2 having the rest. The language model device 3 may also have functional configurations other than those described above. Each functional configuration of the language model device 3 may be realized by software, as described above, or by hardware such as an IC chip, SoC, LSI, or microcomputer.
[0076] <Interview method> An interview method executed by the interview system 1000 will now be described. Fig. 7 is a flowchart showing an example of the first half of the interview method. Fig. 8 is a flowchart showing an example of the second half of the interview method.
[0077] (Step S101) The application unit 232 of the terminal device 2 transmits a request to start an interview to the interview device 1. For example, the interviewee operates the operation unit 25 to click on a specific URL (Uniform Resource Locator). In response to the URL being clicked, the application unit 232 transmits a request to start an interview to the interview device 1. Note that the object operated by the operation unit 25 may not be a URL, but may be software installed on the terminal device 2.
[0078] (Step S102) The acquisition unit 131 of the interview device 1 acquires the interview start request. Next, the interview unit 132 transmits the interview program to the terminal device 2 in response to the acquisition of the interview start request.
[0079] (Step S103) The acquisition unit 231 acquires the interview program. The application unit 232 executes the interview program. When the interview program is executed, the application unit 232 synchronizes with the interview unit 132. When the application unit 232 and the interview unit 132 are synchronized, the interview begins. The interview unit 132 executes the ice-breaking phase of the interview. The application unit 232 generates a screen for conducting the interview. The application unit 232 outputs the generated screen to the display unit 24.
[0080] (Step S104) The application unit 232 generates an AI avatar. The generated AI avatar is displayed on the display unit 24. The application unit 232 controls the AI avatar based on a control instruction from the interview unit 132.
[0081] (Step S105) The acquisition unit 131 acquires topic information 121 from the storage unit 12 or a device capable of communicating via the network N. For example, the acquisition unit 131 may acquire information about the interviewee as topic information 121 from the storage unit 12, or may acquire information about past interviews conducted with the interviewee (e.g., the content of conversations in past interviews, impressions of an AI avatar, etc.). The information about the interviewee may be, for example, information written on the interviewee's resume (e.g., address, name, date of birth, hobbies, qualifications, etc.). A case will be described in which the acquisition unit 131 acquires topic information 121 from a device capable of communicating via the network N. The acquisition unit 131 may acquire any information, such as weather information, date and time, or news, via an API.
[0082] (Step S106) The prompt generation unit 133 generates a prompt based on various information recorded in the storage unit 12. Specifically, the prompt generation unit 133 acquires template information 122 from the storage unit 12. The prompt generation unit 133 inputs the acquired topic information 121 into an input area of the acquired template information 122, thereby generating a prompt 123 that instructs the interviewer to create an initial conversation to be conveyed to the interviewee. The prompt generation unit 133 records the generated prompt 123 in the storage unit 12.
[0083] (Step S107) The output unit 134 inputs the prompt 123 into a predetermined machine learning model, and outputs the content of the dialogue that the AI avatar conveys to the interviewer to the terminal device 2. First, the output unit 134 transmits the generated prompt 123 to the language model device 3. The acquisition unit 331 acquires the prompt 123. The acquisition unit 331 records the acquired prompt 123 in the memory unit 32 as a prompt 321.
[0084] (Step S108) The conversation generation unit 332 generates conversation information 323 based on the prompt 321. Specifically, the conversation generation unit 332 acquires the prompt 321 from the storage unit 32. Next, the conversation generation unit 332 generates the conversation information 323 by inputting the prompt 321 into the large-scale language model 322. The conversation generation unit 332 records the generated conversation information 323 in the storage unit 32.
[0085] (Step S109) The conversation generation unit 332 transmits the conversation information 323 to the interview device 1. The acquisition unit 131 acquires the conversation information 323. The output unit 134 records the acquired conversation information 323 as conversation information 124 in the storage unit 12. Note that the output unit 134 may record the conversation information 124 after performing a predetermined process on the conversation information 323. The predetermined process may be, for example, a process of deleting information other than the conversation, such as an introduction, from the conversation information 323.
[0086] (Step S110) The output unit 134 transmits the conversation information 124 to the terminal device 2. The acquisition unit 231 acquires the conversation information 124 from the interview device 1 and records it in the storage unit 22 as conversation information 221.
[0087] (Step S111) The application unit 232 controls the AI avatar so that it appears as if it is speaking to the interviewer. Specifically, the application unit 232 converts the conversation information 221 recorded in the storage unit 22 into an audio signal. The application unit 232 outputs the converted audio signal to the audio output unit 27. As a result, the content of the conversation indicated by the conversation information 221 is output as audio from the audio output unit 27. At this time, the application unit 232 can control the AI avatar so that it appears as if it is speaking to the interviewer by moving the AI avatar's mouth or body. For example, the application unit 232 can control the AI avatar so that it asks the interviewer, "What are your hobbies?"
[0088] (Step S112) The application unit 232 acquires the content of the interviewee's utterance as utterance information from the input unit 26. The acquired utterance information may include a response to the content of the conversation of the AI avatar. For example, if the AI avatar asks the interviewee, "What are your hobbies?", the utterance information may include a response to the content of the interviewee's utterance, such as, "My hobby is reading."
[0089] (Step S113) The application unit 232 transcribes the acquired utterance information using a known method and transmits it to the interview device 1. Note that the application unit 232 may transmit the acquired utterance information as an audio signal to the interview device 1. The acquisition unit 131 records the acquired utterance information in the memory unit 12 as utterance information 125.
[0090] (Step S114) The prompt generation unit 133 generates a prompt based on various information recorded in the storage unit 12. Specifically, the prompt generation unit 133 acquires template information 122 from the storage unit 12. The prompt generation unit 133 inputs the acquired topic information 121 and utterance information 125 into an input area of the acquired template information 122, thereby generating a prompt 123 that instructs the user to create a response to the utterance. The prompt generation unit 133 records the generated prompt 123 in the storage unit 12. Note that the prompt generation unit 133 may acquire template information 122 different from the template information 122 acquired in step S106. Furthermore, the prompt generation unit 133 may generate the prompt 123 by further inputting other information recorded in the storage unit 12 (e.g., a previously generated prompt 123, previously acquired conversation information 124, previously acquired utterance information 125, etc.) into the input area.
[0091] (Step S115) The output unit 134 inputs the prompt 123 into a predetermined machine learning model, and outputs the content of the dialogue that the AI avatar conveys to the interviewer to the terminal device 2. First, the output unit 134 transmits the generated prompt 123 to the language model device 3. The acquisition unit 331 acquires the prompt 123. The acquisition unit 331 records the acquired prompt 123 in the memory unit 32 as a prompt 321.
[0092] (Step S116) The conversation generation unit 332 generates conversation information 323 based on the prompt 321. Specifically, the conversation generation unit 332 acquires the prompt 321 from the storage unit 32. Next, the conversation generation unit 332 generates the conversation information 323 by inputting the prompt 321 into the large-scale language model 322. The conversation generation unit 332 records the generated conversation information 323 in the storage unit 32. Note that the prompt 321 is generated using the utterance information 125. Therefore, the generated conversation information 323 includes a response content to the content of the interviewer's utterance included in the utterance information 125. For example, if the utterance information 125 includes "My hobby is reading," the conversation information 323 includes a response content to the content of the interviewer's utterance, such as "So reading is your hobby. What's your favorite genre?"
[0093] (Step S117) The conversation generation unit 332 transmits the conversation information 323 to the interview device 1. The acquisition unit 131 acquires the conversation information 323. The output unit 134 records the acquired conversation information 323 as conversation information 124 in the storage unit 12. Note that the output unit 134 may record the conversation information 124 after performing a predetermined process on the conversation information 323. The predetermined process may be, for example, a process of deleting information other than the conversation, such as an introduction, from the conversation information 323.
[0094] (Step S118) The output unit 134 transmits the conversation information 124 to the terminal device 2. The acquisition unit 231 acquires the conversation information 124 from the interview device 1 and records it in the storage unit 22 as conversation information 221.
[0095] (Step S119) The application unit 232 controls the AI avatar so that it appears as if it is speaking to the interviewer. Specifically, the application unit 232 converts the conversation information 221 recorded in the storage unit 22 into an audio signal. The application unit 232 outputs the converted audio signal to the audio output unit 27. As a result, the content of the conversation indicated by the conversation information 221 is output as audio from the audio output unit 27. At this time, the application unit 232 can control the AI avatar by moving the mouth or body of the AI avatar so that it appears as if it is speaking the content of the conversation indicated by the conversation information 124 to the interviewer. The conversation information 221 includes a response content to the content of the interviewer's utterance. Therefore, by having the AI avatar respond to the interviewer with the content included in the conversation information 221, the interviewer can feel as if he or she is having a conversation with the AI avatar.
[0096] (Step S120) The application unit 232 acquires the content of the interviewee's utterance as speech information from the input unit 26. The acquired speech information may include a response to the content of the conversation of the AI avatar. For example, if the AI avatar says to the interviewee, "I see you enjoy reading. What's your favorite genre?", the speech information may include a response to the content of the interviewee's utterance, such as "My favorite genre is science fiction."
[0097] (Step S121) The application unit 232 transcribes the acquired utterance information using a known method and transmits it to the interview device 1. Note that the application unit 232 may transmit the acquired utterance information as an audio signal to the interview device 1. The acquisition unit 131 records the acquired utterance information in the memory unit 12 as utterance information 125.
[0098] (Step S122) The interview unit 132 determines whether the termination condition 126 is satisfied. If the termination condition 126 is not satisfied (step S122: NO), the interview unit 132 transitions to step S123. If the termination condition 126 is satisfied (step S122: YES), the interview unit 132 transitions to step S124.
[0099] For example, if the termination condition 126 is the number of interactions between the AI avatar and the interviewer, the interview unit 132 counts the number of interactions. If the number of interactions exceeds a predetermined number, the interview unit 132 determines that the termination condition 126 is satisfied.
[0100] For example, if the termination condition 126 is the time of a conversation with an interviewer, the interview unit 132 measures the time of the conversation. If the time of the conversation exceeds a predetermined time, the interview unit 132 determines that the termination condition 126 is satisfied.
[0101] For example, if the termination condition 126 is the number of characters processed in the dialogue between the AI avatar and the interviewer, the interview unit 132 counts the number of characters processed in the dialogue between the AI avatar and the interviewer. If the counted number of characters exceeds a predetermined number of characters, the interview unit 132 determines that the termination condition 126 is satisfied.
[0102] If the termination condition 126 is, for example, that the emotional state of the interviewee is in a state where an interview is possible, the emotion estimation unit 135 estimates the emotion of the interviewee. If the estimation result by the emotion estimation unit 135 indicates that the emotional state of the interviewee is in a state where an interview is possible, the interview unit 132 determines that the termination condition 126 is satisfied.
[0103] (Step S123) The interview unit 132 operates the acquisition unit 131, interview unit 132, prompt generation unit 133, and output unit 134 until the termination condition 126 is satisfied. Specifically, the process proceeds to step S114, and the interview system 1000 executes the interview method.
[0104] (Step S124) The interview unit 132 starts an interview with the interviewer. The interview unit 132 transitions from the ice-breaking phase to the interview phase. The interview may be conducted by an AI avatar or a real person. Note that by conducting the ice-breaking phase in advance, the interviewer can conduct the interview in a state where they are accustomed to interviews with an AI avatar. Therefore, the interviewer can respond appropriately to questions in the interview with the AI avatar. Furthermore, by conducting the ice-breaking phase in advance, the interviewer can engage in a conversation without getting bored until the real person appears for the interview. Furthermore, even when the real person appears, the interviewer can conduct the interview in a state where they are accustomed to online interviews. Therefore, the interviewer can respond appropriately to questions in the interview.
[0105] 9 is a diagram showing an example of a display screen of a terminal device during an interview with an AI avatar. In FIG. 9, the display screen includes an AI avatar 201 having an interview with an interviewer, a conversation area 202, a microphone icon 203, and a status icon 204.
[0106] The conversation area 202 is an area where the content of the conversation conveyed by the AI avatar 201 is expressed in natural language. The content of the conversation is output from the audio output unit 27 of the front terminal device 2. According to the conversation area 202, the content of the conversation by the AI avatar 201 is "It's nice weather in the Kanto region."
[0107] Microphone icon 203 is an icon that indicates whether or not the conversation of AI avatar 201 is being output from audio output unit 27. Microphone icon 203 indicates the state when the conversation of AI avatar 201 is being output from audio output unit 27. When the conversation of AI avatar 201 is not being output from audio output unit 27, microphone icon 203 is displayed in a different manner. In this way, by indicating with microphone icon 203 whether or not the conversation of AI avatar 201 is being output from audio output unit 27, the interviewer can distinguish whether it is time for the interviewer or AI avatar 201 to speak.
[0108] The status icon 204 indicates to the interviewer the phase of the interview. The status icon 204 includes icons associated with the ice-breaking phase and the interview phase. The icon shown in FIG. 9 indicates the ice-breaking phase to the interviewer. The icon shown in FIG. 9 displays the text "chatting." The interviewer can know that the interview has not yet started by checking the "chatting" status icon 204. When the status icon 204 is displayed in the interview phase, it changes to an icon associated with the interview phase. For example, the icon indicating the interview phase is an icon displaying the text "interview." Although an example has been shown in which the status icon 204 displays the text "chatting" as the icon indicating the ice-breaking phase and the text "interview" as the icon indicating the interview phase, the icons of the status icon 204 are not limited to these. The status icon 204 may be an image or a combination of an image and text instead of text.
[0109] <Summary> As described above, according to this embodiment, an interview system 1000 can be realized, which includes an acquisition unit 131 that acquires the interviewee's utterance from the interviewee's terminal device 2, a prompt generation unit 133 that generates a prompt 123 that instructs the creation of a response to the utterance based on the utterance and topic information 121 that indicates the topic before the interview with the interviewer, an output unit 134 that outputs the response to the terminal device 2 by inputting the prompt 123 into a machine learning model, and an interview unit 132 that operates the acquisition unit 131, the prompt generation unit 133, and the output unit 134 until the conversation termination condition is met.
[0110] According to the interview system 1000, by inputting the interviewer's utterances and prompts including topic information indicating topics to be discussed before the interview with the interviewer into a machine learning model, it is possible to generate conversation content that will ease the interviewer's tension. By repeating the process of the AI avatar conveying the content of the conversation to the interviewer and the interviewer making further utterances based on the content of the conversation, the interviewer can start the interview in a state of familiarity with the AI avatar before the interview.
[0111] <Modification> In this embodiment, the large-scale language model 332 is recorded in the storage unit 32 of the language model device 3, but it may also be recorded in the storage unit 12 of the interview device 1. In this case, the output unit 134 inputs the prompt 123 into the large-scale language model recorded in the storage unit 12, instead of transmitting the prompt 123 to the language model device 3. Such a large-scale language model may be an existing model, or may be a model developed by the developer of the interview device 1.
[0112] In this embodiment, the prompt generation unit 132 may be configured to change the topic information 121 and generate the prompt 123 during the ice-breaking phase. For example, if the achievement status of the termination condition 126 exceeds 50%, the prompt generation unit 132 may generate the prompt 123 based on different topic information 121. For example, even if the prompt generation unit 132 generates the prompt 123 using weather as the topic information 121 at the start of the ice-breaking phase, the prompt generation unit 132 may generate the prompt 123 using information about the interviewee as the topic information 121 midway through the ice-breaking phase. With this configuration, the interview system 1000 can execute the ice-breaking phase with the AI avatar without boring the interviewee. Therefore, the interview system 1000 allows the interviewee to start the interview in a state where they are more familiar with interviews with the AI avatar.
[0113] The prompt generation unit 132 may be configured to generate the prompt 123 using multiple pieces of topic information 121. For example, the prompt generation unit 132 can generate a prompt 123 related to weather information related to the interviewee's address by using the weather and information (address) about the interviewee as the topic information 121. With this configuration, the interview system 1000 can execute the ice-breaking phase with a topic more relevant to the interviewee. Therefore, the interview system 1000 allows the interviewee to start the interview in a state where they are more familiar with interviews with the AI avatar.
[0114] <Additional Notes> The present embodiment includes the following disclosure.
[0115] (Appendix 1) an acquisition unit that acquires an interviewee's utterance from the interviewee's terminal device; a prompt generation unit that generates a prompt that instructs the user to create a response to the utterance based on the utterance and topic information that indicates a topic that was discussed before the interview with the interviewer; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition of the conversation is satisfied; An interview system that includes:
[0116] (Appendix 2) the acquisition unit acquires the topic information when communication with the terminal device is started; the prompt generation unit generates a prompt that instructs the interviewer to create an initial conversation to be conveyed to the interviewer based on the topic information. The interview system described in Appendix 1.
[0117] (Appendix 3) the prompt generation unit generates the prompt by inputting the utterance and the topic information into predetermined template information. The interview system described in Appendix 1.
[0118] (Appendix 4) the template information includes avoiding inappropriate topics for the interviewee; The interview system described in Appendix 3.
[0119] (Appendix 5) The topic information includes one or more of the following: weather, a calendar, information about the interviewee, and information about past interviews conducted with the interviewee; The interview system described in Appendix 1.
[0120] (Appendix 6) When the termination condition is satisfied, the interview unit conducts an interview with the interviewer. The interview system described in Appendix 1.
[0121] (Appendix 7) The termination condition is satisfied when the number of interactions with the interviewer exceeds a predetermined number. The interview system described in Appendix 1.
[0122] (Appendix 8) The termination condition is met when the duration of the conversation with the interviewer exceeds a predetermined time. The interview system described in Appendix 1.
[0123] (Appendix 9) The termination condition is satisfied when the number of characters processed in the dialogue with the interviewer exceeds a predetermined number of characters. The interview system described in Appendix 1.
[0124] (Appendix 10) The system further includes an emotion estimation unit that estimates an emotional state of the interviewee based on information acquired from the terminal device of the interviewee, The termination condition is satisfied when the emotional state is determined to be an interviewable state. The interview system described in Appendix 1.
[0125] (Appendix 11) the prompt generation unit generates the prompt by further using a past utterance of the interviewee and a response to the past utterance. The interview system according to claim 1.
[0126] (Appendix 12) an acquisition unit that acquires an interviewee's utterance from the interviewee's terminal device; a prompt generation unit that generates a prompt that instructs the user to create a response to the utterance based on the utterance and topic information that indicates a topic that was discussed before the interview with the interviewer; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition of the conversation is satisfied; An interview device comprising:
[0127] (Appendix 13) An interview method executed by an interview system, an acquisition step of acquiring an utterance of an interviewee from a terminal device of the interviewee; a prompt generating step of generating a prompt that instructs the user to create a response to the utterance based on the utterance and topic information indicating a topic that has not yet been discussed with the interviewer; an output step of inputting the prompt into a machine learning model and outputting the response to the terminal device; an interview step of repeating the obtaining step, the prompt generating step, the output step, and the response step until the conversation termination condition is met; An interview method that includes the following.
[0128] (Appendix 14) The interview system an acquisition unit that acquires an interviewee's utterance from the interviewee's terminal device; a prompt generation unit that generates a prompt that instructs the user to create a response to the utterance based on the utterance and topic information that indicates a topic that was discussed before the interview with the interviewer; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition of the conversation is satisfied; An interview program that helps you implement the following.
[0129] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]
[0130] 1: Interview device 2: Terminal device 3: Language model device 131,231,331: Acquisition department 132: Interview Department 133: Prompt generation unit 134: Output section 135: Emotion estimation part 232: Application section 332: Conversation generation unit
Claims
1. an acquisition unit that acquires, from the terminal device of the interviewer, an utterance of the interviewer in response to a conversation by the AI avatar during an ice-breaking phase between the interviewer and the AI avatar displayed on the terminal device of the interviewer before the interview between the interviewer and the AI avatar; a prompt generation unit that generates a prompt that instructs the user to create a response to the utterance based on the utterance and topic information that indicates a topic that was discussed before the interview with the interviewer; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition of the conversation is satisfied; Equipped with When the termination condition is satisfied, the interview section transitions from the ice-breaking phase to an interview phase between the AI avatar and the interviewer. Interview system.
2. the acquisition unit acquires the topic information when communication with the terminal device is started; the prompt generation unit generates a prompt that instructs the interviewer to create an initial conversation to be conveyed to the interviewer based on the topic information. The interview system according to claim 1 .
3. the prompt generation unit generates the prompt by inputting the utterance and the topic information into predetermined template information. The interview system according to claim 1 .
4. the template information includes avoiding inappropriate topics for the interviewee; The interview system according to claim 3 .
5. The topic information includes one or more of the weather, a calendar, information about the interviewee, and information about past interviews conducted with the interviewee. The interview system according to claim 1 .
6. The termination condition is satisfied when the number of interactions with the interviewer exceeds a predetermined number. The interview system according to claim 1 .
7. The termination condition is met when the duration of the conversation with the interviewer exceeds a predetermined time. The interview system according to claim 1 .
8. The termination condition is satisfied when the number of characters processed in the dialogue with the interviewer exceeds a predetermined number of characters. The interview system according to claim 1 .
9. The system further includes an emotion estimation unit that estimates an emotional state of the interviewee based on information acquired from the terminal device of the interviewee, The termination condition is satisfied when the emotional state is determined to be an interviewable state. The interview system according to claim 1 .
10. the prompt generation unit generates the prompt by further using a past utterance of the interviewee and a response to the past utterance. The interview system according to claim 1 .
11. an acquisition unit that acquires, from the terminal device of the interviewer, an utterance of the interviewer in response to a conversation by the AI avatar during an ice-breaking phase between the interviewer and the AI avatar displayed on the terminal device of the interviewer before the interview between the interviewer and the AI avatar; a prompt generation unit that generates a prompt that instructs the user to create a response to the utterance based on the utterance and topic information that indicates a topic that was discussed before the interview with the interviewer; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; an interview unit that operates the acquisition unit, the prompt generation unit, and the output unit until a termination condition of the conversation is satisfied; Equipped with When the termination condition is satisfied, the interview section transitions from the ice-breaking phase to an interview phase between the AI avatar and the interviewer. An interview device comprising:
12. An interview method executed by an interview system, an acquisition step of acquiring, from the terminal device of the interviewer, an utterance of the interviewer in response to a conversation by the AI avatar during an ice-breaking phase between the interviewer and the AI avatar displayed on the terminal device of the interviewer, before the interview between the interviewer and the AI avatar; a prompt generating step of generating a prompt that instructs the user to create a response to the utterance based on the utterance and topic information indicating a topic that has not yet been discussed with the interviewer; an output step of inputting the prompt into a machine learning model and outputting the response to the terminal device; The acquiring step, the prompt generating step, and the like are repeated until the conversation termination condition is met. an interview step of repeating the step and the output step; Equipped with In the interview step, when the termination condition is satisfied, the ice-breaking phase is transitioned to an interview phase between the AI avatar and the interviewer. Interview method.
13. The interview system an acquisition unit that acquires, from the terminal device of the interviewer, an utterance of the interviewer in response to a conversation by the AI avatar during an ice-breaking phase between the interviewer and the AI avatar displayed on the terminal device of the interviewer before the interview between the interviewer and the AI avatar; Based on the utterance and topic information indicating a topic before the interview with the interviewer, a prompt generation unit that generates a prompt that instructs the user to create a response to an utterance; an output unit that inputs the prompt into a machine learning model and outputs the response to the terminal device; The acquiring unit, the prompt generating unit and the output unit operate in unison until the conversation termination condition is met. an interview unit that activates the power unit; Execute When the termination condition is satisfied, the interview section transitions from the ice-breaking phase to an interview phase between the AI avatar and the interviewer. Interview program.
Citation Information
Patent Citations
Interview evaluation method and system based on artificial intelligence
CN117236911A
Resume screening method and related equipment
CN119850161A
Interface testing method and device and storage medium
CN119988235A
Data processing apparatus, data processing method, and program
JP2025015327A
System
JP2025047423A