System
The AI-powered recruitment system addresses high costs and mismatches by automating interviews, generating questions, analyzing responses, and providing feedback, enhancing the efficiency and accuracy of talent selection.
Patent Information
- Application Number
- JP2024123851
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Companies face high recruitment costs and mismatches due to the inefficiencies in processing large numbers of application forms and conducting multiple interviews, which are subjective and time-consuming, leading to early turnover and unnecessary expenses.
A system that uses AI interviewers to acquire entry information, generate questions, conduct interviews via video communication, analyze answers in real-time, and summarize results to create feedback materials, while also analyzing applicants' nervousness to support a natural interview state.
This system reduces recruitment costs and prevents mismatches by efficiently processing applicant information, ensuring the selection of the right talent through streamlined interviews and feedback materials.
Smart Images

Figure 2026022334000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When companies are recruiting, they need to read through a large number of application forms and conduct multiple interviews, which creates a costly problem. Another issue is that mismatches during recruitment can lead to early turnover, resulting in further unnecessary costs. To solve these problems, a method is needed to reduce recruitment costs and ensure the selection of the right talent. [Means for solving the problem]
[0005] To solve this problem, the present invention provides the following measures. First, it provides a means for acquiring entry information. Next, it provides a means for generating questions based on the acquired entry information. Furthermore, it provides a means for asking questions to applicants via video communication and obtaining their answers. It also includes a means for analyzing the obtained answers in real time and generating follow-up questions as needed. Finally, it provides a means for summarizing the interview results and creating feedback materials, thereby preventing mismatches and reducing costs in recruitment activities. It also includes a means for analyzing the applicant's level of nervousness during video communication and speaking to them accordingly, thereby supporting applicants so that they can be interviewed in a natural state. It also includes a means for storing entry information in a database and providing feedback materials to a human resources system based on the analysis, thereby supporting efficient human resources operations.
[0006] "Entry information" refers to information submitted by applicants, such as resumes, job history, and reasons for applying.
[0007] "Means for capturing entry information" refers to a system or method for capturing entry information submitted by applicants through the online recruitment portal.
[0008] "Means for generating questions" refers to a system or method for automatically generating questions to be used in an interview based on acquired entry information.
[0009] "Means for asking questions to applicants and receiving answers via video communication" refers to a system or method for asking questions to applicants using an online video calling tool and receiving answers in real time.
[0010] "Means for analyzing responses in real time and generating follow-up questions" refers to a system or method for analyzing applicant responses on the fly and generating follow-up questions as needed.
[0011] "Means for summarizing interview results and preparing feedback materials" refers to a system or method for summarizing the content and results of interviews and preparing feedback materials in a form that can be checked by human resources personnel.
[0012] "Means for analyzing the applicant's level of tension and speaking to them accordingly" refers to a system or method for analyzing the applicant's voice and facial expressions to estimate their level of tension and speaking to them at the appropriate time to relax them.
[0013] "Means for storing interview data in a database and providing feedback materials to a human resources system based on analysis" refers to a system or method for storing interview data in a database, creating feedback materials based on the analysis results, and using them by human resources personnel. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews during corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results, and creating feedback materials.
[0036] A natural language description of the program's operation
[0037] 1. Acquisition of entry information
[0038] A user (applicant) submits an application form to an online recruitment portal. This information is stored in a database by the server. As a concrete example, consider a scenario in which an applicant uploads a resume, a curriculum vitae, and a statement of reasons for applying. The server receives and stores this application information.
[0039] 2. Question Generation
[0040] The server uses an AI model to generate questions for applicants based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0041] 3. Conducting the interview
[0042] The user selects the interview date and time through an online recruitment portal. At the designated date and time, the user accesses a video call tool such as Zoom through their device (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions.
[0043] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0044] 4. Real-time analysis of responses
[0045] As users answer questions, the server analyzes the answers in real time. For example, if the server determines that an applicant has a deep understanding of a particular technology, it will generate follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0046] 5. Analyzing applicants' tension levels and encouraging them
[0047] The server analyzes the applicant's tone of voice and facial expressions, and if it determines that the applicant is nervous, it will use the AI interviewer to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0048] 6. Summarizing interview results and creating feedback materials
[0049] After the interview is over, the server summarizes the answers given during the interview and creates a feedback document that summarizes the interview results, such as "The candidate has strong technical skills, but lacks details regarding team leadership experience."
[0050] The feedback materials include questions to be asked in the next interview, evaluations of the applicant, concerns, etc. The server stores the generated feedback materials in the HR system and makes them accessible on a terminal (such as the HR person's PC).
[0051] Specific examples
[0052] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates in the interview using Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time and generates further relevant follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0053] In this way, the present invention realizes effective and efficient interviews in companies' recruitment activities, and contributes greatly to reducing recruitment costs and selecting appropriate personnel.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0057] Step 2:
[0058] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0059] Step 3:
[0060] The server analyzes the application information retrieved from the database and generates individual questions using an AI model. For example, if an applicant's academic background is computer science, it generates a question such as, "Tell us about the most interesting technology you have studied."
[0061] Step 4:
[0062] The server sends the generated question list to the terminal (interviewer AI device), where the question list is organized into a format that can be used by the video chat tool.
[0063] Step 5:
[0064] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0065] Step 6:
[0066] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0067] Step 7:
[0068] The server asks the user questions through the AI interviewer based on the submitted question list. For example, the AI interviewer might ask, "Please tell us more about the projects listed in your resume."
[0069] Step 8:
[0070] As users answer questions, the server analyzes the answers in real time using natural language processing techniques.
[0071] Step 9:
[0072] The server generates follow-up questions based on the answers and asks the user additional questions. For example, if the user explains about the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0073] Step 10:
[0074] During the interview, the server analyzes the user's tone of voice and facial expressions, and if it determines that the user is nervous, it will say something like, "Please relax. Please answer calmly" through the AI interviewer.
[0075] Step 11:
[0076] After the interview, the server analyzes the collected response data again and summarizes the interview results, including the applicant's strengths, areas for improvement, and evaluation.
[0077] Step 12:
[0078] The server creates feedback materials based on the summarized interview results and saves them in the personnel system. These materials include questions to be asked in the next interview and evaluation details. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0079] Step 13:
[0080] All interview data is collected by the server and used to improve the accuracy of the AI model, enabling more efficient and accurate question generation and analysis for subsequent interviews.
[0081] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] In corporate recruitment activities, in order to streamline initial interviews and select appropriate candidates, a system is needed that can efficiently process large amounts of applicant information and improve the quality and progress of interviews. However, conventional methods require interviewers to directly interact with the interviewer, which is costly and time-consuming, and is largely dependent on the interviewer's subjective opinion, making it difficult to properly evaluate the interviewer. In addition, there is a lack of ways to appropriately alleviate applicant tension, making it difficult to bring out the true potential of the applicant.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and receiving their answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for generating questions and follow-up questions using a generative AI model, and means for analyzing the applicant's level of nervousness during video communication, thereby making it possible to improve the efficiency of initial interviews and select appropriate candidates.
[0087] "Entry information" refers to information such as a resume, job history, and reason for applying that an applicant submits to a company's recruitment portal.
[0088] "Means for generating questions" refers to a function for creating questions for applicants based on stored entry information.
[0089] "Video communication" is a technology that allows real-time video and audio communication over the Internet and is used in the interview process.
[0090] The "means for analyzing responses in real time" refers to technology that has the function of analyzing applicants' responses in real time, evaluating their content, and generating appropriate follow-up questions.
[0091] "Means for generating follow-up questions" refers to the function of creating additional, detailed questions based on the applicant's answers.
[0092] "Means for summarizing interview results and creating feedback materials" refers to a function for summarizing interview results based on answers given during the interview and creating feedback materials to be provided to human resources personnel.
[0093] "Generative AI model" refers to an artificial intelligence model used to generate questions and follow-up questions.
[0094] The "means for analyzing the level of tension" is a technology that has the function of determining and analyzing the level of tension based on the applicant's voice and video data during video communication.
[0095] This invention is a system that uses AI interviewers to efficiently conduct early-stage interviews (first-stage interviews) in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, and summarizing interview results and creating feedback materials.
[0096] Obtaining entry information
[0097] A user (applicant) submits an application form to an online recruitment portal. This application form includes a resume, a curriculum vitae, and a statement of reasons for applying. This information is received by the server and stored in a database. For example, if a user submits an application form for a "software engineer," the server records information about Applicant A in the database based on this information.
[0098] Question Generation
[0099] The server uses the generative AI model to generate questions for applicants based on the saved entry information, such as "Please tell us about the most difficult experience you had in a past project," and sends these to the terminal (interviewer AI device).
[0100] Conducting interviews
[0101] The user selects the interview date and time through an online recruitment portal. At the specified date and time, the user accesses a video call tool such as Zoom through their terminal (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions. For example, the applicant clicks on the Zoom link at the specified time, and the video call begins. The AI interviewer (server) asks the question, "Tell me about your career history."
[0102] Real-time analysis of responses
[0103] When a user answers a question, the audio and video data are sent to a server in real time. The server then uses an AI analysis model to analyze the answers and evaluate the applicant's skills and experience. For example, if the analysis determines that the applicant has a deep understanding of a particular technology, the server generates questions that dig deeper. It then asks follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0104] Analyzing applicants' tension levels and encouraging them
[0105] The server analyzes the audio and video during the video call to determine the applicant's level of nervousness in real time. If it determines that the applicant is nervous, the server instructs the interviewer AI device to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0106] Summarizing interview results and creating feedback materials
[0107] After the interview is over, the server summarizes the answers given during the interview and creates feedback materials. These materials include an evaluation of the applicant, points to ask in the next interview, and concerns. For example, an evaluation point such as "Applicant A is highly evaluated for his / her strong technical skills and extensive leadership experience" may be written. The feedback materials are saved in the HR system and can be accessed from a terminal (such as the HR manager's PC).
[0108] Specific examples
[0109] When Applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. For example, the applicant participates in an interview via Zoom at a specified time, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answers in real time and generates relevant follow-up questions. After the interview, the data obtained can be used to create feedback materials that HR personnel can use to prepare the next interview.
[0110] Prompt Sentence Examples
[0111] "Tell me about your career so far."
[0112] "What is the most challenging experience you've had on a recent project?"
[0113] "Can you please tell us a bit more about a specific project where you used that technology?"
[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0115] Step 1: Obtaining entry information
[0116] Users access an online recruitment portal and submit their application.
[0117] Specifically, the user enters their resume, job history, and reason for applying into an online form and clicks the submit button.
[0118] Input: Application information such as resume, work history, and reasons for applying
[0119] Output: Entry information stored in the database
[0120] The server receives this information and stores it in a database.
[0121] Step 2: Generate questions
[0122] The server analyzes the entry information stored in the database.
[0123] Specifically, the server uses a generative AI model to receive entry information as input and generate questions based on that information.
[0124] Input: Entry information stored in the database
[0125] Output: Questions to ask the applicant
[0126] For example, a question such as "Please tell us about the most difficult experience you had in a past project" is generated. The generated question is sent to the device.
[0127] Step 3: Set up an interview
[0128] Users select interview dates and times through an online recruitment portal.
[0129] Specifically, the user uses the calendar function to select an available time slot and confirm the reservation.
[0130] Input: Desired interview date and time and contact information
[0131] Output: Notification of confirmed interview schedule
[0132] The server receives this information and confirms the interview schedule.
[0133] Step 4: Conducting the interview
[0134] At the specified date and time, users access video calling tools such as Zoom using their device (PC or mobile device).
[0135] The server detects that a video call has started and begins the interview based on pre-generated questions.
[0136] Input: Access to video call tools such as Zoom, generated questions
[0137] Output: Conducting an interview via video call
[0138] For example, a user clicks on a Zoom link at a specified time to start a video call. The AI interviewer (server) asks, "Tell me about your career history."
[0139] Step 5: Real-time analysis of responses
[0140] As the user answers the questions, audio and video data is transmitted in real time to the server.
[0141] The server analyzes the data using an AI analysis model and evaluates the answers.
[0142] Input: Applicant's audio and video data
[0143] Output: Analysis results and new follow-up questions
[0144] For example, if the analysis reveals that you have a deep understanding of a particular technology, the server will generate a follow-up question such as, "Please tell us more about a specific project using that technology."
[0145] Step 6: Analyze the applicant's level of nervousness and encourage them
[0146] The server analyzes the audio and video during the video call in real time to assess the applicant's level of nervousness.
[0147] Input: Audio and video data during a video call
[0148] Output: Tension analysis results and verbal instructions
[0149] If the server determines that the applicant is nervous, it instructs the AI interviewer device to say something like, "Please relax." Specifically, the AI interviewer uses a voice message to encourage the applicant to relax.
[0150] Step 7: Summarize the interview results and create feedback materials
[0151] After the interview is completed, the server summarizes the answers given during the interview and creates feedback materials.
[0152] Input: Answers given during the interview and analysis results
[0153] Output: Feedback materials and evaluation points for the next interview
[0154] For example, the server can create a summary such as "The applicant has strong technical skills, but lacks details regarding leadership experience" and include it in the feedback material, which is then stored in the HR system and used by HR personnel for the next interview.
[0155] (Application example 1)
[0156] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0157] Until now, human resource costs and time have been major issues in corporate recruitment activities and factory performance evaluations. Furthermore, because the quality of the evaluation depends on human subjectivity, there is a risk of mismatches and unfair evaluations. Furthermore, creating training programs to efficiently improve work performance requires a great deal of effort. There is a need to solve these issues and provide a more efficient and fair evaluation system and training programs suited to each individual worker.
[0158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0159] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants and workers via video communication and receiving answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing interview results and performance evaluation results to create feedback materials, and means for evaluating performance and providing training programs. This makes it possible to efficiently evaluate applicants and workers, achieve fair and error-free evaluations, and provide optimal training programs based on their individual characteristics.
[0160] "Entry information" refers to information such as resumes, work history, and work experience provided by applicants and workers, and is data necessary for evaluation.
[0161] "Means for generating questions" refers to a function that automatically creates questions appropriate for applicants and workers based on entry information and analysis results.
[0162] "Video communication" is a technology that allows for real-time video and audio communication over the Internet, and is used for interviews and job evaluations.
[0163] "Means for obtaining responses" refers to a function for receiving responses from applicants and workers via video communication.
[0164] "Real-time analytical means" refers to the analytical capabilities required to instantly evaluate applicant and worker responses and generate follow-up questions as needed.
[0165] "Means for generating follow-up questions" refers to a function that automatically creates more detailed questions based on the initial answers.
[0166] "Means for summarizing interview results and performance evaluation results to create feedback materials" refers to a device that has the function of summarizing information obtained from interviews and performance evaluations and providing it in a concise, easy-to-understand format.
[0167] "Means for evaluating work performance and providing training programs" refers to a function that evaluates a worker's ability to perform work and assigns an appropriate training program based on the results.
[0168] The present invention is a system for evaluating the work performance of factory workers and providing training programs based on the evaluation results. The main technical elements of this system include obtaining entry information, generating questions, conducting evaluation sessions via video communication, analyzing responses in real time, and creating feedback materials based on the evaluation results.
[0169] Obtaining entry information
[0170] First, the server obtains the application information from the worker through the portal system in the factory. This application information includes the worker's resume, work history, work experience, etc., and this data is stored in a database.
[0171] Question Generation
[0172] Based on the acquired entry information, the server uses an AI model to generate questions for the worker. For example, it generates questions such as "What was the most difficult task in your recent work?" based on the worker's past work history. The generated questions are sent to the terminal.
[0173] Conducting a business evaluation session
[0174] At the designated time, workers access a video chat tool such as Zoom via their device (PC or tablet). The server detects that the video call has started and begins the evaluation session based on pre-generated questions.
[0175] Real-time analysis of responses
[0176] As workers answer questions, the server analyzes their answers in real time, assessing their understanding of specific work processes based on their answers, and generates follow-up questions as needed, such as, "Is there anything in that work process that you feel needs improvement?"
[0177] Analyzing tension levels and encouraging others
[0178] The server analyzes the worker's tone of voice and facial expressions, and if they appear nervous, it will say something like, "Relax." This allows the worker to be evaluated in a natural state.
[0179] Summarizing the assessment results and preparing training materials
[0180] Once the evaluation session is over, the server summarizes the responses and compiles the evaluation results into a feedback document, which includes an evaluation of work performance, areas for improvement, and the next training program to be implemented. For example, a summary might be created such as, "Work speed is high, but quality control needs improvement."
[0181] The hardware used includes servers, databases, PCs, and tablets, while the software includes a portal system, Zoom API, voice analysis engines (e.g., Google Speech-to-Text), natural language processing models (e.g., GPT-3), and facial expression analysis software (e.g., OpenCV).
[0182] Specific examples
[0183] When Worker A submits application information (such as resume, work history, and work experience) to the online portal, the server generates appropriate questions based on this information. For example, it asks questions such as, "What was the most difficult task you recently performed?" or "Is there anything in the work process that you feel needs improvement?" Worker A participates in the evaluation session using a video chat tool, and analysis is carried out in real time. Afterwards, feedback materials summarizing the evaluation results are created, and a training program is provided.
[0184] Prompt Sentence Examples
[0185] "What was the most difficult task you recently undertook? Please explain in detail why."
[0186] "Are there any particular work processes you feel need improvement?"
[0187] "Please tell us about any major problems you have faced in the past and how you dealt with them."
[0188] In this way, the present invention improves the efficiency of work evaluation and training support within a factory, and contributes greatly to improving work performance and preventing mismatches.
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1:
[0191] The server obtains application information from users (workers) through the portal system. Users fill out a self-evaluation sheet and provide details of their resume, work history, and work experience. This application information is stored in a database by the server. The input is the application information provided by the user, and the output is the application information stored in the database.
[0192] Step 2:
[0193] The server uses the AI model with a question generation means to generate appropriate questions for the worker based on the stored entry information. The data input into the AI model includes the worker's resume, work history, and work experience, and specific questions are generated as output. For example, a question such as "What was the most difficult task in your recent work?" is generated.
[0194] Step 3:
[0195] The user accesses a video chat tool such as Zoom through a device (PC or tablet) at a specified date and time. When the server detects the start of the video call, it starts an evaluation session based on the generated questions. The input is access to the video chat tool, and the output is the start of the session. Specifically, when the user clicks the Zoom link and the video call starts, the server begins displaying the questions.
[0196] Step 4:
[0197] As users answer questions, the server analyzes the answers in real time. Specifically, a speech analysis engine (e.g., Google Speech-to-Text) is used to convert the speech data into text. A natural language processing model (e.g., GPT-3) then analyzes the text data to understand the answer. Based on this analysis, follow-up questions are generated as needed. The input is the user's speech answer, and the output is the analyzed text and follow-up questions.
[0198] Step 5:
[0199] The server analyzes the user's tone of voice and facial expression to determine the level of tension. It uses voice analysis software and facial expression analysis software (e.g., OpenCV) to detect whether the user is nervous. If necessary, it will say something like, "Please relax." The input is the user's tone of voice and facial expression data, and the output is the analysis results and a message to encourage the user.
[0200] Step 6:
[0201] After the evaluation session is over, the server summarizes the responses and compiles the evaluation results as feedback materials. It uses a text summarization engine (e.g., BERT) to extract key points and create concise feedback materials. It then evaluates the user's performance based on these materials and provides training programs. The input is the user's complete response data, and the output is the summarized feedback materials and training programs.
[0202] This will enable an efficient and fair evaluation system and training programs based on individual characteristics.
[0203] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0204] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results and creating feedback materials, and an emotion engine that recognizes user emotions.
[0205] A natural language description of the program's operation
[0206] 1. Acquisition of entry information
[0207] A user (applicant) submits an application to an online recruitment portal, including a resume, a job history, and a motivation statement. The server receives this application information and stores it in a database.
[0208] 2. Question Generation
[0209] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0210] 3. Conducting the interview
[0211] The user then accesses the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user connects to a video call tool such as Zoom. The server detects the start of the video call and begins the interview.
[0212] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0213] 4. Real-time analysis of responses
[0214] When a user answers a question, the server analyzes the answer in real time using natural language processing technology. The server also includes an emotion engine that analyzes the user's voice and facial expressions during the answer to collect emotional data.
[0215] 5. Generate follow-up questions
[0216] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0217] 6. Analyzing applicants' tension levels and encouraging them
[0218] If the server determines that the user is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly."
[0219] 7. Summarizing interview results and creating feedback materials
[0220] After the interview is over, the server analyzes the interview responses and emotional data again to summarize the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0221] For example, if the analysis indicates that the applicant is technically strong but has concerns about leadership questions, include that information in the summary.
[0222] 8. Providing Feedback Materials
[0223] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0224] Specific examples
[0225] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates via Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time, using an emotion engine to detect nervousness from the applicant's facial expressions and tone of voice, and generates appropriate follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0226] In this way, by combining emotion engines, companies can provide a more refined and personalized interview experience, enabling them to conduct their recruitment activities more efficiently and effectively.
[0227] The processing flow will be explained below.
[0228] Step 1:
[0229] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0230] Step 2:
[0231] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0232] Step 3:
[0233] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0234] Step 4:
[0235] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0236] Step 5:
[0237] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0238] Step 6:
[0239] The server asks the user questions through the AI interviewer based on the sent question list. For example, the AI interviewer might ask, "Tell me about your career so far."
[0240] Step 7:
[0241] As users answer questions, the server analyzes the answers in real time using natural language processing technology, and the resulting data is instantly evaluated by an emotion engine.
[0242] Step 8:
[0243] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0244] Step 9:
[0245] During the interview, if the server determines that the applicant is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say things like, "Please relax. Please answer calmly." This helps to stabilize the applicant's mental state.
[0246] Step 10:
[0247] After the interview, the server analyzes the collected response data and emotional data again to summarize the interview results. This summary includes the applicant's strengths, weaknesses, and emotional reactions. For example, it may include information such as, "The applicant is technically strong, but is anxious about questions regarding leadership."
[0248] Step 11:
[0249] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0250] Step 12:
[0251] All interview data is collected by the server and used to improve the accuracy of the AI model and emotion engine, enabling more efficient and accurate question generation and analysis for future interviews.
[0252] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0253] Example 2
[0254] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0255] In traditional recruitment activities, a large amount of resources are allocated to conducting initial interviews (stage zero interviews), which often results in mismatches between candidates and candidates. This problem is particularly pronounced in companies conducting large-scale recruitment activities, increasing the burden on human resources personnel and contributing to rising recruitment costs. Furthermore, traditional interview methods make it difficult to accurately grasp the applicant's emotions and level of nervousness in a short amount of time, which makes it difficult for applicants to relax and have an opportunity to present themselves. To solve these problems, a more efficient and objective interview method is required.
[0256] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for acquiring entry information, a means for generating questions using a generative AI model based on the acquired entry information, a means for asking questions to applicants via video communication and obtaining their answers, a means for analyzing the answers in real time and generating follow-up questions using natural language processing technology and an emotion engine, and a means for summarizing the interview results and creating feedback materials. This enables efficient and effective initial interviews, reduces recruitment costs, and prevents mismatches with applicants. Furthermore, by analyzing the applicant's level of nervousness using the emotion engine during video communication and providing appropriate prompts, the applicant can be given an opportunity to relax and promote themselves.
[0257] "Entry information" refers to information submitted by applicants, such as resumes, work history, and reasons for applying.
[0258] "Generative AI model" refers to an artificial intelligence model that uses natural language processing technology to generate individual questions from the input entry information.
[0259] "Server" refers to a central processing unit that performs processes such as obtaining entry information, generating questions, analyzing answers, generating follow-up questions, summarizing interview results, and creating feedback materials.
[0260] "Video communication" refers to a means of communication that uses video calling tools such as Zoom to share audio and video in real time with applicants in remote locations.
[0261] "Natural language processing technology" refers to the technology of analyzing and processing human language using a computer.
[0262] "Emotion engine" refers to a system that analyzes an applicant's emotional state from their tone of voice and facial expressions.
[0263] "Follow-up questions" refer to additional questions that are generated based on the applicant's initial answers or emotional state.
[0264] "Feedback materials" are documents summarizing the results of an interview, including questions for the next interview, evaluations of the applicant, and concerns.
[0265] This invention is a system for efficiently conducting initial interviews using AI interviewers in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The system includes functions for acquiring application information, generating questions, conducting interviews via video communication, analyzing responses in real time, generating follow-up questions, summarizing interview results and creating feedback materials, and an emotion engine for recognizing applicants' emotions.
[0266] Obtaining entry information
[0267] A user (applicant) first accesses an online recruitment portal and submits an application form, which includes a resume, a curriculum vitae, and a statement of reasons for applying. The server receives this application information through an HTTP request and performs the necessary database transactions to store it in a database. Specific software used includes database management systems such as MySQL or PostgreSQL.
[0268] Question Generation
[0269] The server retrieves the user's application information from the database and uses a Python script to call an open-source natural language processing library (e.g., Transformers). A prompt sentence is input into the generative AI model to generate individual questions. The generated questions are sent from the server to the terminal (interviewer AI device) in JSON format. An example of a specific prompt sentence is, "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please use an open-source natural language processing library to generate individual, specific questions."
[0270] Conducting interviews
[0271] The user logs in again to the online recruitment portal and selects the date and time of the interview. At the specified date and time, the user clicks the Zoom link to connect to the video call. The server monitors the connection status through the video call tool's API and detects the user's connection. At the start of the interview, the server sends pre-generated questions to the device, which then conducts the assessment. Specific video call tools used include Zoom and WebEx.
[0272] Real-time analysis of responses
[0273] When a user answers a question, the audio data of the answer is sent to the server in real time via the video chat tool's API. The server converts the audio data into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The converted text is then analyzed using a natural language processing library. Furthermore, an emotion engine analyzes the user's tone of voice and facial expressions to collect emotional data.
[0274] Generate follow-up questions
[0275] The server generates follow-up questions based on the collected emotional data and the content of the answers. For example, if the user is nervous about a particular topic, the server generates follow-up questions such as, "What were the specific benefits of this new technology?" This is again done using a generative AI model.
[0276] Analyzing applicants' tension levels and encouraging them
[0277] If the server determines that the user is nervous based on the data analyzed by the emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly." Specifically, it uses the results of transcription and emotion analysis to display appropriate instructions on the screen and play them back aloud.
[0278] Summarizing interview results and creating feedback materials
[0279] After the interview, the server analyzes the interview responses and emotional data again. For example, it uses a text analysis engine to extract specific keywords and emotional states, and then summarizes the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0280] Providing feedback materials
[0281] The server generates feedback materials in PDF format based on the summarized interview results. The PDF file is then saved in the HR system and a notification is sent. HR personnel can access the materials from their own devices (such as PCs) and prepare for the next interview. This function enables companies to conduct their recruitment activities efficiently and effectively.
[0282] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0283] Step 1: Obtaining entry information
[0284] A user (applicant) accesses an online recruitment portal and submits an application form. The application form includes information such as a resume, a curriculum vitae, and a reason for applying. The submitted application information is sent to the server as an HTTP request. The server receives the application information and stores it in a database using a database management system (e.g., MySQL). Specifically, the operation involves executing an INSERT statement against the database. The input in this process is the application information, and the output is the application information stored in the database.
[0285] Step 2: Generate questions
[0286] The server retrieves the saved application information from the database. Based on the retrieved information, questions are generated using a generative AI model (e.g., Transformers). This process is carried out by executing a Python script and calling a natural language processing library. An example prompt is: "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please generate individual, specific questions using an open-source natural language processing library." The generated questions are sent from the server to the terminal in JSON format. The input in this process is the application information, and the output is the generated questions.
[0287] Step 3: Conducting the interview
[0288] The user logs in to the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user clicks on a link in a video calling tool such as Zoom to connect to the video call. The server monitors the connection status through the video calling tool's API and detects the user's connection. When the interview starts, the server sends pre-generated questions to the terminal and conducts the questions. The input is the interview date and time and the questions, and the output is the video call that has started.
[0289] Step 4: Real-time analysis of responses
[0290] When a user answers a question via video call, the audio of the answer is sent to the server in real time via the video call tool's API. The server then converts the audio data of the answer into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The server then analyzes the answer using a natural language processing library (e.g., Transformers), and the emotion engine uses this information to collect emotional data from the tone of voice and facial expressions. The input to this process is the audio data of the answer, and the output is the analyzed answer text and emotional data.
[0291] Step 5: Generate follow-up questions
[0292] The server generates follow-up questions based on the analyzed answer text and sentiment data. Using the generative AI model, it runs a Python script to generate appropriate follow-up questions. These follow-up questions are also sent to the device in JSON format. For example, if a user is nervous about a particular topic, a follow-up question might be generated: "What were the specific benefits of this new technology?" The input to this process is the analyzed answer text and sentiment data, and the output is a follow-up question.
[0293] Step 6: Analyze the applicant's level of nervousness and encourage them
[0294] The server analyzes the user's level of nervousness based on the data collected by the emotion engine. If it determines that the user is nervous, the server will use the AI interviewer to give appropriate encouragement. For example, it may generate encouragement such as "Please relax. Please answer calmly" and play it back aloud. The input in this process is emotional data, and the output is a encouragement message.
[0295] Step 7: Summarize the interview results and create feedback materials
[0296] After the interview is over, the server re-analyzes the responses and emotional data from the interview. For example, it can use a text analysis engine to extract specific keywords and emotional states and summarize evaluation points. Based on the analysis results, it summarizes the interview results and creates feedback materials in PDF format. The input for this process is the responses and emotional data, and the output is the summarized interview results and feedback materials.
[0297] Step 8: Provide feedback materials
[0298] The server generates feedback materials in PDF format based on the summarized interview results and saves them in the human resources system. The saved feedback materials are notified to human resources personnel so that they can access them from their own devices (such as PCs). The input in this process is the summarized interview results, and the output is the generated feedback materials.
[0299] This system allows companies to carry out their recruitment activities efficiently and effectively.
[0300] (Application example 2)
[0301] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0302] Conventional security systems have had difficulty analyzing visitors' behavior and emotions in real time and responding automatically. While early detection of suspicious individuals and rapid response are required, fully automating this process has been a difficult task. The present invention aims to solve these problems and realize more efficient and accurate security responses.
[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and obtaining answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for acquiring video and audio data and analyzing the visitor's emotions from this information, and means for generating follow-up actions based on the results of the emotion analysis. This makes it possible to analyze the visitor's behavior and emotions in real time and respond quickly and appropriately.
[0304] "Entry information" refers to information submitted by applicants, such as resumes, job history, and reasons for applying.
[0305] The "means for generating a question" is a function for generating an appropriate question based on the acquired entry information.
[0306] "Video communication" is a means of communication that uses video and audio to enable applicants and interviewers to interact online.
[0307] "Real-time analysis" refers to the process of instantly analyzing applicant responses and generating follow-up questions as needed.
[0308] The "means for generating follow-up questions" is a function for automatically generating follow-up questions in response to the applicant's answers.
[0309] The "means of summarizing interview results and creating feedback materials" is a function for summarizing the answers and emotional data given during the interview and creating feedback materials.
[0310] "Video and audio data" refers to visual and audio information obtained through cameras and microphones.
[0311] The "means for analyzing emotions" is a function for analyzing the emotions of visitors from the acquired video and audio data.
[0312] The "means for generating follow-up actions" is a function for automatically taking appropriate action against visitors based on the results of sentiment analysis.
[0313] To realize the present invention, the following hardware and software are used.
[0314] Hardware Configuration
[0315] 1. Security Guard Robot: A robot equipped with a camera and microphone to collect video and audio data (e.g., a typical security guard robot with a camera and microphone).
[0316] 2. Server: A high-performance computer for data processing and analysis (e.g., a typical server).
[0317] Software Configuration
[0318] 1. Natural language processing engine: Technology for analyzing information from captured video and audio data and generating appropriate questions and follow-up actions (e.g., OpenAI GPT-4, Google Cloud Dialogflow).
[0319] 2. Video and audio analysis engine: Technology for analyzing video and audio data and recognizing visitors' emotions (e.g., Microsoft Azure Cognitive Services, Google Cloud Video Intelligence API).
[0320] 3. Emotion recognition engine: Technology for analyzing emotions from collected data (e.g., Affectiva SDK, Microsoft Azure Emotion API).
[0321] Basic operations
[0322] 1. Security guard robots patrol monitored areas such as shopping malls and collect video and audio data in real time.
[0323] 2. The collected data is sent to a server and analyzed using a natural language processing engine and a video and audio analysis engine.
[0324] 3. The server analyzes the visitor's behavior and emotions to detect suspicious behavior and emotions such as tension, anxiety, and anger.
[0325] 4. The server utilizes an emotion recognition engine to generate appropriate follow-up actions and prompts.
[0326] 5. The generated instructions are sent to the security guard robot, which then takes appropriate action against the visitor.
[0327] Specific examples
[0328] For example, a security guard robot patrolling a shopping mall collects video and audio of visitor A. If the server analyzes that visitor A suddenly feels anxious, the robot will say, "Please relax. Is there anything I can help you with?" This information is also reported to the security center, and a human security guard will respond if necessary.
[0329] Prompt Sentence Examples
[0330] "Build a system for a security guard robot patrolling a shopping mall to collect real-time video and audio of visitors and analyze their emotions based on specific actions, facial expressions, and tone of voice. This involves the following steps:
[0331] 1. Video and audio data collection
[0332] 2. Real-time analysis using a sentiment analysis engine
[0333] 3. Take follow-up actions based on sentiment data
[0334] 4. Database recording and alert generation
[0335] In this way, by implementing the present invention, security operations can be automated and highly accurate crisis management can be achieved.
[0336] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0337] Step 1:
[0338] The security guard robot patrols the shopping mall and collects video and audio data in real time. The robot's cameras and microphones capture this data and send it to a server.
[0339] Input: Video and audio data from inside the shopping mall
[0340] Output: Video and audio data sent to the server
[0341] Step 2:
[0342] The server passes the received video and audio data to a natural language processing engine and a video and audio analysis engine for analysis, such as analyzing the visitor's behavior, speech content, tone of voice, and facial expressions to recognize their emotional state (tension, anxiety, anger, etc.).
[0343] Input: Video and audio data sent to the server
[0344] Output: Analyzed behavioral and emotional data
[0345] Step 3:
[0346] Based on the analysis results, the server uses an emotion recognition engine to understand the visitor's emotional state in detail, and in this process, specific emotions such as "tension," "anxiety," and "anger" are identified.
[0347] Input: Parsed behavioral and emotional data
[0348] Output: Detailed emotional state data
[0349] Step 4:
[0350] The server generates follow-up actions based on the detailed emotional state data. For example, if the visitor is feeling anxious, it generates a follow-up prompt such as "Please relax. Is there anything I can help you with?"
[0351] Input: Detailed emotional state data
[0352] Output: Generated follow-up action instructions
[0353] Step 5:
[0354] The generated follow-up action instructions are sent to the security guard robot, and the robot responds to the visitor accordingly. For example, the robot may tell the visitor to "relax."
[0355] Input: Generated follow-up action instructions
[0356] Output: Response action by security guard robot
[0357] Step 6:
[0358] The server records the results of follow-up actions and visitor responses in a database and stores relevant information, which can later be analyzed and used for further security measures.
[0359] Input: Follow-up action results and visitor response data
[0360] Output: Behavioral results and response data recorded in a database
[0361] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0363] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0364] [Second embodiment]
[0365] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0366] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0367] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0368] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0369] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0370] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0371] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0372] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0373] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0374] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0375] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0376] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0377] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews during corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results, and creating feedback materials.
[0378] A natural language description of the program's operation
[0379] 1. Acquisition of entry information
[0380] A user (applicant) submits an application form to an online recruitment portal. This information is stored in a database by the server. As a concrete example, consider a scenario in which an applicant uploads a resume, a curriculum vitae, and a statement of reasons for applying. The server receives and stores this application information.
[0381] 2. Question Generation
[0382] The server uses an AI model to generate questions for applicants based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0383] 3. Conducting the interview
[0384] The user selects the interview date and time through an online recruitment portal. At the designated date and time, the user accesses a video call tool such as Zoom through their device (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions.
[0385] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0386] 4. Real-time analysis of responses
[0387] As users answer questions, the server analyzes the answers in real time. For example, if the server determines that an applicant has a deep understanding of a particular technology, it will generate follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0388] 5. Analyzing applicants' tension levels and encouraging them
[0389] The server analyzes the applicant's tone of voice and facial expressions, and if it determines that the applicant is nervous, it will use the AI interviewer to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0390] 6. Summarizing interview results and creating feedback materials
[0391] After the interview is over, the server summarizes the answers given during the interview and creates a feedback document that summarizes the interview results, such as "The candidate has strong technical skills, but lacks details regarding team leadership experience."
[0392] The feedback materials include questions to be asked in the next interview, evaluations of the applicant, concerns, etc. The server stores the generated feedback materials in the HR system and makes them accessible on a terminal (such as the HR person's PC).
[0393] Specific examples
[0394] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates in the interview using Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time and generates further relevant follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0395] In this way, the present invention realizes effective and efficient interviews in companies' recruitment activities, and contributes greatly to reducing recruitment costs and selecting appropriate personnel.
[0396] The processing flow will be explained below.
[0397] Step 1:
[0398] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0399] Step 2:
[0400] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0401] Step 3:
[0402] The server analyzes the application information retrieved from the database and generates individual questions using an AI model. For example, if an applicant's academic background is computer science, it generates a question such as, "Tell us about the most interesting technology you have studied."
[0403] Step 4:
[0404] The server sends the generated question list to the terminal (interviewer AI device), where the question list is organized into a format that can be used by the video chat tool.
[0405] Step 5:
[0406] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0407] Step 6:
[0408] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0409] Step 7:
[0410] The server asks the user questions through the AI interviewer based on the submitted question list. For example, the AI interviewer might ask, "Please tell us more about the projects listed in your resume."
[0411] Step 8:
[0412] As users answer questions, the server analyzes the answers in real time using natural language processing techniques.
[0413] Step 9:
[0414] The server generates follow-up questions based on the answers and asks the user additional questions. For example, if the user explains about the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0415] Step 10:
[0416] During the interview, the server analyzes the user's tone of voice and facial expressions, and if it determines that the user is nervous, it will say something like, "Please relax. Please answer calmly" through the AI interviewer.
[0417] Step 11:
[0418] After the interview, the server analyzes the collected response data again and summarizes the interview results, including the applicant's strengths, areas for improvement, and evaluation.
[0419] Step 12:
[0420] The server creates feedback materials based on the summarized interview results and saves them in the personnel system. These materials include questions to be asked in the next interview and evaluation details. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0421] Step 13:
[0422] All interview data is collected by the server and used to improve the accuracy of the AI model, enabling more efficient and accurate question generation and analysis for subsequent interviews.
[0423] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0424] Example 1
[0425] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0426] In corporate recruitment activities, in order to streamline initial interviews and select appropriate candidates, a system is needed that can efficiently process large amounts of applicant information and improve the quality and progress of interviews. However, conventional methods require interviewers to directly interact with the interviewer, which is costly and time-consuming, and is largely dependent on the interviewer's subjective opinion, making it difficult to properly evaluate the interviewer. In addition, there is a lack of ways to appropriately alleviate applicant tension, making it difficult to bring out the true potential of the applicant.
[0427] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0428] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and receiving their answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for generating questions and follow-up questions using a generative AI model, and means for analyzing the applicant's level of nervousness during video communication, thereby making it possible to improve the efficiency of initial interviews and select appropriate candidates.
[0429] "Entry information" refers to information such as a resume, job history, and reason for applying that an applicant submits to a company's recruitment portal.
[0430] "Means for generating questions" refers to a function for creating questions for applicants based on stored entry information.
[0431] "Video communication" is a technology that allows real-time video and audio communication over the Internet and is used in the interview process.
[0432] The "means for analyzing responses in real time" refers to technology that has the function of analyzing applicants' responses in real time, evaluating their content, and generating appropriate follow-up questions.
[0433] "Means for generating follow-up questions" refers to the function of creating additional, detailed questions based on the applicant's answers.
[0434] "Means for summarizing interview results and creating feedback materials" refers to a function for summarizing interview results based on answers given during the interview and creating feedback materials to be provided to human resources personnel.
[0435] "Generative AI model" refers to an artificial intelligence model used to generate questions and follow-up questions.
[0436] The "means for analyzing the level of tension" is a technology that has the function of determining and analyzing the level of tension based on the applicant's voice and video data during video communication.
[0437] This invention is a system that uses AI interviewers to efficiently conduct early-stage interviews (first-stage interviews) in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, and summarizing interview results and creating feedback materials.
[0438] Obtaining entry information
[0439] A user (applicant) submits an application form to an online recruitment portal. This application form includes a resume, a curriculum vitae, and a statement of reasons for applying. This information is received by the server and stored in a database. For example, if a user submits an application form for a "software engineer," the server records information about Applicant A in the database based on this information.
[0440] Question Generation
[0441] The server uses the generative AI model to generate questions for applicants based on the saved entry information, such as "Please tell us about the most difficult experience you had in a past project," and sends these to the terminal (interviewer AI device).
[0442] Conducting interviews
[0443] The user selects the interview date and time through an online recruitment portal. At the specified date and time, the user accesses a video call tool such as Zoom through their terminal (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions. For example, the applicant clicks on the Zoom link at the specified time, and the video call begins. The AI interviewer (server) asks the question, "Tell me about your career history."
[0444] Real-time analysis of responses
[0445] When a user answers a question, the audio and video data are sent to a server in real time. The server then uses an AI analysis model to analyze the answers and evaluate the applicant's skills and experience. For example, if the analysis determines that the applicant has a deep understanding of a particular technology, the server generates questions that dig deeper. It then asks follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0446] Analyzing applicants' tension levels and encouraging them
[0447] The server analyzes the audio and video during the video call to determine the applicant's level of nervousness in real time. If it determines that the applicant is nervous, the server instructs the interviewer AI device to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0448] Summarizing interview results and creating feedback materials
[0449] After the interview is over, the server summarizes the answers given during the interview and creates feedback materials. These materials include an evaluation of the applicant, points to ask in the next interview, and concerns. For example, an evaluation point such as "Applicant A is highly evaluated for his / her strong technical skills and extensive leadership experience" may be written. The feedback materials are saved in the HR system and can be accessed from a terminal (such as the HR manager's PC).
[0450] Specific examples
[0451] When Applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. For example, the applicant participates in an interview via Zoom at a specified time, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answers in real time and generates relevant follow-up questions. After the interview, the data obtained can be used to create feedback materials that HR personnel can use to prepare the next interview.
[0452] Prompt Sentence Examples
[0453] "Tell me about your career so far."
[0454] "What is the most challenging experience you've had on a recent project?"
[0455] "Can you please tell us a bit more about a specific project where you used that technology?"
[0456] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0457] Step 1: Obtaining entry information
[0458] Users access an online recruitment portal and submit their application.
[0459] Specifically, the user enters their resume, job history, and reason for applying into an online form and clicks the submit button.
[0460] Input: Application information such as resume, work history, and reasons for applying
[0461] Output: Entry information stored in the database
[0462] The server receives this information and stores it in a database.
[0463] Step 2: Generate questions
[0464] The server analyzes the entry information stored in the database.
[0465] Specifically, the server uses a generative AI model to receive entry information as input and generate questions based on that information.
[0466] Input: Entry information stored in the database
[0467] Output: Questions to ask the applicant
[0468] For example, a question such as "Please tell us about the most difficult experience you had in a past project" is generated. The generated question is sent to the device.
[0469] Step 3: Set up an interview
[0470] Users select interview dates and times through an online recruitment portal.
[0471] Specifically, the user uses the calendar function to select an available time slot and confirm the reservation.
[0472] Input: Desired interview date and time and contact information
[0473] Output: Notification of confirmed interview schedule
[0474] The server receives this information and confirms the interview schedule.
[0475] Step 4: Conducting the interview
[0476] At the specified date and time, users access video calling tools such as Zoom using their device (PC or mobile device).
[0477] The server detects that a video call has started and begins the interview based on pre-generated questions.
[0478] Input: Access to video call tools such as Zoom, generated questions
[0479] Output: Conducting an interview via video call
[0480] For example, a user clicks on a Zoom link at a specified time to start a video call. The AI interviewer (server) asks, "Tell me about your career history."
[0481] Step 5: Real-time analysis of responses
[0482] As the user answers the questions, audio and video data is transmitted in real time to the server.
[0483] The server analyzes the data using an AI analysis model and evaluates the answers.
[0484] Input: Applicant's audio and video data
[0485] Output: Analysis results and new follow-up questions
[0486] For example, if the analysis reveals that you have a deep understanding of a particular technology, the server will generate a follow-up question such as, "Please tell us more about a specific project using that technology."
[0487] Step 6: Analyze the applicant's level of nervousness and encourage them
[0488] The server analyzes the audio and video during the video call in real time to assess the applicant's level of nervousness.
[0489] Input: Audio and video data during a video call
[0490] Output: Tension analysis results and verbal instructions
[0491] If the server determines that the applicant is nervous, it instructs the AI interviewer device to say something like, "Please relax." Specifically, the AI interviewer uses a voice message to encourage the applicant to relax.
[0492] Step 7: Summarize the interview results and create feedback materials
[0493] After the interview is completed, the server summarizes the answers given during the interview and creates feedback materials.
[0494] Input: Answers given during the interview and analysis results
[0495] Output: Feedback materials and evaluation points for the next interview
[0496] For example, the server can create a summary such as "The applicant has strong technical skills, but lacks details regarding leadership experience" and include it in the feedback material, which is then stored in the HR system and used by HR personnel for the next interview.
[0497] (Application example 1)
[0498] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0499] Until now, human resource costs and time have been major issues in corporate recruitment activities and factory performance evaluations. Furthermore, because the quality of the evaluation depends on human subjectivity, there is a risk of mismatches and unfair evaluations. Furthermore, creating training programs to efficiently improve work performance requires a great deal of effort. There is a need to solve these issues and provide a more efficient and fair evaluation system and training programs suited to each individual worker.
[0500] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0501] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants and workers via video communication and receiving answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing interview results and performance evaluation results to create feedback materials, and means for evaluating performance and providing training programs. This makes it possible to efficiently evaluate applicants and workers, achieve fair and error-free evaluations, and provide optimal training programs based on their individual characteristics.
[0502] "Entry information" refers to information such as resumes, work history, and work experience provided by applicants and workers, and is data necessary for evaluation.
[0503] "Means for generating questions" refers to a function that automatically creates questions appropriate for applicants and workers based on entry information and analysis results.
[0504] "Video communication" is a technology that allows for real-time video and audio communication over the Internet, and is used for interviews and job evaluations.
[0505] "Means for obtaining responses" refers to a function for receiving responses from applicants and workers via video communication.
[0506] "Real-time analytical means" refers to the analytical capabilities required to instantly evaluate applicant and worker responses and generate follow-up questions as needed.
[0507] "Means for generating follow-up questions" refers to a function that automatically creates more detailed questions based on the initial answers.
[0508] "Means for summarizing interview results and performance evaluation results to create feedback materials" refers to a device that has the function of summarizing information obtained from interviews and performance evaluations and providing it in a concise, easy-to-understand format.
[0509] "Means for evaluating work performance and providing training programs" refers to a function that evaluates a worker's ability to perform work and assigns an appropriate training program based on the results.
[0510] The present invention is a system for evaluating the work performance of factory workers and providing training programs based on the evaluation results. The main technical elements of this system include obtaining entry information, generating questions, conducting evaluation sessions via video communication, analyzing responses in real time, and creating feedback materials based on the evaluation results.
[0511] Obtaining entry information
[0512] First, the server obtains the application information from the worker through the portal system in the factory. This application information includes the worker's resume, work history, work experience, etc., and this data is stored in a database.
[0513] Question Generation
[0514] Based on the acquired entry information, the server uses an AI model to generate questions for the worker. For example, it generates questions such as "What was the most difficult task in your recent work?" based on the worker's past work history. The generated questions are sent to the terminal.
[0515] Conducting a business evaluation session
[0516] At the designated time, workers access a video chat tool such as Zoom via their device (PC or tablet). The server detects that the video call has started and begins the evaluation session based on pre-generated questions.
[0517] Real-time analysis of responses
[0518] As workers answer questions, the server analyzes their answers in real time, assessing their understanding of specific work processes based on their answers, and generates follow-up questions as needed, such as, "Is there anything in that work process that you feel needs improvement?"
[0519] Analyzing tension levels and encouraging others
[0520] The server analyzes the worker's tone of voice and facial expressions, and if they appear nervous, it will say something like, "Relax." This allows the worker to be evaluated in a natural state.
[0521] Summarizing the assessment results and preparing training materials
[0522] Once the evaluation session is over, the server summarizes the responses and compiles the evaluation results into a feedback document, which includes an evaluation of work performance, areas for improvement, and the next training program to be implemented. For example, a summary might be created such as, "Work speed is high, but quality control needs improvement."
[0523] The hardware used includes servers, databases, PCs, and tablets, while the software includes a portal system, Zoom API, voice analysis engines (e.g., Google Speech-to-Text), natural language processing models (e.g., GPT-3), and facial expression analysis software (e.g., OpenCV).
[0524] Specific examples
[0525] When Worker A submits application information (such as resume, work history, and work experience) to the online portal, the server generates appropriate questions based on this information. For example, it asks questions such as, "What was the most difficult task you recently performed?" or "Is there anything in the work process that you feel needs improvement?" Worker A participates in the evaluation session using a video chat tool, and analysis is carried out in real time. Afterwards, feedback materials summarizing the evaluation results are created, and a training program is provided.
[0526] Prompt Sentence Examples
[0527] "What was the most difficult task you recently undertook? Please explain in detail why."
[0528] "Are there any particular work processes you feel need improvement?"
[0529] "Please tell us about any major problems you have faced in the past and how you dealt with them."
[0530] In this way, the present invention improves the efficiency of work evaluation and training support within a factory, and contributes greatly to improving work performance and preventing mismatches.
[0531] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0532] Step 1:
[0533] The server obtains application information from users (workers) through the portal system. Users fill out a self-evaluation sheet and provide details of their resume, work history, and work experience. This application information is stored in a database by the server. The input is the application information provided by the user, and the output is the application information stored in the database.
[0534] Step 2:
[0535] The server uses the AI model with a question generation means to generate appropriate questions for the worker based on the stored entry information. The data input into the AI model includes the worker's resume, work history, and work experience, and specific questions are generated as output. For example, a question such as "What was the most difficult task in your recent work?" is generated.
[0536] Step 3:
[0537] The user accesses a video chat tool such as Zoom through a device (PC or tablet) at a specified date and time. When the server detects the start of the video call, it starts an evaluation session based on the generated questions. The input is access to the video chat tool, and the output is the start of the session. Specifically, when the user clicks the Zoom link and the video call starts, the server begins displaying the questions.
[0538] Step 4:
[0539] As users answer questions, the server analyzes the answers in real time. Specifically, a speech analysis engine (e.g., Google Speech-to-Text) is used to convert the speech data into text. A natural language processing model (e.g., GPT-3) then analyzes the text data to understand the answer. Based on this analysis, follow-up questions are generated as needed. The input is the user's speech answer, and the output is the analyzed text and follow-up questions.
[0540] Step 5:
[0541] The server analyzes the user's tone of voice and facial expression to determine the level of tension. It uses voice analysis software and facial expression analysis software (e.g., OpenCV) to detect whether the user is nervous. If necessary, it will say something like, "Please relax." The input is the user's tone of voice and facial expression data, and the output is the analysis results and a message to encourage the user.
[0542] Step 6:
[0543] After the evaluation session is over, the server summarizes the responses and compiles the evaluation results as feedback materials. It uses a text summarization engine (e.g., BERT) to extract key points and create concise feedback materials. It then evaluates the user's performance based on these materials and provides training programs. The input is the user's complete response data, and the output is the summarized feedback materials and training programs.
[0544] This will enable an efficient and fair evaluation system and training programs based on individual characteristics.
[0545] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0546] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results and creating feedback materials, and an emotion engine that recognizes user emotions.
[0547] A natural language description of the program's operation
[0548] 1. Acquisition of entry information
[0549] A user (applicant) submits an application to an online recruitment portal, including a resume, a job history, and a motivation statement. The server receives this application information and stores it in a database.
[0550] 2. Question Generation
[0551] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0552] 3. Conducting the interview
[0553] The user then accesses the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user connects to a video call tool such as Zoom. The server detects the start of the video call and begins the interview.
[0554] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0555] 4. Real-time analysis of responses
[0556] When a user answers a question, the server analyzes the answer in real time using natural language processing technology. The server also includes an emotion engine that analyzes the user's voice and facial expressions during the answer to collect emotional data.
[0557] 5. Generate follow-up questions
[0558] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0559] 6. Analyzing applicants' tension levels and encouraging them
[0560] If the server determines that the user is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly."
[0561] 7. Summarizing interview results and creating feedback materials
[0562] After the interview is over, the server analyzes the interview responses and emotional data again to summarize the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0563] For example, if the analysis indicates that the applicant is technically strong but has concerns about leadership questions, include that information in the summary.
[0564] 8. Providing Feedback Materials
[0565] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0566] Specific examples
[0567] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates via Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time, using an emotion engine to detect nervousness from the applicant's facial expressions and tone of voice, and generates appropriate follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0568] In this way, by combining emotion engines, companies can provide a more refined and personalized interview experience, enabling them to conduct their recruitment activities more efficiently and effectively.
[0569] The processing flow will be explained below.
[0570] Step 1:
[0571] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0572] Step 2:
[0573] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0574] Step 3:
[0575] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0576] Step 4:
[0577] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0578] Step 5:
[0579] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0580] Step 6:
[0581] The server asks the user questions through the AI interviewer based on the sent question list. For example, the AI interviewer might ask, "Tell me about your career so far."
[0582] Step 7:
[0583] As users answer questions, the server analyzes the answers in real time using natural language processing technology, and the resulting data is instantly evaluated by an emotion engine.
[0584] Step 8:
[0585] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0586] Step 9:
[0587] During the interview, if the server determines that the applicant is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say things like, "Please relax. Please answer calmly." This helps to stabilize the applicant's mental state.
[0588] Step 10:
[0589] After the interview, the server analyzes the collected response data and emotional data again to summarize the interview results. This summary includes the applicant's strengths, weaknesses, and emotional reactions. For example, it may include information such as, "The applicant is technically strong, but is anxious about questions regarding leadership."
[0590] Step 11:
[0591] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0592] Step 12:
[0593] All interview data is collected by the server and used to improve the accuracy of the AI model and emotion engine, enabling more efficient and accurate question generation and analysis for future interviews.
[0594] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0595] Example 2
[0596] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0597] In traditional recruitment activities, a large amount of resources are allocated to conducting initial interviews (stage zero interviews), which often results in mismatches between candidates and candidates. This problem is particularly pronounced in companies conducting large-scale recruitment activities, increasing the burden on human resources personnel and contributing to rising recruitment costs. Furthermore, traditional interview methods make it difficult to accurately grasp the applicant's emotions and level of nervousness in a short amount of time, which makes it difficult for applicants to relax and have an opportunity to present themselves. To solve these problems, a more efficient and objective interview method is required.
[0598] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for acquiring entry information, a means for generating questions using a generative AI model based on the acquired entry information, a means for asking questions to applicants via video communication and obtaining their answers, a means for analyzing the answers in real time and generating follow-up questions using natural language processing technology and an emotion engine, and a means for summarizing the interview results and creating feedback materials. This enables efficient and effective initial interviews, reduces recruitment costs, and prevents mismatches with applicants. Furthermore, by analyzing the applicant's level of nervousness using the emotion engine during video communication and providing appropriate prompts, the applicant can be given an opportunity to relax and promote themselves.
[0599] "Entry information" refers to information submitted by applicants, such as resumes, work history, and reasons for applying.
[0600] "Generative AI model" refers to an artificial intelligence model that uses natural language processing technology to generate individual questions from the input entry information.
[0601] "Server" refers to a central processing unit that performs processes such as obtaining entry information, generating questions, analyzing answers, generating follow-up questions, summarizing interview results, and creating feedback materials.
[0602] "Video communication" refers to a means of communication that uses video calling tools such as Zoom to share audio and video in real time with applicants in remote locations.
[0603] "Natural language processing technology" refers to the technology of analyzing and processing human language using a computer.
[0604] "Emotion engine" refers to a system that analyzes an applicant's emotional state from their tone of voice and facial expressions.
[0605] "Follow-up questions" refer to additional questions that are generated based on the applicant's initial answers or emotional state.
[0606] "Feedback materials" are documents summarizing the results of an interview, including questions for the next interview, evaluations of the applicant, and concerns.
[0607] This invention is a system for efficiently conducting initial interviews using AI interviewers in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The system includes functions for acquiring application information, generating questions, conducting interviews via video communication, analyzing responses in real time, generating follow-up questions, summarizing interview results and creating feedback materials, and an emotion engine for recognizing applicants' emotions.
[0608] Obtaining entry information
[0609] A user (applicant) first accesses an online recruitment portal and submits an application form, which includes a resume, a curriculum vitae, and a statement of reasons for applying. The server receives this application information through an HTTP request and performs the necessary database transactions to store it in a database. Specific software used includes database management systems such as MySQL or PostgreSQL.
[0610] Question Generation
[0611] The server retrieves the user's application information from the database and uses a Python script to call an open-source natural language processing library (e.g., Transformers). A prompt sentence is input into the generative AI model to generate individual questions. The generated questions are sent from the server to the terminal (interviewer AI device) in JSON format. An example of a specific prompt sentence is, "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please use an open-source natural language processing library to generate individual, specific questions."
[0612] Conducting interviews
[0613] The user logs in again to the online recruitment portal and selects the date and time of the interview. At the specified date and time, the user clicks the Zoom link to connect to the video call. The server monitors the connection status through the video call tool's API and detects the user's connection. At the start of the interview, the server sends pre-generated questions to the device, which then conducts the assessment. Specific video call tools used include Zoom and WebEx.
[0614] Real-time analysis of responses
[0615] When a user answers a question, the audio data of the answer is sent to the server in real time via the video chat tool's API. The server converts the audio data into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The converted text is then analyzed using a natural language processing library. Furthermore, an emotion engine analyzes the user's tone of voice and facial expressions to collect emotional data.
[0616] Generate follow-up questions
[0617] The server generates follow-up questions based on the collected emotional data and the content of the answers. For example, if the user is nervous about a particular topic, the server generates follow-up questions such as, "What were the specific benefits of this new technology?" This is again done using a generative AI model.
[0618] Analyzing applicants' tension levels and encouraging them
[0619] If the server determines that the user is nervous based on the data analyzed by the emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly." Specifically, it uses the results of transcription and emotion analysis to display appropriate instructions on the screen and play them back aloud.
[0620] Summarizing interview results and creating feedback materials
[0621] After the interview, the server analyzes the interview responses and emotional data again. For example, it uses a text analysis engine to extract specific keywords and emotional states, and then summarizes the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0622] Providing feedback materials
[0623] The server generates feedback materials in PDF format based on the summarized interview results. The PDF file is then saved in the HR system and a notification is sent. HR personnel can access the materials from their own devices (such as PCs) and prepare for the next interview. This function enables companies to conduct their recruitment activities efficiently and effectively.
[0624] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0625] Step 1: Obtaining entry information
[0626] A user (applicant) accesses an online recruitment portal and submits an application form. The application form includes information such as a resume, a curriculum vitae, and a reason for applying. The submitted application information is sent to the server as an HTTP request. The server receives the application information and stores it in a database using a database management system (e.g., MySQL). Specifically, the operation involves executing an INSERT statement against the database. The input in this process is the application information, and the output is the application information stored in the database.
[0627] Step 2: Generate questions
[0628] The server retrieves the saved application information from the database. Based on the retrieved information, questions are generated using a generative AI model (e.g., Transformers). This process is carried out by executing a Python script and calling a natural language processing library. An example prompt is: "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please generate individual, specific questions using an open-source natural language processing library." The generated questions are sent from the server to the terminal in JSON format. The input in this process is the application information, and the output is the generated questions.
[0629] Step 3: Conducting the interview
[0630] The user logs in to the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user clicks on a link in a video calling tool such as Zoom to connect to the video call. The server monitors the connection status through the video calling tool's API and detects the user's connection. When the interview starts, the server sends pre-generated questions to the terminal and conducts the questions. The input is the interview date and time and the questions, and the output is the video call that has started.
[0631] Step 4: Real-time analysis of responses
[0632] When a user answers a question via video call, the audio of the answer is sent to the server in real time via the video call tool's API. The server then converts the audio data of the answer into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The server then analyzes the answer using a natural language processing library (e.g., Transformers), and the emotion engine uses this information to collect emotional data from the tone of voice and facial expressions. The input to this process is the audio data of the answer, and the output is the analyzed answer text and emotional data.
[0633] Step 5: Generate follow-up questions
[0634] The server generates follow-up questions based on the analyzed answer text and sentiment data. Using the generative AI model, it runs a Python script to generate appropriate follow-up questions. These follow-up questions are also sent to the device in JSON format. For example, if a user is nervous about a particular topic, a follow-up question might be generated: "What were the specific benefits of this new technology?" The input to this process is the analyzed answer text and sentiment data, and the output is a follow-up question.
[0635] Step 6: Analyze the applicant's level of nervousness and encourage them
[0636] The server analyzes the user's level of nervousness based on the data collected by the emotion engine. If it determines that the user is nervous, the server will use the AI interviewer to give appropriate encouragement. For example, it may generate encouragement such as "Please relax. Please answer calmly" and play it back aloud. The input in this process is emotional data, and the output is a encouragement message.
[0637] Step 7: Summarize the interview results and create feedback materials
[0638] After the interview is over, the server re-analyzes the responses and emotional data from the interview. For example, it can use a text analysis engine to extract specific keywords and emotional states and summarize evaluation points. Based on the analysis results, it summarizes the interview results and creates feedback materials in PDF format. The input for this process is the responses and emotional data, and the output is the summarized interview results and feedback materials.
[0639] Step 8: Provide feedback materials
[0640] The server generates feedback materials in PDF format based on the summarized interview results and saves them in the human resources system. The saved feedback materials are notified to human resources personnel so that they can access them from their own devices (such as PCs). The input in this process is the summarized interview results, and the output is the generated feedback materials.
[0641] This system allows companies to carry out their recruitment activities efficiently and effectively.
[0642] (Application example 2)
[0643] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0644] Conventional security systems have had difficulty analyzing visitors' behavior and emotions in real time and responding automatically. While early detection of suspicious individuals and rapid response are required, fully automating this process has been a difficult task. The present invention aims to solve these problems and realize more efficient and accurate security responses.
[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and obtaining answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for acquiring video and audio data and analyzing the visitor's emotions from this information, and means for generating follow-up actions based on the results of the emotion analysis. This makes it possible to analyze the visitor's behavior and emotions in real time and respond quickly and appropriately.
[0646] "Entry information" refers to information submitted by applicants, such as resumes, job history, and reasons for applying.
[0647] The "means for generating a question" is a function for generating an appropriate question based on the acquired entry information.
[0648] "Video communication" is a means of communication that uses video and audio to enable applicants and interviewers to interact online.
[0649] "Real-time analysis" refers to the process of instantly analyzing applicant responses and generating follow-up questions as needed.
[0650] The "means for generating follow-up questions" is a function for automatically generating follow-up questions in response to the applicant's answers.
[0651] The "means of summarizing interview results and creating feedback materials" is a function for summarizing the answers and emotional data given during the interview and creating feedback materials.
[0652] "Video and audio data" refers to visual and audio information obtained through cameras and microphones.
[0653] The "means for analyzing emotions" is a function for analyzing the emotions of visitors from the acquired video and audio data.
[0654] The "means for generating follow-up actions" is a function for automatically taking appropriate action against visitors based on the results of sentiment analysis.
[0655] To realize the present invention, the following hardware and software are used.
[0656] Hardware Configuration
[0657] 1. Security Guard Robot: A robot equipped with a camera and microphone to collect video and audio data (e.g., a typical security guard robot with a camera and microphone).
[0658] 2. Server: A high-performance computer for data processing and analysis (e.g., a typical server).
[0659] Software Configuration
[0660] 1. Natural language processing engine: Technology for analyzing information from captured video and audio data and generating appropriate questions and follow-up actions (e.g., OpenAI GPT-4, Google Cloud Dialogflow).
[0661] 2. Video and audio analysis engine: Technology for analyzing video and audio data and recognizing visitors' emotions (e.g., Microsoft Azure Cognitive Services, Google Cloud Video Intelligence API).
[0662] 3. Emotion recognition engine: Technology for analyzing emotions from collected data (e.g., Affectiva SDK, Microsoft Azure Emotion API).
[0663] Basic operations
[0664] 1. Security guard robots patrol monitored areas such as shopping malls and collect video and audio data in real time.
[0665] 2. The collected data is sent to a server and analyzed using a natural language processing engine and a video and audio analysis engine.
[0666] 3. The server analyzes the visitor's behavior and emotions to detect suspicious behavior and emotions such as tension, anxiety, and anger.
[0667] 4. The server utilizes an emotion recognition engine to generate appropriate follow-up actions and prompts.
[0668] 5. The generated instructions are sent to the security guard robot, which then takes appropriate action against the visitor.
[0669] Specific examples
[0670] For example, a security guard robot patrolling a shopping mall collects video and audio of visitor A. If the server analyzes that visitor A suddenly feels anxious, the robot will say, "Please relax. Is there anything I can help you with?" This information is also reported to the security center, and a human security guard will respond if necessary.
[0671] Prompt Sentence Examples
[0672] "Build a system for a security guard robot patrolling a shopping mall to collect real-time video and audio of visitors and analyze their emotions based on specific actions, facial expressions, and tone of voice. This involves the following steps:
[0673] 1. Video and audio data collection
[0674] 2. Real-time analysis using a sentiment analysis engine
[0675] 3. Take follow-up actions based on sentiment data
[0676] 4. Database recording and alert generation
[0677] In this way, by implementing the present invention, security operations can be automated and highly accurate crisis management can be achieved.
[0678] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0679] Step 1:
[0680] The security guard robot patrols the shopping mall and collects video and audio data in real time. The robot's cameras and microphones capture this data and send it to a server.
[0681] Input: Video and audio data from inside the shopping mall
[0682] Output: Video and audio data sent to the server
[0683] Step 2:
[0684] The server passes the received video and audio data to a natural language processing engine and a video and audio analysis engine for analysis, such as analyzing the visitor's behavior, speech content, tone of voice, and facial expressions to recognize their emotional state (tension, anxiety, anger, etc.).
[0685] Input: Video and audio data sent to the server
[0686] Output: Analyzed behavioral and emotional data
[0687] Step 3:
[0688] Based on the analysis results, the server uses an emotion recognition engine to understand the visitor's emotional state in detail, and in this process, specific emotions such as "tension," "anxiety," and "anger" are identified.
[0689] Input: Parsed behavioral and emotional data
[0690] Output: Detailed emotional state data
[0691] Step 4:
[0692] The server generates follow-up actions based on the detailed emotional state data. For example, if the visitor is feeling anxious, it generates a follow-up prompt such as "Please relax. Is there anything I can help you with?"
[0693] Input: Detailed emotional state data
[0694] Output: Generated follow-up action instructions
[0695] Step 5:
[0696] The generated follow-up action instructions are sent to the security guard robot, and the robot responds to the visitor accordingly. For example, the robot may tell the visitor to "relax."
[0697] Input: Generated follow-up action instructions
[0698] Output: Response action by security guard robot
[0699] Step 6:
[0700] The server records the results of follow-up actions and visitor responses in a database and stores relevant information, which can later be analyzed and used for further security measures.
[0701] Input: Follow-up action results and visitor response data
[0702] Output: Behavioral results and response data recorded in a database
[0703] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0704] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0705] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0706] [Third embodiment]
[0707] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0708] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0709] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0710] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0711] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0712] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0713] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0714] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0715] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0716] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0717] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0718] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0719] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews during corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results, and creating feedback materials.
[0720] A natural language description of the program's operation
[0721] 1. Acquisition of entry information
[0722] A user (applicant) submits an application form to an online recruitment portal. This information is stored in a database by the server. As a concrete example, consider a scenario in which an applicant uploads a resume, a curriculum vitae, and a statement of reasons for applying. The server receives and stores this application information.
[0723] 2. Question Generation
[0724] The server uses an AI model to generate questions for applicants based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0725] 3. Conducting the interview
[0726] The user selects the interview date and time through an online recruitment portal. At the designated date and time, the user accesses a video call tool such as Zoom through their device (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions.
[0727] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0728] 4. Real-time analysis of responses
[0729] As users answer questions, the server analyzes the answers in real time. For example, if the server determines that an applicant has a deep understanding of a particular technology, it will generate follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0730] 5. Analyzing applicants' tension levels and encouraging them
[0731] The server analyzes the applicant's tone of voice and facial expressions, and if it determines that the applicant is nervous, it will use the AI interviewer to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0732] 6. Summarizing interview results and creating feedback materials
[0733] After the interview is over, the server summarizes the answers given during the interview and creates a feedback document that summarizes the interview results, such as "The candidate has strong technical skills, but lacks details regarding team leadership experience."
[0734] The feedback materials include questions to be asked in the next interview, evaluations of the applicant, concerns, etc. The server stores the generated feedback materials in the HR system and makes them accessible on a terminal (such as the HR person's PC).
[0735] Specific examples
[0736] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates in the interview using Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time and generates further relevant follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0737] In this way, the present invention realizes effective and efficient interviews in companies' recruitment activities, and contributes greatly to reducing recruitment costs and selecting appropriate personnel.
[0738] The processing flow will be explained below.
[0739] Step 1:
[0740] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0741] Step 2:
[0742] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0743] Step 3:
[0744] The server analyzes the application information retrieved from the database and generates individual questions using an AI model. For example, if an applicant's academic background is computer science, it generates a question such as, "Tell us about the most interesting technology you have studied."
[0745] Step 4:
[0746] The server sends the generated question list to the terminal (interviewer AI device), where the question list is organized into a format that can be used by the video chat tool.
[0747] Step 5:
[0748] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0749] Step 6:
[0750] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0751] Step 7:
[0752] The server asks the user questions through the AI interviewer based on the submitted question list. For example, the AI interviewer might ask, "Please tell us more about the projects listed in your resume."
[0753] Step 8:
[0754] As users answer questions, the server analyzes the answers in real time using natural language processing techniques.
[0755] Step 9:
[0756] The server generates follow-up questions based on the answers and asks the user additional questions. For example, if the user explains about the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0757] Step 10:
[0758] During the interview, the server analyzes the user's tone of voice and facial expressions, and if it determines that the user is nervous, it will say something like, "Please relax. Please answer calmly" through the AI interviewer.
[0759] Step 11:
[0760] After the interview, the server analyzes the collected response data again and summarizes the interview results, including the applicant's strengths, areas for improvement, and evaluation.
[0761] Step 12:
[0762] The server creates feedback materials based on the summarized interview results and saves them in the personnel system. These materials include questions to be asked in the next interview and evaluation details. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0763] Step 13:
[0764] All interview data is collected by the server and used to improve the accuracy of the AI model, enabling more efficient and accurate question generation and analysis for subsequent interviews.
[0765] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0766] Example 1
[0767] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0768] In corporate recruitment activities, in order to streamline initial interviews and select appropriate candidates, a system is needed that can efficiently process large amounts of applicant information and improve the quality and progress of interviews. However, conventional methods require interviewers to directly interact with the interviewer, which is costly and time-consuming, and is largely dependent on the interviewer's subjective opinion, making it difficult to properly evaluate the interviewer. In addition, there is a lack of ways to appropriately alleviate applicant tension, making it difficult to bring out the true potential of the applicant.
[0769] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0770] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and receiving their answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for generating questions and follow-up questions using a generative AI model, and means for analyzing the applicant's level of nervousness during video communication, thereby making it possible to improve the efficiency of initial interviews and select appropriate candidates.
[0771] "Entry information" refers to information such as a resume, job history, and reason for applying that an applicant submits to a company's recruitment portal.
[0772] "Means for generating questions" refers to a function for creating questions for applicants based on stored entry information.
[0773] "Video communication" is a technology that allows real-time video and audio communication over the Internet and is used in the interview process.
[0774] The "means for analyzing responses in real time" refers to technology that has the function of analyzing applicants' responses in real time, evaluating their content, and generating appropriate follow-up questions.
[0775] "Means for generating follow-up questions" refers to the function of creating additional, detailed questions based on the applicant's answers.
[0776] "Means for summarizing interview results and creating feedback materials" refers to a function for summarizing interview results based on answers given during the interview and creating feedback materials to be provided to human resources personnel.
[0777] "Generative AI model" refers to an artificial intelligence model used to generate questions and follow-up questions.
[0778] The "means for analyzing the level of tension" is a technology that has the function of determining and analyzing the level of tension based on the applicant's voice and video data during video communication.
[0779] This invention is a system that uses AI interviewers to efficiently conduct early-stage interviews (first-stage interviews) in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, and summarizing interview results and creating feedback materials.
[0780] Obtaining entry information
[0781] A user (applicant) submits an application form to an online recruitment portal. This application form includes a resume, a curriculum vitae, and a statement of reasons for applying. This information is received by the server and stored in a database. For example, if a user submits an application form for a "software engineer," the server records information about Applicant A in the database based on this information.
[0782] Question Generation
[0783] The server uses the generative AI model to generate questions for applicants based on the saved entry information, such as "Please tell us about the most difficult experience you had in a past project," and sends these to the terminal (interviewer AI device).
[0784] Conducting interviews
[0785] The user selects the interview date and time through an online recruitment portal. At the specified date and time, the user accesses a video call tool such as Zoom through their terminal (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions. For example, the applicant clicks on the Zoom link at the specified time, and the video call begins. The AI interviewer (server) asks the question, "Tell me about your career history."
[0786] Real-time analysis of responses
[0787] When a user answers a question, the audio and video data are sent to a server in real time. The server then uses an AI analysis model to analyze the answers and evaluate the applicant's skills and experience. For example, if the analysis determines that the applicant has a deep understanding of a particular technology, the server generates questions that dig deeper. It then asks follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[0788] Analyzing applicants' tension levels and encouraging them
[0789] The server analyzes the audio and video during the video call to determine the applicant's level of nervousness in real time. If it determines that the applicant is nervous, the server instructs the interviewer AI device to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[0790] Summarizing interview results and creating feedback materials
[0791] After the interview is over, the server summarizes the answers given during the interview and creates feedback materials. These materials include an evaluation of the applicant, points to ask in the next interview, and concerns. For example, an evaluation point such as "Applicant A is highly evaluated for his / her strong technical skills and extensive leadership experience" may be written. The feedback materials are saved in the HR system and can be accessed from a terminal (such as the HR manager's PC).
[0792] Specific examples
[0793] When Applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. For example, the applicant participates in an interview via Zoom at a specified time, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answers in real time and generates relevant follow-up questions. After the interview, the data obtained can be used to create feedback materials that HR personnel can use to prepare the next interview.
[0794] Prompt Sentence Examples
[0795] "Tell me about your career so far."
[0796] "What is the most challenging experience you've had on a recent project?"
[0797] "Can you please tell us a bit more about a specific project where you used that technology?"
[0798] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0799] Step 1: Obtaining entry information
[0800] Users access an online recruitment portal and submit their application.
[0801] Specifically, the user enters their resume, job history, and reason for applying into an online form and clicks the submit button.
[0802] Input: Application information such as resume, work history, and reasons for applying
[0803] Output: Entry information stored in the database
[0804] The server receives this information and stores it in a database.
[0805] Step 2: Generate questions
[0806] The server analyzes the entry information stored in the database.
[0807] Specifically, the server uses a generative AI model to receive entry information as input and generate questions based on that information.
[0808] Input: Entry information stored in the database
[0809] Output: Questions to ask the applicant
[0810] For example, a question such as "Please tell us about the most difficult experience you had in a past project" is generated. The generated question is sent to the device.
[0811] Step 3: Set up an interview
[0812] Users select interview dates and times through an online recruitment portal.
[0813] Specifically, the user uses the calendar function to select an available time slot and confirm the reservation.
[0814] Input: Desired interview date and time and contact information
[0815] Output: Notification of confirmed interview schedule
[0816] The server receives this information and confirms the interview schedule.
[0817] Step 4: Conducting the interview
[0818] At the specified date and time, users access video calling tools such as Zoom using their device (PC or mobile device).
[0819] The server detects that a video call has started and begins the interview based on pre-generated questions.
[0820] Input: Access to video call tools such as Zoom, generated questions
[0821] Output: Conducting an interview via video call
[0822] For example, a user clicks on a Zoom link at a specified time to start a video call. The AI interviewer (server) asks, "Tell me about your career history."
[0823] Step 5: Real-time analysis of responses
[0824] As the user answers the questions, audio and video data is transmitted in real time to the server.
[0825] The server analyzes the data using an AI analysis model and evaluates the answers.
[0826] Input: Applicant's audio and video data
[0827] Output: Analysis results and new follow-up questions
[0828] For example, if the analysis reveals that you have a deep understanding of a particular technology, the server will generate a follow-up question such as, "Please tell us more about a specific project using that technology."
[0829] Step 6: Analyze the applicant's level of nervousness and encourage them
[0830] The server analyzes the audio and video during the video call in real time to assess the applicant's level of nervousness.
[0831] Input: Audio and video data during a video call
[0832] Output: Tension analysis results and verbal instructions
[0833] If the server determines that the applicant is nervous, it instructs the AI interviewer device to say something like, "Please relax." Specifically, the AI interviewer uses a voice message to encourage the applicant to relax.
[0834] Step 7: Summarize the interview results and create feedback materials
[0835] After the interview is completed, the server summarizes the answers given during the interview and creates feedback materials.
[0836] Input: Answers given during the interview and analysis results
[0837] Output: Feedback materials and evaluation points for the next interview
[0838] For example, the server can create a summary such as "The applicant has strong technical skills, but lacks details regarding leadership experience" and include it in the feedback material, which is then stored in the HR system and used by HR personnel for the next interview.
[0839] (Application example 1)
[0840] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0841] Until now, human resource costs and time have been major issues in corporate recruitment activities and factory performance evaluations. Furthermore, because the quality of the evaluation depends on human subjectivity, there is a risk of mismatches and unfair evaluations. Furthermore, creating training programs to efficiently improve work performance requires a great deal of effort. There is a need to solve these issues and provide a more efficient and fair evaluation system and training programs suited to each individual worker.
[0842] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0843] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants and workers via video communication and receiving answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing interview results and performance evaluation results to create feedback materials, and means for evaluating performance and providing training programs. This makes it possible to efficiently evaluate applicants and workers, achieve fair and error-free evaluations, and provide optimal training programs based on their individual characteristics.
[0844] "Entry information" refers to information such as resumes, work history, and work experience provided by applicants and workers, and is data necessary for evaluation.
[0845] "Means for generating questions" refers to a function that automatically creates questions appropriate for applicants and workers based on entry information and analysis results.
[0846] "Video communication" is a technology that allows for real-time video and audio communication over the Internet, and is used for interviews and job evaluations.
[0847] "Means for obtaining responses" refers to a function for receiving responses from applicants and workers via video communication.
[0848] "Real-time analytical means" refers to the analytical capabilities required to instantly evaluate applicant and worker responses and generate follow-up questions as needed.
[0849] "Means for generating follow-up questions" refers to a function that automatically creates more detailed questions based on the initial answers.
[0850] "Means for summarizing interview results and performance evaluation results to create feedback materials" refers to a device that has the function of summarizing information obtained from interviews and performance evaluations and providing it in a concise, easy-to-understand format.
[0851] "Means for evaluating work performance and providing training programs" refers to a function that evaluates a worker's ability to perform work and assigns an appropriate training program based on the results.
[0852] The present invention is a system for evaluating the work performance of factory workers and providing training programs based on the evaluation results. The main technical elements of this system include obtaining entry information, generating questions, conducting evaluation sessions via video communication, analyzing responses in real time, and creating feedback materials based on the evaluation results.
[0853] Obtaining entry information
[0854] First, the server obtains the application information from the worker through the portal system in the factory. This application information includes the worker's resume, work history, work experience, etc., and this data is stored in a database.
[0855] Question Generation
[0856] Based on the acquired entry information, the server uses an AI model to generate questions for the worker. For example, it generates questions such as "What was the most difficult task in your recent work?" based on the worker's past work history. The generated questions are sent to the terminal.
[0857] Conducting a business evaluation session
[0858] At the designated time, workers access a video chat tool such as Zoom via their device (PC or tablet). The server detects that the video call has started and begins the evaluation session based on pre-generated questions.
[0859] Real-time analysis of responses
[0860] As workers answer questions, the server analyzes their answers in real time, assessing their understanding of specific work processes based on their answers, and generates follow-up questions as needed, such as, "Is there anything in that work process that you feel needs improvement?"
[0861] Analyzing tension levels and encouraging others
[0862] The server analyzes the worker's tone of voice and facial expressions, and if they appear nervous, it will say something like, "Relax." This allows the worker to be evaluated in a natural state.
[0863] Summarizing the assessment results and preparing training materials
[0864] Once the evaluation session is over, the server summarizes the responses and compiles the evaluation results into a feedback document, which includes an evaluation of work performance, areas for improvement, and the next training program to be implemented. For example, a summary might be created such as, "Work speed is high, but quality control needs improvement."
[0865] The hardware used includes servers, databases, PCs, and tablets, while the software includes a portal system, Zoom API, voice analysis engines (e.g., Google Speech-to-Text), natural language processing models (e.g., GPT-3), and facial expression analysis software (e.g., OpenCV).
[0866] Specific examples
[0867] When Worker A submits application information (such as resume, work history, and work experience) to the online portal, the server generates appropriate questions based on this information. For example, it asks questions such as, "What was the most difficult task you recently performed?" or "Is there anything in the work process that you feel needs improvement?" Worker A participates in the evaluation session using a video chat tool, and analysis is carried out in real time. Afterwards, feedback materials summarizing the evaluation results are created, and a training program is provided.
[0868] Prompt Sentence Examples
[0869] "What was the most difficult task you recently undertook? Please explain in detail why."
[0870] "Are there any particular work processes you feel need improvement?"
[0871] "Please tell us about any major problems you have faced in the past and how you dealt with them."
[0872] In this way, the present invention improves the efficiency of work evaluation and training support within a factory, and contributes greatly to improving work performance and preventing mismatches.
[0873] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0874] Step 1:
[0875] The server obtains application information from users (workers) through the portal system. Users fill out a self-evaluation sheet and provide details of their resume, work history, and work experience. This application information is stored in a database by the server. The input is the application information provided by the user, and the output is the application information stored in the database.
[0876] Step 2:
[0877] The server uses the AI model with a question generation means to generate appropriate questions for the worker based on the stored entry information. The data input into the AI model includes the worker's resume, work history, and work experience, and specific questions are generated as output. For example, a question such as "What was the most difficult task in your recent work?" is generated.
[0878] Step 3:
[0879] The user accesses a video chat tool such as Zoom through a device (PC or tablet) at a specified date and time. When the server detects the start of the video call, it starts an evaluation session based on the generated questions. The input is access to the video chat tool, and the output is the start of the session. Specifically, when the user clicks the Zoom link and the video call starts, the server begins displaying the questions.
[0880] Step 4:
[0881] As users answer questions, the server analyzes the answers in real time. Specifically, a speech analysis engine (e.g., Google Speech-to-Text) is used to convert the speech data into text. A natural language processing model (e.g., GPT-3) then analyzes the text data to understand the answer. Based on this analysis, follow-up questions are generated as needed. The input is the user's speech answer, and the output is the analyzed text and follow-up questions.
[0882] Step 5:
[0883] The server analyzes the user's tone of voice and facial expression to determine the level of tension. It uses voice analysis software and facial expression analysis software (e.g., OpenCV) to detect whether the user is nervous. If necessary, it will say something like, "Please relax." The input is the user's tone of voice and facial expression data, and the output is the analysis results and a message to encourage the user.
[0884] Step 6:
[0885] After the evaluation session is over, the server summarizes the responses and compiles the evaluation results as feedback materials. It uses a text summarization engine (e.g., BERT) to extract key points and create concise feedback materials. It then evaluates the user's performance based on these materials and provides training programs. The input is the user's complete response data, and the output is the summarized feedback materials and training programs.
[0886] This will enable an efficient and fair evaluation system and training programs based on individual characteristics.
[0887] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0888] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results and creating feedback materials, and an emotion engine that recognizes user emotions.
[0889] A natural language description of the program's operation
[0890] 1. Acquisition of entry information
[0891] A user (applicant) submits an application to an online recruitment portal, including a resume, a job history, and a motivation statement. The server receives this application information and stores it in a database.
[0892] 2. Question Generation
[0893] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0894] 3. Conducting the interview
[0895] The user then accesses the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user connects to a video call tool such as Zoom. The server detects the start of the video call and begins the interview.
[0896] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[0897] 4. Real-time analysis of responses
[0898] When a user answers a question, the server analyzes the answer in real time using natural language processing technology. The server also includes an emotion engine that analyzes the user's voice and facial expressions during the answer to collect emotional data.
[0899] 5. Generate follow-up questions
[0900] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0901] 6. Analyzing applicants' tension levels and encouraging them
[0902] If the server determines that the user is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly."
[0903] 7. Summarizing interview results and creating feedback materials
[0904] After the interview is over, the server analyzes the interview responses and emotional data again to summarize the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0905] For example, if the analysis indicates that the applicant is technically strong but has concerns about leadership questions, include that information in the summary.
[0906] 8. Providing Feedback Materials
[0907] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0908] Specific examples
[0909] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates via Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time, using an emotion engine to detect nervousness from the applicant's facial expressions and tone of voice, and generates appropriate follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[0910] In this way, by combining emotion engines, companies can provide a more refined and personalized interview experience, enabling them to conduct their recruitment activities more efficiently and effectively.
[0911] The processing flow will be explained below.
[0912] Step 1:
[0913] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[0914] Step 2:
[0915] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[0916] Step 3:
[0917] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[0918] Step 4:
[0919] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[0920] Step 5:
[0921] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[0922] Step 6:
[0923] The server asks the user questions through the AI interviewer based on the sent question list. For example, the AI interviewer might ask, "Tell me about your career so far."
[0924] Step 7:
[0925] As users answer questions, the server analyzes the answers in real time using natural language processing technology, and the resulting data is instantly evaluated by an emotion engine.
[0926] Step 8:
[0927] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[0928] Step 9:
[0929] During the interview, if the server determines that the applicant is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say things like, "Please relax. Please answer calmly." This helps to stabilize the applicant's mental state.
[0930] Step 10:
[0931] After the interview, the server analyzes the collected response data and emotional data again to summarize the interview results. This summary includes the applicant's strengths, weaknesses, and emotional reactions. For example, it may include information such as, "The applicant is technically strong, but is anxious about questions regarding leadership."
[0932] Step 11:
[0933] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[0934] Step 12:
[0935] All interview data is collected by the server and used to improve the accuracy of the AI model and emotion engine, enabling more efficient and accurate question generation and analysis for future interviews.
[0936] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[0937] Example 2
[0938] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0939] In traditional recruitment activities, a large amount of resources are allocated to conducting initial interviews (stage zero interviews), which often results in mismatches between candidates and candidates. This problem is particularly pronounced in companies conducting large-scale recruitment activities, increasing the burden on human resources personnel and contributing to rising recruitment costs. Furthermore, traditional interview methods make it difficult to accurately grasp the applicant's emotions and level of nervousness in a short amount of time, which makes it difficult for applicants to relax and have an opportunity to present themselves. To solve these problems, a more efficient and objective interview method is required.
[0940] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for acquiring entry information, a means for generating questions using a generative AI model based on the acquired entry information, a means for asking questions to applicants via video communication and obtaining their answers, a means for analyzing the answers in real time and generating follow-up questions using natural language processing technology and an emotion engine, and a means for summarizing the interview results and creating feedback materials. This enables efficient and effective initial interviews, reduces recruitment costs, and prevents mismatches with applicants. Furthermore, by analyzing the applicant's level of nervousness using the emotion engine during video communication and providing appropriate prompts, the applicant can be given an opportunity to relax and promote themselves.
[0941] "Entry information" refers to information submitted by applicants, such as resumes, work history, and reasons for applying.
[0942] "Generative AI model" refers to an artificial intelligence model that uses natural language processing technology to generate individual questions from the input entry information.
[0943] "Server" refers to a central processing unit that performs processes such as obtaining entry information, generating questions, analyzing answers, generating follow-up questions, summarizing interview results, and creating feedback materials.
[0944] "Video communication" refers to a means of communication that uses video calling tools such as Zoom to share audio and video in real time with applicants in remote locations.
[0945] "Natural language processing technology" refers to the technology of analyzing and processing human language using a computer.
[0946] "Emotion engine" refers to a system that analyzes an applicant's emotional state from their tone of voice and facial expressions.
[0947] "Follow-up questions" refer to additional questions that are generated based on the applicant's initial answers or emotional state.
[0948] "Feedback materials" are documents summarizing the results of an interview, including questions for the next interview, evaluations of the applicant, and concerns.
[0949] This invention is a system for efficiently conducting initial interviews using AI interviewers in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The system includes functions for acquiring application information, generating questions, conducting interviews via video communication, analyzing responses in real time, generating follow-up questions, summarizing interview results and creating feedback materials, and an emotion engine for recognizing applicants' emotions.
[0950] Obtaining entry information
[0951] A user (applicant) first accesses an online recruitment portal and submits an application form, which includes a resume, a curriculum vitae, and a statement of reasons for applying. The server receives this application information through an HTTP request and performs the necessary database transactions to store it in a database. Specific software used includes database management systems such as MySQL or PostgreSQL.
[0952] Question Generation
[0953] The server retrieves the user's application information from the database and uses a Python script to call an open-source natural language processing library (e.g., Transformers). A prompt sentence is input into the generative AI model to generate individual questions. The generated questions are sent from the server to the terminal (interviewer AI device) in JSON format. An example of a specific prompt sentence is, "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please use an open-source natural language processing library to generate individual, specific questions."
[0954] Conducting interviews
[0955] The user logs in again to the online recruitment portal and selects the date and time of the interview. At the specified date and time, the user clicks the Zoom link to connect to the video call. The server monitors the connection status through the video call tool's API and detects the user's connection. At the start of the interview, the server sends pre-generated questions to the device, which then conducts the assessment. Specific video call tools used include Zoom and WebEx.
[0956] Real-time analysis of responses
[0957] When a user answers a question, the audio data of the answer is sent to the server in real time via the video chat tool's API. The server converts the audio data into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The converted text is then analyzed using a natural language processing library. Furthermore, an emotion engine analyzes the user's tone of voice and facial expressions to collect emotional data.
[0958] Generate follow-up questions
[0959] The server generates follow-up questions based on the collected emotional data and the content of the answers. For example, if the user is nervous about a particular topic, the server generates follow-up questions such as, "What were the specific benefits of this new technology?" This is again done using a generative AI model.
[0960] Analyzing applicants' tension levels and encouraging them
[0961] If the server determines that the user is nervous based on the data analyzed by the emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly." Specifically, it uses the results of transcription and emotion analysis to display appropriate instructions on the screen and play them back aloud.
[0962] Summarizing interview results and creating feedback materials
[0963] After the interview, the server analyzes the interview responses and emotional data again. For example, it uses a text analysis engine to extract specific keywords and emotional states, and then summarizes the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[0964] Providing feedback materials
[0965] The server generates feedback materials in PDF format based on the summarized interview results. The PDF file is then saved in the HR system and a notification is sent. HR personnel can access the materials from their own devices (such as PCs) and prepare for the next interview. This function enables companies to conduct their recruitment activities efficiently and effectively.
[0966] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0967] Step 1: Obtaining entry information
[0968] A user (applicant) accesses an online recruitment portal and submits an application form. The application form includes information such as a resume, a curriculum vitae, and a reason for applying. The submitted application information is sent to the server as an HTTP request. The server receives the application information and stores it in a database using a database management system (e.g., MySQL). Specifically, the operation involves executing an INSERT statement against the database. The input in this process is the application information, and the output is the application information stored in the database.
[0969] Step 2: Generate questions
[0970] The server retrieves the saved application information from the database. Based on the retrieved information, questions are generated using a generative AI model (e.g., Transformers). This process is carried out by executing a Python script and calling a natural language processing library. An example prompt is: "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please generate individual, specific questions using an open-source natural language processing library." The generated questions are sent from the server to the terminal in JSON format. The input in this process is the application information, and the output is the generated questions.
[0971] Step 3: Conducting the interview
[0972] The user logs in to the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user clicks on a link in a video calling tool such as Zoom to connect to the video call. The server monitors the connection status through the video calling tool's API and detects the user's connection. When the interview starts, the server sends pre-generated questions to the terminal and conducts the questions. The input is the interview date and time and the questions, and the output is the video call that has started.
[0973] Step 4: Real-time analysis of responses
[0974] When a user answers a question via video call, the audio of the answer is sent to the server in real time via the video call tool's API. The server then converts the audio data of the answer into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The server then analyzes the answer using a natural language processing library (e.g., Transformers), and the emotion engine uses this information to collect emotional data from the tone of voice and facial expressions. The input to this process is the audio data of the answer, and the output is the analyzed answer text and emotional data.
[0975] Step 5: Generate follow-up questions
[0976] The server generates follow-up questions based on the analyzed answer text and sentiment data. Using the generative AI model, it runs a Python script to generate appropriate follow-up questions. These follow-up questions are also sent to the device in JSON format. For example, if a user is nervous about a particular topic, a follow-up question might be generated: "What were the specific benefits of this new technology?" The input to this process is the analyzed answer text and sentiment data, and the output is a follow-up question.
[0977] Step 6: Analyze the applicant's level of nervousness and encourage them
[0978] The server analyzes the user's level of nervousness based on the data collected by the emotion engine. If it determines that the user is nervous, the server will use the AI interviewer to give appropriate encouragement. For example, it may generate encouragement such as "Please relax. Please answer calmly" and play it back aloud. The input in this process is emotional data, and the output is a encouragement message.
[0979] Step 7: Summarize the interview results and create feedback materials
[0980] After the interview is over, the server re-analyzes the responses and emotional data from the interview. For example, it can use a text analysis engine to extract specific keywords and emotional states and summarize evaluation points. Based on the analysis results, it summarizes the interview results and creates feedback materials in PDF format. The input for this process is the responses and emotional data, and the output is the summarized interview results and feedback materials.
[0981] Step 8: Provide feedback materials
[0982] The server generates feedback materials in PDF format based on the summarized interview results and saves them in the human resources system. The saved feedback materials are notified to human resources personnel so that they can access them from their own devices (such as PCs). The input in this process is the summarized interview results, and the output is the generated feedback materials.
[0983] This system allows companies to carry out their recruitment activities efficiently and effectively.
[0984] (Application example 2)
[0985] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0986] Conventional security systems have had difficulty analyzing visitors' behavior and emotions in real time and responding automatically. While early detection of suspicious individuals and rapid response are required, fully automating this process has been a difficult task. The present invention aims to solve these problems and realize more efficient and accurate security responses.
[0987] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and obtaining answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for acquiring video and audio data and analyzing the visitor's emotions from this information, and means for generating follow-up actions based on the results of the emotion analysis. This makes it possible to analyze the visitor's behavior and emotions in real time and respond quickly and appropriately.
[0988] "Entry information" refers to information submitted by applicants, such as resumes, job history, and reasons for applying.
[0989] The "means for generating a question" is a function for generating an appropriate question based on the acquired entry information.
[0990] "Video communication" is a means of communication that uses video and audio to enable applicants and interviewers to interact online.
[0991] "Real-time analysis" refers to the process of instantly analyzing applicant responses and generating follow-up questions as needed.
[0992] The "means for generating follow-up questions" is a function for automatically generating follow-up questions in response to the applicant's answers.
[0993] The "means of summarizing interview results and creating feedback materials" is a function for summarizing the answers and emotional data given during the interview and creating feedback materials.
[0994] "Video and audio data" refers to visual and audio information obtained through cameras and microphones.
[0995] The "means for analyzing emotions" is a function for analyzing the emotions of visitors from the acquired video and audio data.
[0996] The "means for generating follow-up actions" is a function for automatically taking appropriate action against visitors based on the results of sentiment analysis.
[0997] To realize the present invention, the following hardware and software are used.
[0998] Hardware Configuration
[0999] 1. Security Guard Robot: A robot equipped with a camera and microphone to collect video and audio data (e.g., a typical security guard robot with a camera and microphone).
[1000] 2. Server: A high-performance computer for data processing and analysis (e.g., a typical server).
[1001] Software Configuration
[1002] 1. Natural language processing engine: Technology for analyzing information from captured video and audio data and generating appropriate questions and follow-up actions (e.g., OpenAI GPT-4, Google Cloud Dialogflow).
[1003] 2. Video and audio analysis engine: Technology for analyzing video and audio data and recognizing visitors' emotions (e.g., Microsoft Azure Cognitive Services, Google Cloud Video Intelligence API).
[1004] 3. Emotion recognition engine: Technology for analyzing emotions from collected data (e.g., Affectiva SDK, Microsoft Azure Emotion API).
[1005] Basic operations
[1006] 1. Security guard robots patrol monitored areas such as shopping malls and collect video and audio data in real time.
[1007] 2. The collected data is sent to a server and analyzed using a natural language processing engine and a video and audio analysis engine.
[1008] 3. The server analyzes the visitor's behavior and emotions to detect suspicious behavior and emotions such as tension, anxiety, and anger.
[1009] 4. The server utilizes an emotion recognition engine to generate appropriate follow-up actions and prompts.
[1010] 5. The generated instructions are sent to the security guard robot, which then takes appropriate action against the visitor.
[1011] Specific examples
[1012] For example, a security guard robot patrolling a shopping mall collects video and audio of visitor A. If the server analyzes that visitor A suddenly feels anxious, the robot will say, "Please relax. Is there anything I can help you with?" This information is also reported to the security center, and a human security guard will respond if necessary.
[1013] Prompt Sentence Examples
[1014] "Build a system for a security guard robot patrolling a shopping mall to collect real-time video and audio of visitors and analyze their emotions based on specific actions, facial expressions, and tone of voice. This involves the following steps:
[1015] 1. Video and audio data collection
[1016] 2. Real-time analysis using a sentiment analysis engine
[1017] 3. Take follow-up actions based on sentiment data
[1018] 4. Database recording and alert generation
[1019] In this way, by implementing the present invention, security operations can be automated and highly accurate crisis management can be achieved.
[1020] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1021] Step 1:
[1022] The security guard robot patrols the shopping mall and collects video and audio data in real time. The robot's cameras and microphones capture this data and send it to a server.
[1023] Input: Video and audio data from inside the shopping mall
[1024] Output: Video and audio data sent to the server
[1025] Step 2:
[1026] The server passes the received video and audio data to a natural language processing engine and a video and audio analysis engine for analysis, such as analyzing the visitor's behavior, speech content, tone of voice, and facial expressions to recognize their emotional state (tension, anxiety, anger, etc.).
[1027] Input: Video and audio data sent to the server
[1028] Output: Analyzed behavioral and emotional data
[1029] Step 3:
[1030] Based on the analysis results, the server uses an emotion recognition engine to understand the visitor's emotional state in detail, and in this process, specific emotions such as "tension," "anxiety," and "anger" are identified.
[1031] Input: Parsed behavioral and emotional data
[1032] Output: Detailed emotional state data
[1033] Step 4:
[1034] The server generates follow-up actions based on the detailed emotional state data. For example, if the visitor is feeling anxious, it generates a follow-up prompt such as "Please relax. Is there anything I can help you with?"
[1035] Input: Detailed emotional state data
[1036] Output: Generated follow-up action instructions
[1037] Step 5:
[1038] The generated follow-up action instructions are sent to the security guard robot, and the robot responds to the visitor accordingly. For example, the robot may tell the visitor to "relax."
[1039] Input: Generated follow-up action instructions
[1040] Output: Response action by security guard robot
[1041] Step 6:
[1042] The server records the results of follow-up actions and visitor responses in a database and stores relevant information, which can later be analyzed and used for further security measures.
[1043] Input: Follow-up action results and visitor response data
[1044] Output: Behavioral results and response data recorded in a database
[1045] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1046] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1047] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1048] [Fourth embodiment]
[1049] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1050] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1051] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1052] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1053] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1054] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1055] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1056] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1057] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1058] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1059] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1060] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1061] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1062] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews during corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results, and creating feedback materials.
[1063] A natural language description of the program's operation
[1064] 1. Acquisition of entry information
[1065] A user (applicant) submits an application form to an online recruitment portal. This information is stored in a database by the server. As a concrete example, consider a scenario in which an applicant uploads a resume, a curriculum vitae, and a statement of reasons for applying. The server receives and stores this application information.
[1066] 2. Question Generation
[1067] The server uses an AI model to generate questions for applicants based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[1068] 3. Conducting the interview
[1069] The user selects the interview date and time through an online recruitment portal. At the designated date and time, the user accesses a video call tool such as Zoom through their device (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions.
[1070] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[1071] 4. Real-time analysis of responses
[1072] As users answer questions, the server analyzes the answers in real time. For example, if the server determines that an applicant has a deep understanding of a particular technology, it will generate follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[1073] 5. Analyzing applicants' tension levels and encouraging them
[1074] The server analyzes the applicant's tone of voice and facial expressions, and if it determines that the applicant is nervous, it will use the AI interviewer to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[1075] 6. Summarizing interview results and creating feedback materials
[1076] After the interview is over, the server summarizes the answers given during the interview and creates a feedback document that summarizes the interview results, such as "The candidate has strong technical skills, but lacks details regarding team leadership experience."
[1077] The feedback materials include questions to be asked in the next interview, evaluations of the applicant, concerns, etc. The server stores the generated feedback materials in the HR system and makes them accessible on a terminal (such as the HR person's PC).
[1078] Specific examples
[1079] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates in the interview using Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time and generates further relevant follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[1080] In this way, the present invention realizes effective and efficient interviews in companies' recruitment activities, and contributes greatly to reducing recruitment costs and selecting appropriate personnel.
[1081] The processing flow will be explained below.
[1082] Step 1:
[1083] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[1084] Step 2:
[1085] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[1086] Step 3:
[1087] The server analyzes the application information retrieved from the database and generates individual questions using an AI model. For example, if an applicant's academic background is computer science, it generates a question such as, "Tell us about the most interesting technology you have studied."
[1088] Step 4:
[1089] The server sends the generated question list to the terminal (interviewer AI device), where the question list is organized into a format that can be used by the video chat tool.
[1090] Step 5:
[1091] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[1092] Step 6:
[1093] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[1094] Step 7:
[1095] The server asks the user questions through the AI interviewer based on the submitted question list. For example, the AI interviewer might ask, "Please tell us more about the projects listed in your resume."
[1096] Step 8:
[1097] As users answer questions, the server analyzes the answers in real time using natural language processing techniques.
[1098] Step 9:
[1099] The server generates follow-up questions based on the answers and asks the user additional questions. For example, if the user explains about the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[1100] Step 10:
[1101] During the interview, the server analyzes the user's tone of voice and facial expressions, and if it determines that the user is nervous, it will say something like, "Please relax. Please answer calmly" through the AI interviewer.
[1102] Step 11:
[1103] After the interview, the server analyzes the collected response data again and summarizes the interview results, including the applicant's strengths, areas for improvement, and evaluation.
[1104] Step 12:
[1105] The server creates feedback materials based on the summarized interview results and saves them in the personnel system. These materials include questions to be asked in the next interview and evaluation details. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[1106] Step 13:
[1107] All interview data is collected by the server and used to improve the accuracy of the AI model, enabling more efficient and accurate question generation and analysis for subsequent interviews.
[1108] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[1109] Example 1
[1110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1111] In corporate recruitment activities, in order to streamline initial interviews and select appropriate candidates, a system is needed that can efficiently process large amounts of applicant information and improve the quality and progress of interviews. However, conventional methods require interviewers to directly interact with the interviewer, which is costly and time-consuming, and is largely dependent on the interviewer's subjective opinion, making it difficult to properly evaluate the interviewer. In addition, there is a lack of ways to appropriately alleviate applicant tension, making it difficult to bring out the true potential of the applicant.
[1112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1113] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and receiving their answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for generating questions and follow-up questions using a generative AI model, and means for analyzing the applicant's level of nervousness during video communication, thereby making it possible to improve the efficiency of initial interviews and select appropriate candidates.
[1114] "Entry information" refers to information such as a resume, job history, and reason for applying that an applicant submits to a company's recruitment portal.
[1115] "Means for generating questions" refers to a function for creating questions for applicants based on stored entry information.
[1116] "Video communication" is a technology that allows real-time video and audio communication over the Internet and is used in the interview process.
[1117] The "means for analyzing responses in real time" refers to technology that has the function of analyzing applicants' responses in real time, evaluating their content, and generating appropriate follow-up questions.
[1118] "Means for generating follow-up questions" refers to the function of creating additional, detailed questions based on the applicant's answers.
[1119] "Means for summarizing interview results and creating feedback materials" refers to a function for summarizing interview results based on answers given during the interview and creating feedback materials to be provided to human resources personnel.
[1120] "Generative AI model" refers to an artificial intelligence model used to generate questions and follow-up questions.
[1121] The "means for analyzing the level of tension" is a technology that has the function of determining and analyzing the level of tension based on the applicant's voice and video data during video communication.
[1122] This invention is a system that uses AI interviewers to efficiently conduct early-stage interviews (first-stage interviews) in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, and summarizing interview results and creating feedback materials.
[1123] Obtaining entry information
[1124] A user (applicant) submits an application form to an online recruitment portal. This application form includes a resume, a curriculum vitae, and a statement of reasons for applying. This information is received by the server and stored in a database. For example, if a user submits an application form for a "software engineer," the server records information about Applicant A in the database based on this information.
[1125] Question Generation
[1126] The server uses the generative AI model to generate questions for applicants based on the saved entry information, such as "Please tell us about the most difficult experience you had in a past project," and sends these to the terminal (interviewer AI device).
[1127] Conducting interviews
[1128] The user selects the interview date and time through an online recruitment portal. At the specified date and time, the user accesses a video call tool such as Zoom through their terminal (PC or mobile device). The server detects that the video call has started and begins the interview based on pre-generated questions. For example, the applicant clicks on the Zoom link at the specified time, and the video call begins. The AI interviewer (server) asks the question, "Tell me about your career history."
[1129] Real-time analysis of responses
[1130] When a user answers a question, the audio and video data are sent to a server in real time. The server then uses an AI analysis model to analyze the answers and evaluate the applicant's skills and experience. For example, if the analysis determines that the applicant has a deep understanding of a particular technology, the server generates questions that dig deeper. It then asks follow-up questions such as, "Please tell us in more detail about a specific project that used that technology."
[1131] Analyzing applicants' tension levels and encouraging them
[1132] The server analyzes the audio and video during the video call to determine the applicant's level of nervousness in real time. If it determines that the applicant is nervous, the server instructs the interviewer AI device to say things like, "Please relax." This allows the user to approach the interview in a more natural state.
[1133] Summarizing interview results and creating feedback materials
[1134] After the interview is over, the server summarizes the answers given during the interview and creates feedback materials. These materials include an evaluation of the applicant, points to ask in the next interview, and concerns. For example, an evaluation point such as "Applicant A is highly evaluated for his / her strong technical skills and extensive leadership experience" may be written. The feedback materials are saved in the HR system and can be accessed from a terminal (such as the HR manager's PC).
[1135] Specific examples
[1136] When Applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. For example, the applicant participates in an interview via Zoom at a specified time, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answers in real time and generates relevant follow-up questions. After the interview, the data obtained can be used to create feedback materials that HR personnel can use to prepare the next interview.
[1137] Prompt Sentence Examples
[1138] "Tell me about your career so far."
[1139] "What is the most challenging experience you've had on a recent project?"
[1140] "Can you please tell us a bit more about a specific project where you used that technology?"
[1141] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1142] Step 1: Obtaining entry information
[1143] Users access an online recruitment portal and submit their application.
[1144] Specifically, the user enters their resume, job history, and reason for applying into an online form and clicks the submit button.
[1145] Input: Application information such as resume, work history, and reasons for applying
[1146] Output: Entry information stored in the database
[1147] The server receives this information and stores it in a database.
[1148] Step 2: Generate questions
[1149] The server analyzes the entry information stored in the database.
[1150] Specifically, the server uses a generative AI model to receive entry information as input and generate questions based on that information.
[1151] Input: Entry information stored in the database
[1152] Output: Questions to ask the applicant
[1153] For example, a question such as "Please tell us about the most difficult experience you had in a past project" is generated. The generated question is sent to the device.
[1154] Step 3: Set up an interview
[1155] Users select interview dates and times through an online recruitment portal.
[1156] Specifically, the user uses the calendar function to select an available time slot and confirm the reservation.
[1157] Input: Desired interview date and time and contact information
[1158] Output: Notification of confirmed interview schedule
[1159] The server receives this information and confirms the interview schedule.
[1160] Step 4: Conducting the interview
[1161] At the specified date and time, users access video calling tools such as Zoom using their device (PC or mobile device).
[1162] The server detects that a video call has started and begins the interview based on pre-generated questions.
[1163] Input: Access to video call tools such as Zoom, generated questions
[1164] Output: Conducting an interview via video call
[1165] For example, a user clicks on a Zoom link at a specified time to start a video call. The AI interviewer (server) asks, "Tell me about your career history."
[1166] Step 5: Real-time analysis of responses
[1167] As the user answers the questions, audio and video data is transmitted in real time to the server.
[1168] The server analyzes the data using an AI analysis model and evaluates the answers.
[1169] Input: Applicant's audio and video data
[1170] Output: Analysis results and new follow-up questions
[1171] For example, if the analysis reveals that you have a deep understanding of a particular technology, the server will generate a follow-up question such as, "Please tell us more about a specific project using that technology."
[1172] Step 6: Analyze the applicant's level of nervousness and encourage them
[1173] The server analyzes the audio and video during the video call in real time to assess the applicant's level of nervousness.
[1174] Input: Audio and video data during a video call
[1175] Output: Tension analysis results and verbal instructions
[1176] If the server determines that the applicant is nervous, it instructs the AI interviewer device to say something like, "Please relax." Specifically, the AI interviewer uses a voice message to encourage the applicant to relax.
[1177] Step 7: Summarize the interview results and create feedback materials
[1178] After the interview is completed, the server summarizes the answers given during the interview and creates feedback materials.
[1179] Input: Answers given during the interview and analysis results
[1180] Output: Feedback materials and evaluation points for the next interview
[1181] For example, the server can create a summary such as "The applicant has strong technical skills, but lacks details regarding leadership experience" and include it in the feedback material, which is then stored in the HR system and used by HR personnel for the next interview.
[1182] (Application example 1)
[1183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1184] Until now, human resource costs and time have been major issues in corporate recruitment activities and factory performance evaluations. Furthermore, because the quality of the evaluation depends on human subjectivity, there is a risk of mismatches and unfair evaluations. Furthermore, creating training programs to efficiently improve work performance requires a great deal of effort. There is a need to solve these issues and provide a more efficient and fair evaluation system and training programs suited to each individual worker.
[1185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1186] In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants and workers via video communication and receiving answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing interview results and performance evaluation results to create feedback materials, and means for evaluating performance and providing training programs. This makes it possible to efficiently evaluate applicants and workers, achieve fair and error-free evaluations, and provide optimal training programs based on their individual characteristics.
[1187] "Entry information" refers to information such as resumes, work history, and work experience provided by applicants and workers, and is data necessary for evaluation.
[1188] "Means for generating questions" refers to a function that automatically creates questions appropriate for applicants and workers based on entry information and analysis results.
[1189] "Video communication" is a technology that allows for real-time video and audio communication over the Internet, and is used for interviews and job evaluations.
[1190] "Means for obtaining responses" refers to a function for receiving responses from applicants and workers via video communication.
[1191] "Real-time analytical means" refers to the analytical capabilities required to instantly evaluate applicant and worker responses and generate follow-up questions as needed.
[1192] "Means for generating follow-up questions" refers to a function that automatically creates more detailed questions based on the initial answers.
[1193] "Means for summarizing interview results and performance evaluation results to create feedback materials" refers to a device that has the function of summarizing information obtained from interviews and performance evaluations and providing it in a concise, easy-to-understand format.
[1194] "Means for evaluating work performance and providing training programs" refers to a function that evaluates a worker's ability to perform work and assigns an appropriate training program based on the results.
[1195] The present invention is a system for evaluating the work performance of factory workers and providing training programs based on the evaluation results. The main technical elements of this system include obtaining entry information, generating questions, conducting evaluation sessions via video communication, analyzing responses in real time, and creating feedback materials based on the evaluation results.
[1196] Obtaining entry information
[1197] First, the server obtains the application information from the worker through the portal system in the factory. This application information includes the worker's resume, work history, work experience, etc., and this data is stored in a database.
[1198] Question Generation
[1199] Based on the acquired entry information, the server uses an AI model to generate questions for the worker. For example, it generates questions such as "What was the most difficult task in your recent work?" based on the worker's past work history. The generated questions are sent to the terminal.
[1200] Conducting a business evaluation session
[1201] At the designated time, workers access a video chat tool such as Zoom via their device (PC or tablet). The server detects that the video call has started and begins the evaluation session based on pre-generated questions.
[1202] Real-time analysis of responses
[1203] As workers answer questions, the server analyzes their answers in real time, assessing their understanding of specific work processes based on their answers, and generates follow-up questions as needed, such as, "Is there anything in that work process that you feel needs improvement?"
[1204] Analyzing tension levels and encouraging others
[1205] The server analyzes the worker's tone of voice and facial expressions, and if they appear nervous, it will say something like, "Relax." This allows the worker to be evaluated in a natural state.
[1206] Summarizing the assessment results and preparing training materials
[1207] Once the evaluation session is over, the server summarizes the responses and compiles the evaluation results into a feedback document, which includes an evaluation of work performance, areas for improvement, and the next training program to be implemented. For example, a summary might be created such as, "Work speed is high, but quality control needs improvement."
[1208] The hardware used includes servers, databases, PCs, and tablets, while the software includes a portal system, Zoom API, voice analysis engines (e.g., Google Speech-to-Text), natural language processing models (e.g., GPT-3), and facial expression analysis software (e.g., OpenCV).
[1209] Specific examples
[1210] When Worker A submits application information (such as resume, work history, and work experience) to the online portal, the server generates appropriate questions based on this information. For example, it asks questions such as, "What was the most difficult task you recently performed?" or "Is there anything in the work process that you feel needs improvement?" Worker A participates in the evaluation session using a video chat tool, and analysis is carried out in real time. Afterwards, feedback materials summarizing the evaluation results are created, and a training program is provided.
[1211] Prompt Sentence Examples
[1212] "What was the most difficult task you recently undertook? Please explain in detail why."
[1213] "Are there any particular work processes you feel need improvement?"
[1214] "Please tell us about any major problems you have faced in the past and how you dealt with them."
[1215] In this way, the present invention improves the efficiency of work evaluation and training support within a factory, and contributes greatly to improving work performance and preventing mismatches.
[1216] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1217] Step 1:
[1218] The server obtains application information from users (workers) through the portal system. Users fill out a self-evaluation sheet and provide details of their resume, work history, and work experience. This application information is stored in a database by the server. The input is the application information provided by the user, and the output is the application information stored in the database.
[1219] Step 2:
[1220] The server uses the AI model with a question generation means to generate appropriate questions for the worker based on the stored entry information. The data input into the AI model includes the worker's resume, work history, and work experience, and specific questions are generated as output. For example, a question such as "What was the most difficult task in your recent work?" is generated.
[1221] Step 3:
[1222] The user accesses a video chat tool such as Zoom through a device (PC or tablet) at a specified date and time. When the server detects the start of the video call, it starts an evaluation session based on the generated questions. The input is access to the video chat tool, and the output is the start of the session. Specifically, when the user clicks the Zoom link and the video call starts, the server begins displaying the questions.
[1223] Step 4:
[1224] As users answer questions, the server analyzes the answers in real time. Specifically, a speech analysis engine (e.g., Google Speech-to-Text) is used to convert the speech data into text. A natural language processing model (e.g., GPT-3) then analyzes the text data to understand the answer. Based on this analysis, follow-up questions are generated as needed. The input is the user's speech answer, and the output is the analyzed text and follow-up questions.
[1225] Step 5:
[1226] The server analyzes the user's tone of voice and facial expression to determine the level of tension. It uses voice analysis software and facial expression analysis software (e.g., OpenCV) to detect whether the user is nervous. If necessary, it will say something like, "Please relax." The input is the user's tone of voice and facial expression data, and the output is the analysis results and a message to encourage the user.
[1227] Step 6:
[1228] After the evaluation session is over, the server summarizes the responses and compiles the evaluation results as feedback materials. It uses a text summarization engine (e.g., BERT) to extract key points and create concise feedback materials. It then evaluates the user's performance based on these materials and provides training programs. The input is the user's complete response data, and the output is the summarized feedback materials and training programs.
[1229] This will enable an efficient and fair evaluation system and training programs based on individual characteristics.
[1230] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1231] This invention is a system that uses AI interviewers to efficiently conduct preliminary interviews in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The main technical elements of this system include obtaining application information, generating questions, conducting interviews via video communication, analyzing answers in real time, summarizing interview results and creating feedback materials, and an emotion engine that recognizes user emotions.
[1232] A natural language description of the program's operation
[1233] 1. Acquisition of entry information
[1234] A user (applicant) submits an application to an online recruitment portal, including a resume, a job history, and a motivation statement. The server receives this application information and stores it in a database.
[1235] 2. Question Generation
[1236] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[1237] 3. Conducting the interview
[1238] The user then accesses the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user connects to a video call tool such as Zoom. The server detects the start of the video call and begins the interview.
[1239] For example, an applicant clicks on a Zoom link at the designated time, and a video call begins. The AI interviewer (server) asks questions such as, "Tell me about your career history."
[1240] 4. Real-time analysis of responses
[1241] When a user answers a question, the server analyzes the answer in real time using natural language processing technology. The server also includes an emotion engine that analyzes the user's voice and facial expressions during the answer to collect emotional data.
[1242] 5. Generate follow-up questions
[1243] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[1244] 6. Analyzing applicants' tension levels and encouraging them
[1245] If the server determines that the user is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly."
[1246] 7. Summarizing interview results and creating feedback materials
[1247] After the interview is over, the server analyzes the interview responses and emotional data again to summarize the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[1248] For example, if the analysis indicates that the applicant is technically strong but has concerns about leadership questions, include that information in the summary.
[1249] 8. Providing Feedback Materials
[1250] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[1251] Specific examples
[1252] For example, if applicant A submits an application form to an online recruitment portal and the application information includes "a degree in computer science and work experience in software development," the server generates appropriate questions based on this information. On the day of the interview, the user participates via Zoom, and the AI interviewer asks, "What was the most difficult thing about your most recent project?" The server analyzes the answer in real time, using an emotion engine to detect nervousness from the applicant's facial expressions and tone of voice, and generates appropriate follow-up questions. After the interview, feedback materials can be created based on the data obtained, which can be used by human resources personnel for the next interview.
[1253] In this way, by combining emotion engines, companies can provide a more refined and personalized interview experience, enabling them to conduct their recruitment activities more efficiently and effectively.
[1254] The processing flow will be explained below.
[1255] Step 1:
[1256] Users (applicants) access an online recruitment portal and submit an application form, which includes a resume, a job history, and a statement of motivation for applying.
[1257] Step 2:
[1258] The server retrieves application form information submitted through an online recruitment portal and stores it in a database, including the applicant's name, educational background, work history, skills, etc.
[1259] Step 3:
[1260] The server uses an AI model to generate individual questions based on the saved application information. For example, it generates questions such as "What was the most difficult experience you had in a past project?" based on the applicant's educational background and work history. The generated questions are sent from the server to the terminal (interviewer AI device).
[1261] Step 4:
[1262] The user accesses the online recruitment portal again and selects the date and time of the interview. The information on the confirmed date and time of the interview is saved on the server.
[1263] Step 5:
[1264] At the specified date and time, the user connects to a video call tool such as Zoom using a terminal (PC or mobile device). The server detects the start of the video call and begins the interview.
[1265] Step 6:
[1266] The server asks the user questions through the AI interviewer based on the sent question list. For example, the AI interviewer might ask, "Tell me about your career so far."
[1267] Step 7:
[1268] As users answer questions, the server analyzes the answers in real time using natural language processing technology, and the resulting data is instantly evaluated by an emotion engine.
[1269] Step 8:
[1270] The server generates follow-up questions based on the answers and emotional data. For example, if a user feels nervous when explaining the introduction of a new technology, the server asks a follow-up question such as, "What were the specific benefits of that new technology?"
[1271] Step 9:
[1272] During the interview, if the server determines that the applicant is nervous based on data analyzed by the user's emotion engine, it will use the AI interviewer to say things like, "Please relax. Please answer calmly." This helps to stabilize the applicant's mental state.
[1273] Step 10:
[1274] After the interview, the server analyzes the collected response data and emotional data again to summarize the interview results. This summary includes the applicant's strengths, weaknesses, and emotional reactions. For example, it may include information such as, "The applicant is technically strong, but is anxious about questions regarding leadership."
[1275] Step 11:
[1276] The server creates feedback materials based on the summarized interview results and stores them in the personnel system. These materials include questions to ask in the next interview, evaluations of the applicant, and areas of concern. The saved feedback materials can be accessed from a terminal (such as the personnel manager's PC).
[1277] Step 12:
[1278] All interview data is collected by the server and used to improve the accuracy of the AI model and emotion engine, enabling more efficient and accurate question generation and analysis for future interviews.
[1279] In this way, through the specific actions taken at each step, this system streamlines a company's recruitment process and helps them select the right talent.
[1280] Example 2
[1281] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1282] In traditional recruitment activities, a large amount of resources are allocated to conducting initial interviews (stage zero interviews), which often results in mismatches between candidates and candidates. This problem is particularly pronounced in companies conducting large-scale recruitment activities, increasing the burden on human resources personnel and contributing to rising recruitment costs. Furthermore, traditional interview methods make it difficult to accurately grasp the applicant's emotions and level of nervousness in a short amount of time, which makes it difficult for applicants to relax and have an opportunity to present themselves. To solve these problems, a more efficient and objective interview method is required.
[1283] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for acquiring entry information, a means for generating questions using a generative AI model based on the acquired entry information, a means for asking questions to applicants via video communication and obtaining their answers, a means for analyzing the answers in real time and generating follow-up questions using natural language processing technology and an emotion engine, and a means for summarizing the interview results and creating feedback materials. This enables efficient and effective initial interviews, reduces recruitment costs, and prevents mismatches with applicants. Furthermore, by analyzing the applicant's level of nervousness using the emotion engine during video communication and providing appropriate prompts, the applicant can be given an opportunity to relax and promote themselves.
[1284] "Entry information" refers to information submitted by applicants, such as resumes, work history, and reasons for applying.
[1285] "Generative AI model" refers to an artificial intelligence model that uses natural language processing technology to generate individual questions from the input entry information.
[1286] "Server" refers to a central processing unit that performs processes such as obtaining entry information, generating questions, analyzing answers, generating follow-up questions, summarizing interview results, and creating feedback materials.
[1287] "Video communication" refers to a means of communication that uses video calling tools such as Zoom to share audio and video in real time with applicants in remote locations.
[1288] "Natural language processing technology" refers to the technology of analyzing and processing human language using a computer.
[1289] "Emotion engine" refers to a system that analyzes an applicant's emotional state from their tone of voice and facial expressions.
[1290] "Follow-up questions" refer to additional questions that are generated based on the applicant's initial answers or emotional state.
[1291] "Feedback materials" are documents summarizing the results of an interview, including questions for the next interview, evaluations of the applicant, and concerns.
[1292] This invention is a system for efficiently conducting initial interviews using AI interviewers in corporate recruitment activities, reducing recruitment costs and preventing mismatches. The system includes functions for acquiring application information, generating questions, conducting interviews via video communication, analyzing responses in real time, generating follow-up questions, summarizing interview results and creating feedback materials, and an emotion engine for recognizing applicants' emotions.
[1293] Obtaining entry information
[1294] A user (applicant) first accesses an online recruitment portal and submits an application form, which includes a resume, a curriculum vitae, and a statement of reasons for applying. The server receives this application information through an HTTP request and performs the necessary database transactions to store it in a database. Specific software used includes database management systems such as MySQL or PostgreSQL.
[1295] Question Generation
[1296] The server retrieves the user's application information from the database and uses a Python script to call an open-source natural language processing library (e.g., Transformers). A prompt sentence is input into the generative AI model to generate individual questions. The generated questions are sent from the server to the terminal (interviewer AI device) in JSON format. An example of a specific prompt sentence is, "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please use an open-source natural language processing library to generate individual, specific questions."
[1297] Conducting interviews
[1298] The user logs in again to the online recruitment portal and selects the date and time of the interview. At the specified date and time, the user clicks the Zoom link to connect to the video call. The server monitors the connection status through the video call tool's API and detects the user's connection. At the start of the interview, the server sends pre-generated questions to the device, which then conducts the assessment. Specific video call tools used include Zoom and WebEx.
[1299] Real-time analysis of responses
[1300] When a user answers a question, the audio data of the answer is sent to the server in real time via the video chat tool's API. The server converts the audio data into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The converted text is then analyzed using a natural language processing library. Furthermore, an emotion engine analyzes the user's tone of voice and facial expressions to collect emotional data.
[1301] Generate follow-up questions
[1302] The server generates follow-up questions based on the collected emotional data and the content of the answers. For example, if the user is nervous about a particular topic, the server generates follow-up questions such as, "What were the specific benefits of this new technology?" This is again done using a generative AI model.
[1303] Analyzing applicants' tension levels and encouraging them
[1304] If the server determines that the user is nervous based on the data analyzed by the emotion engine, it will use the AI interviewer to say something like, "Please relax. Please answer calmly." Specifically, it uses the results of transcription and emotion analysis to display appropriate instructions on the screen and play them back aloud.
[1305] Summarizing interview results and creating feedback materials
[1306] After the interview, the server analyzes the interview responses and emotional data again. For example, it uses a text analysis engine to extract specific keywords and emotional states, and then summarizes the interview results, including the applicant's strengths, weaknesses, and emotional responses.
[1307] Providing feedback materials
[1308] The server generates feedback materials in PDF format based on the summarized interview results. The PDF file is then saved in the HR system and a notification is sent. HR personnel can access the materials from their own devices (such as PCs) and prepare for the next interview. This function enables companies to conduct their recruitment activities efficiently and effectively.
[1309] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1310] Step 1: Obtaining entry information
[1311] A user (applicant) accesses an online recruitment portal and submits an application form. The application form includes information such as a resume, a curriculum vitae, and a reason for applying. The submitted application information is sent to the server as an HTTP request. The server receives the application information and stores it in a database using a database management system (e.g., MySQL). Specifically, the operation involves executing an INSERT statement against the database. The input in this process is the application information, and the output is the application information stored in the database.
[1312] Step 2: Generate questions
[1313] The server retrieves the saved application information from the database. Based on the retrieved information, questions are generated using a generative AI model (e.g., Transformers). This process is carried out by executing a Python script and calling a natural language processing library. An example prompt is: "Based on the applicant's application information (educational background, work history, motivation for applying, etc.), please generate individual, specific questions using an open-source natural language processing library." The generated questions are sent from the server to the terminal in JSON format. The input in this process is the application information, and the output is the generated questions.
[1314] Step 3: Conducting the interview
[1315] The user logs in to the online recruitment portal again and selects the date and time of the interview. At the specified date and time, the user clicks on a link in a video calling tool such as Zoom to connect to the video call. The server monitors the connection status through the video calling tool's API and detects the user's connection. When the interview starts, the server sends pre-generated questions to the terminal and conducts the questions. The input is the interview date and time and the questions, and the output is the video call that has started.
[1316] Step 4: Real-time analysis of responses
[1317] When a user answers a question via video call, the audio of the answer is sent to the server in real time via the video call tool's API. The server then converts the audio data of the answer into text using a transcription service (e.g., Google Cloud Speech-to-Text API). The server then analyzes the answer using a natural language processing library (e.g., Transformers), and the emotion engine uses this information to collect emotional data from the tone of voice and facial expressions. The input to this process is the audio data of the answer, and the output is the analyzed answer text and emotional data.
[1318] Step 5: Generate follow-up questions
[1319] The server generates follow-up questions based on the analyzed answer text and sentiment data. Using the generative AI model, it runs a Python script to generate appropriate follow-up questions. These follow-up questions are also sent to the device in JSON format. For example, if a user is nervous about a particular topic, a follow-up question might be generated: "What were the specific benefits of this new technology?" The input to this process is the analyzed answer text and sentiment data, and the output is a follow-up question.
[1320] Step 6: Analyze the applicant's level of nervousness and encourage them
[1321] The server analyzes the user's level of nervousness based on the data collected by the emotion engine. If it determines that the user is nervous, the server will use the AI interviewer to give appropriate encouragement. For example, it may generate encouragement such as "Please relax. Please answer calmly" and play it back aloud. The input in this process is emotional data, and the output is a encouragement message.
[1322] Step 7: Summarize the interview results and create feedback materials
[1323] After the interview is over, the server re-analyzes the responses and emotional data from the interview. For example, it can use a text analysis engine to extract specific keywords and emotional states and summarize evaluation points. Based on the analysis results, it summarizes the interview results and creates feedback materials in PDF format. The input for this process is the responses and emotional data, and the output is the summarized interview results and feedback materials.
[1324] Step 8: Provide feedback materials
[1325] The server generates feedback materials in PDF format based on the summarized interview results and saves them in the human resources system. The saved feedback materials are notified to human resources personnel so that they can access them from their own devices (such as PCs). The input in this process is the summarized interview results, and the output is the generated feedback materials.
[1326] This system allows companies to carry out their recruitment activities efficiently and effectively.
[1327] (Application example 2)
[1328] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1329] Conventional security systems have had difficulty analyzing visitors' behavior and emotions in real time and responding automatically. While early detection of suspicious individuals and rapid response are required, fully automating this process has been a difficult task. The present invention aims to solve these problems and realize more efficient and accurate security responses.
[1330] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring entry information, means for generating questions based on the acquired entry information, means for asking questions to applicants via video communication and obtaining answers, means for analyzing the answers in real time and generating follow-up questions, means for summarizing the interview results and creating feedback materials, means for acquiring video and audio data and analyzing the visitor's emotions from this information, and means for generating follow-up actions based on the results of the emotion analysis. This makes it possible to analyze the visitor's behavior and emotions in real time and respond quickly and appropriately.
[1331] "Entry information" refers to information submitted by applicants, such as resumes, job history, and reasons for applying.
[1332] The "means for generating a question" is a function for generating an appropriate question based on the acquired entry information.
[1333] "Video communication" is a means of communication that uses video and audio to enable applicants and interviewers to interact online.
[1334] "Real-time analysis" refers to the process of instantly analyzing applicant responses and generating follow-up questions as needed.
[1335] The "means for generating follow-up questions" is a function for automatically generating follow-up questions in response to the applicant's answers.
[1336] The "means of summarizing interview results and creating feedback materials" is a function for summarizing the answers and emotional data given during the interview and creating feedback materials.
[1337] "Video and audio data" refers to visual and audio information obtained through cameras and microphones.
[1338] The "means for analyzing emotions" is a function for analyzing the emotions of visitors from the acquired video and audio data.
[1339] The "means for generating follow-up actions" is a function for automatically taking appropriate action against visitors based on the results of sentiment analysis.
[1340] To realize the present invention, the following hardware and software are used.
[1341] Hardware Configuration
[1342] 1. Security Guard Robot: A robot equipped with a camera and microphone to collect video and audio data (e.g., a typical security guard robot with a camera and microphone).
[1343] 2. Server: A high-performance computer for data processing and analysis (e.g., a typical server).
[1344] Software Configuration
[1345] 1. Natural language processing engine: Technology for analyzing information from captured video and audio data and generating appropriate questions and follow-up actions (e.g., OpenAI GPT-4, Google Cloud Dialogflow).
[1346] 2. Video and audio analysis engine: Technology for analyzing video and audio data and recognizing visitors' emotions (e.g., Microsoft Azure Cognitive Services, Google Cloud Video Intelligence API).
[1347] 3. Emotion recognition engine: Technology for analyzing emotions from collected data (e.g., Affectiva SDK, Microsoft Azure Emotion API).
[1348] Basic operations
[1349] 1. Security guard robots patrol monitored areas such as shopping malls and collect video and audio data in real time.
[1350] 2. The collected data is sent to a server and analyzed using a natural language processing engine and a video and audio analysis engine.
[1351] 3. The server analyzes the visitor's behavior and emotions to detect suspicious behavior and emotions such as tension, anxiety, and anger.
[1352] 4. The server utilizes an emotion recognition engine to generate appropriate follow-up actions and prompts.
[1353] 5. The generated instructions are sent to the security guard robot, which then takes appropriate action against the visitor.
[1354] Specific examples
[1355] For example, a security guard robot patrolling a shopping mall collects video and audio of visitor A. If the server analyzes that visitor A suddenly feels anxious, the robot will say, "Please relax. Is there anything I can help you with?" This information is also reported to the security center, and a human security guard will respond if necessary.
[1356] Prompt Sentence Examples
[1357] "Build a system for a security guard robot patrolling a shopping mall to collect real-time video and audio of visitors and analyze their emotions based on specific actions, facial expressions, and tone of voice. This involves the following steps:
[1358] 1. Video and audio data collection
[1359] 2. Real-time analysis using a sentiment analysis engine
[1360] 3. Take follow-up actions based on sentiment data
[1361] 4. Database recording and alert generation
[1362] In this way, by implementing the present invention, security operations can be automated and highly accurate crisis management can be achieved.
[1363] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1364] Step 1:
[1365] The security guard robot patrols the shopping mall and collects video and audio data in real time. The robot's cameras and microphones capture this data and send it to a server.
[1366] Input: Video and audio data from inside the shopping mall
[1367] Output: Video and audio data sent to the server
[1368] Step 2:
[1369] The server passes the received video and audio data to a natural language processing engine and a video and audio analysis engine for analysis, such as analyzing the visitor's behavior, speech content, tone of voice, and facial expressions to recognize their emotional state (tension, anxiety, anger, etc.).
[1370] Input: Video and audio data sent to the server
[1371] Output: Analyzed behavioral and emotional data
[1372] Step 3:
[1373] Based on the analysis results, the server uses an emotion recognition engine to understand the visitor's emotional state in detail, and in this process, specific emotions such as "tension," "anxiety," and "anger" are identified.
[1374] Input: Parsed behavioral and emotional data
[1375] Output: Detailed emotional state data
[1376] Step 4:
[1377] The server generates follow-up actions based on the detailed emotional state data. For example, if the visitor is feeling anxious, it generates a follow-up prompt such as "Please relax. Is there anything I can help you with?"
[1378] Input: Detailed emotional state data
[1379] Output: Generated follow-up action instructions
[1380] Step 5:
[1381] The generated follow-up action instructions are sent to the security guard robot, and the robot responds to the visitor accordingly. For example, the robot may tell the visitor to "relax."
[1382] Input: Generated follow-up action instructions
[1383] Output: Response action by security guard robot
[1384] Step 6:
[1385] The server records the results of follow-up actions and visitor responses in a database and stores relevant information, which can later be analyzed and used for further security measures.
[1386] Input: Follow-up action results and visitor response data
[1387] Output: Behavioral results and response data recorded in a database
[1388] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1389] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1390] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1391] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1392] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1393] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1394] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1395] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1396] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1397] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1398] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1399] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1400] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1401] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1402] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1403] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1404] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1405] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1406] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1407] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1408] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1409] The following is further disclosed regarding the above embodiment.
[1410] (Claim 1)
[1411] A means for obtaining entry information;
[1412] A means for generating a question based on the acquired entry information;
[1413] A means to ask questions and receive answers from applicants via video communication;
[1414] a means of analyzing responses in real time and generating follow-up questions;
[1415] A means of summarizing interview results and creating feedback materials;
[1416] A system including:
[1417] (Claim 2)
[1418] 2. The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness and speaking to the applicant accordingly when said video communication is carried out.
[1419] (Claim 3)
[1420] 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis.
[1421] "Example 1"
[1422] (Claim 1)
[1423] A means for obtaining entry information;
[1424] A means for generating a question based on the acquired entry information;
[1425] A means to ask questions and receive answers from applicants via video communication;
[1426] a means of analyzing responses in real time and generating follow-up questions;
[1427] A means of summarizing interview results and creating feedback materials;
[1428] a means for generating questions and follow-up questions using a generative AI model;
[1429] A means for analyzing the applicant's nervousness during the video communication;
[1430] A system including:
[1431] (Claim 2)
[1432] 2. The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness and speaking to the applicant accordingly when said video communication is carried out.
[1433] (Claim 3)
[1434] 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis.
[1435] "Application Example 1"
[1436] (Claim 1)
[1437] A means for obtaining entry information;
[1438] A means for generating a question based on the acquired entry information;
[1439] A means to ask questions and receive answers from applicants via video communication;
[1440] a means of analyzing responses in real time and generating follow-up questions;
[1441] A means of summarizing interview results and creating feedback materials;
[1442] a means of evaluating job performance and providing training programs;
[1443] A system including:
[1444] (Claim 2)
[1445] 2. The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness and speaking to the applicant accordingly when said video communication is carried out.
[1446] (Claim 3)
[1447] 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis.
[1448] "Example 2: Combining Emotion Engines"
[1449] (Claim 1)
[1450] A means for obtaining entry information;
[1451] A means for generating questions using a generative AI model based on the acquired entry information;
[1452] A means to ask questions and receive answers from applicants via video communication;
[1453] a means for analyzing the responses in real time and generating follow-up questions using natural language processing techniques and an emotion engine;
[1454] A means of summarizing interview results and creating feedback materials;
[1455] A system including:
[1456] (Claim 2)
[1457] The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness using an emotion engine and speaking to the applicant accordingly when conducting video communication.
[1458] (Claim 3)
[1459] 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis.
[1460] "Application example 2 when combining emotion engines"
[1461] (Claim 1)
[1462] A means for obtaining entry information;
[1463] A means for generating a question based on the acquired entry information;
[1464] A means to ask questions and receive answers from applicants via video communication;
[1465] a means of analyzing responses in real time and generating follow-up questions;
[1466] A means of summarizing interview results and creating feedback materials;
[1467] means for acquiring video and audio data and analyzing visitor sentiment from this information;
[1468] means for generating follow-up actions based on the results of the sentiment analysis;
[1469] A system including:
[1470] (Claim 2)
[1471] 2. The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness and speaking to the applicant accordingly when said video communication is carried out.
[1472] (Claim 3)
[1473] 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis. [Explanation of symbols]
[1474] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for obtaining entry information; A means for generating a question based on the acquired entry information; A means to ask questions and receive answers from applicants via video communication; a means of analyzing responses in real time and generating follow-up questions; A means of summarizing interview results and creating feedback materials; A system including:
2. 2. The system according to claim 1, further comprising means for analyzing the applicant's level of nervousness when said video communication is being carried out and for speaking to the applicant in accordance with said level of nervousness.
3. 10. The system of claim 1, further comprising means for storing the entry information in a database and providing feedback material to a human resources system based on the analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A