system

The system addresses recruitment inefficiencies by analyzing application forms and generating voice questions for automated evaluation, enhancing efficiency and fairness in candidate assessment.

JP2026041367APending Publication Date: 2026-03-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Recruitment processes are costly, time-consuming, and prone to subjective judgment, leading to mismatches and inefficiencies in candidate evaluation.

Method used

A system that analyzes application forms using natural language processing, generates voice questions with generative AI, records and evaluates responses, and provides feedback to improve efficiency and fairness.

Benefits of technology

Enables efficient scrutiny of large numbers of application forms, reduces recruitment costs, and ensures fairer, more objective evaluations by automating the initial interview process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041367000001_ABST
    Figure 2026041367000001_ABST
Patent Text Reader

Abstract

Provide a system. A means for receiving an entry form; A means for analyzing the contents of the application form and extracting characteristics of the applicant; means for generating a query based on said characteristics; means for converting the question into speech; means for recording the applicant's responses to said audio questions; means for analyzing the responses and generating a reputation score; means for storing said evaluation scores and generating feedback; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Recruiting activities require scrutinizing a large number of application forms and conducting multiple interviews, which is extremely costly and time-consuming. It also makes it difficult for recruitment teams to determine the suitability of candidates, resulting in the risk of early turnover and mismatches. Furthermore, traditional interview methods make it difficult to ensure fairness, and there is a risk of subjective judgment. A new and effective system is needed to solve these issues. [Means for solving the problem]

[0005] The present invention aims to improve the efficiency and accuracy of the recruitment process by providing a system that analyzes application forms, generates questions and synthesizes voice using AI, and records and analyzes responses.

[0006] Specifically, the system includes a means for receiving application forms, a means for analyzing the contents of the application form to extract the applicant's characteristics, a means for generating questions using generative AI based on the characteristics, a means for converting the questions into voice, a means for recording the applicant's responses to the voice questions, a means for analyzing the voice responses to generate an evaluation score, and a means for saving the evaluation score and generating feedback. This system enables the efficient scrutiny of large numbers of application forms, achieving fairer and more objective evaluations. It also contributes to reducing recruitment costs and preventing mismatches.

[0007] An "application form" is a written or electronic document submitted by a job applicant, which contains information such as personal information, educational background, work history, skills, and reasons for applying.

[0008] "Analysis" is the process of extracting information from input data (such as application forms) and organizing and evaluating it according to a specific purpose.

[0009] "Applicant" means an individual applying for a position whose information and responses are processed by the System.

[0010] "Characteristics" refers to characteristic information such as the applicant's skills, experience, and inclinations extracted through analysis.

[0011] "Generative AI" refers to algorithms or systems that generate questions and ideas without human intervention, especially those that use natural language processing to generate written or conversational information.

[0012] "Questions" are interactive text or voice messages generated to assess an applicant's aptitudes and skills.

[0013] "Speech synthesis" is a technology that converts text information into voice data, generating natural-sounding speech.

[0014] "Recording" is the process of digitally recording an applicant's voice responses.

[0015] "Response" refers to the answer provided by an Applicant to an interview question, either by voice or text.

[0016] The "evaluation score" is a numerical representation of the applicant's aptitude and skill level, based on an analysis of the content of the responses.

[0017] "Feedback" is information provided to an applicant based on the evaluation results, including evaluation information and instructions regarding next steps. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The system of the present invention is composed of users (applicants), a server, and terminals, and each component operates in cooperation with the others as follows.

[0040] Submit your application:

[0041] The user submits an application form.

[0042] The user accesses a recruitment website and uses a dedicated application form to enter the required information (personal information, educational background, work history, skills, reasons for applying, etc.). Once the information is complete, the user clicks the "Submit" button to send the application form to the server.

[0043] Application Form Analysis:

[0044] The server analyzes the application form.

[0045] The server checks the received application form and analyzes its contents using natural language processing (NLP) technology. Specifically, it extracts the applicant's skills, experience, and motivation for applying from the text, and extracts the necessary characteristic information.

[0046] Question generation:

[0047] The server generates questions using generative AI.

[0048] Based on the analyzed characteristics, the server uses generative AI to generate appropriate questions for the applicant, such as questions about details related to the applicant's work history and technical skills, to assess the applicant's aptitude and skills.

[0049] Text-to-Speech:

[0050] The server uses a voice generation AI to synthesize the interviewer's voice.

[0051] The server converts the generated questions from text to audio files. Using speech generation AI, it generates questions as audio files in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0052] Interview Conducted:

[0053] The user responds to the voice questions.

[0054] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to each question in their own voice and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0055] Response analysis:

[0056] The server analyzes the response.

[0057] The server converts the user's voice response into text and analyzes it again using natural language processing technology. Specifically, it extracts elements from the applicant's response that evaluate their technical skills, problem-solving ability, and communication ability, and generates an evaluation score.

[0058] Grading and feedback:

[0059] The server scores the responses and stores the results.

[0060] The server calculates the evaluation points for each question and generates an overall score. The evaluation results and feedback are notified to the applicant, and the applicant is guided to the next step. The generated feedback is also provided to the recruiter and used as reference material for the subsequent selection process.

[0061] Examples:

[0062] Taking the hiring process for software engineers as an example, a user would list "3+ years of development experience in Python" as a skill on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a project using Python and how you solved them," and then perform voice synthesis. The user would then respond to the questions by voice, and the response flow would be sent to the server for analysis and evaluation. This process ensures a fair and efficient initial interview.

[0063] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0064] The processing flow will be explained below.

[0065] Step 1:

[0066] The user submits an application form.

[0067] The user accesses a recruitment website and enters the required information into the application form, such as personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0068] Step 2:

[0069] The server analyzes the application form.

[0070] The server receives the submitted application form data and uses natural language processing (NLP) technology to analyze the content of the application form and extract characteristic information such as the applicant's work history, skills, and motivation for applying.

[0071] Step 3:

[0072] The server generates questions using generative AI.

[0073] The server uses generative AI to generate appropriate questions based on the analyzed characteristics, such as specific questions about past project experience and problem-solving ability based on the applicant's work history.

[0074] Step 4:

[0075] The server converts the question into speech.

[0076] The server converts the generated question text into an audio file using speech generation AI. Through this process, a naturally spoken question is prepared as audio data.

[0077] Step 5:

[0078] The user responds to the voice questions.

[0079] An interview-specific application is launched on the device (user's PC or smartphone). The server sends an audio file to the device, which then plays the audio questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0080] Step 6:

[0081] The server analyzes the response.

[0082] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score.

[0083] Step 7:

[0084] The server scores the responses and stores the results.

[0085] The server calculates the evaluation points for each question and generates an overall score. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who uses them in the next selection step.

[0086] Through the above processing steps, the system of the present invention realizes efficient scrutiny of application forms and highly accurate initial interviews.

[0087] Example 1

[0088] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0089] In the conventional hiring process, reviewing applicants' application forms and conducting interviews takes a great deal of time and effort, and there are issues with inconsistent evaluation criteria and a high degree of reliance on subjective judgment. Another problem is that the quality of questions posed to applicants varies depending on the interviewer, resulting in a lack of fairness. A new system is needed to solve these issues and realize an efficient and fair selection process.

[0090] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0091] In this invention, the server includes means for receiving application forms, means for analyzing the contents of the application form using natural language processing technology and extracting characteristic information about the applicant, means for generating questions using generative AI technology based on the characteristic information, means for converting the questions into voice using speech synthesis technology, means for the applicant to record a response to the voice question, means for converting the response voice into text using speech recognition technology and re-analyzing the content to generate an evaluation score, and means for saving the evaluation score and generating feedback. This automates the confirmation of application form contents and the implementation of interviews, enabling an efficient and fair hiring process.

[0092] An "entry sheet" is a document in which an applicant writes down their personal information, educational background, work history, skills, reasons for applying, etc.

[0093] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language.

[0094] "Characteristic information" refers to information such as skills, work history, and motivation for applying that is extracted from the applicant's application form.

[0095] "Generative AI technology" is a technology in which artificial intelligence generates new text or data based on given information.

[0096] "Questions" are a series of questions that are generated based on the applicant's characteristic information and used during the interview.

[0097] "Speech synthesis technology" is a technology that converts text data into speech format.

[0098] A "response" is an oral response given by an applicant to a question during an interview.

[0099] "Speech recognition technology" is a technology that converts voice data into text data.

[0100] The "evaluation score" is a score calculated based on a set of criteria by analyzing the applicant's responses.

[0101] "Feedback" refers to information such as evaluation results and comments provided to applicants based on their evaluation scores.

[0102] The system of the present invention is composed of a user, a server, and a terminal, which operate in cooperation with each other. A specific embodiment of the system will be described below.

[0103] User operations

[0104] First, a user accesses a recruitment website and fills out an application form using a dedicated application form entry form. The application form includes personal information, educational background, work history, skills, and reasons for applying. Once the user has completed the entry, they click the "Submit" button to send the application form data to the server.

[0105] Server Processing

[0106] The server receives the application form and analyzes its contents using natural language processing technology, specifically using Python natural language processing libraries such as spaCy and NLTK. This allows the server to extract characteristic information such as the applicant's skills, work history, and motivation for applying.

[0107] Next, the server generates questions based on the extracted characteristic information using generative AI technology. For generative AI technology, we use OpenAI's GPT-3 model. For example, we create prompt sentences like the following:

[0108] "Describe a challenge you faced while working on a Python project and how you solved it."

[0109] The generated questions are converted into audio files by the server using speech synthesis technology, such as Google® Text-to-Speech (GCP TTS) or Amazon Polly. The generated audio questions are then used in the user's interview application.

[0110] User response

[0111] The user launches the interview application and connects to the server. The server sequentially transfers audio files to the user's device and plays back the questions. The user responds to each question verbally and records their responses using the device's microphone. Once recording is complete, the response data is uploaded from the user's device to the server.

[0112] Server response analysis

[0113] The server receives the response speech and converts it into text using speech recognition technology, such as Google Cloud Speech-to-Text or IBM Watson® Speech to Text. The converted text is then analyzed again using natural language processing technology.

[0114] Finally, the server evaluates the applicant's technical skills, problem-solving ability, and communication ability based on the responses and generates an evaluation score. The evaluation results and feedback are notified to the applicant and provided to the recruiter, enabling efficient and fair initial interviews.

[0115] Specific examples

[0116] In the software engineer recruitment process, a user writes "3+ years of development experience in Python" on an application form. The server analyzes this information and uses generative AI technology to generate a question such as "Please explain the challenges you faced in a project using Python and how you solved it," which is then converted into an audio file using speech synthesis technology. The user responds to the question by voice, and the response data is sent to the server, where it is analyzed and evaluated. This ensures a fair and efficient initial interview.

[0117] The above is a specific embodiment of the present invention.

[0118] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0119] Step 1: Submit your application form

[0120] The user creates an entry form and sends it to the server.

[0121] Input: The user accesses a recruitment website and enters the necessary information, such as personal information, educational background, work history, skills, and reasons for applying, into a dedicated application form.

[0122] How it works: When the user clicks the "Submit" button, the application form data is sent to the server via an HTTP POST request.

[0123] Output: The application form data is saved on the server.

[0124] Step 2: Analyzing the application form

[0125] The server analyzes the received application form data and extracts characteristic information.

[0126] Input: Entry sheet data received by the server.

[0127] How it works: The server uses Python to analyze the application form text with a natural language processing library (e.g., spaCy) and extracts characteristic information such as the applicant's skills, work history, and motivation for applying. Specifically, it uses NLP techniques such as morphological analysis, named entity extraction, and keyword identification.

[0128] Output: The extracted characteristic information is stored in a database.

[0129] Step 3: Question Generation

[0130] The server generates questions based on the characteristic information using generative AI technology.

[0131] Input: The characteristic information parsed by the server.

[0132] How it works: The server uses generative AI technology (e.g., OpenAI GPT-3 model) to generate prompts based on the characteristics. Specifically, the prompt generates a question asking for details about the applicant's work history.

[0133] For example: "Describe a challenge you faced while working on a project using Python and how you solved it."

[0134] Output: The generated questions are saved on the server.

[0135] Step 4: Text-to-speech questions

[0136] The server generates a question and converts it into an audio file.

[0137] Input: Server-generated question text.

[0138] How it works: The server uses speech synthesis technology (e.g., Google Text-to-Speech or Amazon Polly) to convert the question text into an audio file. Specifically, it calls a speech synthesis API and obtains the generated audio data.

[0139] Output: The generated audio file is saved on the server.

[0140] Step 5: Conduct the interview

[0141] The user responds to the voice questions.

[0142] Input: Server-generated audio file.

[0143] Operation: The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device, and the application plays them.

[0144] Output: The user responds verbally and records their response using the device's microphone. The recording is uploaded from the user's device to the server.

[0145] Step 6: Analyzing the response

[0146] The server analyzes the response voice and generates an evaluation score.

[0147] Input: Response audio data uploaded by the user.

[0148] How it works: The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text or IBM Watson Speech to Text) to convert the voice data into text, and then uses natural language processing technology again to analyze the response and extract characteristic information.

[0149] Output: The analyzed response data is saved and a rating score is calculated.

[0150] Step 7: Marking and feedback

[0151] The server scores the responses, stores the results, and generates feedback.

[0152] Input: Parsed response data.

[0153] How it works: The server calculates evaluation points based on the evaluation criteria and generates an overall score. The generated evaluation score and feedback are stored in a database and notified to the applicant via email or other means. Feedback is also provided to recruiters.

[0154] Output: Generation, storage and communication of assessment results and feedback.

[0155] (Application example 1)

[0156] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0157] Traditional hiring processes have had the problem of requiring a lot of time and effort to efficiently select applicants. Furthermore, particularly for technical positions, specialized questions are required to properly evaluate applicants' technical skills and experience, making it difficult to ensure fairness and appropriateness in interviews. To solve these problems and streamline talent acquisition, an automated interview system was needed.

[0158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0159] In this invention, the server includes a means for receiving an application form, a means for analyzing the contents of the application form to extract characteristics of the applicant, and a means for generating questions based on the characteristics. This automates the interview process in the early stages of recruitment, enabling efficient and fair evaluation of applicants.

[0160] Furthermore, by utilizing generative AI models and speech recognition technology, the generated questions are converted into natural-sounding speech, and the applicant's voice responses are analyzed using speech recognition technology, enabling highly accurate technical skill evaluations. This will enable efficient selection of factory robot operators and maintenance personnel, and ensure the right talent is secured.

[0161] The "entry form receiving means" is a means for receiving the entry form sent by the applicant into the server.

[0162] The "means for analyzing the contents of an application form" is a means for analyzing the contents of a received application form and extracting the characteristics of the applicant.

[0163] The "question generation means" is a means for generating appropriate questions based on the analyzed characteristics of the applicant.

[0164] The "speech conversion means" is a means for converting the generated question from text to speech.

[0165] The "answer recording means" is a means for recording the applicant's response to the voice questions.

[0166] The "response analysis means" is a means for analyzing the recorded responses and generating an evaluation score.

[0167] The "evaluation score storage means" is a means for storing the generated evaluation scores and generating feedback.

[0168] A "generative AI model" is an artificial intelligence model that generates technical questions based on analyzed data.

[0169] "Voice recognition technology" is a technology for transcribing applicants' voice responses and analyzing their content.

[0170] The present invention provides an automated interview system for effectively and efficiently evaluating an applicant's technical skills and experience. Specific embodiments will now be described.

[0171] 1. Submit your application form

[0172] Users submit application forms via application sites or dedicated applications. Devices used for this purpose include PCs, smartphones, tablets, etc. This application form includes personal information, educational background, work history, technical skills, and reasons for applying.

[0173] 2. Analysis of application forms

[0174] The server receives the application form and analyzes its contents using natural language processing technology. This analysis extracts characteristic information such as the applicant's skills, experience, and motivation for applying. The software used is natural language processing technology such as OpenAI API.

[0175] 3. Question generation

[0176] Based on the analyzed applicant characteristics, the server uses a generative AI model to generate appropriate questions. The generative AI model creates questions to elicit technical details and the applicant's experience. The generated questions are in text format.

[0177] 4. Speech Synthesis

[0178] The textual questions are then further processed and converted into audio using speech synthesis technologies such as gTTS (Google Text-to-Speech), which produces natural-spoken questions.

[0179] 5. Interview

[0180] The user answers the questions by voice. In this section, the questions are played on the applicant's device and the responses are recorded via a microphone. The recorded responses are then sent to the server.

[0181] 6. Analysis of response content

[0182] The server converts the received voice response into text using the Google Speech-to-Text API, and then uses natural language processing technology to analyze the response and evaluate the participant's technical skills, problem-solving ability, and communication ability.

[0183] 7. Ratings and Feedback

[0184] The server generates an evaluation score based on the analysis results and stores it in a database. It also generates feedback for the applicant and notifies them with instructions on next steps. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0185] Examples:

[0186] If a factory automation engineer applicant lists "3+ years of PLC programming experience," the generative AI model generates questions such as "Tell us about the challenges you faced in PLC programming projects and how you solved them." These questions are converted into audio and presented to the applicant, and the applicant's responses are analyzed.

[0187] Example prompt for a generative AI model:

[0188] Please analyze the following application form and extract your key skills and experience.

[0189] Application Form: "I have over three years of experience in PLC programming and have participated in multiple production line automation projects."

[0190] Generate specific questions based on your analysis.

[0191] Analysis result: "Please tell us about the challenges you faced in your PLC programming projects and how you solved them."

[0192] In this way, the present invention evaluates the technical skills of applicants with high accuracy, realizing an efficient and fair hiring process.

[0193] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0194] Step 1:

[0195] The user submits an application form through an application site or a dedicated application. The input data includes personal information, educational background, work history, technical skills, and motivation for applying. This data is received by the server. The server stores the received data in a database and proceeds to the next analysis step.

[0196] Step 2:

[0197] The server analyzes the received application form using natural language processing (NLP) technology. Specifically, it uses the OpenAI API to extract characteristic information such as the applicant's skills, experience, and motivation for applying from the text. The input is the application form data received in step 1, and the output is the analyzed characteristic information. This information is used to proceed to the next step of question generation.

[0198] Step 3:

[0199] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. The input is the characteristic information obtained in step 2, and the output is a text version of technical questions to ask the applicant. The generative AI model uses a pre-trained question generation model, which generates detailed questions.

[0200] Step 4:

[0201] The server converts the generated questions into speech. The input is the text of the questions generated in step 3, and the output is an audio file. This speech synthesis uses speech generation technology such as gTTS (Google Text-to-Speech). The audio file is then transferred to the applicant's device.

[0202] Step 5:

[0203] The applicant plays the audio questions on the device and responds verbally through the microphone. The input is an audio file transferred from the server, and the output is the applicant's voice response. The device records this audio and sends it to the server. The server saves the received audio data and proceeds to the next analysis step.

[0204] Step 6:

[0205] The server converts the applicant's voice response into text using the Google Speech-to-Text API. The input is the audio file obtained in step 5, and the output is text data. Furthermore, the response content is analyzed using natural language processing technology to evaluate the applicant's technical skills, problem-solving ability, and communication ability.

[0206] Step 7:

[0207] The server generates an evaluation score based on the analysis results and stores it in a database. The input is the analysis data obtained in step 6, and the output is the evaluation score and feedback. The evaluation score is also notified to the applicant, along with instructions on the next step. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0208] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0209] The system of the present invention is composed of users (applicants), a server, a terminal, and an emotion engine, and each component operates in cooperation with the others as follows.

[0210] Submit your application:

[0211] The user submits an application form.

[0212] The user accesses a recruitment website and enters the necessary information such as personal information, work history, educational background, and skills into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0213] Application Form Analysis:

[0214] The server analyzes the application form.

[0215] The server checks the received application form data and analyzes the content using natural language processing (NLP) technology, specifically extracting characteristic information such as the applicant's work history, skills, and motivation for applying.

[0216] Question generation:

[0217] The server generates questions using generative AI.

[0218] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, based on the applicant's work history, it creates specific questions about past project experience and problem-solving ability. These questions are intended to assess the applicant's aptitude and skills.

[0219] Text-to-Speech:

[0220] The server uses a voice generation AI to synthesize the interviewer's voice.

[0221] The server converts the generated question text into an audio file using speech generation AI. This process prepares the question as audio data in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0222] Interview Conducted:

[0223] The user responds to the voice questions.

[0224] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0225] Response analysis:

[0226] The server analyzes the response.

[0227] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score. The server also uses an emotion engine to analyze the emotions contained in the voice response. For example, the server can extract evaluation points based on the applicant's emotional state, such as confidence, calmness, or nervousness, expressed in their response.

[0228] Grading and feedback:

[0229] The server scores the responses and stores the results.

[0230] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the emotion engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0231] Examples:

[0232] For example, when hiring a software engineer, a user would list "5+ years of development experience in Java (registered trademark)" as one of their skills on an application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would respond to the questions verbally, and the device would record and send the responses to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[0233] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0234] The processing flow will be explained below.

[0235] Step 1:

[0236] The user submits an application form.

[0237] Users access a dedicated recruitment website and view an application form, enter required information such as personal information, work history, educational background, and skill set, and click the "Submit" button to send the application form data to the server.

[0238] Step 2:

[0239] The server analyzes the application form.

[0240] The server reviews the received application form and uses natural language processing (NLP) technology to extract and analyze data from each field, such as the applicant's past work experience or specific technical skills, and stores the data in a database.

[0241] Step 3:

[0242] The server generates questions using generative AI.

[0243] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, if an applicant has experience developing Java, the server generates a question such as, "Please tell us about the challenges you faced in Java projects and how you solved them."

[0244] Step 4:

[0245] The server converts the question into speech.

[0246] The server converts the generated question text into an audio file using speech generation AI. This audio file is generated in a natural tone and is used to ask the user questions on behalf of the interviewer.

[0247] Step 5:

[0248] The user responds to the voice questions.

[0249] The interview application installed on the device (user's PC or smartphone) is launched. The server sends voice questions to the device, and the application plays them back to the user. The user responds to the questions by voice, and the responses are recorded by the device's microphone. The recorded response data is immediately uploaded to the server.

[0250] Step 6:

[0251] The server analyzes the response.

[0252] The server converts the received voice response into text using speech recognition technology, and then analyzes the text using natural language processing. At this stage, the server evaluates the user's technical skills, problem-solving ability, and logical thinking. At the same time, it also analyzes the user's emotional state using an emotion engine, extracting emotional elements such as confidence and nervousness as evaluation points.

[0253] Step 7:

[0254] The server scores the responses and stores the results.

[0255] The server calculates evaluation points based on the analysis results for each question and generates an overall score. At this time, the emotion analysis results from the emotion engine are also incorporated into the score. The evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0256] The above processing steps realize a system that generates appropriate questions from the user's application information, conducts the interview process by voice, and performs a comprehensive evaluation including analysis of the user's emotional state. This system is designed to efficiently and accurately carry out the initial stage of the hiring process.

[0257] Example 2

[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0259] In the traditional recruitment process, reviewing application forms and conducting interviews to evaluate candidates were all done manually, which not only took time and effort, but also made the evaluations subjective. This made it difficult to conduct efficient and fair evaluations, and there was a high possibility of overlooking suitable candidates. In addition, interviewers' questions and evaluations were inconsistent, making it difficult to accurately evaluate skills and characteristics.

[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0261] In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting characteristics of the applicant, means for generating questions using a generative artificial intelligence model based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for converting the responses into text data using speech recognition technology, analyzing the content of the responses using a sentiment analysis engine, and generating an evaluation score, and means for saving the evaluation score and generating feedback. This increases the efficiency and fairness of the hiring process and enables accurate evaluation of skills and characteristics.

[0262] The "means for receiving an entry form" refers to a device or program that has the function of transmitting the entry form data entered by the applicant to a server and receiving it.

[0263] "Means for analyzing the contents of the application form and extracting the characteristics of the applicant" refers to a device or program that has the function of analyzing the contents of the received application form using natural language processing technology and extracting necessary characteristic information such as the applicant's work history, skills, and motivation for applying.

[0264] A "generative artificial intelligence model" is a machine learning model that can generate appropriate questions by inputting a prompt sentence.

[0265] The "means for generating questions" refers to a device or program that has the function of automatically generating questions based on the characteristics of applicants using a generative artificial intelligence model.

[0266] The "means for converting a question into speech" is a device or program that has the function of converting text into speech in order to output the generated question as speech data.

[0267] The "means for recording the applicant's response to the voice question" is a device or program having the function of recording the voice of the applicant's response to the voice question.

[0268] The "means for converting responses into text data using voice recognition technology" refers to a device or program that has the function of using voice recognition technology to convert the applicant's voice responses into text data.

[0269] An "emotion analysis engine" is a machine learning algorithm or program that analyzes emotional states from voice or text data and outputs the analysis results.

[0270] "Means for generating an evaluation score" refers to a device or program that has the function of evaluating the skills and aptitude of an applicant based on the converted character data and the analysis results of the emotion analysis engine, and generating a numerical score.

[0271] The "means for storing evaluation scores and generating feedback" refers to a device or program that has the function of storing the generated evaluation scores in a storage device and generating feedback messages to be provided to applicants and recruiters.

[0272] The present invention provides a system that streamlines the hiring process and performs fair and accurate evaluations by linking users, servers, terminals, and an emotion analysis engine. Specific embodiments of this system are described in detail below.

[0273] Submitting an application form

[0274] The user submits an application form.

[0275] A user accesses a recruitment website and enters the required information, such as personal information, work history, educational background, and skills, into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server. This process is carried out using a web browser.

[0276] Application form analysis

[0277] The server analyzes the application form.

[0278] The server checks the data in the received application form and analyzes the content using natural language processing (NLP) technology. Specifically, it uses the Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying. For example, "more than five years of development experience in Java" is extracted as characteristic information. The results of this analysis are stored in a database.

[0279] question generation

[0280] The server generates questions using a generative artificial intelligence model.

[0281] Based on the analyzed characteristic information, the server uses a generative artificial intelligence model to generate appropriate questions. For example, OpenAI's GPT-3 is used for this model. Based on the analysis results, a prompt is entered, "Please generate a question for the applicant with more than five years of development experience in Java," and a question is generated using the GPT-3 API. The generated question is stored in a database.

[0282] Speech synthesis

[0283] The server uses a voice generation AI to synthesize the interviewer's voice.

[0284] The server converts the generated question text into an audio file using speech generation AI (e.g., Google Text-to-Speech API), which is then stored in a database and used during the interview.

[0285] Interview

[0286] The user responds to the voice questions.

[0287] The user launches an application specifically for interviews and connects to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0288] Analysis of response content

[0289] The server analyzes the response.

[0290] The server receives the response and converts it into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). It then performs text analysis using NLP technology. Furthermore, it uses an emotion analysis engine to extract the emotional state contained in the response as an evaluation point. For example, it evaluates the confidence, calmness, or nervousness displayed by the applicant in their response.

[0291] Grading and feedback

[0292] The server scores the responses and stores the results.

[0293] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the sentiment analysis engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved in a database, and the user is notified of the feedback content. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0294] Specific examples

[0295] For example, when hiring a software engineer, a user would list "more than five years of Java development experience" as one of their skills on an application form. The server would analyze this information and use a generative artificial intelligence model to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[0296] An example of a prompt for the generative AI model would be, "Generate questions for applicants with more than five years of development experience in Java."

[0297] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0298] Step 1:

[0299] Submitting an application form

[0300] A user submits an application form. The user accesses a recruitment website and enters the necessary information into the application form, including personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button. This operation sends the entered data to the server as form-data. The server receives this data and stores it in the application form database.

[0301] Input: Data entered into the application form by the user

[0302] Output: Application form data saved on the server

[0303] Step 2:

[0304] Application form analysis

[0305] The server analyzes the application form. It checks the received form data and formats the content. It uses the natural language processing (NLP) technology Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying from the text. The analysis results are saved as characteristic data in a characteristic database.

[0306] Input: Application form data

[0307] Output: Characteristic data

[0308] Step 3:

[0309] question generation

[0310] The server generates a prompt using a generative artificial intelligence model (e.g., OpenAI's GPT-3) based on the characteristic information stored in the characteristic database. The prompt, "Please generate a question for when the applicant has more than five years of development experience in Java," is input into the GPT-3 API, and the generated question is retrieved. The created question is stored in the question database.

[0311] Input: characteristic data, prompt statement

[0312] Output: Question data

[0313] Step 4:

[0314] Speech synthesis

[0315] The server reads the question data and uses the Google Text-to-Speech API to convert the generated question text into an audio file, which is then stored in a speech database.

[0316] Input: Question data

[0317] Output: Audio file

[0318] Step 5:

[0319] Interview

[0320] The user responds to the audio questions. The user launches an interview application and connects to the server. The server sends the audio file to the user's device, and the interview application plays the audio file. The user responds to the questions by voice and records the response audio using the device's microphone. Once recording is complete, the response audio file is uploaded from the device to the server.

[0321] Input: Audio file

[0322] Output: Response audio file

[0323] Step 6:

[0324] Analysis of response content

[0325] The server receives the response audio file and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is then stored in an analysis results database. It is then analyzed using NLP technology and further analyzed for emotional state using an emotion analysis engine. The analysis results of the emotional state are stored as evaluation data.

[0326] Input: Response audio file

[0327] Output: Analysis result data, emotion evaluation data

[0328] Step 7:

[0329] Grading and feedback

[0330] The server calculates an evaluation score based on the analysis result data and the sentiment evaluation data. The server generates an evaluation score and stores it in a score database. It also generates a feedback message and notifies the user and recruiter. Notifications are sent via the email system.

[0331] Input: Analysis result data, emotion evaluation data

[0332] Output: Evaluation score, feedback message

[0333] (Application example 2)

[0334] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0335] Traditional hiring processes have the problem of requiring a lot of time and effort to accurately evaluate applicants' aptitude and skills. Furthermore, the interviewer's subjectivity can sometimes affect the evaluation, potentially resulting in a lack of fairness. Especially when hiring workers in factories, efficient and accurate skill evaluation is required. Therefore, an automated system is needed to quickly and fairly evaluate applicants and find the right talent.

[0336] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting the applicant's characteristics, means for generating questions based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for analyzing the responses and generating an evaluation score, means for saving the evaluation score and generating feedback, means for converting the applicant's responses into text using speech recognition technology, means for analyzing the emotions of the responses, means for providing feedback based on the analysis results, and means for using a generative AI model to generate questions for evaluating the applicant's aptitude and skills. This enables the applicant's skills and characteristics to be evaluated accurately and efficiently through an automated process, enabling a fair and reliable hiring process.

[0337] The "means for receiving the application form" refers to a communication means for transferring the information on the application form submitted by the applicant to the server.

[0338] "Means for analyzing the contents of application forms and extracting the characteristics of applicants" refers to means that include natural language processing technology for analyzing the information in application forms submitted by applicants and extracting the skills, experience, and other characteristics of the applicants.

[0339] The "means for generating questions" refers to a means for using a generative AI model to generate appropriate interview questions based on the analyzed applicant's characteristic information.

[0340] The "means for converting a question into voice" refers to a means that utilizes voice synthesis technology to convert the text data of the generated question into voice data.

[0341] "Means for recording applicant responses" refers to a means for recording the applicant's voice responses and saving them as digital data.

[0342] The "means for analyzing responses and generating an evaluation score" refers to a means for converting the recorded responses of applicants into text using voice recognition technology, analyzing the content of the text, and generating an evaluation score.

[0343] The "means for storing evaluation scores and generating feedback" refers to a means for storing the generated evaluation scores in a database or the like and providing them as feedback to applicants and hiring managers.

[0344] "Means of converting applicants' responses into text using voice recognition technology" refers to the use of voice recognition software to convert recorded voice data into text data.

[0345] "Means for analyzing emotion from response voice" refers to means for analyzing voice data of an applicant and using an emotion analysis engine to evaluate the applicant's emotional state (e.g., confident, calm, nervous).

[0346] The "means for providing feedback based on the analysis results" refers to a means for generating detailed feedback including an evaluation score and an emotional evaluation based on the results of analyzing the applicant's responses, and providing the feedback to the applicant and the recruiter.

[0347] "Means of using a generative AI model to generate questions to assess aptitudes and skills" refers to means of using generative AI technology, such as a machine learning model, to generate appropriate interview questions based on the characteristics of applicants.

[0348] The system of the present invention is composed of a user, a server, a terminal, and an emotion engine, and these elements work in cooperation with each other. Specific embodiments will be described below.

[0349] Submitting an application form

[0350] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server.

[0351] Application form analysis

[0352] The server analyzes the received application form data and uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. This analysis uses Python's NLP library and machine learning models.

[0353] question generation

[0354] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used.

[0355] Speech synthesis

[0356] The server converts the generated question text into audio data. A speech synthesis AI is used to convert the text data into natural spoken language. This audio file is saved in an appropriate format and sent to the user's device.

[0357] Interview

[0358] The user starts up the interview application installed on the terminal and conducts the interview. The terminal sequentially plays back the voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server.

[0359] Analysis of response content

[0360] The server receives the response data and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the response. For example, emotions expressed by the user in their response, such as confidence, calmness, or nervousness, can be extracted as evaluation points.

[0361] Grading and feedback

[0362] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The generated evaluation results and feedback are saved in a database, and the feedback is notified to the user. The evaluation results and feedback are also sent to the recruiter, who uses them in the next selection step.

[0363] Specific examples

[0364] For example, when hiring a software engineer, a user would write "5+ years of Java development experience" on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, and the device would record and send the responses to the server. The server would then analyze the responses and evaluate the user's emotional state, including whether they were confident or nervous. This allows for a fair and effective initial interview.

[0365] Prompt Sentence Examples

[0366] "Based on the technical skills (e.g. Java, ROS) that applicants have listed on their application form, generate questions to assess their specific experience and problem-solving abilities."

[0367] The above is a specific embodiment for carrying out the present invention. Through this system, applicant characteristics can be accurately evaluated, and an efficient and fair hiring process can be realized.

[0368] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0369] Step 1:

[0370] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server. The input data includes the application form information in text format, and the output data is the application form data that has reached the server.

[0371] Step 2:

[0372] The server analyzes the received application form data. It uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. Specifically, it uses a Python NLP library (e.g., NLTK, Spacy, etc.) to analyze the text data. The input data includes the application form data, and the output data is the extracted characteristic information.

[0373] Step 3:

[0374] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used. The input data includes characteristic information, and the generated question text is obtained as output data.

[0375] Step 4:

[0376] The server converts the generated question text into audio data. A speech synthesis AI (e.g., Google Text-to-Speech) is used to generate the audio, converting the text data into natural spoken language. This audio file is saved in an appropriate format (e.g., MP3, WAV, etc.). The input data includes the question text, and the output data is an audio file.

[0377] Step 5:

[0378] The user starts a dedicated interview application installed on the terminal and conducts the interview. The terminal sequentially plays back voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server. The input data includes the voice questions and the user's responses, and the recorded response data is obtained as output data.

[0379] Step 6:

[0380] The server receives the response data and converts the response to text using speech recognition technology. Specifically, it uses speech recognition software (e.g., Google Speech-to-Text) to convert the recorded voice data to text. The input data includes the voice response, and the output data is the text response.

[0381] Step 7:

[0382] The server analyzes the text responses to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the responses. For example, it uses voice analysis software (e.g., IBM Watson Tone Analyzer) to evaluate the emotional state (e.g., confident, calm, nervous). Input data includes the text responses and voice data, and output data includes the emotion analysis results and an evaluation score.

[0383] Step 8:

[0384] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The evaluation results and feedback are saved in a database and notified to the user as feedback. The evaluation results and feedback are also sent to the hiring manager. The input data includes the evaluation score and emotion analysis results, and the output data is an overall evaluation score and feedback.

[0385] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0386] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0387] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0388] [Second embodiment]

[0389] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0390] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0391] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0392] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0393] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0394] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0395] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0396] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0397] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0398] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0399] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0400] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0401] The system of the present invention is composed of users (applicants), a server, and terminals, and each component operates in cooperation with the others as follows.

[0402] Submit your application:

[0403] The user submits an application form.

[0404] The user accesses a recruitment website and uses a dedicated application form to enter the required information (personal information, educational background, work history, skills, reasons for applying, etc.). Once the information is complete, the user clicks the "Submit" button to send the application form to the server.

[0405] Application Form Analysis:

[0406] The server analyzes the application form.

[0407] The server checks the received application form and analyzes its contents using natural language processing (NLP) technology. Specifically, it extracts the applicant's skills, experience, and motivation for applying from the text, and extracts the necessary characteristic information.

[0408] Question generation:

[0409] The server generates questions using generative AI.

[0410] Based on the analyzed characteristics, the server uses generative AI to generate appropriate questions for the applicant, such as questions about details related to the applicant's work history and technical skills, to assess the applicant's aptitude and skills.

[0411] Text-to-Speech:

[0412] The server uses a voice generation AI to synthesize the interviewer's voice.

[0413] The server converts the generated questions from text to audio files. Using speech generation AI, it generates questions as audio files in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0414] Interview Conducted:

[0415] The user responds to the voice questions.

[0416] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to each question in their own voice and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0417] Response analysis:

[0418] The server analyzes the response.

[0419] The server converts the user's voice response into text and analyzes it again using natural language processing technology. Specifically, it extracts elements from the applicant's response that evaluate their technical skills, problem-solving ability, and communication ability, and generates an evaluation score.

[0420] Grading and feedback:

[0421] The server scores the responses and stores the results.

[0422] The server calculates the evaluation points for each question and generates an overall score. The evaluation results and feedback are notified to the applicant, and the applicant is guided to the next step. The generated feedback is also provided to the recruiter and used as reference material for the subsequent selection process.

[0423] Examples:

[0424] Taking the hiring process for software engineers as an example, a user would list "3+ years of development experience in Python" as a skill on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a project using Python and how you solved them," and then perform voice synthesis. The user would then respond to the questions by voice, and the response flow would be sent to the server for analysis and evaluation. This process ensures a fair and efficient initial interview.

[0425] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0426] The processing flow will be explained below.

[0427] Step 1:

[0428] The user submits an application form.

[0429] The user accesses a recruitment website and enters the required information into the application form, such as personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0430] Step 2:

[0431] The server analyzes the application form.

[0432] The server receives the submitted application form data and uses natural language processing (NLP) technology to analyze the content of the application form and extract characteristic information such as the applicant's work history, skills, and motivation for applying.

[0433] Step 3:

[0434] The server generates questions using generative AI.

[0435] The server uses generative AI to generate appropriate questions based on the analyzed characteristics, such as specific questions about past project experience and problem-solving ability based on the applicant's work history.

[0436] Step 4:

[0437] The server converts the question into speech.

[0438] The server converts the generated question text into an audio file using speech generation AI. Through this process, a naturally spoken question is prepared as audio data.

[0439] Step 5:

[0440] The user responds to the voice questions.

[0441] An interview-specific application is launched on the device (user's PC or smartphone). The server sends an audio file to the device, which then plays the audio questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0442] Step 6:

[0443] The server analyzes the response.

[0444] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score.

[0445] Step 7:

[0446] The server scores the responses and stores the results.

[0447] The server calculates the evaluation points for each question and generates an overall score. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who uses them in the next selection step.

[0448] Through the above processing steps, the system of the present invention realizes efficient scrutiny of application forms and highly accurate initial interviews.

[0449] Example 1

[0450] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0451] In the conventional hiring process, reviewing applicants' application forms and conducting interviews takes a great deal of time and effort, and there are issues with inconsistent evaluation criteria and a high degree of reliance on subjective judgment. Another problem is that the quality of questions posed to applicants varies depending on the interviewer, resulting in a lack of fairness. A new system is needed to solve these issues and realize an efficient and fair selection process.

[0452] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0453] In this invention, the server includes means for receiving application forms, means for analyzing the contents of the application form using natural language processing technology and extracting characteristic information about the applicant, means for generating questions using generative AI technology based on the characteristic information, means for converting the questions into voice using speech synthesis technology, means for the applicant to record a response to the voice question, means for converting the response voice into text using speech recognition technology and re-analyzing the content to generate an evaluation score, and means for saving the evaluation score and generating feedback. This automates the confirmation of application form contents and the implementation of interviews, enabling an efficient and fair hiring process.

[0454] An "entry sheet" is a document in which an applicant writes down their personal information, educational background, work history, skills, reasons for applying, etc.

[0455] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language.

[0456] "Characteristic information" refers to information such as skills, work history, and motivation for applying that is extracted from the applicant's application form.

[0457] "Generative AI technology" is a technology in which artificial intelligence generates new text or data based on given information.

[0458] "Questions" are a series of questions that are generated based on the applicant's characteristic information and used during the interview.

[0459] "Speech synthesis technology" is a technology that converts text data into speech format.

[0460] A "response" is an oral response given by an applicant to a question during an interview.

[0461] "Speech recognition technology" is a technology that converts voice data into text data.

[0462] The "evaluation score" is a score calculated based on a set of criteria by analyzing the applicant's responses.

[0463] "Feedback" refers to information such as evaluation results and comments provided to applicants based on their evaluation scores.

[0464] The system of the present invention is composed of a user, a server, and a terminal, which operate in cooperation with each other. A specific embodiment of the system will be described below.

[0465] User operations

[0466] First, a user accesses a recruitment website and fills out an application form using a dedicated application form entry form. The application form includes personal information, educational background, work history, skills, and reasons for applying. Once the user has completed the entry, they click the "Submit" button to send the application form data to the server.

[0467] Server Processing

[0468] The server receives the application form and analyzes its contents using natural language processing technology, specifically using Python natural language processing libraries such as spaCy and NLTK. This allows the server to extract characteristic information such as the applicant's skills, work history, and motivation for applying.

[0469] Next, the server generates questions based on the extracted characteristic information using generative AI technology. For generative AI technology, we use OpenAI's GPT-3 model. For example, we create prompt sentences like the following:

[0470] "Describe a challenge you faced while working on a Python project and how you solved it."

[0471] The generated questions are converted into audio files by the server using speech synthesis technology, such as Google Text-to-Speech (GCP TTS) or Amazon Polly. The generated audio questions are then used in the user's interview application.

[0472] User response

[0473] The user launches the interview application and connects to the server. The server sequentially transfers audio files to the user's device and plays back the questions. The user responds to each question verbally and records their responses using the device's microphone. Once recording is complete, the response data is uploaded from the user's device to the server.

[0474] Server response analysis

[0475] The server receives the response voice and converts it into text using speech recognition technology, such as Google Cloud Speech-to-Text or IBM Watson Speech to Text. The converted text response is then analyzed again using natural language processing technology.

[0476] Finally, the server evaluates the applicant's technical skills, problem-solving ability, and communication ability based on the responses and generates an evaluation score. The evaluation results and feedback are notified to the applicant and provided to the recruiter, enabling efficient and fair initial interviews.

[0477] Specific examples

[0478] In the software engineer recruitment process, a user writes "3+ years of development experience in Python" on an application form. The server analyzes this information and uses generative AI technology to generate a question such as "Please explain the challenges you faced in a project using Python and how you solved it," which is then converted into an audio file using speech synthesis technology. The user responds to the question by voice, and the response data is sent to the server, where it is analyzed and evaluated. This ensures a fair and efficient initial interview.

[0479] The above is a specific embodiment of the present invention.

[0480] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0481] Step 1: Submit your application form

[0482] The user creates an entry form and sends it to the server.

[0483] Input: The user accesses a recruitment website and enters the necessary information, such as personal information, educational background, work history, skills, and reasons for applying, into a dedicated application form.

[0484] How it works: When the user clicks the "Submit" button, the application form data is sent to the server via an HTTP POST request.

[0485] Output: The application form data is saved on the server.

[0486] Step 2: Analyzing the application form

[0487] The server analyzes the received application form data and extracts characteristic information.

[0488] Input: Entry sheet data received by the server.

[0489] How it works: The server uses Python to analyze the application form text with a natural language processing library (e.g., spaCy) and extracts characteristic information such as the applicant's skills, work history, and motivation for applying. Specifically, it uses NLP techniques such as morphological analysis, named entity extraction, and keyword identification.

[0490] Output: The extracted characteristic information is stored in a database.

[0491] Step 3: Question Generation

[0492] The server generates questions based on the characteristic information using generative AI technology.

[0493] Input: The characteristic information parsed by the server.

[0494] How it works: The server uses generative AI technology (e.g., OpenAI GPT-3 model) to generate prompts based on the characteristics. Specifically, the prompt generates a question asking for details about the applicant's work history.

[0495] For example: "Describe a challenge you faced while working on a project using Python and how you solved it."

[0496] Output: The generated questions are saved on the server.

[0497] Step 4: Text-to-speech questions

[0498] The server generates a question and converts it into an audio file.

[0499] Input: Server-generated question text.

[0500] How it works: The server uses speech synthesis technology (e.g., Google Text-to-Speech or Amazon Polly) to convert the question text into an audio file. Specifically, it calls a speech synthesis API and obtains the generated audio data.

[0501] Output: The generated audio file is saved on the server.

[0502] Step 5: Conduct the interview

[0503] The user responds to the voice questions.

[0504] Input: Server-generated audio file.

[0505] Operation: The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device, and the application plays them.

[0506] Output: The user responds verbally and records their response using the device's microphone. The recording is uploaded from the user's device to the server.

[0507] Step 6: Analyzing the response

[0508] The server analyzes the response voice and generates an evaluation score.

[0509] Input: Response audio data uploaded by the user.

[0510] How it works: The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text or IBM Watson Speech to Text) to convert the voice data into text, and then uses natural language processing technology again to analyze the response and extract characteristic information.

[0511] Output: The analyzed response data is saved and a rating score is calculated.

[0512] Step 7: Marking and feedback

[0513] The server scores the responses, stores the results, and generates feedback.

[0514] Input: Parsed response data.

[0515] How it works: The server calculates evaluation points based on the evaluation criteria and generates an overall score. The generated evaluation score and feedback are stored in a database and notified to the applicant via email or other means. Feedback is also provided to recruiters.

[0516] Output: Generation, storage and communication of assessment results and feedback.

[0517] (Application example 1)

[0518] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0519] Traditional hiring processes have had the problem of requiring a lot of time and effort to efficiently select applicants. Furthermore, particularly for technical positions, specialized questions are required to properly evaluate applicants' technical skills and experience, making it difficult to ensure fairness and appropriateness in interviews. To solve these problems and streamline talent acquisition, an automated interview system was needed.

[0520] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0521] In this invention, the server includes a means for receiving an application form, a means for analyzing the contents of the application form to extract characteristics of the applicant, and a means for generating questions based on the characteristics. This automates the interview process in the early stages of recruitment, enabling efficient and fair evaluation of applicants.

[0522] Furthermore, by utilizing generative AI models and speech recognition technology, the generated questions are converted into natural-sounding speech, and the applicant's voice responses are analyzed using speech recognition technology, enabling highly accurate technical skill evaluations. This will enable efficient selection of factory robot operators and maintenance personnel, and ensure the right talent is secured.

[0523] The "entry form receiving means" is a means for receiving the entry form sent by the applicant into the server.

[0524] The "means for analyzing the contents of an application form" is a means for analyzing the contents of a received application form and extracting the characteristics of the applicant.

[0525] The "question generation means" is a means for generating appropriate questions based on the analyzed characteristics of the applicant.

[0526] The "speech conversion means" is a means for converting the generated question from text to speech.

[0527] The "answer recording means" is a means for recording the applicant's response to the voice questions.

[0528] The "response analysis means" is a means for analyzing the recorded responses and generating an evaluation score.

[0529] The "evaluation score storage means" is a means for storing the generated evaluation scores and generating feedback.

[0530] A "generative AI model" is an artificial intelligence model that generates technical questions based on analyzed data.

[0531] "Voice recognition technology" is a technology for transcribing applicants' voice responses and analyzing their content.

[0532] The present invention provides an automated interview system for effectively and efficiently evaluating an applicant's technical skills and experience. Specific embodiments will now be described.

[0533] 1. Submit your application form

[0534] Users submit application forms via application sites or dedicated applications. Devices used for this purpose include PCs, smartphones, tablets, etc. This application form includes personal information, educational background, work history, technical skills, and reasons for applying.

[0535] 2. Analysis of application forms

[0536] The server receives the application form and analyzes its contents using natural language processing technology. This analysis extracts characteristic information such as the applicant's skills, experience, and motivation for applying. The software used is natural language processing technology such as OpenAI API.

[0537] 3. Question generation

[0538] Based on the analyzed applicant characteristics, the server uses a generative AI model to generate appropriate questions. The generative AI model creates questions to elicit technical details and the applicant's experience. The generated questions are in text format.

[0539] 4. Speech Synthesis

[0540] The textual questions are then further processed and converted into audio using speech synthesis technologies such as gTTS (Google Text-to-Speech), which produces natural-spoken questions.

[0541] 5. Interview

[0542] The user answers the questions by voice. In this section, the questions are played on the applicant's device and the responses are recorded via a microphone. The recorded responses are then sent to the server.

[0543] 6. Analysis of response content

[0544] The server converts the received voice response into text using the Google Speech-to-Text API, and then uses natural language processing technology to analyze the response and evaluate the participant's technical skills, problem-solving ability, and communication ability.

[0545] 7. Ratings and Feedback

[0546] The server generates an evaluation score based on the analysis results and stores it in a database. It also generates feedback for the applicant and notifies them with instructions on next steps. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0547] Examples:

[0548] If a factory automation engineer applicant lists "3+ years of PLC programming experience," the generative AI model generates questions such as "Tell us about the challenges you faced in PLC programming projects and how you solved them." These questions are converted into audio and presented to the applicant, and the applicant's responses are analyzed.

[0549] Example prompt for a generative AI model:

[0550] Please analyze the following application form and extract your key skills and experience.

[0551] Application Form: "I have over three years of experience in PLC programming and have participated in multiple production line automation projects."

[0552] Generate specific questions based on your analysis.

[0553] Analysis result: "Please tell us about the challenges you faced in your PLC programming projects and how you solved them."

[0554] In this way, the present invention evaluates the technical skills of applicants with high accuracy, realizing an efficient and fair hiring process.

[0555] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0556] Step 1:

[0557] The user submits an application form through an application site or a dedicated application. The input data includes personal information, educational background, work history, technical skills, and motivation for applying. This data is received by the server. The server stores the received data in a database and proceeds to the next analysis step.

[0558] Step 2:

[0559] The server analyzes the received application form using natural language processing (NLP) technology. Specifically, it uses the OpenAI API to extract characteristic information such as the applicant's skills, experience, and motivation for applying from the text. The input is the application form data received in step 1, and the output is the analyzed characteristic information. This information is used to proceed to the next step of question generation.

[0560] Step 3:

[0561] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. The input is the characteristic information obtained in step 2, and the output is a text version of technical questions to ask the applicant. The generative AI model uses a pre-trained question generation model, which generates detailed questions.

[0562] Step 4:

[0563] The server converts the generated questions into speech. The input is the text of the questions generated in step 3, and the output is an audio file. This speech synthesis uses speech generation technology such as gTTS (Google Text-to-Speech). The audio file is then transferred to the applicant's device.

[0564] Step 5:

[0565] The applicant plays the audio questions on the device and responds verbally through the microphone. The input is an audio file transferred from the server, and the output is the applicant's voice response. The device records this audio and sends it to the server. The server saves the received audio data and proceeds to the next analysis step.

[0566] Step 6:

[0567] The server converts the applicant's voice response into text using the Google Speech-to-Text API. The input is the audio file obtained in step 5, and the output is text data. Furthermore, the response content is analyzed using natural language processing technology to evaluate the applicant's technical skills, problem-solving ability, and communication ability.

[0568] Step 7:

[0569] The server generates an evaluation score based on the analysis results and stores it in a database. The input is the analysis data obtained in step 6, and the output is the evaluation score and feedback. The evaluation score is also notified to the applicant, along with instructions on the next step. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0570] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0571] The system of the present invention is composed of users (applicants), a server, a terminal, and an emotion engine, and each component operates in cooperation with the others as follows.

[0572] Submit your application:

[0573] The user submits an application form.

[0574] The user accesses a recruitment website and enters the necessary information such as personal information, work history, educational background, and skills into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0575] Application Form Analysis:

[0576] The server analyzes the application form.

[0577] The server checks the received application form data and analyzes the content using natural language processing (NLP) technology, specifically extracting characteristic information such as the applicant's work history, skills, and motivation for applying.

[0578] Question generation:

[0579] The server generates questions using generative AI.

[0580] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, based on the applicant's work history, it creates specific questions about past project experience and problem-solving ability. These questions are intended to assess the applicant's aptitude and skills.

[0581] Text-to-Speech:

[0582] The server uses a voice generation AI to synthesize the interviewer's voice.

[0583] The server converts the generated question text into an audio file using speech generation AI. This process prepares the question as audio data in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0584] Interview Conducted:

[0585] The user responds to the voice questions.

[0586] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0587] Response analysis:

[0588] The server analyzes the response.

[0589] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score. The server also uses an emotion engine to analyze the emotions contained in the voice response. For example, the server can extract evaluation points based on the applicant's emotional state, such as confidence, calmness, or nervousness, expressed in their response.

[0590] Grading and feedback:

[0591] The server scores the responses and stores the results.

[0592] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the emotion engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0593] Examples:

[0594] For example, when hiring a software engineer, a user would list "5+ years of Java development experience" as one of their skills on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[0595] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0596] The processing flow will be explained below.

[0597] Step 1:

[0598] The user submits an application form.

[0599] Users access a dedicated recruitment website and view an application form, enter required information such as personal information, work history, educational background, and skill set, and click the "Submit" button to send the application form data to the server.

[0600] Step 2:

[0601] The server analyzes the application form.

[0602] The server reviews the received application form and uses natural language processing (NLP) technology to extract and analyze data from each field, such as the applicant's past work experience or specific technical skills, and stores the data in a database.

[0603] Step 3:

[0604] The server generates questions using generative AI.

[0605] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, if an applicant has experience developing Java, the server generates a question such as, "Please tell us about the challenges you faced in Java projects and how you solved them."

[0606] Step 4:

[0607] The server converts the question into speech.

[0608] The server converts the generated question text into an audio file using speech generation AI. This audio file is generated in a natural tone and is used to ask the user questions on behalf of the interviewer.

[0609] Step 5:

[0610] The user responds to the voice questions.

[0611] The interview application installed on the device (user's PC or smartphone) is launched. The server sends voice questions to the device, and the application plays them back to the user. The user responds to the questions by voice, and the responses are recorded by the device's microphone. The recorded response data is immediately uploaded to the server.

[0612] Step 6:

[0613] The server analyzes the response.

[0614] The server converts the received voice response into text using speech recognition technology, and then analyzes the text using natural language processing. At this stage, the server evaluates the user's technical skills, problem-solving ability, and logical thinking. At the same time, it also analyzes the user's emotional state using an emotion engine, extracting emotional elements such as confidence and nervousness as evaluation points.

[0615] Step 7:

[0616] The server scores the responses and stores the results.

[0617] The server calculates evaluation points based on the analysis results for each question and generates an overall score. At this time, the emotion analysis results from the emotion engine are also incorporated into the score. The evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0618] The above processing steps realize a system that generates appropriate questions from the user's application information, conducts the interview process by voice, and performs a comprehensive evaluation including analysis of the user's emotional state. This system is designed to efficiently and accurately carry out the initial stage of the hiring process.

[0619] Example 2

[0620] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0621] In the traditional recruitment process, reviewing application forms and conducting interviews to evaluate candidates were all done manually, which not only took time and effort, but also made the evaluations subjective. This made it difficult to conduct efficient and fair evaluations, and there was a high possibility of overlooking suitable candidates. In addition, interviewers' questions and evaluations were inconsistent, making it difficult to accurately evaluate skills and characteristics.

[0622] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0623] In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting characteristics of the applicant, means for generating questions using a generative artificial intelligence model based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for converting the responses into text data using speech recognition technology, analyzing the content of the responses using a sentiment analysis engine, and generating an evaluation score, and means for saving the evaluation score and generating feedback. This increases the efficiency and fairness of the hiring process and enables accurate evaluation of skills and characteristics.

[0624] The "means for receiving an entry form" refers to a device or program that has the function of transmitting the entry form data entered by the applicant to a server and receiving it.

[0625] "Means for analyzing the contents of the application form and extracting the characteristics of the applicant" refers to a device or program that has the function of analyzing the contents of the received application form using natural language processing technology and extracting necessary characteristic information such as the applicant's work history, skills, and motivation for applying.

[0626] A "generative artificial intelligence model" is a machine learning model that can generate appropriate questions by inputting a prompt sentence.

[0627] The "means for generating questions" refers to a device or program that has the function of automatically generating questions based on the characteristics of applicants using a generative artificial intelligence model.

[0628] The "means for converting a question into speech" is a device or program that has the function of converting text into speech in order to output the generated question as speech data.

[0629] The "means for recording the applicant's response to the voice question" is a device or program having the function of recording the voice of the applicant's response to the voice question.

[0630] The "means for converting responses into text data using voice recognition technology" refers to a device or program that has the function of using voice recognition technology to convert the applicant's voice responses into text data.

[0631] An "emotion analysis engine" is a machine learning algorithm or program that analyzes emotional states from voice or text data and outputs the analysis results.

[0632] "Means for generating an evaluation score" refers to a device or program that has the function of evaluating the skills and aptitude of an applicant based on the converted character data and the analysis results of the emotion analysis engine, and generating a numerical score.

[0633] The "means for storing evaluation scores and generating feedback" refers to a device or program that has the function of storing the generated evaluation scores in a storage device and generating feedback messages to be provided to applicants and recruiters.

[0634] The present invention provides a system that streamlines the hiring process and performs fair and accurate evaluations by linking users, servers, terminals, and an emotion analysis engine. Specific embodiments of this system are described in detail below.

[0635] Submitting an application form

[0636] The user submits an application form.

[0637] A user accesses a recruitment website and enters the required information, such as personal information, work history, educational background, and skills, into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server. This process is carried out using a web browser.

[0638] Application form analysis

[0639] The server analyzes the application form.

[0640] The server checks the data in the received application form and analyzes the content using natural language processing (NLP) technology. Specifically, it uses the Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying. For example, "more than five years of development experience in Java" is extracted as characteristic information. The results of this analysis are stored in a database.

[0641] question generation

[0642] The server generates questions using a generative artificial intelligence model.

[0643] Based on the analyzed characteristic information, the server uses a generative artificial intelligence model to generate appropriate questions. For example, OpenAI's GPT-3 is used for this model. Based on the analysis results, a prompt is entered, "Please generate a question for the applicant with more than five years of development experience in Java," and a question is generated using the GPT-3 API. The generated question is stored in a database.

[0644] Speech synthesis

[0645] The server uses a voice generation AI to synthesize the interviewer's voice.

[0646] The server converts the generated question text into an audio file using speech generation AI (e.g., Google Text-to-Speech API), which is then stored in a database and used during the interview.

[0647] Interview

[0648] The user responds to the voice questions.

[0649] The user launches an application specifically for interviews and connects to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0650] Analysis of response content

[0651] The server analyzes the response.

[0652] The server receives the response and converts it into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). It then performs text analysis using NLP technology. Furthermore, it uses an emotion analysis engine to extract the emotional state contained in the response as an evaluation point. For example, it evaluates the confidence, calmness, or nervousness displayed by the applicant in their response.

[0653] Grading and feedback

[0654] The server scores the responses and stores the results.

[0655] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the sentiment analysis engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved in a database, and the user is notified of the feedback content. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0656] Specific examples

[0657] For example, when hiring a software engineer, a user would list "more than five years of Java development experience" as one of their skills on an application form. The server would analyze this information and use a generative artificial intelligence model to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[0658] An example of a prompt for the generative AI model would be, "Generate questions for applicants with more than five years of development experience in Java."

[0659] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0660] Step 1:

[0661] Submitting an application form

[0662] A user submits an application form. The user accesses a recruitment website and enters the necessary information into the application form, including personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button. This operation sends the entered data to the server as form-data. The server receives this data and stores it in the application form database.

[0663] Input: Data entered into the application form by the user

[0664] Output: Application form data saved on the server

[0665] Step 2:

[0666] Application form analysis

[0667] The server analyzes the application form. It checks the received form data and formats the content. It uses the natural language processing (NLP) technology Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying from the text. The analysis results are saved as characteristic data in a characteristic database.

[0668] Input: Application form data

[0669] Output: Characteristic data

[0670] Step 3:

[0671] question generation

[0672] The server generates a prompt using a generative artificial intelligence model (e.g., OpenAI's GPT-3) based on the characteristic information stored in the characteristic database. The prompt, "Please generate a question for when the applicant has more than five years of development experience in Java," is input into the GPT-3 API, and the generated question is retrieved. The created question is stored in the question database.

[0673] Input: characteristic data, prompt statement

[0674] Output: Question data

[0675] Step 4:

[0676] Speech synthesis

[0677] The server reads the question data and uses the Google Text-to-Speech API to convert the generated question text into an audio file, which is then stored in a speech database.

[0678] Input: Question data

[0679] Output: Audio file

[0680] Step 5:

[0681] Interview

[0682] The user responds to the audio questions. The user launches an interview application and connects to the server. The server sends the audio file to the user's device, and the interview application plays the audio file. The user responds to the questions by voice and records the response audio using the device's microphone. Once recording is complete, the response audio file is uploaded from the device to the server.

[0683] Input: Audio file

[0684] Output: Response audio file

[0685] Step 6:

[0686] Analysis of response content

[0687] The server receives the response audio file and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is then stored in an analysis results database. It is then analyzed using NLP technology and further analyzed for emotional state using an emotion analysis engine. The analysis results of the emotional state are stored as evaluation data.

[0688] Input: Response audio file

[0689] Output: Analysis result data, emotion evaluation data

[0690] Step 7:

[0691] Grading and feedback

[0692] The server calculates an evaluation score based on the analysis result data and the sentiment evaluation data. The server generates an evaluation score and stores it in a score database. It also generates a feedback message and notifies the user and recruiter. Notifications are sent via the email system.

[0693] Input: Analysis result data, emotion evaluation data

[0694] Output: Evaluation score, feedback message

[0695] (Application example 2)

[0696] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0697] Traditional hiring processes have the problem of requiring a lot of time and effort to accurately evaluate applicants' aptitude and skills. Furthermore, the interviewer's subjectivity can sometimes affect the evaluation, potentially resulting in a lack of fairness. Especially when hiring workers in factories, efficient and accurate skill evaluation is required. Therefore, an automated system is needed to quickly and fairly evaluate applicants and find the right talent.

[0698] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting the applicant's characteristics, means for generating questions based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for analyzing the responses and generating an evaluation score, means for saving the evaluation score and generating feedback, means for converting the applicant's responses into text using speech recognition technology, means for analyzing the emotions of the responses, means for providing feedback based on the analysis results, and means for using a generative AI model to generate questions for evaluating the applicant's aptitude and skills. This enables the applicant's skills and characteristics to be evaluated accurately and efficiently through an automated process, enabling a fair and reliable hiring process.

[0699] The "means for receiving the application form" refers to a communication means for transferring the information on the application form submitted by the applicant to the server.

[0700] "Means for analyzing the contents of application forms and extracting the characteristics of applicants" refers to means that include natural language processing technology for analyzing the information in application forms submitted by applicants and extracting the skills, experience, and other characteristics of the applicants.

[0701] The "means for generating questions" refers to a means for using a generative AI model to generate appropriate interview questions based on the analyzed applicant's characteristic information.

[0702] The "means for converting a question into voice" refers to a means that utilizes voice synthesis technology to convert the text data of the generated question into voice data.

[0703] "Means for recording applicant responses" refers to a means for recording the applicant's voice responses and saving them as digital data.

[0704] The "means for analyzing responses and generating an evaluation score" refers to a means for converting the recorded responses of applicants into text using voice recognition technology, analyzing the content of the text, and generating an evaluation score.

[0705] The "means for storing evaluation scores and generating feedback" refers to a means for storing the generated evaluation scores in a database or the like and providing them as feedback to applicants and hiring managers.

[0706] "Means of converting applicants' responses into text using voice recognition technology" refers to the use of voice recognition software to convert recorded voice data into text data.

[0707] "Means for analyzing emotion from response voice" refers to means for analyzing voice data of an applicant and using an emotion analysis engine to evaluate the applicant's emotional state (e.g., confident, calm, nervous).

[0708] The "means for providing feedback based on the analysis results" refers to a means for generating detailed feedback including an evaluation score and an emotional evaluation based on the results of analyzing the applicant's responses, and providing the feedback to the applicant and the recruiter.

[0709] "Means of using a generative AI model to generate questions to assess aptitudes and skills" refers to means of using generative AI technology, such as a machine learning model, to generate appropriate interview questions based on the characteristics of applicants.

[0710] The system of the present invention is composed of a user, a server, a terminal, and an emotion engine, and these elements work in cooperation with each other. Specific embodiments will be described below.

[0711] Submitting an application form

[0712] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server.

[0713] Application form analysis

[0714] The server analyzes the received application form data and uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. This analysis uses Python's NLP library and machine learning models.

[0715] question generation

[0716] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used.

[0717] Speech synthesis

[0718] The server converts the generated question text into audio data. A speech synthesis AI is used to convert the text data into natural spoken language. This audio file is saved in an appropriate format and sent to the user's device.

[0719] Interview

[0720] The user starts up the interview application installed on the terminal and conducts the interview. The terminal sequentially plays back the voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server.

[0721] Analysis of response content

[0722] The server receives the response data and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the response. For example, emotions expressed by the user in their response, such as confidence, calmness, or nervousness, can be extracted as evaluation points.

[0723] Grading and feedback

[0724] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The generated evaluation results and feedback are saved in a database, and the feedback is notified to the user. The evaluation results and feedback are also sent to the recruiter, who uses them in the next selection step.

[0725] Specific examples

[0726] For example, when hiring a software engineer, a user would write "5+ years of Java development experience" on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, and the device would record and send the responses to the server. The server would then analyze the responses and evaluate the user's emotional state, including whether they were confident or nervous. This allows for a fair and effective initial interview.

[0727] Prompt Sentence Examples

[0728] "Based on the technical skills (e.g. Java, ROS) that applicants have listed on their application form, generate questions to assess their specific experience and problem-solving abilities."

[0729] The above is a specific embodiment for carrying out the present invention. Through this system, applicant characteristics can be accurately evaluated, and an efficient and fair hiring process can be realized.

[0730] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0731] Step 1:

[0732] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server. The input data includes the application form information in text format, and the output data is the application form data that has reached the server.

[0733] Step 2:

[0734] The server analyzes the received application form data. It uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. Specifically, it uses a Python NLP library (e.g., NLTK, Spacy, etc.) to analyze the text data. The input data includes the application form data, and the output data is the extracted characteristic information.

[0735] Step 3:

[0736] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used. The input data includes characteristic information, and the generated question text is obtained as output data.

[0737] Step 4:

[0738] The server converts the generated question text into audio data. A speech synthesis AI (e.g., Google Text-to-Speech) is used to generate the audio, converting the text data into natural spoken language. This audio file is saved in an appropriate format (e.g., MP3, WAV, etc.). The input data includes the question text, and the output data is an audio file.

[0739] Step 5:

[0740] The user starts a dedicated interview application installed on the terminal and conducts the interview. The terminal sequentially plays back voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server. The input data includes the voice questions and the user's responses, and the recorded response data is obtained as output data.

[0741] Step 6:

[0742] The server receives the response data and converts the response to text using speech recognition technology. Specifically, it uses speech recognition software (e.g., Google Speech-to-Text) to convert the recorded voice data to text. The input data includes the voice response, and the output data is the text response.

[0743] Step 7:

[0744] The server analyzes the text responses to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the responses. For example, it uses voice analysis software (e.g., IBM Watson Tone Analyzer) to evaluate the emotional state (e.g., confident, calm, nervous). Input data includes the text responses and voice data, and output data includes the emotion analysis results and an evaluation score.

[0745] Step 8:

[0746] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The evaluation results and feedback are saved in a database and notified to the user as feedback. The evaluation results and feedback are also sent to the hiring manager. The input data includes the evaluation score and emotion analysis results, and the output data is an overall evaluation score and feedback.

[0747] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0748] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0749] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0750] [Third embodiment]

[0751] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0752] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0753] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0754] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0755] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0756] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0757] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0758] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0759] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0760] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0761] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0762] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0763] The system of the present invention is composed of users (applicants), a server, and terminals, and each component operates in cooperation with the others as follows.

[0764] Submit your application:

[0765] The user submits an application form.

[0766] The user accesses a recruitment website and uses a dedicated application form to enter the required information (personal information, educational background, work history, skills, reasons for applying, etc.). Once the information is complete, the user clicks the "Submit" button to send the application form to the server.

[0767] Application Form Analysis:

[0768] The server analyzes the application form.

[0769] The server checks the received application form and analyzes its contents using natural language processing (NLP) technology. Specifically, it extracts the applicant's skills, experience, and motivation for applying from the text, and extracts the necessary characteristic information.

[0770] Question generation:

[0771] The server generates questions using generative AI.

[0772] Based on the analyzed characteristics, the server uses generative AI to generate appropriate questions for the applicant, such as questions about details related to the applicant's work history and technical skills, to assess the applicant's aptitude and skills.

[0773] Text-to-Speech:

[0774] The server uses a voice generation AI to synthesize the interviewer's voice.

[0775] The server converts the generated questions from text to audio files. Using speech generation AI, it generates questions as audio files in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0776] Interview Conducted:

[0777] The user responds to the voice questions.

[0778] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to each question in their own voice and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0779] Response analysis:

[0780] The server analyzes the response.

[0781] The server converts the user's voice response into text and analyzes it again using natural language processing technology. Specifically, it extracts elements from the applicant's response that evaluate their technical skills, problem-solving ability, and communication ability, and generates an evaluation score.

[0782] Grading and feedback:

[0783] The server scores the responses and stores the results.

[0784] The server calculates the evaluation points for each question and generates an overall score. The evaluation results and feedback are notified to the applicant, and the applicant is guided to the next step. The generated feedback is also provided to the recruiter and used as reference material for the subsequent selection process.

[0785] Examples:

[0786] Taking the hiring process for software engineers as an example, a user would list "3+ years of development experience in Python" as a skill on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a project using Python and how you solved them," and then perform voice synthesis. The user would then respond to the questions by voice, and the response flow would be sent to the server for analysis and evaluation. This process ensures a fair and efficient initial interview.

[0787] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0788] The processing flow will be explained below.

[0789] Step 1:

[0790] The user submits an application form.

[0791] The user accesses a recruitment website and enters the required information into the application form, such as personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0792] Step 2:

[0793] The server analyzes the application form.

[0794] The server receives the submitted application form data and uses natural language processing (NLP) technology to analyze the content of the application form and extract characteristic information such as the applicant's work history, skills, and motivation for applying.

[0795] Step 3:

[0796] The server generates questions using generative AI.

[0797] The server uses generative AI to generate appropriate questions based on the analyzed characteristics, such as specific questions about past project experience and problem-solving ability based on the applicant's work history.

[0798] Step 4:

[0799] The server converts the question into speech.

[0800] The server converts the generated question text into an audio file using speech generation AI. Through this process, a naturally spoken question is prepared as audio data.

[0801] Step 5:

[0802] The user responds to the voice questions.

[0803] An interview-specific application is launched on the device (user's PC or smartphone). The server sends an audio file to the device, which then plays the audio questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0804] Step 6:

[0805] The server analyzes the response.

[0806] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score.

[0807] Step 7:

[0808] The server scores the responses and stores the results.

[0809] The server calculates the evaluation points for each question and generates an overall score. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who uses them in the next selection step.

[0810] Through the above processing steps, the system of the present invention realizes efficient scrutiny of application forms and highly accurate initial interviews.

[0811] Example 1

[0812] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0813] In the conventional hiring process, reviewing applicants' application forms and conducting interviews takes a great deal of time and effort, and there are issues with inconsistent evaluation criteria and a high degree of reliance on subjective judgment. Another problem is that the quality of questions posed to applicants varies depending on the interviewer, resulting in a lack of fairness. A new system is needed to solve these issues and realize an efficient and fair selection process.

[0814] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0815] In this invention, the server includes means for receiving application forms, means for analyzing the contents of the application form using natural language processing technology and extracting characteristic information about the applicant, means for generating questions using generative AI technology based on the characteristic information, means for converting the questions into voice using speech synthesis technology, means for the applicant to record a response to the voice question, means for converting the response voice into text using speech recognition technology and re-analyzing the content to generate an evaluation score, and means for saving the evaluation score and generating feedback. This automates the confirmation of application form contents and the implementation of interviews, enabling an efficient and fair hiring process.

[0816] An "entry sheet" is a document in which an applicant writes down their personal information, educational background, work history, skills, reasons for applying, etc.

[0817] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language.

[0818] "Characteristic information" refers to information such as skills, work history, and motivation for applying that is extracted from the applicant's application form.

[0819] "Generative AI technology" is a technology in which artificial intelligence generates new text or data based on given information.

[0820] "Questions" are a series of questions that are generated based on the applicant's characteristic information and used during the interview.

[0821] "Speech synthesis technology" is a technology that converts text data into speech format.

[0822] A "response" is an oral response given by an applicant to a question during an interview.

[0823] "Speech recognition technology" is a technology that converts voice data into text data.

[0824] The "evaluation score" is a score calculated based on a set of criteria by analyzing the applicant's responses.

[0825] "Feedback" refers to information such as evaluation results and comments provided to applicants based on their evaluation scores.

[0826] The system of the present invention is composed of a user, a server, and a terminal, which operate in cooperation with each other. A specific embodiment of the system will be described below.

[0827] User operations

[0828] First, a user accesses a recruitment website and fills out an application form using a dedicated application form entry form. The application form includes personal information, educational background, work history, skills, and reasons for applying. Once the user has completed the entry, they click the "Submit" button to send the application form data to the server.

[0829] Server Processing

[0830] The server receives the application form and analyzes its contents using natural language processing technology, specifically using Python natural language processing libraries such as spaCy and NLTK. This allows the server to extract characteristic information such as the applicant's skills, work history, and motivation for applying.

[0831] Next, the server generates questions based on the extracted characteristic information using generative AI technology. For generative AI technology, we use OpenAI's GPT-3 model. For example, we create prompt sentences like the following:

[0832] "Describe a challenge you faced while working on a Python project and how you solved it."

[0833] The generated questions are converted into audio files by the server using speech synthesis technology, such as Google Text-to-Speech (GCP TTS) or Amazon Polly. The generated audio questions are then used in the user's interview application.

[0834] User response

[0835] The user launches the interview application and connects to the server. The server sequentially transfers audio files to the user's device and plays back the questions. The user responds to each question verbally and records their responses using the device's microphone. Once recording is complete, the response data is uploaded from the user's device to the server.

[0836] Server response analysis

[0837] The server receives the response voice and converts it into text using speech recognition technology, such as Google Cloud Speech-to-Text or IBM Watson Speech to Text. The converted text response is then analyzed again using natural language processing technology.

[0838] Finally, the server evaluates the applicant's technical skills, problem-solving ability, and communication ability based on the responses and generates an evaluation score. The evaluation results and feedback are notified to the applicant and provided to the recruiter, enabling efficient and fair initial interviews.

[0839] Specific examples

[0840] In the software engineer recruitment process, a user writes "3+ years of development experience in Python" on an application form. The server analyzes this information and uses generative AI technology to generate a question such as "Please explain the challenges you faced in a project using Python and how you solved it," which is then converted into an audio file using speech synthesis technology. The user responds to the question by voice, and the response data is sent to the server, where it is analyzed and evaluated. This ensures a fair and efficient initial interview.

[0841] The above is a specific embodiment of the present invention.

[0842] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0843] Step 1: Submit your application form

[0844] The user creates an entry form and sends it to the server.

[0845] Input: The user accesses a recruitment website and enters the necessary information, such as personal information, educational background, work history, skills, and reasons for applying, into a dedicated application form.

[0846] How it works: When the user clicks the "Submit" button, the application form data is sent to the server via an HTTP POST request.

[0847] Output: The application form data is saved on the server.

[0848] Step 2: Analyzing the application form

[0849] The server analyzes the received application form data and extracts characteristic information.

[0850] Input: Entry sheet data received by the server.

[0851] How it works: The server uses Python to analyze the application form text with a natural language processing library (e.g., spaCy) and extracts characteristic information such as the applicant's skills, work history, and motivation for applying. Specifically, it uses NLP techniques such as morphological analysis, named entity extraction, and keyword identification.

[0852] Output: The extracted characteristic information is stored in a database.

[0853] Step 3: Question Generation

[0854] The server generates questions based on the characteristic information using generative AI technology.

[0855] Input: The characteristic information parsed by the server.

[0856] How it works: The server uses generative AI technology (e.g., OpenAI GPT-3 model) to generate prompts based on the characteristics. Specifically, the prompt generates a question asking for details about the applicant's work history.

[0857] For example: "Describe a challenge you faced while working on a project using Python and how you solved it."

[0858] Output: The generated questions are saved on the server.

[0859] Step 4: Text-to-speech questions

[0860] The server generates a question and converts it into an audio file.

[0861] Input: Server-generated question text.

[0862] How it works: The server uses speech synthesis technology (e.g., Google Text-to-Speech or Amazon Polly) to convert the question text into an audio file. Specifically, it calls a speech synthesis API and obtains the generated audio data.

[0863] Output: The generated audio file is saved on the server.

[0864] Step 5: Conduct the interview

[0865] The user responds to the voice questions.

[0866] Input: Server-generated audio file.

[0867] Operation: The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device, and the application plays them.

[0868] Output: The user responds verbally and records their response using the device's microphone. The recording is uploaded from the user's device to the server.

[0869] Step 6: Analyzing the response

[0870] The server analyzes the response voice and generates an evaluation score.

[0871] Input: Response audio data uploaded by the user.

[0872] How it works: The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text or IBM Watson Speech to Text) to convert the voice data into text, and then uses natural language processing technology again to analyze the response and extract characteristic information.

[0873] Output: The analyzed response data is saved and a rating score is calculated.

[0874] Step 7: Marking and feedback

[0875] The server scores the responses, stores the results, and generates feedback.

[0876] Input: Parsed response data.

[0877] How it works: The server calculates evaluation points based on the evaluation criteria and generates an overall score. The generated evaluation score and feedback are stored in a database and notified to the applicant via email or other means. Feedback is also provided to recruiters.

[0878] Output: Generation, storage and communication of assessment results and feedback.

[0879] (Application example 1)

[0880] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0881] Traditional hiring processes have had the problem of requiring a lot of time and effort to efficiently select applicants. Furthermore, particularly for technical positions, specialized questions are required to properly evaluate applicants' technical skills and experience, making it difficult to ensure fairness and appropriateness in interviews. To solve these problems and streamline talent acquisition, an automated interview system was needed.

[0882] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0883] In this invention, the server includes a means for receiving an application form, a means for analyzing the contents of the application form to extract characteristics of the applicant, and a means for generating questions based on the characteristics. This automates the interview process in the early stages of recruitment, enabling efficient and fair evaluation of applicants.

[0884] Furthermore, by utilizing generative AI models and speech recognition technology, the generated questions are converted into natural-sounding speech, and the applicant's voice responses are analyzed using speech recognition technology, enabling highly accurate technical skill evaluations. This will enable efficient selection of factory robot operators and maintenance personnel, and ensure the right talent is secured.

[0885] The "entry form receiving means" is a means for receiving the entry form sent by the applicant into the server.

[0886] The "means for analyzing the contents of an application form" is a means for analyzing the contents of a received application form and extracting the characteristics of the applicant.

[0887] The "question generation means" is a means for generating appropriate questions based on the analyzed characteristics of the applicant.

[0888] The "speech conversion means" is a means for converting the generated question from text to speech.

[0889] The "answer recording means" is a means for recording the applicant's response to the voice questions.

[0890] The "response analysis means" is a means for analyzing the recorded responses and generating an evaluation score.

[0891] The "evaluation score storage means" is a means for storing the generated evaluation scores and generating feedback.

[0892] A "generative AI model" is an artificial intelligence model that generates technical questions based on analyzed data.

[0893] "Voice recognition technology" is a technology for transcribing applicants' voice responses and analyzing their content.

[0894] The present invention provides an automated interview system for effectively and efficiently evaluating an applicant's technical skills and experience. Specific embodiments will now be described.

[0895] 1. Submit your application form

[0896] Users submit application forms via application sites or dedicated applications. Devices used for this purpose include PCs, smartphones, tablets, etc. This application form includes personal information, educational background, work history, technical skills, and reasons for applying.

[0897] 2. Analysis of application forms

[0898] The server receives the application form and analyzes its contents using natural language processing technology. This analysis extracts characteristic information such as the applicant's skills, experience, and motivation for applying. The software used is natural language processing technology such as OpenAI API.

[0899] 3. Question generation

[0900] Based on the analyzed applicant characteristics, the server uses a generative AI model to generate appropriate questions. The generative AI model creates questions to elicit technical details and the applicant's experience. The generated questions are in text format.

[0901] 4. Speech Synthesis

[0902] The textual questions are then further processed and converted into audio using speech synthesis technologies such as gTTS (Google Text-to-Speech), which produces natural-spoken questions.

[0903] 5. Interview

[0904] The user answers the questions by voice. In this section, the questions are played on the applicant's device and the responses are recorded via a microphone. The recorded responses are then sent to the server.

[0905] 6. Analysis of response content

[0906] The server converts the received voice response into text using the Google Speech-to-Text API, and then uses natural language processing technology to analyze the response and evaluate the participant's technical skills, problem-solving ability, and communication ability.

[0907] 7. Ratings and Feedback

[0908] The server generates an evaluation score based on the analysis results and stores it in a database. It also generates feedback for the applicant and notifies them with instructions on next steps. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0909] Examples:

[0910] If a factory automation engineer applicant lists "3+ years of PLC programming experience," the generative AI model generates questions such as "Tell us about the challenges you faced in PLC programming projects and how you solved them." These questions are converted into audio and presented to the applicant, and the applicant's responses are analyzed.

[0911] Example prompt for a generative AI model:

[0912] Please analyze the following application form and extract your key skills and experience.

[0913] Application Form: "I have over three years of experience in PLC programming and have participated in multiple production line automation projects."

[0914] Generate specific questions based on your analysis.

[0915] Analysis result: "Please tell us about the challenges you faced in your PLC programming projects and how you solved them."

[0916] In this way, the present invention evaluates the technical skills of applicants with high accuracy, realizing an efficient and fair hiring process.

[0917] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0918] Step 1:

[0919] The user submits an application form through an application site or a dedicated application. The input data includes personal information, educational background, work history, technical skills, and motivation for applying. This data is received by the server. The server stores the received data in a database and proceeds to the next analysis step.

[0920] Step 2:

[0921] The server analyzes the received application form using natural language processing (NLP) technology. Specifically, it uses the OpenAI API to extract characteristic information such as the applicant's skills, experience, and motivation for applying from the text. The input is the application form data received in step 1, and the output is the analyzed characteristic information. This information is used to proceed to the next step of question generation.

[0922] Step 3:

[0923] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. The input is the characteristic information obtained in step 2, and the output is a text version of technical questions to ask the applicant. The generative AI model uses a pre-trained question generation model, which generates detailed questions.

[0924] Step 4:

[0925] The server converts the generated questions into speech. The input is the text of the questions generated in step 3, and the output is an audio file. This speech synthesis uses speech generation technology such as gTTS (Google Text-to-Speech). The audio file is then transferred to the applicant's device.

[0926] Step 5:

[0927] The applicant plays the audio questions on the device and responds verbally through the microphone. The input is an audio file transferred from the server, and the output is the applicant's voice response. The device records this audio and sends it to the server. The server saves the received audio data and proceeds to the next analysis step.

[0928] Step 6:

[0929] The server converts the applicant's voice response into text using the Google Speech-to-Text API. The input is the audio file obtained in step 5, and the output is text data. Furthermore, the response content is analyzed using natural language processing technology to evaluate the applicant's technical skills, problem-solving ability, and communication ability.

[0930] Step 7:

[0931] The server generates an evaluation score based on the analysis results and stores it in a database. The input is the analysis data obtained in step 6, and the output is the evaluation score and feedback. The evaluation score is also notified to the applicant, along with instructions on the next step. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[0932] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0933] The system of the present invention is composed of users (applicants), a server, a terminal, and an emotion engine, and each component operates in cooperation with the others as follows.

[0934] Submit your application:

[0935] The user submits an application form.

[0936] The user accesses a recruitment website and enters the necessary information such as personal information, work history, educational background, and skills into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[0937] Application Form Analysis:

[0938] The server analyzes the application form.

[0939] The server checks the received application form data and analyzes the content using natural language processing (NLP) technology, specifically extracting characteristic information such as the applicant's work history, skills, and motivation for applying.

[0940] Question generation:

[0941] The server generates questions using generative AI.

[0942] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, based on the applicant's work history, it creates specific questions about past project experience and problem-solving ability. These questions are intended to assess the applicant's aptitude and skills.

[0943] Text-to-Speech:

[0944] The server uses a voice generation AI to synthesize the interviewer's voice.

[0945] The server converts the generated question text into an audio file using speech generation AI. This process prepares the question as audio data in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[0946] Interview Conducted:

[0947] The user responds to the voice questions.

[0948] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[0949] Response analysis:

[0950] The server analyzes the response.

[0951] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score. The server also uses an emotion engine to analyze the emotions contained in the voice response. For example, the server can extract evaluation points based on the applicant's emotional state, such as confidence, calmness, or nervousness, expressed in their response.

[0952] Grading and feedback:

[0953] The server scores the responses and stores the results.

[0954] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the emotion engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0955] Examples:

[0956] For example, when hiring a software engineer, a user would list "5+ years of Java development experience" as one of their skills on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[0957] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[0958] The processing flow will be explained below.

[0959] Step 1:

[0960] The user submits an application form.

[0961] Users access a dedicated recruitment website and view an application form, enter required information such as personal information, work history, educational background, and skill set, and click the "Submit" button to send the application form data to the server.

[0962] Step 2:

[0963] The server analyzes the application form.

[0964] The server reviews the received application form and uses natural language processing (NLP) technology to extract and analyze data from each field, such as the applicant's past work experience or specific technical skills, and stores the data in a database.

[0965] Step 3:

[0966] The server generates questions using generative AI.

[0967] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, if an applicant has experience developing Java, the server generates a question such as, "Please tell us about the challenges you faced in Java projects and how you solved them."

[0968] Step 4:

[0969] The server converts the question into speech.

[0970] The server converts the generated question text into an audio file using speech generation AI. This audio file is generated in a natural tone and is used to ask the user questions on behalf of the interviewer.

[0971] Step 5:

[0972] The user responds to the voice questions.

[0973] The interview application installed on the device (user's PC or smartphone) is launched. The server sends voice questions to the device, and the application plays them back to the user. The user responds to the questions by voice, and the responses are recorded by the device's microphone. The recorded response data is immediately uploaded to the server.

[0974] Step 6:

[0975] The server analyzes the response.

[0976] The server converts the received voice response into text using speech recognition technology, and then analyzes the text using natural language processing. At this stage, the server evaluates the user's technical skills, problem-solving ability, and logical thinking. At the same time, it also analyzes the user's emotional state using an emotion engine, extracting emotional elements such as confidence and nervousness as evaluation points.

[0977] Step 7:

[0978] The server scores the responses and stores the results.

[0979] The server calculates evaluation points based on the analysis results for each question and generates an overall score. At this time, the emotion analysis results from the emotion engine are also incorporated into the score. The evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[0980] The above processing steps realize a system that generates appropriate questions from the user's application information, conducts the interview process by voice, and performs a comprehensive evaluation including analysis of the user's emotional state. This system is designed to efficiently and accurately carry out the initial stage of the hiring process.

[0981] Example 2

[0982] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0983] In the traditional recruitment process, reviewing application forms and conducting interviews to evaluate candidates were all done manually, which not only took time and effort, but also made the evaluations subjective. This made it difficult to conduct efficient and fair evaluations, and there was a high possibility of overlooking suitable candidates. In addition, interviewers' questions and evaluations were inconsistent, making it difficult to accurately evaluate skills and characteristics.

[0984] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0985] In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting characteristics of the applicant, means for generating questions using a generative artificial intelligence model based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for converting the responses into text data using speech recognition technology, analyzing the content of the responses using a sentiment analysis engine, and generating an evaluation score, and means for saving the evaluation score and generating feedback. This increases the efficiency and fairness of the hiring process and enables accurate evaluation of skills and characteristics.

[0986] The "means for receiving an entry form" refers to a device or program that has the function of transmitting the entry form data entered by the applicant to a server and receiving it.

[0987] "Means for analyzing the contents of the application form and extracting the characteristics of the applicant" refers to a device or program that has the function of analyzing the contents of the received application form using natural language processing technology and extracting necessary characteristic information such as the applicant's work history, skills, and motivation for applying.

[0988] A "generative artificial intelligence model" is a machine learning model that can generate appropriate questions by inputting a prompt sentence.

[0989] The "means for generating questions" refers to a device or program that has the function of automatically generating questions based on the characteristics of applicants using a generative artificial intelligence model.

[0990] The "means for converting a question into speech" is a device or program that has the function of converting text into speech in order to output the generated question as speech data.

[0991] The "means for recording the applicant's response to the voice question" is a device or program having the function of recording the voice of the applicant's response to the voice question.

[0992] The "means for converting responses into text data using voice recognition technology" refers to a device or program that has the function of using voice recognition technology to convert the applicant's voice responses into text data.

[0993] An "emotion analysis engine" is a machine learning algorithm or program that analyzes emotional states from voice or text data and outputs the analysis results.

[0994] "Means for generating an evaluation score" refers to a device or program that has the function of evaluating the skills and aptitude of an applicant based on the converted character data and the analysis results of the emotion analysis engine, and generating a numerical score.

[0995] The "means for storing evaluation scores and generating feedback" refers to a device or program that has the function of storing the generated evaluation scores in a storage device and generating feedback messages to be provided to applicants and recruiters.

[0996] The present invention provides a system that streamlines the hiring process and performs fair and accurate evaluations by linking users, servers, terminals, and an emotion analysis engine. Specific embodiments of this system are described in detail below.

[0997] Submitting an application form

[0998] The user submits an application form.

[0999] A user accesses a recruitment website and enters the required information, such as personal information, work history, educational background, and skills, into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server. This process is carried out using a web browser.

[1000] Application form analysis

[1001] The server analyzes the application form.

[1002] The server checks the data in the received application form and analyzes the content using natural language processing (NLP) technology. Specifically, it uses the Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying. For example, "more than five years of development experience in Java" is extracted as characteristic information. The results of this analysis are stored in a database.

[1003] question generation

[1004] The server generates questions using a generative artificial intelligence model.

[1005] Based on the analyzed characteristic information, the server uses a generative artificial intelligence model to generate appropriate questions. For example, OpenAI's GPT-3 is used for this model. Based on the analysis results, a prompt is entered, "Please generate a question for the applicant with more than five years of development experience in Java," and a question is generated using the GPT-3 API. The generated question is stored in a database.

[1006] Speech synthesis

[1007] The server uses a voice generation AI to synthesize the interviewer's voice.

[1008] The server converts the generated question text into an audio file using speech generation AI (e.g., Google Text-to-Speech API), which is then stored in a database and used during the interview.

[1009] Interview

[1010] The user responds to the voice questions.

[1011] The user launches an application specifically for interviews and connects to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[1012] Analysis of response content

[1013] The server analyzes the response.

[1014] The server receives the response and converts it into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). It then performs text analysis using NLP technology. Furthermore, it uses an emotion analysis engine to extract the emotional state contained in the response as an evaluation point. For example, it evaluates the confidence, calmness, or nervousness displayed by the applicant in their response.

[1015] Grading and feedback

[1016] The server scores the responses and stores the results.

[1017] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the sentiment analysis engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved in a database, and the user is notified of the feedback content. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[1018] Specific examples

[1019] For example, when hiring a software engineer, a user would list "more than five years of Java development experience" as one of their skills on an application form. The server would analyze this information and use a generative artificial intelligence model to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[1020] An example of a prompt for the generative AI model would be, "Generate questions for applicants with more than five years of development experience in Java."

[1021] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1022] Step 1:

[1023] Submitting an application form

[1024] A user submits an application form. The user accesses a recruitment website and enters the necessary information into the application form, including personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button. This operation sends the entered data to the server as form-data. The server receives this data and stores it in the application form database.

[1025] Input: Data entered into the application form by the user

[1026] Output: Application form data saved on the server

[1027] Step 2:

[1028] Application form analysis

[1029] The server analyzes the application form. It checks the received form data and formats the content. It uses the natural language processing (NLP) technology Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying from the text. The analysis results are saved as characteristic data in a characteristic database.

[1030] Input: Application form data

[1031] Output: Characteristic data

[1032] Step 3:

[1033] question generation

[1034] The server generates a prompt using a generative artificial intelligence model (e.g., OpenAI's GPT-3) based on the characteristic information stored in the characteristic database. The prompt, "Please generate a question for when the applicant has more than five years of development experience in Java," is input into the GPT-3 API, and the generated question is retrieved. The created question is stored in the question database.

[1035] Input: characteristic data, prompt statement

[1036] Output: Question data

[1037] Step 4:

[1038] Speech synthesis

[1039] The server reads the question data and uses the Google Text-to-Speech API to convert the generated question text into an audio file, which is then stored in a speech database.

[1040] Input: Question data

[1041] Output: Audio file

[1042] Step 5:

[1043] Interview

[1044] The user responds to the audio questions. The user launches an interview application and connects to the server. The server sends the audio file to the user's device, and the interview application plays the audio file. The user responds to the questions by voice and records the response audio using the device's microphone. Once recording is complete, the response audio file is uploaded from the device to the server.

[1045] Input: Audio file

[1046] Output: Response audio file

[1047] Step 6:

[1048] Analysis of response content

[1049] The server receives the response audio file and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is then stored in an analysis results database. It is then analyzed using NLP technology and further analyzed for emotional state using an emotion analysis engine. The analysis results of the emotional state are stored as evaluation data.

[1050] Input: Response audio file

[1051] Output: Analysis result data, emotion evaluation data

[1052] Step 7:

[1053] Grading and feedback

[1054] The server calculates an evaluation score based on the analysis result data and the sentiment evaluation data. The server generates an evaluation score and stores it in a score database. It also generates a feedback message and notifies the user and recruiter. Notifications are sent via the email system.

[1055] Input: Analysis result data, emotion evaluation data

[1056] Output: Evaluation score, feedback message

[1057] (Application example 2)

[1058] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1059] Traditional hiring processes have the problem of requiring a lot of time and effort to accurately evaluate applicants' aptitude and skills. Furthermore, the interviewer's subjectivity can sometimes affect the evaluation, potentially resulting in a lack of fairness. Especially when hiring workers in factories, efficient and accurate skill evaluation is required. Therefore, an automated system is needed to quickly and fairly evaluate applicants and find the right talent.

[1060] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting the applicant's characteristics, means for generating questions based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for analyzing the responses and generating an evaluation score, means for saving the evaluation score and generating feedback, means for converting the applicant's responses into text using speech recognition technology, means for analyzing the emotions of the responses, means for providing feedback based on the analysis results, and means for using a generative AI model to generate questions for evaluating the applicant's aptitude and skills. This enables the applicant's skills and characteristics to be evaluated accurately and efficiently through an automated process, enabling a fair and reliable hiring process.

[1061] The "means for receiving the application form" refers to a communication means for transferring the information on the application form submitted by the applicant to the server.

[1062] "Means for analyzing the contents of application forms and extracting the characteristics of applicants" refers to means that include natural language processing technology for analyzing the information in application forms submitted by applicants and extracting the skills, experience, and other characteristics of the applicants.

[1063] The "means for generating questions" refers to a means for using a generative AI model to generate appropriate interview questions based on the analyzed applicant's characteristic information.

[1064] The "means for converting a question into voice" refers to a means that utilizes voice synthesis technology to convert the text data of the generated question into voice data.

[1065] "Means for recording applicant responses" refers to a means for recording the applicant's voice responses and saving them as digital data.

[1066] The "means for analyzing responses and generating an evaluation score" refers to a means for converting the recorded responses of applicants into text using voice recognition technology, analyzing the content of the text, and generating an evaluation score.

[1067] The "means for storing evaluation scores and generating feedback" refers to a means for storing the generated evaluation scores in a database or the like and providing them as feedback to applicants and hiring managers.

[1068] "Means of converting applicants' responses into text using voice recognition technology" refers to the use of voice recognition software to convert recorded voice data into text data.

[1069] "Means for analyzing emotion from response voice" refers to means for analyzing voice data of an applicant and using an emotion analysis engine to evaluate the applicant's emotional state (e.g., confident, calm, nervous).

[1070] The "means for providing feedback based on the analysis results" refers to a means for generating detailed feedback including an evaluation score and an emotional evaluation based on the results of analyzing the applicant's responses, and providing the feedback to the applicant and the recruiter.

[1071] "Means of using a generative AI model to generate questions to assess aptitudes and skills" refers to means of using generative AI technology, such as a machine learning model, to generate appropriate interview questions based on the characteristics of applicants.

[1072] The system of the present invention is composed of a user, a server, a terminal, and an emotion engine, and these elements work in cooperation with each other. Specific embodiments will be described below.

[1073] Submitting an application form

[1074] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server.

[1075] Application form analysis

[1076] The server analyzes the received application form data and uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. This analysis uses Python's NLP library and machine learning models.

[1077] question generation

[1078] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used.

[1079] Speech synthesis

[1080] The server converts the generated question text into audio data. A speech synthesis AI is used to convert the text data into natural spoken language. This audio file is saved in an appropriate format and sent to the user's device.

[1081] Interview

[1082] The user starts up the interview application installed on the terminal and conducts the interview. The terminal sequentially plays back the voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server.

[1083] Analysis of response content

[1084] The server receives the response data and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the response. For example, emotions expressed by the user in their response, such as confidence, calmness, or nervousness, can be extracted as evaluation points.

[1085] Grading and feedback

[1086] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The generated evaluation results and feedback are saved in a database, and the feedback is notified to the user. The evaluation results and feedback are also sent to the recruiter, who uses them in the next selection step.

[1087] Specific examples

[1088] For example, when hiring a software engineer, a user would write "5+ years of Java development experience" on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, and the device would record and send the responses to the server. The server would then analyze the responses and evaluate the user's emotional state, including whether they were confident or nervous. This allows for a fair and effective initial interview.

[1089] Prompt Sentence Examples

[1090] "Based on the technical skills (e.g. Java, ROS) that applicants have listed on their application form, generate questions to assess their specific experience and problem-solving abilities."

[1091] The above is a specific embodiment for carrying out the present invention. Through this system, applicant characteristics can be accurately evaluated, and an efficient and fair hiring process can be realized.

[1092] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1093] Step 1:

[1094] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server. The input data includes the application form information in text format, and the output data is the application form data that has reached the server.

[1095] Step 2:

[1096] The server analyzes the received application form data. It uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. Specifically, it uses a Python NLP library (e.g., NLTK, Spacy, etc.) to analyze the text data. The input data includes the application form data, and the output data is the extracted characteristic information.

[1097] Step 3:

[1098] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used. The input data includes characteristic information, and the generated question text is obtained as output data.

[1099] Step 4:

[1100] The server converts the generated question text into audio data. A speech synthesis AI (e.g., Google Text-to-Speech) is used to generate the audio, converting the text data into natural spoken language. This audio file is saved in an appropriate format (e.g., MP3, WAV, etc.). The input data includes the question text, and the output data is an audio file.

[1101] Step 5:

[1102] The user starts a dedicated interview application installed on the terminal and conducts the interview. The terminal sequentially plays back voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server. The input data includes the voice questions and the user's responses, and the recorded response data is obtained as output data.

[1103] Step 6:

[1104] The server receives the response data and converts the response to text using speech recognition technology. Specifically, it uses speech recognition software (e.g., Google Speech-to-Text) to convert the recorded voice data to text. The input data includes the voice response, and the output data is the text response.

[1105] Step 7:

[1106] The server analyzes the text responses to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the responses. For example, it uses voice analysis software (e.g., IBM Watson Tone Analyzer) to evaluate the emotional state (e.g., confident, calm, nervous). Input data includes the text responses and voice data, and output data includes the emotion analysis results and an evaluation score.

[1107] Step 8:

[1108] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The evaluation results and feedback are saved in a database and notified to the user as feedback. The evaluation results and feedback are also sent to the hiring manager. The input data includes the evaluation score and emotion analysis results, and the output data is an overall evaluation score and feedback.

[1109] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1110] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1111] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1112] [Fourth embodiment]

[1113] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1114] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1115] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1116] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1117] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1119] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1120] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1121] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1122] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1123] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1124] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1125] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1126] The system of the present invention is composed of users (applicants), a server, and terminals, and each component operates in cooperation with the others as follows.

[1127] Submit your application:

[1128] The user submits an application form.

[1129] The user accesses a recruitment website and uses a dedicated application form to enter the required information (personal information, educational background, work history, skills, reasons for applying, etc.). Once the information is complete, the user clicks the "Submit" button to send the application form to the server.

[1130] Application Form Analysis:

[1131] The server analyzes the application form.

[1132] The server checks the received application form and analyzes its contents using natural language processing (NLP) technology. Specifically, it extracts the applicant's skills, experience, and motivation for applying from the text, and extracts the necessary characteristic information.

[1133] Question generation:

[1134] The server generates questions using generative AI.

[1135] Based on the analyzed characteristics, the server uses generative AI to generate appropriate questions for the applicant, such as questions about details related to the applicant's work history and technical skills, to assess the applicant's aptitude and skills.

[1136] Text-to-Speech:

[1137] The server uses a voice generation AI to synthesize the interviewer's voice.

[1138] The server converts the generated questions from text to audio files. Using speech generation AI, it generates questions as audio files in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[1139] Interview Conducted:

[1140] The user responds to the voice questions.

[1141] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to each question in their own voice and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[1142] Response analysis:

[1143] The server analyzes the response.

[1144] The server converts the user's voice response into text and analyzes it again using natural language processing technology. Specifically, it extracts elements from the applicant's response that evaluate their technical skills, problem-solving ability, and communication ability, and generates an evaluation score.

[1145] Grading and feedback:

[1146] The server scores the responses and stores the results.

[1147] The server calculates the evaluation points for each question and generates an overall score. The evaluation results and feedback are notified to the applicant, and the applicant is guided to the next step. The generated feedback is also provided to the recruiter and used as reference material for the subsequent selection process.

[1148] Examples:

[1149] Taking the hiring process for software engineers as an example, a user would list "3+ years of development experience in Python" as a skill on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a project using Python and how you solved them," and then perform voice synthesis. The user would then respond to the questions by voice, and the response flow would be sent to the server for analysis and evaluation. This process ensures a fair and efficient initial interview.

[1150] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[1151] The processing flow will be explained below.

[1152] Step 1:

[1153] The user submits an application form.

[1154] The user accesses a recruitment website and enters the required information into the application form, such as personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[1155] Step 2:

[1156] The server analyzes the application form.

[1157] The server receives the submitted application form data and uses natural language processing (NLP) technology to analyze the content of the application form and extract characteristic information such as the applicant's work history, skills, and motivation for applying.

[1158] Step 3:

[1159] The server generates questions using generative AI.

[1160] The server uses generative AI to generate appropriate questions based on the analyzed characteristics, such as specific questions about past project experience and problem-solving ability based on the applicant's work history.

[1161] Step 4:

[1162] The server converts the question into speech.

[1163] The server converts the generated question text into an audio file using speech generation AI. Through this process, a naturally spoken question is prepared as audio data.

[1164] Step 5:

[1165] The user responds to the voice questions.

[1166] An interview-specific application is launched on the device (user's PC or smartphone). The server sends an audio file to the device, which then plays the audio questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[1167] Step 6:

[1168] The server analyzes the response.

[1169] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score.

[1170] Step 7:

[1171] The server scores the responses and stores the results.

[1172] The server calculates the evaluation points for each question and generates an overall score. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who uses them in the next selection step.

[1173] Through the above processing steps, the system of the present invention realizes efficient scrutiny of application forms and highly accurate initial interviews.

[1174] Example 1

[1175] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1176] In the conventional hiring process, reviewing applicants' application forms and conducting interviews takes a great deal of time and effort, and there are issues with inconsistent evaluation criteria and a high degree of reliance on subjective judgment. Another problem is that the quality of questions posed to applicants varies depending on the interviewer, resulting in a lack of fairness. A new system is needed to solve these issues and realize an efficient and fair selection process.

[1177] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1178] In this invention, the server includes means for receiving application forms, means for analyzing the contents of the application form using natural language processing technology and extracting characteristic information about the applicant, means for generating questions using generative AI technology based on the characteristic information, means for converting the questions into voice using speech synthesis technology, means for the applicant to record a response to the voice question, means for converting the response voice into text using speech recognition technology and re-analyzing the content to generate an evaluation score, and means for saving the evaluation score and generating feedback. This automates the confirmation of application form contents and the implementation of interviews, enabling an efficient and fair hiring process.

[1179] An "entry sheet" is a document in which an applicant writes down their personal information, educational background, work history, skills, reasons for applying, etc.

[1180] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language.

[1181] "Characteristic information" refers to information such as skills, work history, and motivation for applying that is extracted from the applicant's application form.

[1182] "Generative AI technology" is a technology in which artificial intelligence generates new text or data based on given information.

[1183] "Questions" are a series of questions that are generated based on the applicant's characteristic information and used during the interview.

[1184] "Speech synthesis technology" is a technology that converts text data into speech format.

[1185] A "response" is an oral response given by an applicant to a question during an interview.

[1186] "Speech recognition technology" is a technology that converts voice data into text data.

[1187] The "evaluation score" is a score calculated based on a set of criteria by analyzing the applicant's responses.

[1188] "Feedback" refers to information such as evaluation results and comments provided to applicants based on their evaluation scores.

[1189] The system of the present invention is composed of a user, a server, and a terminal, which operate in cooperation with each other. A specific embodiment of the system will be described below.

[1190] User operations

[1191] First, a user accesses a recruitment website and fills out an application form using a dedicated application form entry form. The application form includes personal information, educational background, work history, skills, and reasons for applying. Once the user has completed the entry, they click the "Submit" button to send the application form data to the server.

[1192] Server Processing

[1193] The server receives the application form and analyzes its contents using natural language processing technology, specifically using Python natural language processing libraries such as spaCy and NLTK. This allows the server to extract characteristic information such as the applicant's skills, work history, and motivation for applying.

[1194] Next, the server generates questions based on the extracted characteristic information using generative AI technology. For generative AI technology, we use OpenAI's GPT-3 model. For example, we create prompt sentences like the following:

[1195] "Describe a challenge you faced while working on a Python project and how you solved it."

[1196] The generated questions are converted into audio files by the server using speech synthesis technology, such as Google Text-to-Speech (GCP TTS) or Amazon Polly. The generated audio questions are then used in the user's interview application.

[1197] User response

[1198] The user launches the interview application and connects to the server. The server sequentially transfers audio files to the user's device and plays back the questions. The user responds to each question verbally and records their responses using the device's microphone. Once recording is complete, the response data is uploaded from the user's device to the server.

[1199] Server response analysis

[1200] The server receives the response voice and converts it into text using speech recognition technology, such as Google Cloud Speech-to-Text or IBM Watson Speech to Text. The converted text response is then analyzed again using natural language processing technology.

[1201] Finally, the server evaluates the applicant's technical skills, problem-solving ability, and communication ability based on the responses and generates an evaluation score. The evaluation results and feedback are notified to the applicant and provided to the recruiter, enabling efficient and fair initial interviews.

[1202] Specific examples

[1203] In the software engineer recruitment process, a user writes "3+ years of development experience in Python" on an application form. The server analyzes this information and uses generative AI technology to generate a question such as "Please explain the challenges you faced in a project using Python and how you solved it," which is then converted into an audio file using speech synthesis technology. The user responds to the question by voice, and the response data is sent to the server, where it is analyzed and evaluated. This ensures a fair and efficient initial interview.

[1204] The above is a specific embodiment of the present invention.

[1205] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1206] Step 1: Submit your application form

[1207] The user creates an entry form and sends it to the server.

[1208] Input: The user accesses a recruitment website and enters the necessary information, such as personal information, educational background, work history, skills, and reasons for applying, into a dedicated application form.

[1209] How it works: When the user clicks the "Submit" button, the application form data is sent to the server via an HTTP POST request.

[1210] Output: The application form data is saved on the server.

[1211] Step 2: Analyzing the application form

[1212] The server analyzes the received application form data and extracts characteristic information.

[1213] Input: Entry sheet data received by the server.

[1214] How it works: The server uses Python to analyze the application form text with a natural language processing library (e.g., spaCy) and extracts characteristic information such as the applicant's skills, work history, and motivation for applying. Specifically, it uses NLP techniques such as morphological analysis, named entity extraction, and keyword identification.

[1215] Output: The extracted characteristic information is stored in a database.

[1216] Step 3: Question Generation

[1217] The server generates questions based on the characteristic information using generative AI technology.

[1218] Input: The characteristic information parsed by the server.

[1219] How it works: The server uses generative AI technology (e.g., OpenAI GPT-3 model) to generate prompts based on the characteristics. Specifically, the prompt generates a question asking for details about the applicant's work history.

[1220] For example: "Describe a challenge you faced while working on a project using Python and how you solved it."

[1221] Output: The generated questions are saved on the server.

[1222] Step 4: Text-to-speech questions

[1223] The server generates a question and converts it into an audio file.

[1224] Input: Server-generated question text.

[1225] How it works: The server uses speech synthesis technology (e.g., Google Text-to-Speech or Amazon Polly) to convert the question text into an audio file. Specifically, it calls a speech synthesis API and obtains the generated audio data.

[1226] Output: The generated audio file is saved on the server.

[1227] Step 5: Conduct the interview

[1228] The user responds to the voice questions.

[1229] Input: Server-generated audio file.

[1230] Operation: The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device, and the application plays them.

[1231] Output: The user responds verbally and records their response using the device's microphone. The recording is uploaded from the user's device to the server.

[1232] Step 6: Analyzing the response

[1233] The server analyzes the response voice and generates an evaluation score.

[1234] Input: Response audio data uploaded by the user.

[1235] How it works: The server uses speech recognition technology (e.g., Google Cloud Speech-to-Text or IBM Watson Speech to Text) to convert the voice data into text, and then uses natural language processing technology again to analyze the response and extract characteristic information.

[1236] Output: The analyzed response data is saved and a rating score is calculated.

[1237] Step 7: Marking and feedback

[1238] The server scores the responses, stores the results, and generates feedback.

[1239] Input: Parsed response data.

[1240] How it works: The server calculates evaluation points based on the evaluation criteria and generates an overall score. The generated evaluation score and feedback are stored in a database and notified to the applicant via email or other means. Feedback is also provided to recruiters.

[1241] Output: Generation, storage and communication of assessment results and feedback.

[1242] (Application example 1)

[1243] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1244] Traditional hiring processes have had the problem of requiring a lot of time and effort to efficiently select applicants. Furthermore, particularly for technical positions, specialized questions are required to properly evaluate applicants' technical skills and experience, making it difficult to ensure fairness and appropriateness in interviews. To solve these problems and streamline talent acquisition, an automated interview system was needed.

[1245] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1246] In this invention, the server includes a means for receiving an application form, a means for analyzing the contents of the application form to extract characteristics of the applicant, and a means for generating questions based on the characteristics. This automates the interview process in the early stages of recruitment, enabling efficient and fair evaluation of applicants.

[1247] Furthermore, by utilizing generative AI models and speech recognition technology, the generated questions are converted into natural-sounding speech, and the applicant's voice responses are analyzed using speech recognition technology, enabling highly accurate technical skill evaluations. This will enable efficient selection of factory robot operators and maintenance personnel, and ensure the right talent is secured.

[1248] The "entry form receiving means" is a means for receiving the entry form sent by the applicant into the server.

[1249] The "means for analyzing the contents of an application form" is a means for analyzing the contents of a received application form and extracting the characteristics of the applicant.

[1250] The "question generation means" is a means for generating appropriate questions based on the analyzed characteristics of the applicant.

[1251] The "speech conversion means" is a means for converting the generated question from text to speech.

[1252] The "answer recording means" is a means for recording the applicant's response to the voice questions.

[1253] The "response analysis means" is a means for analyzing the recorded responses and generating an evaluation score.

[1254] The "evaluation score storage means" is a means for storing the generated evaluation scores and generating feedback.

[1255] A "generative AI model" is an artificial intelligence model that generates technical questions based on analyzed data.

[1256] "Voice recognition technology" is a technology for transcribing applicants' voice responses and analyzing their content.

[1257] The present invention provides an automated interview system for effectively and efficiently evaluating an applicant's technical skills and experience. Specific embodiments will now be described.

[1258] 1. Submit your application form

[1259] Users submit application forms via application sites or dedicated applications. Devices used for this purpose include PCs, smartphones, tablets, etc. This application form includes personal information, educational background, work history, technical skills, and reasons for applying.

[1260] 2. Analysis of application forms

[1261] The server receives the application form and analyzes its contents using natural language processing technology. This analysis extracts characteristic information such as the applicant's skills, experience, and motivation for applying. The software used is natural language processing technology such as OpenAI API.

[1262] 3. Question generation

[1263] Based on the analyzed applicant characteristics, the server uses a generative AI model to generate appropriate questions. The generative AI model creates questions to elicit technical details and the applicant's experience. The generated questions are in text format.

[1264] 4. Speech Synthesis

[1265] The textual questions are then further processed and converted into audio using speech synthesis technologies such as gTTS (Google Text-to-Speech), which produces natural-spoken questions.

[1266] 5. Interview

[1267] The user answers the questions by voice. In this section, the questions are played on the applicant's device and the responses are recorded via a microphone. The recorded responses are then sent to the server.

[1268] 6. Analysis of response content

[1269] The server converts the received voice response into text using the Google Speech-to-Text API, and then uses natural language processing technology to analyze the response and evaluate the participant's technical skills, problem-solving ability, and communication ability.

[1270] 7. Ratings and Feedback

[1271] The server generates an evaluation score based on the analysis results and stores it in a database. It also generates feedback for the applicant and notifies them with instructions on next steps. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[1272] Examples:

[1273] If a factory automation engineer applicant lists "3+ years of PLC programming experience," the generative AI model generates questions such as "Tell us about the challenges you faced in PLC programming projects and how you solved them." These questions are converted into audio and presented to the applicant, and the applicant's responses are analyzed.

[1274] Example prompt for a generative AI model:

[1275] Please analyze the following application form and extract your key skills and experience.

[1276] Application Form: "I have over three years of experience in PLC programming and have participated in multiple production line automation projects."

[1277] Generate specific questions based on your analysis.

[1278] Analysis result: "Please tell us about the challenges you faced in your PLC programming projects and how you solved them."

[1279] In this way, the present invention evaluates the technical skills of applicants with high accuracy, realizing an efficient and fair hiring process.

[1280] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1281] Step 1:

[1282] The user submits an application form through an application site or a dedicated application. The input data includes personal information, educational background, work history, technical skills, and motivation for applying. This data is received by the server. The server stores the received data in a database and proceeds to the next analysis step.

[1283] Step 2:

[1284] The server analyzes the received application form using natural language processing (NLP) technology. Specifically, it uses the OpenAI API to extract characteristic information such as the applicant's skills, experience, and motivation for applying from the text. The input is the application form data received in step 1, and the output is the analyzed characteristic information. This information is used to proceed to the next step of question generation.

[1285] Step 3:

[1286] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. The input is the characteristic information obtained in step 2, and the output is a text version of technical questions to ask the applicant. The generative AI model uses a pre-trained question generation model, which generates detailed questions.

[1287] Step 4:

[1288] The server converts the generated questions into speech. The input is the text of the questions generated in step 3, and the output is an audio file. This speech synthesis uses speech generation technology such as gTTS (Google Text-to-Speech). The audio file is then transferred to the applicant's device.

[1289] Step 5:

[1290] The applicant plays the audio questions on the device and responds verbally through the microphone. The input is an audio file transferred from the server, and the output is the applicant's voice response. The device records this audio and sends it to the server. The server saves the received audio data and proceeds to the next analysis step.

[1291] Step 6:

[1292] The server converts the applicant's voice response into text using the Google Speech-to-Text API. The input is the audio file obtained in step 5, and the output is text data. Furthermore, the response content is analyzed using natural language processing technology to evaluate the applicant's technical skills, problem-solving ability, and communication ability.

[1293] Step 7:

[1294] The server generates an evaluation score based on the analysis results and stores it in a database. The input is the analysis data obtained in step 6, and the output is the evaluation score and feedback. The evaluation score is also notified to the applicant, along with instructions on the next step. The evaluation results are also provided to recruiters and used in the subsequent selection process.

[1295] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1296] The system of the present invention is composed of users (applicants), a server, a terminal, and an emotion engine, and each component operates in cooperation with the others as follows.

[1297] Submit your application:

[1298] The user submits an application form.

[1299] The user accesses a recruitment website and enters the necessary information such as personal information, work history, educational background, and skills into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server.

[1300] Application Form Analysis:

[1301] The server analyzes the application form.

[1302] The server checks the received application form data and analyzes the content using natural language processing (NLP) technology, specifically extracting characteristic information such as the applicant's work history, skills, and motivation for applying.

[1303] Question generation:

[1304] The server generates questions using generative AI.

[1305] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, based on the applicant's work history, it creates specific questions about past project experience and problem-solving ability. These questions are intended to assess the applicant's aptitude and skills.

[1306] Text-to-Speech:

[1307] The server uses a voice generation AI to synthesize the interviewer's voice.

[1308] The server converts the generated question text into an audio file using speech generation AI. This process prepares the question as audio data in a natural speech format. For example, the generated question "What is the most difficult project you have ever faced?" is converted into an audio file.

[1309] Interview Conducted:

[1310] The user responds to the voice questions.

[1311] The interview application installed on the user's device is launched and connected to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[1312] Response analysis:

[1313] The server analyzes the response.

[1314] The server receives the voice response and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to extract elements that evaluate the applicant's technical depth and communication skills, generating an evaluation score. The server also uses an emotion engine to analyze the emotions contained in the voice response. For example, the server can extract evaluation points based on the applicant's emotional state, such as confidence, calmness, or nervousness, expressed in their response.

[1315] Grading and feedback:

[1316] The server scores the responses and stores the results.

[1317] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the emotion engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[1318] Examples:

[1319] For example, when hiring a software engineer, a user would list "5+ years of Java development experience" as one of their skills on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[1320] The above is a specific embodiment of the system of the present invention, which realizes efficiency and accuracy improvement in the early stages of the recruitment process.

[1321] The processing flow will be explained below.

[1322] Step 1:

[1323] The user submits an application form.

[1324] Users access a dedicated recruitment website and view an application form, enter required information such as personal information, work history, educational background, and skill set, and click the "Submit" button to send the application form data to the server.

[1325] Step 2:

[1326] The server analyzes the application form.

[1327] The server reviews the received application form and uses natural language processing (NLP) technology to extract and analyze data from each field, such as the applicant's past work experience or specific technical skills, and stores the data in a database.

[1328] Step 3:

[1329] The server generates questions using generative AI.

[1330] The server uses generative AI to generate appropriate questions based on the analyzed characteristics. For example, if an applicant has experience developing Java, the server generates a question such as, "Please tell us about the challenges you faced in Java projects and how you solved them."

[1331] Step 4:

[1332] The server converts the question into speech.

[1333] The server converts the generated question text into an audio file using speech generation AI. This audio file is generated in a natural tone and is used to ask the user questions on behalf of the interviewer.

[1334] Step 5:

[1335] The user responds to the voice questions.

[1336] The interview application installed on the device (user's PC or smartphone) is launched. The server sends voice questions to the device, and the application plays them back to the user. The user responds to the questions by voice, and the responses are recorded by the device's microphone. The recorded response data is immediately uploaded to the server.

[1337] Step 6:

[1338] The server analyzes the response.

[1339] The server converts the received voice response into text using speech recognition technology, and then analyzes the text using natural language processing. At this stage, the server evaluates the user's technical skills, problem-solving ability, and logical thinking. At the same time, it also analyzes the user's emotional state using an emotion engine, extracting emotional elements such as confidence and nervousness as evaluation points.

[1340] Step 7:

[1341] The server scores the responses and stores the results.

[1342] The server calculates evaluation points based on the analysis results for each question and generates an overall score. At this time, the emotion analysis results from the emotion engine are also incorporated into the score. The evaluation results and feedback are saved and the user is notified of the feedback. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[1343] The above processing steps realize a system that generates appropriate questions from the user's application information, conducts the interview process by voice, and performs a comprehensive evaluation including analysis of the user's emotional state. This system is designed to efficiently and accurately carry out the initial stage of the hiring process.

[1344] Example 2

[1345] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1346] In the traditional recruitment process, reviewing application forms and conducting interviews to evaluate candidates were all done manually, which not only took time and effort, but also made the evaluations subjective. This made it difficult to conduct efficient and fair evaluations, and there was a high possibility of overlooking suitable candidates. In addition, interviewers' questions and evaluations were inconsistent, making it difficult to accurately evaluate skills and characteristics.

[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1348] In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting characteristics of the applicant, means for generating questions using a generative artificial intelligence model based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for converting the responses into text data using speech recognition technology, analyzing the content of the responses using a sentiment analysis engine, and generating an evaluation score, and means for saving the evaluation score and generating feedback. This increases the efficiency and fairness of the hiring process and enables accurate evaluation of skills and characteristics.

[1349] The "means for receiving an entry form" refers to a device or program that has the function of transmitting the entry form data entered by the applicant to a server and receiving it.

[1350] "Means for analyzing the contents of the application form and extracting the characteristics of the applicant" refers to a device or program that has the function of analyzing the contents of the received application form using natural language processing technology and extracting necessary characteristic information such as the applicant's work history, skills, and motivation for applying.

[1351] A "generative artificial intelligence model" is a machine learning model that can generate appropriate questions by inputting a prompt sentence.

[1352] The "means for generating questions" refers to a device or program that has the function of automatically generating questions based on the characteristics of applicants using a generative artificial intelligence model.

[1353] The "means for converting a question into speech" is a device or program that has the function of converting text into speech in order to output the generated question as speech data.

[1354] The "means for recording the applicant's response to the voice question" is a device or program having the function of recording the voice of the applicant's response to the voice question.

[1355] The "means for converting responses into text data using voice recognition technology" refers to a device or program that has the function of using voice recognition technology to convert the applicant's voice responses into text data.

[1356] An "emotion analysis engine" is a machine learning algorithm or program that analyzes emotional states from voice or text data and outputs the analysis results.

[1357] "Means for generating an evaluation score" refers to a device or program that has the function of evaluating the skills and aptitude of an applicant based on the converted character data and the analysis results of the emotion analysis engine, and generating a numerical score.

[1358] The "means for storing evaluation scores and generating feedback" refers to a device or program that has the function of storing the generated evaluation scores in a storage device and generating feedback messages to be provided to applicants and recruiters.

[1359] The present invention provides a system that streamlines the hiring process and performs fair and accurate evaluations by linking users, servers, terminals, and an emotion analysis engine. Specific embodiments of this system are described in detail below.

[1360] Submitting an application form

[1361] The user submits an application form.

[1362] A user accesses a recruitment website and enters the required information, such as personal information, work history, educational background, and skills, into the application form. After completing the entry, the user clicks the "Submit" button, and the application form data is sent to the server. This process is carried out using a web browser.

[1363] Application form analysis

[1364] The server analyzes the application form.

[1365] The server checks the data in the received application form and analyzes the content using natural language processing (NLP) technology. Specifically, it uses the Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying. For example, "more than five years of development experience in Java" is extracted as characteristic information. The results of this analysis are stored in a database.

[1366] question generation

[1367] The server generates questions using a generative artificial intelligence model.

[1368] Based on the analyzed characteristic information, the server uses a generative artificial intelligence model to generate appropriate questions. For example, OpenAI's GPT-3 is used for this model. Based on the analysis results, a prompt is entered, "Please generate a question for the applicant with more than five years of development experience in Java," and a question is generated using the GPT-3 API. The generated question is stored in a database.

[1369] Speech synthesis

[1370] The server uses a voice generation AI to synthesize the interviewer's voice.

[1371] The server converts the generated question text into an audio file using speech generation AI (e.g., Google Text-to-Speech API), which is then stored in a database and used during the interview.

[1372] Interview

[1373] The user responds to the voice questions.

[1374] The user launches an application specifically for interviews and connects to the server. The server sequentially transfers audio files to the user's device and plays questions to the user. The user responds to the questions verbally and records their responses using the device's microphone. Once the recording is complete, the response data is uploaded to the server.

[1375] Analysis of response content

[1376] The server analyzes the response.

[1377] The server receives the response and converts it into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API). It then performs text analysis using NLP technology. Furthermore, it uses an emotion analysis engine to extract the emotional state contained in the response as an evaluation point. For example, it evaluates the confidence, calmness, or nervousness displayed by the applicant in their response.

[1378] Grading and feedback

[1379] The server scores the responses and stores the results.

[1380] The server calculates the evaluation points for each question and generates an overall score. At this time, the analysis results of the sentiment analysis engine are also incorporated into the evaluation score to provide a more comprehensive evaluation. The generated evaluation results and feedback are saved in a database, and the user is notified of the feedback content. At the same time, the evaluation results and feedback are sent to the recruiter, who will use them in the next selection step.

[1381] Specific examples

[1382] For example, when hiring a software engineer, a user would list "more than five years of Java development experience" as one of their skills on an application form. The server would analyze this information and use a generative artificial intelligence model to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, which would then be recorded on the device and sent to the server. The server would then analyze the responses and incorporate their emotional state (e.g., whether they were confident or nervous) into the evaluation. This would enable a fair and effective initial interview and optimize the hiring process.

[1383] An example of a prompt for the generative AI model would be, "Generate questions for applicants with more than five years of development experience in Java."

[1384] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1385] Step 1:

[1386] Submitting an application form

[1387] A user submits an application form. The user accesses a recruitment website and enters the necessary information into the application form, including personal information, work history, educational background, and skills. After completing the entry, the user clicks the "Submit" button. This operation sends the entered data to the server as form-data. The server receives this data and stores it in the application form database.

[1388] Input: Data entered into the application form by the user

[1389] Output: Application form data saved on the server

[1390] Step 2:

[1391] Application form analysis

[1392] The server analyzes the application form. It checks the received form data and formats the content. It uses the natural language processing (NLP) technology Python library spaCy to extract characteristic information such as the applicant's work history, skills, and motivation for applying from the text. The analysis results are saved as characteristic data in a characteristic database.

[1393] Input: Application form data

[1394] Output: Characteristic data

[1395] Step 3:

[1396] question generation

[1397] The server generates a prompt using a generative artificial intelligence model (e.g., OpenAI's GPT-3) based on the characteristic information stored in the characteristic database. The prompt, "Please generate a question for when the applicant has more than five years of development experience in Java," is input into the GPT-3 API, and the generated question is retrieved. The created question is stored in the question database.

[1398] Input: characteristic data, prompt statement

[1399] Output: Question data

[1400] Step 4:

[1401] Speech synthesis

[1402] The server reads the question data and uses the Google Text-to-Speech API to convert the generated question text into an audio file, which is then stored in a speech database.

[1403] Input: Question data

[1404] Output: Audio file

[1405] Step 5:

[1406] Interview

[1407] The user responds to the audio questions. The user launches an interview application and connects to the server. The server sends the audio file to the user's device, and the interview application plays the audio file. The user responds to the questions by voice and records the response audio using the device's microphone. Once recording is complete, the response audio file is uploaded from the device to the server.

[1408] Input: Audio file

[1409] Output: Response audio file

[1410] Step 6:

[1411] Analysis of response content

[1412] The server receives the response audio file and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is then stored in an analysis results database. It is then analyzed using NLP technology and further analyzed for emotional state using an emotion analysis engine. The analysis results of the emotional state are stored as evaluation data.

[1413] Input: Response audio file

[1414] Output: Analysis result data, emotion evaluation data

[1415] Step 7:

[1416] Grading and feedback

[1417] The server calculates an evaluation score based on the analysis result data and the sentiment evaluation data. The server generates an evaluation score and stores it in a score database. It also generates a feedback message and notifies the user and recruiter. Notifications are sent via the email system.

[1418] Input: Analysis result data, emotion evaluation data

[1419] Output: Evaluation score, feedback message

[1420] (Application example 2)

[1421] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1422] Traditional hiring processes have the problem of requiring a lot of time and effort to accurately evaluate applicants' aptitude and skills. Furthermore, the interviewer's subjectivity can sometimes affect the evaluation, potentially resulting in a lack of fairness. Especially when hiring workers in factories, efficient and accurate skill evaluation is required. Therefore, an automated system is needed to quickly and fairly evaluate applicants and find the right talent.

[1423] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an application form, means for analyzing the contents of the application form and extracting the applicant's characteristics, means for generating questions based on the characteristics, means for converting the questions into speech, means for recording the applicant's responses to the speech questions, means for analyzing the responses and generating an evaluation score, means for saving the evaluation score and generating feedback, means for converting the applicant's responses into text using speech recognition technology, means for analyzing the emotions of the responses, means for providing feedback based on the analysis results, and means for using a generative AI model to generate questions for evaluating the applicant's aptitude and skills. This enables the applicant's skills and characteristics to be evaluated accurately and efficiently through an automated process, enabling a fair and reliable hiring process.

[1424] The "means for receiving the application form" refers to a communication means for transferring the information on the application form submitted by the applicant to the server.

[1425] "Means for analyzing the contents of application forms and extracting the characteristics of applicants" refers to means that include natural language processing technology for analyzing the information in application forms submitted by applicants and extracting the skills, experience, and other characteristics of the applicants.

[1426] The "means for generating questions" refers to a means for using a generative AI model to generate appropriate interview questions based on the analyzed applicant's characteristic information.

[1427] The "means for converting a question into voice" refers to a means that utilizes voice synthesis technology to convert the text data of the generated question into voice data.

[1428] "Means for recording applicant responses" refers to a means for recording the applicant's voice responses and saving them as digital data.

[1429] The "means for analyzing responses and generating an evaluation score" refers to a means for converting the recorded responses of applicants into text using voice recognition technology, analyzing the content of the text, and generating an evaluation score.

[1430] The "means for storing evaluation scores and generating feedback" refers to a means for storing the generated evaluation scores in a database or the like and providing them as feedback to applicants and hiring managers.

[1431] "Means of converting applicants' responses into text using voice recognition technology" refers to the use of voice recognition software to convert recorded voice data into text data.

[1432] "Means for analyzing emotion from response voice" refers to means for analyzing voice data of an applicant and using an emotion analysis engine to evaluate the applicant's emotional state (e.g., confident, calm, nervous).

[1433] The "means for providing feedback based on the analysis results" refers to a means for generating detailed feedback including an evaluation score and an emotional evaluation based on the results of analyzing the applicant's responses, and providing the feedback to the applicant and the recruiter.

[1434] "Means of using a generative AI model to generate questions to assess aptitudes and skills" refers to means of using generative AI technology, such as a machine learning model, to generate appropriate interview questions based on the characteristics of applicants.

[1435] The system of the present invention is composed of a user, a server, a terminal, and an emotion engine, and these elements work in cooperation with each other. Specific embodiments will be described below.

[1436] Submitting an application form

[1437] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server.

[1438] Application form analysis

[1439] The server analyzes the received application form data and uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. This analysis uses Python's NLP library and machine learning models.

[1440] question generation

[1441] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used.

[1442] Speech synthesis

[1443] The server converts the generated question text into audio data. A speech synthesis AI is used to convert the text data into natural spoken language. This audio file is saved in an appropriate format and sent to the user's device.

[1444] Interview

[1445] The user starts up the interview application installed on the terminal and conducts the interview. The terminal sequentially plays back the voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server.

[1446] Analysis of response content

[1447] The server receives the response data and converts it into text using speech recognition technology. It then analyzes the text using NLP technology to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the response. For example, emotions expressed by the user in their response, such as confidence, calmness, or nervousness, can be extracted as evaluation points.

[1448] Grading and feedback

[1449] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The generated evaluation results and feedback are saved in a database, and the feedback is notified to the user. The evaluation results and feedback are also sent to the recruiter, who uses them in the next selection step.

[1450] Specific examples

[1451] For example, when hiring a software engineer, a user would write "5+ years of Java development experience" on their application form. The server would analyze this information and use generative AI to generate questions such as "Please explain the challenges you faced in a Java project and how you solved them," and then perform speech synthesis. The user would then respond to the questions verbally, and the device would record and send the responses to the server. The server would then analyze the responses and evaluate the user's emotional state, including whether they were confident or nervous. This allows for a fair and effective initial interview.

[1452] Prompt Sentence Examples

[1453] "Based on the technical skills (e.g. Java, ROS) that applicants have listed on their application form, generate questions to assess their specific experience and problem-solving abilities."

[1454] The above is a specific embodiment for carrying out the present invention. Through this system, applicant characteristics can be accurately evaluated, and an efficient and fair hiring process can be realized.

[1455] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1456] Step 1:

[1457] Users submit application forms using a dedicated website or application. They enter the necessary information, such as their personal information, work history, educational background, and skills, and click the submit button. The entered application form data is sent to the server. The input data includes the application form information in text format, and the output data is the application form data that has reached the server.

[1458] Step 2:

[1459] The server analyzes the received application form data. It uses natural language processing (NLP) technology to extract characteristic information such as the user's work history, skills, and motivation for applying. Specifically, it uses a Python NLP library (e.g., NLTK, Spacy, etc.) to analyze the text data. The input data includes the application form data, and the output data is the extracted characteristic information.

[1460] Step 3:

[1461] The server uses a generative AI model to generate appropriate questions based on the analyzed characteristic information. For example, specific questions about "past project experience and problem-solving ability" are created based on the applicant's work history. Generative AI such as GPT (Generative Pre-trained Transformer) is used. The input data includes characteristic information, and the generated question text is obtained as output data.

[1462] Step 4:

[1463] The server converts the generated question text into audio data. A speech synthesis AI (e.g., Google Text-to-Speech) is used to generate the audio, converting the text data into natural spoken language. This audio file is saved in an appropriate format (e.g., MP3, WAV, etc.). The input data includes the question text, and the output data is an audio file.

[1464] Step 5:

[1465] The user starts a dedicated interview application installed on the terminal and conducts the interview. The terminal sequentially plays back voice questions transferred from the server, and the user responds to the questions by voice. The terminal records the responses and uploads the response data to the server. The input data includes the voice questions and the user's responses, and the recorded response data is obtained as output data.

[1466] Step 6:

[1467] The server receives the response data and converts the response to text using speech recognition technology. Specifically, it uses speech recognition software (e.g., Google Speech-to-Text) to convert the recorded voice data to text. The input data includes the voice response, and the output data is the text response.

[1468] Step 7:

[1469] The server analyzes the text responses to evaluate the applicant's technical depth and communication skills. It also uses an emotion engine to analyze the emotional state contained in the responses. For example, it uses voice analysis software (e.g., IBM Watson Tone Analyzer) to evaluate the emotional state (e.g., confident, calm, nervous). Input data includes the text responses and voice data, and output data includes the emotion analysis results and an evaluation score.

[1470] Step 8:

[1471] The server calculates the evaluation points for each question and generates an overall score. At this time, a comprehensive evaluation is performed, including the analysis results of the emotion engine. The evaluation results and feedback are saved in a database and notified to the user as feedback. The evaluation results and feedback are also sent to the hiring manager. The input data includes the evaluation score and emotion analysis results, and the output data is an overall evaluation score and feedback.

[1472] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1473] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1474] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1475] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1476] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1477] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1478] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1479] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1480] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1481] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1482] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1483] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1484] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1485] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1486] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1487] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1488] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1489] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1490] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1491] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1492] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1493] The following is further disclosed regarding the above embodiment.

[1494] (Claim 1)

[1495] A means for receiving the application form;

[1496] A means for analyzing the contents of the application form and extracting characteristics of the applicant;

[1497] means for generating a query based on said characteristics;

[1498] means for converting the question into speech;

[1499] means for recording the applicant's responses to said audio questions;

[1500] means for analyzing the responses and generating a reputation score;

[1501] means for storing said evaluation scores and generating feedback;

[1502] A system including:

[1503] (Claim 2)

[1504] 10. The system of claim 1, wherein the means for generating questions utilizes natural language processing techniques to generate questions.

[1505] (Claim 3)

[1506] 2. The system of claim 1, wherein the means for analyzing the applicant's responses to the voice questions uses voice recognition technology to transcribe the responses and analyze their content.

[1507] "Example 1"

[1508] (Claim 1)

[1509] A means for receiving the application form;

[1510] A means for analyzing the contents of the application form using natural language processing technology and extracting characteristic information of the applicant;

[1511] A means for generating questions using AI generation technology based on the characteristic information;

[1512] means for converting the question into speech using speech synthesis technology;

[1513] means for the applicant to record responses to said audio questions;

[1514] a means for converting the response voice into text using a voice recognition technology and analyzing the content of the text again to generate an evaluation score;

[1515] means for storing said evaluation scores and generating feedback;

[1516] A system including:

[1517] (Claim 2)

[1518] The system of claim 1, wherein the means for generating questions using the generative AI technology uses natural language processing technology to create prompt sentences and generate questions based on the prompt sentences.

[1519] (Claim 3)

[1520] 2. The system according to claim 1, wherein the response is converted into text using the voice recognition technology, and the content of the text is analyzed to extract characteristic information about the applicant.

[1521] "Application Example 1"

[1522] (Claim 1)

[1523] A means for receiving the application form;

[1524] A means for analyzing the contents of the application form and extracting characteristics of the applicant;

[1525] means for generating a query based on said characteristics;

[1526] means for converting the question into speech;

[1527] means for recording the applicant's responses to said audio questions;

[1528] means for analyzing the responses and generating a reputation score;

[1529] means for storing said evaluation scores and generating feedback;

[1530] a generative AI model that generates technical questions based on the analyzed data; and

[1531] a means for transcribing the applicant's responses to the voice questions using speech recognition technology and then analyzing them again using natural language processing technology;

[1532] A system including:

[1533] (Claim 2)

[1534] 10. The system of claim 1, wherein the means for generating questions utilizes natural language processing techniques to generate questions.

[1535] (Claim 3)

[1536] 2. The system of claim 1, wherein the means for analyzing the applicant's responses to the voice questions uses voice recognition technology to transcribe the responses and analyze their content.

[1537] "Example 2: Combining Emotion Engines"

[1538] (Claim 1)

[1539] A means for receiving the application form;

[1540] A means for analyzing the contents of the application form and extracting characteristics of the applicant;

[1541] means for generating questions using a generative artificial intelligence model based on the characteristics;

[1542] means for converting the question into speech;

[1543] means for recording the applicant's responses to said audio questions;

[1544] a means for converting the response into text data using a voice recognition technology, analyzing the content of the text data using a sentiment analysis engine, and generating an evaluation score;

[1545] means for storing said evaluation scores and generating feedback;

[1546] A system including:

[1547] (Claim 2)

[1548] 2. The system of claim 1, wherein the means for generating a question uses an input prompt sentence to generate a question using a generative artificial intelligence model.

[1549] (Claim 3)

[1550] 10. The system of claim 1, wherein the means for analyzing the response utilizes a sentiment analysis engine to evaluate the emotional state of the response.

[1551] "Application example 2 when combining emotion engines"

[1552] (Claim 1)

[1553] A means for receiving the application form;

[1554] A means for analyzing the contents of the application form and extracting characteristics of the applicant;

[1555] means for generating a query based on said characteristics;

[1556] means for converting the question into speech;

[1557] means for recording the applicant's responses to said audio questions;

[1558] means for analyzing the responses and generating a reputation score;

[1559] means for storing said evaluation scores and generating feedback;

[1560] A means for converting applicants' responses into text using voice recognition technology;

[1561] A means for analyzing emotions in the response voice;

[1562] means for providing feedback based on the analysis results;

[1563] using a generative AI model to generate questions to assess the aptitudes and skills of applicants;

[1564] A system including:

[1565] (Claim 2)

[1566] 10. The system of claim 1, wherein the means for generating questions utilizes natural language processing techniques to generate questions.

[1567] (Claim 3)

[1568] 2. The system of claim 1, wherein the means for analyzing the applicant's responses to the voice questions uses voice recognition technology to transcribe the responses and analyze their content. [Explanation of symbols]

[1569] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving the application form; A means for analyzing the contents of the application form and extracting characteristics of the applicant; means for generating a query based on said characteristics; means for converting the question into speech; means for recording the applicant's responses to said audio questions; means for analyzing the responses and generating a reputation score; means for storing said evaluation scores and generating feedback; A system including:

2. 10. The system of claim 1, wherein the means for generating questions utilizes natural language processing techniques to generate questions.

3. 2. The system of claim 1, wherein the means for analyzing the applicant's responses to the voice questions uses voice recognition technology to transcribe the responses and analyze their content.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A