system

The system addresses subjective bias in interviews by converting audio to text, analyzing for evaluation metrics, and generating scores, improving the fairness and accuracy of personnel selection through customized criteria.

JP2026100550APending Publication Date: 2026-06-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-09
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Conventional employment interviews face challenges in conducting fair and accurate personnel evaluations due to subjective judgment and bias, leading to mismatches in personnel selection and increased evaluation burden, with a lack of standardized processes hindering continuous improvement.

Method used

A system that utilizes speech recognition technology to convert interview audio into text data, analyzes it for evaluation metrics, and generates scores, allowing for objective evaluation by presenting these scores to interviewers, with the ability to customize evaluation criteria based on company needs.

Benefits of technology

Enhances the fairness and accuracy of personnel evaluations by reducing bias and enabling tailored metrics, facilitating the selection of suitable candidates aligned with organizational requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026100550000001_ABST
    Figure 2026100550000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for receiving interview data, The means for converting the aforementioned interview data into text data using speech recognition technology, A means for analyzing the aforementioned text data and generating scores for each interview evaluation metric, A means for presenting the generated score to the interviewer, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional employment interviews, there is a problem that it is difficult to conduct a fair and accurate personnel evaluation due to the subjective judgment and bias of interviewers. For this reason, it is impossible to efficiently employ personnel suitable for the needs of the company, and there is a possibility of a mismatch of personnel. Furthermore, the evaluation burden on interviewers is large, and there is no standardized interview evaluation process, so continuous improvement is difficult.

Means for Solving the Problems

[0005] This invention provides a system that receives interview data, converts it into text data using speech recognition technology, analyzes it, and generates scores for each evaluation metric of the interview. This allows the system to present each score to the interviewer, providing them with a basis for their final evaluation decision, thereby enhancing the fairness of the evaluation. Furthermore, the system improves analysis accuracy by training the evaluation criteria using past interview record data. In addition, it aims to improve the suitability of personnel within companies by allowing the setting and customization of evaluation metrics according to the specific needs of each company.

[0006] "Interview data" refers to audio files and related text-based information collected during the interview.

[0007] "Speech recognition technology" refers to the process of analyzing speech data and converting it into corresponding text data.

[0008] "Text data" refers to data that is generated by speech recognition technology and expressed as textual information.

[0009] "Evaluation metrics" refer to the criteria used to quantify and evaluate specific skills and abilities during an interview.

[0010] A "score" refers to a numerical representation of performance in each evaluation metric based on the analysis results.

[0011] An "interviewer" refers to a person responsible for conducting and evaluating interviews.

[0012] A "generative AI model" refers to an algorithm or program that learns from past data to support the decision-making process in interview evaluations.

[0013] "Customization" refers to adjusting a system or evaluation metrics to suit specific usage purposes or conditions. [Brief explanation of the drawing]

[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a system for reducing bias and inequality in evaluations during job interviews. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation metric. To achieve this objective, the system operates as follows:

[0036] Data reception and conversion

[0037] The user provides the system with audio data collected during the interview.

[0038] The server receives this audio data and converts it into text data using speech recognition technology.

[0039] Data analysis and evaluation

[0040] The server analyzes the generated text data and extracts key evaluation metrics that should be assessed during the interview.

[0041] The server generates interview scores based on evaluation metrics, including communication skills and problem-solving abilities.

[0042] Presentation of results and feedback

[0043] The device displays the generated score and its evaluation criteria on the screen and presents them to the user acting as the interviewer.

[0044] Users can use this information to make their final hiring decisions.

[0045] Specific example

[0046] For example, in a company interview, a user collects audio data from five candidates. The server converts this data into text in real time and analyzes it for each candidate. The analysis results display scores for each evaluation metric, such as "Communication Skills: 8 / 10" for Candidate A and "Problem-Solving Ability: 7 / 10" for Candidate B. The user can use this information to help select the most suitable candidate for the company.

[0047] This system can improve the accuracy of its analysis by having the model learn evaluation metrics using past interview data. The system also allows for customized evaluation criteria tailored to each company's needs, thus supporting the selection of personnel who are a better fit for the company.

[0048] The following describes the processing flow.

[0049] Step 1:

[0050] The user initiates the interview and collects audio data to record the interaction with the candidate. The audio data is transmitted to the system via a pre-configured, dedicated interface.

[0051] Step 2:

[0052] The server processes the audio data received from the user in real time using a speech recognition engine and converts it into text data. This text conversion utilizes natural language processing technology to generate accurate text.

[0053] Step 3:

[0054] The server analyzes the converted text data and extracts key evaluation metrics that should be assessed during the interview. These metrics are based on criteria such as communication skills and technical abilities.

[0055] Step 4:

[0056] The server generates quantified scores for each evaluation metric. This scoring system utilizes insights and criteria accumulated from past interview data to provide accurate assessments.

[0057] Step 5:

[0058] The terminal displays the generated score and the underlying analytical data on its screen, presenting it to the user acting as the interviewer. The displayed information is summarized visually in an easy-to-understand format using graphs and numerical data.

[0059] Step 6:

[0060] Users make their final hiring decisions based on the presented scores and related evaluation information. This allows for the efficient selection of candidates who best meet the company's needs.

[0061] (Example 1)

[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0063] During the interview process, subjective evaluations are often biased, making it difficult to select appropriate candidates. Furthermore, the lack of standardized evaluation criteria means that evaluations differ from organization to organization, posing a challenge. There is also a need to improve the accuracy of data analysis to enhance the objectivity of evaluation results.

[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0065] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using a speech recognition method, means for analyzing the text information and generating evaluation values ​​for each interview evaluation criterion, and means for generating prompt sentences and inputting them to a generation AI model. This reduces evaluation bias, enables highly objective evaluations based on unified criteria, and allows for the selection of personnel suitable for each organization.

[0066] "Interview information" refers to digital information, including audio, text, and related data, collected during the interview process.

[0067] "Speech recognition methods" refer to the technology that converts speech data into text data, and the process of making human speech understandable to machines.

[0068] "Textual information" refers to information in text format obtained as a result of conversion using speech recognition methods.

[0069] "Analysis" refers to the process of analyzing textual information to extract meaningful evaluation indicators, and is a process for making specific evaluations based on that information.

[0070] "Evaluation criteria" refer to the scales or indicators used to assess a candidate's abilities and suitability during an interview.

[0071] An "evaluation value" is a numerical or quantitative result obtained through analysis, and is a score calculated based on evaluation criteria.

[0072] A "generative AI model" is an artificial intelligence model that uses machine learning or deep learning algorithms to train itself to solve specific problems.

[0073] A "prompt" is a text-based instruction used to tell a generative AI model to input information or perform an action.

[0074] This invention relates to a system that reduces bias and prejudice in human evaluations during interviews, enabling objective evaluation. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation criterion.

[0075] The user collects the interviewee's statements during the interview using an audio recording device. This data is uploaded to a server in digital format. The audio data collected by the user is transmitted to the server via the internet.

[0076] The server converts the received audio data into text using speech recognition software (e.g., a common speech recognition API). Based on this converted text information, the server performs further analysis. This analysis uses natural language processing libraries (such as NLTK or spaCy). The purpose of this analysis is to extract important evaluation criteria that should be assessed in an interview, such as communication skills and problem-solving abilities.

[0077] Based on the analysis results, the server generates evaluation values ​​for each evaluation criterion. Machine learning models and generative AI models are used to generate these evaluation values. For example, a prompt such as "Evaluate this candidate's communication skills on a scale of 1 to 10" can be input into a generative AI model to perform appropriate scoring.

[0078] The device visually displays the generated evaluation scores and presents them to the user acting as the interviewer. This allows the user to eliminate individual bias and evaluate candidates from an objective perspective.

[0079] As a concrete example, in an interview process at a certain organization, the user collects audio data from multiple candidates, and a server scores them. The results are displayed on the device, for example, Candidate A might have a score of "Communication Skills: 8 / 10," and Candidate B might have a score of "Problem-Solving Ability: 7 / 10." Based on this information, the user can select the person best suited to the organization.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The user records the candidate's statements during the interview and saves the audio data in digital format. This audio data is then uploaded to a server via the internet. The input is the audio data, and the output is the status indicating that the data transfer to the server is complete.

[0083] Step 2:

[0084] The server converts the received audio data into text using speech recognition technology. In this process, the speech recognition API takes the audio data as input and outputs the resulting text. Checks are also performed to verify the accuracy of the converted text.

[0085] Step 3:

[0086] The server analyzes the text information obtained through speech recognition and extracts keywords and phrases related to the evaluation criteria. A natural language processing library analyzes the text information as input, identifies the necessary evaluation criteria, and outputs them. The specific actions performed in this step are text analysis and information extraction.

[0087] Step 4:

[0088] The server generates prompt sentences based on the information extracted through analysis and inputs them into the AI ​​model. At this point, the machine learning model generates an evaluation score. This process takes the results of text analysis as input and outputs evaluation values ​​for each evaluation criterion. Specifically, the process involves the formation of prompt sentences and their input into the AI ​​model.

[0089] Step 5:

[0090] The terminal visually displays the evaluation score and its details received from the server and presents them to the user. The output is a visualized result of the evaluation score. The specific function is to enable the user to make a quick decision by displaying the evaluation score on the screen.

[0091] Step 6:

[0092] The user makes a final hiring decision based on the displayed evaluation scores. In this process, the data displayed on the screen is used as input, and the output is a decision based on the interview results. Specifically, the user compares and reviews the interview results and selects the best candidate.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] Traditional interview systems often rely on the subjective opinions of interviewers, leading to bias and inaccuracies. Furthermore, the difficulty in visually comparing candidate evaluations can hinder rational decision-making. Additionally, interview evaluation criteria may not align with the needs of the company or organization, posing challenges in selecting the most suitable candidates.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes means for receiving interview data, means for converting the interview data into text data using speech recognition technology, means for analyzing the text data and generating scores for each interview evaluation metric, and means for making the evaluation results of candidates visually comparable. This reduces bias in interview evaluations and enables a fair and rational interview process. Furthermore, by setting customized evaluation criteria according to the needs of companies and organizations, it becomes possible to select appropriate personnel.

[0098] "Interview data" refers to data that includes audio information exchanged between the candidate and the interviewer.

[0099] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.

[0100] "Text data" refers to data that contains string information converted using speech recognition technology.

[0101] "Analysis" is a method of extracting specific evaluation metrics from text data and performing evaluations based on those metrics.

[0102] "Evaluation metrics" are standards used to measure specific abilities and skills that are considered important during an interview.

[0103] A "score" represents an evaluation result that has been quantified based on evaluation indicators.

[0104] A "terminal" is an electronic device used to display evaluation results.

[0105] "Comparable" means that by comparing multiple evaluation results side-by-side, characteristics and abilities can be easily compared.

[0106] "Organizational needs" refer to the specific skill sets and aptitude requirements that a particular company or organization seeks.

[0107] "Customized performance metrics" are evaluation criteria that have been tailored to the specific requirements of a particular organization.

[0108] The system implementing this invention is configured to receive interview data and convert it into text data using speech recognition technology. The server uses the "SpeechRecognition" library as the speech recognition technology to effectively convert the audio information into text.

[0109] The server uses natural language processing libraries such as "NLTK" or "spaCy" to analyze this text data. During the analysis, it extracts important elements based on interview evaluation metrics and generates scores for them. A generative AI model is used for the analysis, leveraging patterns learned through past interview data.

[0110] The device presents the generated score to the user. At this time, it displays the score in a visually comparable format, such as a graph or list, to allow the user to intuitively understand the evaluation. It also has a function to set customized evaluation metrics based on the needs of the company or organization.

[0111] As a concrete example, when a user opens the app on their smartphone and presses the record button, the conversation with the candidate is automatically recorded and sent to the server. The results, analyzed by the server, are displayed as a score on the smartphone screen in real time.

[0112] An example of a prompt for a generative AI model is, "Design a prompt to identify the necessary evaluation metrics in an interview and generate a score based on the candidate's voice data." This allows users to create a fairer and more efficient interview process.

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The user begins the interview and starts recording audio data using an application on the device. The device receives and records the interview audio data via the microphone. It takes the user's audio data as input and generates an audio file saved to local storage as output.

[0116] Step 2:

[0117] The terminal sends the recorded audio file to the server. The server receives this audio data and converts it into text data using speech recognition technology. Specifically, it analyzes the audio waveform using the "SpeechRecognition" library and generates the corresponding text. It takes an audio file as input and generates text data in string format as output.

[0118] Step 3:

[0119] The server analyzes the generated text data using a natural language processing library ("NLTK" or "spaCy"). Here, keywords and phrases related to the evaluation metrics are extracted. The generative AI model identifies features for score generation based on prompts derived from past data. The input is text data, and the output is data with identified evaluation metrics.

[0120] Step 4:

[0121] The server calculates a score based on the identified evaluation metrics. The calculation uses machine learning algorithms (e.g., scikit-learn or TENSORFLOW®) to generate numerical scores for each metric. The input is the evaluation metric data, and the output is the numerical score.

[0122] Step 5:

[0123] The terminal receives scores sent from the server and visualizes them through a user interface. Graphs and lists are used to allow for visual comparison of scores among multiple candidates. The input is a numerical score, and the output is a visualized evaluation result.

[0124] Step 6:

[0125] Users compare and evaluate candidates based on the provided evaluation results and make hiring decisions. By comparing candidates, users can identify their relative strengths and weaknesses, enabling more rational decision-making. The input is the visualized evaluation results, and the output is the final hiring decision.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention provides a system that enables more precise talent evaluation by combining an emotion engine to recognize and incorporate the emotional state of candidates during job interviews into the evaluation process. The system includes processing to detect emotions in addition to analyzing voice data.

[0128] Data reception and conversion

[0129] The user records the conversation with the candidate as audio during the interview and sends the data to the system.

[0130] Recognition of voice and emotion

[0131] The server processes the received audio data using a speech recognition engine to convert it into text data. Furthermore, it uses an emotion engine to analyze the tone and tempo of the voice and recognize the candidate's emotional state.

[0132] Data analysis and evaluation

[0133] The server generates scores for each evaluation metric based on the content of the text data and the recognized sentiment information. Sentiment information is incorporated into the evaluation of communication skills and personal characteristics.

[0134] Presentation of results and feedback

[0135] The device displays emotional information, along with the generated evaluation score, as part of the analysis results on the screen and presents it to the user acting as the interviewer. This enables a comprehensive evaluation that includes emotional aspects.

[0136] Specific example

[0137] In a typical interview, a user collects audio data from five candidates. The server converts the audio data into text and uses an emotion engine to analyze the emotions each candidate expressed. For example, it evaluates the level of confidence and nervousness that candidate A displayed when answering questions, and reflects these factors influencing their communication skills in a score. The device then graphs this information, providing the interviewer with an intuitive understanding. The user can leverage this detailed feedback to make informed decisions about selecting the best candidates.

[0138] Thus, by integrating an emotion engine, the system of this invention grasps the latent characteristics of candidates that cannot be captured by conventional language analysis alone, and supports talent evaluation that is more suitable for companies.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] After the interview begins, the user uses a dedicated recording device to collect audio data of the conversation with the candidate. This data is transmitted to the system in real time via a secure connection.

[0142] Step 2:

[0143] The server inputs the audio data sent by the user into the speech recognition engine and converts it into text data sequentially. This conversion process includes pre-processing such as background noise reduction and speaker separation.

[0144] Step 3:

[0145] The server simultaneously inputs voice data into the emotion engine and identifies the candidate's emotional state by analyzing features such as tone, volume, and speed of voice. Emotional information is classified into categories such as joy, tension, and surprise.

[0146] Step 4:

[0147] The server integrates the generated text data and sentiment data, inputs it into the evaluation model, and calculates scores for each evaluation metric. Specifically, it scores metrics such as communication skills, cooperativeness, and adaptability.

[0148] Step 5:

[0149] The device displays the evaluation score and sentiment analysis results for each candidate as visualized data on the screen. The display uses graphs, charts, and numerical data, providing information in a visually easy-to-understand format.

[0150] Step 6:

[0151] Users deepen their thinking based on the information presented and make a comprehensive evaluation of candidates from a more objective perspective. This improves the transparency and rationality of the selection process, enabling the selection of the most suitable talent.

[0152] (Example 2)

[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0154] In recruitment interviews, traditional evaluation methods struggle to adequately grasp candidates' nonverbal characteristics and emotions, resulting in difficulties in selecting appropriate personnel. Furthermore, reliance on the interviewer's subjective evaluation leads to a lack of consistency and objectivity.

[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0156] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using speech analysis technology, and means for recognizing the emotional state from the text information and generating numerical values ​​for each evaluation criterion using the results. This makes it possible to reflect the candidate's emotional state during the interview in the evaluation, enabling a more objective and consistent personnel evaluation.

[0157] "Interview information" refers to audio data and related data obtained from conversations and interactions with candidates during job interviews.

[0158] "Speech analysis technology" is a technology that processes speech data and converts it into text information, and in particular, it uses a speech recognition engine for analysis.

[0159] "Textual information" refers to text data that represents the content of audio data, converted using speech analysis technology.

[0160] "Emotional state" refers to the emotional characteristics and reactions that a candidate exhibits during a conversation, and is analyzed from their voice and tone.

[0161] "Evaluation criteria" are standards or scales used to measure a candidate's abilities and aptitudes, and are set according to the organization's needs and objectives.

[0162] "Numerical values" refer to data that quantitatively represents a candidate's abilities and characteristics, calculated based on evaluation criteria.

[0163] A "decision-maker" refers to an individual or organization that has the authority to make decisions regarding personnel selection based on the results of job interviews.

[0164] This invention is a system for recognizing and evaluating the emotional state of candidates during job interviews, and is primarily composed of audio data processing. This system operates as follows:

[0165] Data collection and transmission

[0166] The user records the conversation with the candidate during the interview as audio data using a recording device. This data is then sent to a server.

[0167] Analysis of audio data

[0168] The server analyzes the received audio data. Specifically, it uses a speech recognition engine (e.g., a common cloud-based speech recognition API) as an audio analysis technology to convert the audio data into text information. It also uses an emotion engine (e.g., commercially available emotion recognition software) to analyze the tone and speed of the voice and recognize the candidate's emotional state.

[0169] Generation and display of evaluation scores

[0170] The server uses a generative AI model to generate numerical values ​​for each evaluation criterion, based on the converted text information and recognized emotional states.

[0171] The terminal visualizes numerical data and sentiment information generated on the server, displaying it intuitively through graphs and charts. This information is provided to the user, who is the decision-maker, to assist in the overall evaluation of candidates.

[0172] As a concrete example, suppose a user collects audio data from five candidates during an interview session. The server processes this audio, performing speech recognition and sentiment analysis. For instance, it evaluates candidate A's confidence and nervousness, reflecting these in a numerical value corresponding to their communication skills. The terminal then intuitively displays these evaluation results, allowing the user to easily compare the characteristics of each candidate. This detailed information is extremely useful in making hiring decisions.

[0173] An example of a prompt might be: "Please describe a method for analyzing candidate voice data during interviews to quantify the influence of emotions on the evaluation of communication skills."

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] The user records the conversation with the candidate during the job interview as audio data using a recording device. The recorded audio data is sent to the server in its original format. In this step, the input is the candidate's audio data, and the output is the transmission of the audio data to the server. Specifically, the user presses the record button and continues recording until the audio data is of sufficient length.

[0177] Step 2:

[0178] The server converts received audio data into text data through a speech recognition engine. A cloud-based speech recognition API is used for this process. The input is audio data, and the output is the corresponding text data. Specifically, the server batch processes the audio data for analysis and generates recognition results in real time.

[0179] Step 3:

[0180] The server uses an emotion engine to analyze the converted text data and the original audio data to recognize the candidate's emotional state. In this step, text and audio data are taken as input, and information about the emotional state is obtained as output. Specifically, the server analyzes the tone and speed of the voice and extracts emotional characteristics.

[0181] Step 4:

[0182] The server utilizes a generative AI model to quantify the obtained text data and sentiment information based on evaluation criteria. The input consists of text data and sentiment information, and the output is a numerical value generated according to the evaluation criteria. The server processes the data using algorithms and generates numerical values ​​for each evaluation metric.

[0183] Step 5:

[0184] The terminal visualizes numerical data and emotional states transmitted from the server, displaying them on the screen as graphs and charts. Inputs are numerical data and emotional state information, while output is a visual evaluation report. Specifically, the terminal updates data in real time and displays it in a dashboard format for intuitive user understanding.

[0185] Step 6:

[0186] The user makes an overall evaluation of the candidates based on the data displayed on the device. The input is the displayed evaluation results, and the output is the final hiring decision. Specifically, the user reviews the evaluation results on the device and compares the candidates to reach a conclusion.

[0187] (Application Example 2)

[0188] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0189] In collaborative work between humans and robots in factories, it is not easy to grasp the emotional state of workers in real time and for robots to take appropriate actions and provide feedback in response to those emotions. Therefore, there is a need for systems that can improve work efficiency and safety.

[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0191] In this invention, the server includes means for receiving voice data, means for converting it into text data using speech recognition technology, and means for analyzing voice tone and tempo to recognize emotional state. This enables real-time understanding of the emotional state of workers and allows for appropriate feedback and decision-making regarding actions.

[0192] "Interview data" refers to all information collected during an interview, including audio and video.

[0193] "Speech recognition technology" is a technology that converts speech data into text data.

[0194] "Text data" refers to data that has been converted using speech recognition technology and is represented as textual information.

[0195] "Tone" refers to elements that describe characteristics such as pitch and volume of sound.

[0196] "Tempo" refers to the speed and rhythm of speech.

[0197] "Emotional state" refers to the psychological state analyzed from audio data, and includes emotions such as joy, anger, sadness, and happiness.

[0198] "Evaluation indicators" refer to the criteria and perspectives used to evaluate individuals during interviews or work assignments.

[0199] A "score" is a rating or numerical value calculated based on evaluation metrics.

[0200] "User" is a general term referring to anyone who uses a system or receives a service.

[0201] This system is designed to facilitate collaborative work with factory workers. The server receives voice data from workers in real time and converts it to text using speech recognition technology. Here, speech recognition software such as Google® Cloud Speech-to-Text API is used. Subsequently, an emotion engine, such as Microsoft® Azure® Emotion API, is used to analyze the tone and tempo of the voice and recognize the worker's emotional state.

[0202] Based on the information obtained through this emotion analysis, the server generates feedback and instructions tailored to the work situation and displays them on the terminal. The terminal refers to a smartphone or head-mounted display, providing information in a way that the worker can intuitively understand. For example, if the server determines that the worker is feeling anxious, the terminal will display a message such as, "Take your time, there's no need to rush."

[0203] In this system, the user refers to a worker, and by receiving feedback, improvements in work efficiency and safety can be expected. For example, if an incorrect operation is detected while handling specialized tools, the robot will suggest appropriate corrective steps.

[0204] An example of a prompt for a generative AI model is: "Analyze the worker's current emotions from this audio data and generate instructions for the robot to provide appropriate feedback."

[0205] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0206] Step 1:

[0207] Users record spoken audio in real time at the work site using a smartphone or head-mounted display and send the audio data to a server. The input is audio data, and the output is the raw audio data sent to the server. This process allows workers' voices to be stored as data.

[0208] Step 2:

[0209] The server converts the received audio data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose. The input is raw audio data, and the output is text data extracted from the audio. This process converts the audio content into a format that can be parsed as text information.

[0210] Step 3:

[0211] The server analyzes the converted text data and the tone and tempo of the speech, and uses the Microsoft Azure Emotion API to recognize the emotional state. The input is text data and speech feature information, and the output is the recognized emotional state. This process makes it possible to determine the user's psychological state.

[0212] Step 4:

[0213] The server determines appropriate feedback or action plans based on the worker's emotional state and work progress. Inputs are emotional state and individual work status data, while outputs are feedback content and action plans. This allows for appropriate advice and warnings to be given to the worker.

[0214] Step 5:

[0215] The device displays the determined feedback or action plan to the user. The input is the feedback content, and the output is visual or audio information presented to the user. This operation allows the user to intuitively receive guidance for improvement or correction while working.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] This invention provides a system for reducing bias and inequality in evaluations during job interviews. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation metric. To achieve this objective, the system operates as follows:

[0233] Data reception and conversion

[0234] The user provides the system with audio data collected during the interview.

[0235] The server receives this audio data and converts it into text data using speech recognition technology.

[0236] Data analysis and evaluation

[0237] The server analyzes the generated text data and extracts key evaluation metrics that should be assessed during the interview.

[0238] The server generates interview scores based on evaluation metrics, including communication skills and problem-solving abilities.

[0239] Presentation of results and feedback

[0240] The device displays the generated score and its evaluation criteria on the screen and presents them to the user acting as the interviewer.

[0241] Users can use this information to make their final hiring decisions.

[0242] Specific example

[0243] For example, in a company interview, a user collects audio data from five candidates. The server converts this data into text in real time and analyzes it for each candidate. The analysis results display scores for each evaluation metric, such as "Communication Skills: 8 / 10" for Candidate A and "Problem-Solving Ability: 7 / 10" for Candidate B. The user can use this information to help select the most suitable candidate for the company.

[0244] This system can improve the accuracy of its analysis by having the model learn evaluation metrics using past interview data. The system also allows for customized evaluation criteria tailored to each company's needs, thus supporting the selection of personnel who are a better fit for the company.

[0245] The following describes the processing flow.

[0246] Step 1:

[0247] The user initiates the interview and collects audio data to record the interaction with the candidate. The audio data is transmitted to the system via a pre-configured, dedicated interface.

[0248] Step 2:

[0249] The server processes the audio data received from the user in real time using a speech recognition engine and converts it into text data. This text conversion utilizes natural language processing technology to generate accurate text.

[0250] Step 3:

[0251] The server analyzes the converted text data and extracts key evaluation metrics that should be assessed during the interview. These metrics are based on criteria such as communication skills and technical abilities.

[0252] Step 4:

[0253] The server generates quantified scores for each evaluation metric. This scoring system utilizes insights and criteria accumulated from past interview data to provide accurate assessments.

[0254] Step 5:

[0255] The terminal displays the generated score and the underlying analytical data on its screen, presenting it to the user acting as the interviewer. The displayed information is summarized visually in an easy-to-understand format using graphs and numerical data.

[0256] Step 6:

[0257] Users make their final hiring decisions based on the presented scores and related evaluation information. This allows for the efficient selection of candidates who best meet the company's needs.

[0258] (Example 1)

[0259] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0260] During the interview process, subjective evaluations are often biased, making it difficult to select appropriate candidates. Furthermore, the lack of standardized evaluation criteria means that evaluations differ from organization to organization, posing a challenge. There is also a need to improve the accuracy of data analysis to enhance the objectivity of evaluation results.

[0261] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0262] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using a speech recognition method, means for analyzing the text information and generating evaluation values ​​for each interview evaluation criterion, and means for generating prompt sentences and inputting them to a generation AI model. This reduces evaluation bias, enables highly objective evaluations based on unified criteria, and allows for the selection of personnel suitable for each organization.

[0263] "Interview information" refers to digital information, including audio, text, and related data, collected during the interview process.

[0264] "Speech recognition methods" refer to the technology that converts speech data into text data, and the process of making human speech understandable to machines.

[0265] "Textual information" refers to information in text format obtained as a result of conversion using speech recognition methods.

[0266] "Analysis" refers to the process of analyzing textual information to extract meaningful evaluation indicators, and is a process for making specific evaluations based on that information.

[0267] "Evaluation criteria" refer to the scales or indicators used to assess a candidate's abilities and suitability during an interview.

[0268] An "evaluation value" is a numerical or quantitative result obtained through analysis, and is a score calculated based on evaluation criteria.

[0269] A "generative AI model" is an artificial intelligence model that uses machine learning or deep learning algorithms to train itself to solve specific problems.

[0270] A "prompt" is a text-based instruction used to tell a generative AI model to input information or perform an action.

[0271] This invention relates to a system that reduces bias and prejudice in human evaluations during interviews, enabling objective evaluation. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation criterion.

[0272] The user collects the interviewee's statements during the interview using an audio recording device. This data is uploaded to a server in digital format. The audio data collected by the user is transmitted to the server via the internet.

[0273] The server converts the received audio data into text using speech recognition software (e.g., a common speech recognition API). Based on this converted text information, the server performs further analysis. This analysis uses natural language processing libraries (such as NLTK or spaCy). The purpose of this analysis is to extract important evaluation criteria that should be assessed in an interview, such as communication skills and problem-solving abilities.

[0274] Based on the analysis results, the server generates evaluation values ​​for each evaluation criterion. Machine learning models and generative AI models are used to generate these evaluation values. For example, a prompt such as "Evaluate this candidate's communication skills on a scale of 1 to 10" can be input into a generative AI model to perform appropriate scoring.

[0275] The device visually displays the generated evaluation scores and presents them to the user acting as the interviewer. This allows the user to eliminate individual bias and evaluate candidates from an objective perspective.

[0276] As a concrete example, in an interview process at a certain organization, the user collects audio data from multiple candidates, and a server scores them. The results are displayed on the device, for example, Candidate A might have a score of "Communication Skills: 8 / 10," and Candidate B might have a score of "Problem-Solving Ability: 7 / 10." Based on this information, the user can select the person best suited to the organization.

[0277] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0278] Step 1:

[0279] The user records the candidate's statements during the interview and saves the audio data in digital format. This audio data is then uploaded to a server via the internet. The input is the audio data, and the output is the status indicating that the data transfer to the server is complete.

[0280] Step 2:

[0281] The server converts the received audio data into text using speech recognition technology. In this process, the speech recognition API takes the audio data as input and outputs the resulting text. Checks are also performed to verify the accuracy of the converted text.

[0282] Step 3:

[0283] The server analyzes the character information obtained by voice recognition and extracts keywords and phrases related to the evaluation criteria. The natural language processing library analyzes the character information as input and identifies and outputs the necessary evaluation criteria. The specific operations performed in this step are text analysis and information extraction.

[0284] Step 4:

[0285] Based on the information extracted by the analysis, the server generates a prompt sentence and inputs it into the generation AI model. At this time, the machine learning model generates an evaluation score. In this process, the result of text analysis is used as input, and an evaluation value for each evaluation criterion is output. The specific operations are the formation of the prompt sentence and the input to the AI model.

[0286] Step 5:

[0287] The terminal visually displays the evaluation score and its details received from the server and presents them to the user. The output is the visualized result of the evaluation score. The specific operation is to display the evaluation score on the screen so that the user can make a quick judgment.

[0288] Step 6:

[0289] Based on the displayed evaluation score, the user makes a final adoption decision. In this process, the data displayed on the screen is used as input, and a decision based on the interview results is made as output. As a specific example, the user compares and considers the interview results and selects the optimal candidate.

[0290] (Application Example 1)

[0291] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0292] Traditional interview systems often rely on the subjective opinions of interviewers, leading to bias and inaccuracies. Furthermore, the difficulty in visually comparing candidate evaluations can hinder rational decision-making. Additionally, interview evaluation criteria may not align with the needs of the company or organization, posing challenges in selecting the most suitable candidates.

[0293] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0294] In this invention, the server includes means for receiving interview data, means for converting the interview data into text data using speech recognition technology, means for analyzing the text data and generating scores for each interview evaluation metric, and means for making the evaluation results of candidates visually comparable. This reduces bias in interview evaluations and enables a fair and rational interview process. Furthermore, by setting customized evaluation criteria according to the needs of companies and organizations, it becomes possible to select appropriate personnel.

[0295] "Interview data" refers to data that includes audio information exchanged between the candidate and the interviewer.

[0296] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.

[0297] "Text data" refers to data that contains string information converted using speech recognition technology.

[0298] "Analysis" is a method of extracting specific evaluation metrics from text data and performing evaluations based on those metrics.

[0299] "Evaluation metrics" are standards used to measure specific abilities and skills that are considered important during an interview.

[0300] A "score" represents an evaluation result that has been quantified based on evaluation indicators.

[0301] The "terminal" is an electronic device for displaying evaluation results.

[0302] "Comparable" means that characteristics and capabilities can be easily compared by arranging multiple evaluation results side by side.

[0303] "Organizational needs" refer to specific skill sets and appropriate requirements sought by specific companies or organizations.

[0304] "Customized evaluation criteria" are evaluation criteria adjusted according to the requirements of a specific organization.

[0305] The system for implementing this invention is configured to receive interview data and convert it into text data using speech recognition technology. The server uses the "SpeechRecognition" library as speech recognition technology to effectively convert speech information into text.

[0306] The server uses a natural language processing library such as "NLTK" or "spaCy" to analyze this text data. During the analysis process, important elements are extracted based on the interview evaluation criteria, and scores for these elements are generated. A generative AI model is used for the analysis, leveraging patterns learned through past interview data.

[0307] The terminal presents the generated scores to the user. At this time, by presenting them in a visually comparable format, such as a graph or list, the user can intuitively understand the evaluation. It also has a function to set customized evaluation criteria based on the needs of companies or organizations.

[0308] As a specific example, when the user opens the app on their smartphone and presses the recording button, the conversation with the candidate is automatically recorded and sent to the server. The results analyzed by the server are displayed as scores on the smartphone screen in real time.

[0309] An example of a prompt for a generative AI model is, "Design a prompt to identify the necessary evaluation metrics in an interview and generate a score based on the candidate's voice data." This allows users to create a fairer and more efficient interview process.

[0310] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0311] Step 1:

[0312] The user begins the interview and starts recording audio data using an application on the device. The device receives and records the interview audio data via the microphone. It takes the user's audio data as input and generates an audio file saved to local storage as output.

[0313] Step 2:

[0314] The terminal sends the recorded audio file to the server. The server receives this audio data and converts it into text data using speech recognition technology. Specifically, it analyzes the audio waveform using the "SpeechRecognition" library and generates the corresponding text. It takes an audio file as input and generates text data in string format as output.

[0315] Step 3:

[0316] The server analyzes the generated text data using a natural language processing library ("NLTK" or "spaCy"). Here, keywords and phrases related to the evaluation metrics are extracted. The generative AI model identifies features for score generation based on prompts derived from past data. The input is text data, and the output is data with identified evaluation metrics.

[0317] Step 4:

[0318] The server calculates a score based on the identified evaluation metrics. The calculation uses machine learning algorithms (e.g., scikit-learn or TensorFlow) to generate numerical scores for each metric. The input is the evaluation metric data, and the output is the numerical score.

[0319] Step 5:

[0320] The terminal receives scores sent from the server and visualizes them through a user interface. Graphs and lists are used to allow for visual comparison of scores among multiple candidates. The input is a numerical score, and the output is a visualized evaluation result.

[0321] Step 6:

[0322] Users compare and evaluate candidates based on the provided evaluation results and make hiring decisions. By comparing candidates, users can identify their relative strengths and weaknesses, enabling more rational decision-making. The input is the visualized evaluation results, and the output is the final hiring decision.

[0323] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0324] This invention provides a system that enables more precise talent evaluation by combining an emotion engine to recognize and incorporate the emotional state of candidates during job interviews into the evaluation process. The system includes processing to detect emotions in addition to analyzing voice data.

[0325] Data reception and conversion

[0326] The user records the conversation with the candidate as audio during the interview and sends the data to the system.

[0327] Recognition of voice and emotion

[0328] The server processes the received audio data using a speech recognition engine to convert it into text data. Furthermore, it uses an emotion engine to analyze the tone and tempo of the voice and recognize the candidate's emotional state.

[0329] Data analysis and evaluation

[0330] The server generates scores for each evaluation metric based on the content of the text data and the recognized sentiment information. Sentiment information is incorporated into the evaluation of communication skills and personal characteristics.

[0331] Presentation of results and feedback

[0332] The device displays emotional information, along with the generated evaluation score, as part of the analysis results on the screen and presents it to the user acting as the interviewer. This enables a comprehensive evaluation that includes emotional aspects.

[0333] Specific example

[0334] In a typical interview, a user collects audio data from five candidates. The server converts the audio data into text and uses an emotion engine to analyze the emotions each candidate expressed. For example, it evaluates the level of confidence and nervousness that candidate A displayed when answering questions, and reflects these factors influencing their communication skills in a score. The device then graphs this information, providing the interviewer with an intuitive understanding. The user can leverage this detailed feedback to make informed decisions about selecting the best candidates.

[0335] Thus, by integrating an emotion engine, the system of this invention grasps the latent characteristics of candidates that cannot be captured by conventional language analysis alone, and supports talent evaluation that is more suitable for companies.

[0336] The following describes the processing flow.

[0337] Step 1:

[0338] After the interview begins, the user uses a dedicated recording device to collect audio data of the conversation with the candidate. This data is transmitted to the system in real time via a secure connection.

[0339] Step 2:

[0340] The server inputs the audio data sent by the user into the speech recognition engine and converts it into text data sequentially. This conversion process includes pre-processing such as background noise reduction and speaker separation.

[0341] Step 3:

[0342] The server simultaneously inputs voice data into the emotion engine and identifies the candidate's emotional state by analyzing features such as tone, volume, and speed of voice. Emotional information is classified into categories such as joy, tension, and surprise.

[0343] Step 4:

[0344] The server integrates the generated text data and sentiment data, inputs it into the evaluation model, and calculates scores for each evaluation metric. Specifically, it scores metrics such as communication skills, cooperativeness, and adaptability.

[0345] Step 5:

[0346] The device displays the evaluation score and sentiment analysis results for each candidate as visualized data on the screen. The display uses graphs, charts, and numerical data, providing information in a visually easy-to-understand format.

[0347] Step 6:

[0348] Users deepen their thinking based on the information presented and make a comprehensive evaluation of candidates from a more objective perspective. This improves the transparency and rationality of the selection process, enabling the selection of the most suitable talent.

[0349] (Example 2)

[0350] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0351] In recruitment interviews, traditional evaluation methods struggle to adequately grasp candidates' nonverbal characteristics and emotions, resulting in difficulties in selecting appropriate personnel. Furthermore, reliance on the interviewer's subjective evaluation leads to a lack of consistency and objectivity.

[0352] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0353] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using speech analysis technology, and means for recognizing the emotional state from the text information and generating numerical values ​​for each evaluation criterion using the results. This makes it possible to reflect the candidate's emotional state during the interview in the evaluation, enabling a more objective and consistent personnel evaluation.

[0354] "Interview information" refers to audio data and related data obtained from conversations and interactions with candidates during job interviews.

[0355] "Speech analysis technology" is a technology that processes speech data and converts it into text information, and in particular, it uses a speech recognition engine for analysis.

[0356] "Textual information" refers to text data that represents the content of audio data, converted using speech analysis technology.

[0357] "Emotional state" refers to the emotional characteristics and reactions that a candidate exhibits during a conversation, and is analyzed from their voice and tone.

[0358] "Evaluation criteria" are standards or scales used to measure a candidate's abilities and aptitudes, and are set according to the organization's needs and objectives.

[0359] "Numerical values" refer to data that quantitatively represents a candidate's abilities and characteristics, calculated based on evaluation criteria.

[0360] A "decision-maker" refers to an individual or organization that has the authority to make decisions regarding personnel selection based on the results of job interviews.

[0361] This invention is a system for recognizing and evaluating the emotional state of candidates during job interviews, and is primarily composed of audio data processing. This system operates as follows:

[0362] Data collection and transmission

[0363] The user records the conversation with the candidate during the interview as audio data using a recording device. This data is then sent to a server.

[0364] Analysis of audio data

[0365] The server analyzes the received audio data. Specifically, it uses a speech recognition engine (e.g., a common cloud-based speech recognition API) as an audio analysis technology to convert the audio data into text information. It also uses an emotion engine (e.g., commercially available emotion recognition software) to analyze the tone and speed of the voice and recognize the candidate's emotional state.

[0366] Generation and display of evaluation scores

[0367] The server uses a generative AI model to generate numerical values ​​for each evaluation criterion, based on the converted text information and recognized emotional states.

[0368] The terminal visualizes numerical data and sentiment information generated on the server, displaying it intuitively through graphs and charts. This information is provided to the user, who is the decision-maker, to assist in the overall evaluation of candidates.

[0369] As a concrete example, suppose a user collects audio data from five candidates during an interview session. The server processes this audio, performing speech recognition and sentiment analysis. For instance, it evaluates candidate A's confidence and nervousness, reflecting these in a numerical value corresponding to their communication skills. The terminal then intuitively displays these evaluation results, allowing the user to easily compare the characteristics of each candidate. This detailed information is extremely useful in making hiring decisions.

[0370] An example of a prompt might be: "Please describe a method for analyzing candidate voice data during interviews to quantify the influence of emotions on the evaluation of communication skills."

[0371] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0372] Step 1:

[0373] The user records the conversation with the candidate during the job interview as audio data using a recording device. The recorded audio data is sent to the server in its original format. In this step, the input is the candidate's audio data, and the output is the transmission of the audio data to the server. Specifically, the user presses the record button and continues recording until the audio data is of sufficient length.

[0374] Step 2:

[0375] The server converts received audio data into text data through a speech recognition engine. A cloud-based speech recognition API is used for this process. The input is audio data, and the output is the corresponding text data. Specifically, the server batch processes the audio data for analysis and generates recognition results in real time.

[0376] Step 3:

[0377] The server uses an emotion engine to analyze the converted text data and the original audio data to recognize the candidate's emotional state. In this step, text and audio data are taken as input, and information about the emotional state is obtained as output. Specifically, the server analyzes the tone and speed of the voice and extracts emotional characteristics.

[0378] Step 4:

[0379] The server utilizes a generative AI model to quantify the obtained text data and sentiment information based on evaluation criteria. The input consists of text data and sentiment information, and the output is a numerical value generated according to the evaluation criteria. The server processes the data using algorithms and generates numerical values ​​for each evaluation metric.

[0380] Step 5:

[0381] The terminal visualizes numerical data and emotional states transmitted from the server, displaying them on the screen as graphs and charts. Inputs are numerical data and emotional state information, while output is a visual evaluation report. Specifically, the terminal updates data in real time and displays it in a dashboard format for intuitive user understanding.

[0382] Step 6:

[0383] The user makes an overall evaluation of the candidates based on the data displayed on the device. The input is the displayed evaluation results, and the output is the final hiring decision. Specifically, the user reviews the evaluation results on the device and compares the candidates to reach a conclusion.

[0384] (Application Example 2)

[0385] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0386] In collaborative work between humans and robots in factories, it is not easy to grasp the emotional state of workers in real time and for robots to take appropriate actions and provide feedback in response to those emotions. Therefore, there is a need for systems that can improve work efficiency and safety.

[0387] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0388] In this invention, the server includes means for receiving voice data, means for converting it into text data using speech recognition technology, and means for analyzing voice tone and tempo to recognize emotional state. This enables real-time understanding of the emotional state of workers and allows for appropriate feedback and decision-making regarding actions.

[0389] "Interview data" refers to all information collected during an interview, including audio and video.

[0390] "Speech recognition technology" is a technology that converts speech data into text data.

[0391] "Text data" refers to data that has been converted using speech recognition technology and is represented as textual information.

[0392] "Tone" refers to elements that describe characteristics such as pitch and volume of sound.

[0393] "Tempo" refers to the speed and rhythm of speech.

[0394] "Emotional state" refers to the psychological state analyzed from audio data, and includes emotions such as joy, anger, sadness, and happiness.

[0395] "Evaluation indicators" refer to the criteria and perspectives used to evaluate individuals during interviews or work assignments.

[0396] A "score" is a rating or numerical value calculated based on evaluation metrics.

[0397] "User" is a general term referring to anyone who uses a system or receives a service.

[0398] This system is designed to facilitate collaborative work with factory workers. The server receives voice data from workers in real time and converts it to text using speech recognition technology. Here, speech recognition software such as Google Cloud Speech-to-Text API is used. Subsequently, an emotion engine, such as Microsoft Azure Emotion API, is used to analyze the tone and tempo of the voice and recognize the worker's emotional state.

[0399] Based on the information obtained through this emotion analysis, the server generates feedback and instructions tailored to the work situation and displays them on the terminal. The terminal refers to a smartphone or head-mounted display, providing information in a way that the worker can intuitively understand. For example, if the server determines that the worker is feeling anxious, the terminal will display a message such as, "Take your time, there's no need to rush."

[0400] In this system, the user refers to a worker, and by receiving feedback, improvements in work efficiency and safety can be expected. For example, if an incorrect operation is detected while handling specialized tools, the robot will suggest appropriate corrective steps.

[0401] An example of a prompt for a generative AI model is: "Analyze the worker's current emotions from this audio data and generate instructions for the robot to provide appropriate feedback."

[0402] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0403] Step 1:

[0404] Users record spoken audio in real time at the work site using a smartphone or head-mounted display and send the audio data to a server. The input is audio data, and the output is the raw audio data sent to the server. This process allows workers' voices to be stored as data.

[0405] Step 2:

[0406] The server converts the received audio data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose. The input is raw audio data, and the output is text data extracted from the audio. This process converts the audio content into a format that can be parsed as text information.

[0407] Step 3:

[0408] The server analyzes the converted text data and the tone and tempo of the speech, and uses the Microsoft Azure Emotion API to recognize the emotional state. The input is text data and speech feature information, and the output is the recognized emotional state. This process makes it possible to determine the user's psychological state.

[0409] Step 4:

[0410] The server determines appropriate feedback or action plans based on the worker's emotional state and work progress. Inputs are emotional state and individual work status data, while outputs are feedback content and action plans. This allows for appropriate advice and warnings to be given to the worker.

[0411] Step 5:

[0412] The device displays the determined feedback or action plan to the user. The input is the feedback content, and the output is visual or audio information presented to the user. This operation allows the user to intuitively receive guidance for improvement or correction while working.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0416] [Third Embodiment]

[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0429] This invention provides a system for reducing bias and inequality in evaluations during job interviews. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation metric. To achieve this objective, the system operates as follows:

[0430] Data reception and conversion

[0431] The user provides the system with audio data collected during the interview.

[0432] The server receives this audio data and converts it into text data using speech recognition technology.

[0433] Data analysis and evaluation

[0434] The server analyzes the generated text data and extracts key evaluation metrics that should be assessed during the interview.

[0435] The server generates interview scores based on evaluation metrics, including communication skills and problem-solving abilities.

[0436] Presentation of results and feedback

[0437] The device displays the generated score and its evaluation criteria on the screen and presents them to the user acting as the interviewer.

[0438] Users can use this information to make their final hiring decisions.

[0439] Specific example

[0440] For example, in a company interview, a user collects audio data from five candidates. The server converts this data into text in real time and analyzes it for each candidate. The analysis results display scores for each evaluation metric, such as "Communication Skills: 8 / 10" for Candidate A and "Problem-Solving Ability: 7 / 10" for Candidate B. The user can use this information to help select the most suitable candidate for the company.

[0441] This system can improve the accuracy of its analysis by having the model learn evaluation metrics using past interview data. The system also allows for customized evaluation criteria tailored to each company's needs, thus supporting the selection of personnel who are a better fit for the company.

[0442] The following describes the processing flow.

[0443] Step 1:

[0444] The user initiates the interview and collects audio data to record the interaction with the candidate. The audio data is transmitted to the system via a pre-configured, dedicated interface.

[0445] Step 2:

[0446] The server processes the audio data received from the user in real time using a speech recognition engine and converts it into text data. This text conversion utilizes natural language processing technology to generate accurate text.

[0447] Step 3:

[0448] The server analyzes the converted text data and extracts key evaluation metrics that should be assessed during the interview. These metrics are based on criteria such as communication skills and technical abilities.

[0449] Step 4:

[0450] The server generates quantified scores for each evaluation metric. This scoring system utilizes insights and criteria accumulated from past interview data to provide accurate assessments.

[0451] Step 5:

[0452] The terminal displays the generated score and the underlying analytical data on its screen, presenting it to the user acting as the interviewer. The displayed information is summarized visually in an easy-to-understand format using graphs and numerical data.

[0453] Step 6:

[0454] Users make their final hiring decisions based on the presented scores and related evaluation information. This allows for the efficient selection of candidates who best meet the company's needs.

[0455] (Example 1)

[0456] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0457] During the interview process, subjective evaluations are often biased, making it difficult to select appropriate candidates. Furthermore, the lack of standardized evaluation criteria means that evaluations differ from organization to organization, posing a challenge. There is also a need to improve the accuracy of data analysis to enhance the objectivity of evaluation results.

[0458] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0459] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using a speech recognition method, means for analyzing the text information and generating evaluation values ​​for each interview evaluation criterion, and means for generating prompt sentences and inputting them to a generation AI model. This reduces evaluation bias, enables highly objective evaluations based on unified criteria, and allows for the selection of personnel suitable for each organization.

[0460] "Interview information" refers to digital information, including audio, text, and related data, collected during the interview process.

[0461] "Speech recognition methods" refer to the technology that converts speech data into text data, and the process of making human speech understandable to machines.

[0462] "Textual information" refers to information in text format obtained as a result of conversion using speech recognition methods.

[0463] "Analysis" refers to the process of analyzing textual information to extract meaningful evaluation indicators, and is a process for making specific evaluations based on that information.

[0464] "Evaluation criteria" refer to the scales or indicators used to assess a candidate's abilities and suitability during an interview.

[0465] An "evaluation value" is a numerical or quantitative result obtained through analysis, and is a score calculated based on evaluation criteria.

[0466] A "generative AI model" is an artificial intelligence model that uses machine learning or deep learning algorithms to train itself to solve specific problems.

[0467] A "prompt" is a text-based instruction used to tell a generative AI model to input information or perform an action.

[0468] This invention relates to a system that reduces bias and prejudice in human evaluations during interviews, enabling objective evaluation. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation criterion.

[0469] The user collects the interviewee's statements during the interview using an audio recording device. This data is uploaded to a server in digital format. The audio data collected by the user is transmitted to the server via the internet.

[0470] The server converts the received audio data into text using speech recognition software (e.g., a common speech recognition API). Based on this converted text information, the server performs further analysis. This analysis uses natural language processing libraries (such as NLTK or spaCy). The purpose of this analysis is to extract important evaluation criteria that should be assessed in an interview, such as communication skills and problem-solving abilities.

[0471] Based on the analysis results, the server generates evaluation values ​​for each evaluation criterion. Machine learning models and generative AI models are used to generate these evaluation values. For example, a prompt such as "Evaluate this candidate's communication skills on a scale of 1 to 10" can be input into a generative AI model to perform appropriate scoring.

[0472] The device visually displays the generated evaluation scores and presents them to the user acting as the interviewer. This allows the user to eliminate individual bias and evaluate candidates from an objective perspective.

[0473] As a concrete example, in an interview process at a certain organization, the user collects audio data from multiple candidates, and a server scores them. The results are displayed on the device, for example, Candidate A might have a score of "Communication Skills: 8 / 10," and Candidate B might have a score of "Problem-Solving Ability: 7 / 10." Based on this information, the user can select the person best suited to the organization.

[0474] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0475] Step 1:

[0476] The user records the candidate's statements during the interview and saves the audio data in digital format. This audio data is then uploaded to a server via the internet. The input is the audio data, and the output is the status indicating that the data transfer to the server is complete.

[0477] Step 2:

[0478] The server converts the received audio data into text using speech recognition technology. In this process, the speech recognition API takes the audio data as input and outputs the resulting text. Checks are also performed to verify the accuracy of the converted text.

[0479] Step 3:

[0480] The server analyzes the text information obtained through speech recognition and extracts keywords and phrases related to the evaluation criteria. A natural language processing library analyzes the text information as input, identifies the necessary evaluation criteria, and outputs them. The specific actions performed in this step are text analysis and information extraction.

[0481] Step 4:

[0482] The server generates prompt sentences based on the information extracted through analysis and inputs them into the AI ​​model. At this point, the machine learning model generates an evaluation score. This process takes the results of text analysis as input and outputs evaluation values ​​for each evaluation criterion. Specifically, the process involves the formation of prompt sentences and their input into the AI ​​model.

[0483] Step 5:

[0484] The terminal visually displays the evaluation score and its details received from the server and presents them to the user. The output is a visualized result of the evaluation score. The specific function is to enable the user to make a quick decision by displaying the evaluation score on the screen.

[0485] Step 6:

[0486] The user makes a final hiring decision based on the displayed evaluation scores. In this process, the data displayed on the screen is used as input, and the output is a decision based on the interview results. Specifically, the user compares and reviews the interview results and selects the best candidate.

[0487] (Application Example 1)

[0488] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0489] Traditional interview systems often rely on the subjective opinions of interviewers, leading to bias and inaccuracies. Furthermore, the difficulty in visually comparing candidate evaluations can hinder rational decision-making. Additionally, interview evaluation criteria may not align with the needs of the company or organization, posing challenges in selecting the most suitable candidates.

[0490] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0491] In this invention, the server includes means for receiving interview data, means for converting the interview data into text data using speech recognition technology, means for analyzing the text data and generating scores for each interview evaluation metric, and means for making the evaluation results of candidates visually comparable. This reduces bias in interview evaluations and enables a fair and rational interview process. Furthermore, by setting customized evaluation criteria according to the needs of companies and organizations, it becomes possible to select appropriate personnel.

[0492] "Interview data" refers to data that includes audio information exchanged between the candidate and the interviewer.

[0493] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.

[0494] "Text data" refers to data that contains string information converted using speech recognition technology.

[0495] "Analysis" is a method of extracting specific evaluation metrics from text data and performing evaluations based on those metrics.

[0496] "Evaluation metrics" are standards used to measure specific abilities and skills that are considered important during an interview.

[0497] A "score" represents an evaluation result that has been quantified based on evaluation indicators.

[0498] A "terminal" is an electronic device used to display evaluation results.

[0499] "Comparable" means that by comparing multiple evaluation results side-by-side, characteristics and abilities can be easily compared.

[0500] "Organizational needs" refer to the specific skill sets and aptitude requirements that a particular company or organization seeks.

[0501] "Customized performance metrics" are evaluation criteria that have been tailored to the specific requirements of a particular organization.

[0502] The system implementing this invention is configured to receive interview data and convert it into text data using speech recognition technology. The server uses the "SpeechRecognition" library as the speech recognition technology to effectively convert the audio information into text.

[0503] The server uses natural language processing libraries such as "NLTK" or "spaCy" to analyze this text data. During the analysis, it extracts important elements based on interview evaluation metrics and generates scores for them. A generative AI model is used for the analysis, leveraging patterns learned through past interview data.

[0504] The device presents the generated score to the user. At this time, it displays the score in a visually comparable format, such as a graph or list, to allow the user to intuitively understand the evaluation. It also has a function to set customized evaluation metrics based on the needs of the company or organization.

[0505] As a concrete example, when a user opens the app on their smartphone and presses the record button, the conversation with the candidate is automatically recorded and sent to the server. The results, analyzed by the server, are displayed as a score on the smartphone screen in real time.

[0506] An example of a prompt for a generative AI model is, "Design a prompt to identify the necessary evaluation metrics in an interview and generate a score based on the candidate's voice data." This allows users to create a fairer and more efficient interview process.

[0507] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0508] Step 1:

[0509] The user begins the interview and starts recording audio data using an application on the device. The device receives and records the interview audio data via the microphone. It takes the user's audio data as input and generates an audio file saved to local storage as output.

[0510] Step 2:

[0511] The terminal sends the recorded audio file to the server. The server receives this audio data and converts it into text data using speech recognition technology. Specifically, it analyzes the audio waveform using the "SpeechRecognition" library and generates the corresponding text. It takes an audio file as input and generates text data in string format as output.

[0512] Step 3:

[0513] The server analyzes the generated text data using a natural language processing library ("NLTK" or "spaCy"). Here, keywords and phrases related to the evaluation metrics are extracted. The generative AI model identifies features for score generation based on prompts derived from past data. The input is text data, and the output is data with identified evaluation metrics.

[0514] Step 4:

[0515] The server calculates a score based on the identified evaluation metrics. The calculation uses machine learning algorithms (e.g., scikit-learn or TensorFlow) to generate numerical scores for each metric. The input is the evaluation metric data, and the output is the numerical score.

[0516] Step 5:

[0517] The terminal receives scores sent from the server and visualizes them through a user interface. Graphs and lists are used to allow for visual comparison of scores among multiple candidates. The input is a numerical score, and the output is a visualized evaluation result.

[0518] Step 6:

[0519] Users compare and evaluate candidates based on the provided evaluation results and make hiring decisions. By comparing candidates, users can identify their relative strengths and weaknesses, enabling more rational decision-making. The input is the visualized evaluation results, and the output is the final hiring decision.

[0520] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0521] This invention provides a system that enables more precise talent evaluation by combining an emotion engine to recognize and incorporate the emotional state of candidates during job interviews into the evaluation process. The system includes processing to detect emotions in addition to analyzing voice data.

[0522] Data reception and conversion

[0523] The user records the conversation with the candidate as audio during the interview and sends the data to the system.

[0524] Recognition of voice and emotion

[0525] The server processes the received audio data using a speech recognition engine to convert it into text data. Furthermore, it uses an emotion engine to analyze the tone and tempo of the voice and recognize the candidate's emotional state.

[0526] Data analysis and evaluation

[0527] The server generates scores for each evaluation metric based on the content of the text data and the recognized sentiment information. Sentiment information is incorporated into the evaluation of communication skills and personal characteristics.

[0528] Presentation of results and feedback

[0529] The device displays emotional information, along with the generated evaluation score, as part of the analysis results on the screen and presents it to the user acting as the interviewer. This enables a comprehensive evaluation that includes emotional aspects.

[0530] Specific example

[0531] In a typical interview, a user collects audio data from five candidates. The server converts the audio data into text and uses an emotion engine to analyze the emotions each candidate expressed. For example, it evaluates the level of confidence and nervousness that candidate A displayed when answering questions, and reflects these factors influencing their communication skills in a score. The device then graphs this information, providing the interviewer with an intuitive understanding. The user can leverage this detailed feedback to make informed decisions about selecting the best candidates.

[0532] Thus, by integrating an emotion engine, the system of this invention grasps the latent characteristics of candidates that cannot be captured by conventional language analysis alone, and supports talent evaluation that is more suitable for companies.

[0533] The following describes the processing flow.

[0534] Step 1:

[0535] After the interview begins, the user uses a dedicated recording device to collect audio data of the conversation with the candidate. This data is transmitted to the system in real time via a secure connection.

[0536] Step 2:

[0537] The server inputs the audio data sent by the user into the speech recognition engine and converts it into text data sequentially. This conversion process includes pre-processing such as background noise reduction and speaker separation.

[0538] Step 3:

[0539] The server simultaneously inputs voice data into the emotion engine and identifies the candidate's emotional state by analyzing features such as tone, volume, and speed of voice. Emotional information is classified into categories such as joy, tension, and surprise.

[0540] Step 4:

[0541] The server integrates the generated text data and sentiment data, inputs it into the evaluation model, and calculates scores for each evaluation metric. Specifically, it scores metrics such as communication skills, cooperativeness, and adaptability.

[0542] Step 5:

[0543] The device displays the evaluation score and sentiment analysis results for each candidate as visualized data on the screen. The display uses graphs, charts, and numerical data, providing information in a visually easy-to-understand format.

[0544] Step 6:

[0545] Users deepen their thinking based on the information presented and make a comprehensive evaluation of candidates from a more objective perspective. This improves the transparency and rationality of the selection process, enabling the selection of the most suitable talent.

[0546] (Example 2)

[0547] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0548] In recruitment interviews, traditional evaluation methods struggle to adequately grasp candidates' nonverbal characteristics and emotions, resulting in difficulties in selecting appropriate personnel. Furthermore, reliance on the interviewer's subjective evaluation leads to a lack of consistency and objectivity.

[0549] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0550] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using speech analysis technology, and means for recognizing the emotional state from the text information and generating numerical values ​​for each evaluation criterion using the results. This makes it possible to reflect the candidate's emotional state during the interview in the evaluation, enabling a more objective and consistent personnel evaluation.

[0551] "Interview information" refers to audio data and related data obtained from conversations and interactions with candidates during job interviews.

[0552] "Speech analysis technology" is a technology that processes speech data and converts it into text information, and in particular, it uses a speech recognition engine for analysis.

[0553] "Textual information" refers to text data that represents the content of audio data, converted using speech analysis technology.

[0554] "Emotional state" refers to the emotional characteristics and reactions that a candidate exhibits during a conversation, and is analyzed from their voice and tone.

[0555] "Evaluation criteria" are standards or scales used to measure a candidate's abilities and aptitudes, and are set according to the organization's needs and objectives.

[0556] "Numerical values" refer to data that quantitatively represents a candidate's abilities and characteristics, calculated based on evaluation criteria.

[0557] A "decision-maker" refers to an individual or organization that has the authority to make decisions regarding personnel selection based on the results of job interviews.

[0558] This invention is a system for recognizing and evaluating the emotional state of candidates during job interviews, and is primarily composed of audio data processing. This system operates as follows:

[0559] Data collection and transmission

[0560] The user records the conversation with the candidate during the interview as audio data using a recording device. This data is then sent to a server.

[0561] Analysis of audio data

[0562] The server analyzes the received audio data. Specifically, it uses a speech recognition engine (e.g., a common cloud-based speech recognition API) as an audio analysis technology to convert the audio data into text information. It also uses an emotion engine (e.g., commercially available emotion recognition software) to analyze the tone and speed of the voice and recognize the candidate's emotional state.

[0563] Generation and display of evaluation scores

[0564] The server uses a generative AI model to generate numerical values ​​for each evaluation criterion, based on the converted text information and recognized emotional states.

[0565] The terminal visualizes numerical data and sentiment information generated on the server, displaying it intuitively through graphs and charts. This information is provided to the user, who is the decision-maker, to assist in the overall evaluation of candidates.

[0566] As a concrete example, suppose a user collects audio data from five candidates during an interview session. The server processes this audio, performing speech recognition and sentiment analysis. For instance, it evaluates candidate A's confidence and nervousness, reflecting these in a numerical value corresponding to their communication skills. The terminal then intuitively displays these evaluation results, allowing the user to easily compare the characteristics of each candidate. This detailed information is extremely useful in making hiring decisions.

[0567] An example of a prompt might be: "Please describe a method for analyzing candidate voice data during interviews to quantify the influence of emotions on the evaluation of communication skills."

[0568] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0569] Step 1:

[0570] The user records the conversation with the candidate during the job interview as audio data using a recording device. The recorded audio data is sent to the server in its original format. In this step, the input is the candidate's audio data, and the output is the transmission of the audio data to the server. Specifically, the user presses the record button and continues recording until the audio data is of sufficient length.

[0571] Step 2:

[0572] The server converts received audio data into text data through a speech recognition engine. A cloud-based speech recognition API is used for this process. The input is audio data, and the output is the corresponding text data. Specifically, the server batch processes the audio data for analysis and generates recognition results in real time.

[0573] Step 3:

[0574] The server uses an emotion engine to analyze the converted text data and the original audio data to recognize the candidate's emotional state. In this step, text and audio data are taken as input, and information about the emotional state is obtained as output. Specifically, the server analyzes the tone and speed of the voice and extracts emotional characteristics.

[0575] Step 4:

[0576] The server utilizes a generative AI model to quantify the obtained text data and sentiment information based on evaluation criteria. The input consists of text data and sentiment information, and the output is a numerical value generated according to the evaluation criteria. The server processes the data using algorithms and generates numerical values ​​for each evaluation metric.

[0577] Step 5:

[0578] The terminal visualizes numerical data and emotional states transmitted from the server, displaying them on the screen as graphs and charts. Inputs are numerical data and emotional state information, while output is a visual evaluation report. Specifically, the terminal updates data in real time and displays it in a dashboard format for intuitive user understanding.

[0579] Step 6:

[0580] The user makes an overall evaluation of the candidates based on the data displayed on the device. The input is the displayed evaluation results, and the output is the final hiring decision. Specifically, the user reviews the evaluation results on the device and compares the candidates to reach a conclusion.

[0581] (Application Example 2)

[0582] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0583] In collaborative work between humans and robots in factories, it is not easy to grasp the emotional state of workers in real time and for robots to take appropriate actions and provide feedback in response to those emotions. Therefore, there is a need for systems that can improve work efficiency and safety.

[0584] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0585] In this invention, the server includes means for receiving voice data, means for converting it into text data using speech recognition technology, and means for analyzing voice tone and tempo to recognize emotional state. This enables real-time understanding of the emotional state of workers and allows for appropriate feedback and decision-making regarding actions.

[0586] "Interview data" refers to all information collected during an interview, including audio and video.

[0587] "Speech recognition technology" is a technology that converts speech data into text data.

[0588] "Text data" refers to data that has been converted using speech recognition technology and is represented as textual information.

[0589] "Tone" refers to elements that describe characteristics such as pitch and volume of sound.

[0590] "Tempo" refers to the speed and rhythm of speech.

[0591] "Emotional state" refers to the psychological state analyzed from audio data, and includes emotions such as joy, anger, sadness, and happiness.

[0592] "Evaluation indicators" refer to the criteria and perspectives used to evaluate individuals during interviews or work assignments.

[0593] A "score" is a rating or numerical value calculated based on evaluation metrics.

[0594] "User" is a general term referring to anyone who uses a system or receives a service.

[0595] This system is designed to facilitate collaborative work with factory workers. The server receives voice data from workers in real time and converts it to text using speech recognition technology. Here, speech recognition software such as Google Cloud Speech-to-Text API is used. Subsequently, an emotion engine, such as Microsoft Azure Emotion API, is used to analyze the tone and tempo of the voice and recognize the worker's emotional state.

[0596] Based on the information obtained through this emotion analysis, the server generates feedback and instructions tailored to the work situation and displays them on the terminal. The terminal refers to a smartphone or head-mounted display, providing information in a way that the worker can intuitively understand. For example, if the server determines that the worker is feeling anxious, the terminal will display a message such as, "Take your time, there's no need to rush."

[0597] In this system, the user refers to a worker, and by receiving feedback, improvements in work efficiency and safety can be expected. For example, if an incorrect operation is detected while handling specialized tools, the robot will suggest appropriate corrective steps.

[0598] An example of a prompt for a generative AI model is: "Analyze the worker's current emotions from this audio data and generate instructions for the robot to provide appropriate feedback."

[0599] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0600] Step 1:

[0601] Users record spoken audio in real time at the work site using a smartphone or head-mounted display and send the audio data to a server. The input is audio data, and the output is the raw audio data sent to the server. This process allows workers' voices to be stored as data.

[0602] Step 2:

[0603] The server converts the received audio data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose. The input is raw audio data, and the output is text data extracted from the audio. This process converts the audio content into a format that can be parsed as text information.

[0604] Step 3:

[0605] The server analyzes the converted text data and the tone and tempo of the speech, and uses the Microsoft Azure Emotion API to recognize the emotional state. The input is text data and speech feature information, and the output is the recognized emotional state. This process makes it possible to determine the user's psychological state.

[0606] Step 4:

[0607] The server determines appropriate feedback or action plans based on the worker's emotional state and work progress. Inputs are emotional state and individual work status data, while outputs are feedback content and action plans. This allows for appropriate advice and warnings to be given to the worker.

[0608] Step 5:

[0609] The device displays the determined feedback or action plan to the user. The input is the feedback content, and the output is visual or audio information presented to the user. This operation allows the user to intuitively receive guidance for improvement or correction while working.

[0610] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0611] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0612] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0613] [Fourth Embodiment]

[0614] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0615] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0616] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0617] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0618] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0619] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0620] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0621] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0622] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0623] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0624] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0625] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0626] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0627] This invention provides a system for reducing bias and inequality in evaluations during job interviews. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation metric. To achieve this objective, the system operates as follows:

[0628] Data reception and conversion

[0629] The user provides the system with audio data collected during the interview.

[0630] The server receives this audio data and converts it into text data using speech recognition technology.

[0631] Data analysis and evaluation

[0632] The server analyzes the generated text data and extracts key evaluation metrics that should be assessed during the interview.

[0633] The server generates interview scores based on evaluation metrics, including communication skills and problem-solving abilities.

[0634] Presentation of results and feedback

[0635] The device displays the generated score and its evaluation criteria on the screen and presents them to the user acting as the interviewer.

[0636] Users can use this information to make their final hiring decisions.

[0637] Specific example

[0638] For example, in a company interview, a user collects audio data from five candidates. The server converts this data into text in real time and analyzes it for each candidate. The analysis results display scores for each evaluation metric, such as "Communication Skills: 8 / 10" for Candidate A and "Problem-Solving Ability: 7 / 10" for Candidate B. The user can use this information to help select the most suitable candidate for the company.

[0639] This system can improve the accuracy of its analysis by having the model learn evaluation metrics using past interview data. The system also allows for customized evaluation criteria tailored to each company's needs, thus supporting the selection of personnel who are a better fit for the company.

[0640] The following describes the processing flow.

[0641] Step 1:

[0642] The user initiates the interview and collects audio data to record the interaction with the candidate. The audio data is transmitted to the system via a pre-configured, dedicated interface.

[0643] Step 2:

[0644] The server processes the audio data received from the user in real time using a speech recognition engine and converts it into text data. This text conversion utilizes natural language processing technology to generate accurate text.

[0645] Step 3:

[0646] The server analyzes the converted text data and extracts key evaluation metrics that should be assessed during the interview. These metrics are based on criteria such as communication skills and technical abilities.

[0647] Step 4:

[0648] The server generates quantified scores for each evaluation metric. This scoring system utilizes insights and criteria accumulated from past interview data to provide accurate assessments.

[0649] Step 5:

[0650] The terminal displays the generated score and the underlying analytical data on its screen, presenting it to the user acting as the interviewer. The displayed information is summarized visually in an easy-to-understand format using graphs and numerical data.

[0651] Step 6:

[0652] Users make their final hiring decisions based on the presented scores and related evaluation information. This allows for the efficient selection of candidates who best meet the company's needs.

[0653] (Example 1)

[0654] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0655] During the interview process, subjective evaluations are often biased, making it difficult to select appropriate candidates. Furthermore, the lack of standardized evaluation criteria means that evaluations differ from organization to organization, posing a challenge. There is also a need to improve the accuracy of data analysis to enhance the objectivity of evaluation results.

[0656] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0657] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using a speech recognition method, means for analyzing the text information and generating evaluation values ​​for each interview evaluation criterion, and means for generating prompt sentences and inputting them to a generation AI model. This reduces evaluation bias, enables highly objective evaluations based on unified criteria, and allows for the selection of personnel suitable for each organization.

[0658] "Interview information" refers to digital information, including audio, text, and related data, collected during the interview process.

[0659] "Speech recognition methods" refer to the technology that converts speech data into text data, and the process of making human speech understandable to machines.

[0660] "Textual information" refers to information in text format obtained as a result of conversion using speech recognition methods.

[0661] "Analysis" refers to the process of analyzing textual information to extract meaningful evaluation indicators, and is a process for making specific evaluations based on that information.

[0662] "Evaluation criteria" refer to the scales or indicators used to assess a candidate's abilities and suitability during an interview.

[0663] An "evaluation value" is a numerical or quantitative result obtained through analysis, and is a score calculated based on evaluation criteria.

[0664] A "generative AI model" is an artificial intelligence model that uses machine learning or deep learning algorithms to train itself to solve specific problems.

[0665] A "prompt" is a text-based instruction used to tell a generative AI model to input information or perform an action.

[0666] This invention relates to a system that reduces bias and prejudice in human evaluations during interviews, enabling objective evaluation. The system includes a process of collecting and analyzing audio data and generating scores for each evaluation criterion.

[0667] The user collects the interviewee's statements during the interview using an audio recording device. This data is uploaded to a server in digital format. The audio data collected by the user is transmitted to the server via the internet.

[0668] The server converts the received audio data into text using speech recognition software (e.g., a common speech recognition API). Based on this converted text information, the server performs further analysis. This analysis uses natural language processing libraries (such as NLTK or spaCy). The purpose of this analysis is to extract important evaluation criteria that should be assessed in an interview, such as communication skills and problem-solving abilities.

[0669] Based on the analysis results, the server generates evaluation values ​​for each evaluation criterion. Machine learning models and generative AI models are used to generate these evaluation values. For example, a prompt such as "Evaluate this candidate's communication skills on a scale of 1 to 10" can be input into a generative AI model to perform appropriate scoring.

[0670] The device visually displays the generated evaluation scores and presents them to the user acting as the interviewer. This allows the user to eliminate individual bias and evaluate candidates from an objective perspective.

[0671] As a concrete example, in an interview process at a certain organization, the user collects audio data from multiple candidates, and a server scores them. The results are displayed on the device, for example, Candidate A might have a score of "Communication Skills: 8 / 10," and Candidate B might have a score of "Problem-Solving Ability: 7 / 10." Based on this information, the user can select the person best suited to the organization.

[0672] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0673] Step 1:

[0674] The user records the candidate's statements during the interview and saves the audio data in digital format. This audio data is then uploaded to a server via the internet. The input is the audio data, and the output is the status indicating that the data transfer to the server is complete.

[0675] Step 2:

[0676] The server converts the received audio data into text using speech recognition technology. In this process, the speech recognition API takes the audio data as input and outputs the resulting text. Checks are also performed to verify the accuracy of the converted text.

[0677] Step 3:

[0678] The server analyzes the text information obtained through speech recognition and extracts keywords and phrases related to the evaluation criteria. A natural language processing library analyzes the text information as input, identifies the necessary evaluation criteria, and outputs them. The specific actions performed in this step are text analysis and information extraction.

[0679] Step 4:

[0680] The server generates prompt sentences based on the information extracted through analysis and inputs them into the AI ​​model. At this point, the machine learning model generates an evaluation score. This process takes the results of text analysis as input and outputs evaluation values ​​for each evaluation criterion. Specifically, the process involves the formation of prompt sentences and their input into the AI ​​model.

[0681] Step 5:

[0682] The terminal visually displays the evaluation score and its details received from the server and presents them to the user. The output is a visualized result of the evaluation score. The specific function is to enable the user to make a quick decision by displaying the evaluation score on the screen.

[0683] Step 6:

[0684] The user makes a final hiring decision based on the displayed evaluation scores. In this process, the data displayed on the screen is used as input, and the output is a decision based on the interview results. Specifically, the user compares and reviews the interview results and selects the best candidate.

[0685] (Application Example 1)

[0686] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0687] Traditional interview systems often rely on the subjective opinions of interviewers, leading to bias and inaccuracies. Furthermore, the difficulty in visually comparing candidate evaluations can hinder rational decision-making. Additionally, interview evaluation criteria may not align with the needs of the company or organization, posing challenges in selecting the most suitable candidates.

[0688] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0689] In this invention, the server includes means for receiving interview data, means for converting the interview data into text data using speech recognition technology, means for analyzing the text data and generating scores for each interview evaluation metric, and means for making the evaluation results of candidates visually comparable. This reduces bias in interview evaluations and enables a fair and rational interview process. Furthermore, by setting customized evaluation criteria according to the needs of companies and organizations, it becomes possible to select appropriate personnel.

[0690] "Interview data" refers to data that includes audio information exchanged between the candidate and the interviewer.

[0691] "Speech recognition technology" is a technology that analyzes speech data and converts it into corresponding text data.

[0692] "Text data" refers to data that contains string information converted using speech recognition technology.

[0693] "Analysis" is a method of extracting specific evaluation metrics from text data and performing evaluations based on those metrics.

[0694] "Evaluation metrics" are standards used to measure specific abilities and skills that are considered important during an interview.

[0695] A "score" represents an evaluation result that has been quantified based on evaluation indicators.

[0696] A "terminal" is an electronic device used to display evaluation results.

[0697] "Comparable" means that by comparing multiple evaluation results side-by-side, characteristics and abilities can be easily compared.

[0698] "Organizational needs" refer to the specific skill sets and aptitude requirements that a particular company or organization seeks.

[0699] "Customized performance metrics" are evaluation criteria that have been tailored to the specific requirements of a particular organization.

[0700] The system implementing this invention is configured to receive interview data and convert it into text data using speech recognition technology. The server uses the "SpeechRecognition" library as the speech recognition technology to effectively convert the audio information into text.

[0701] The server uses natural language processing libraries such as "NLTK" or "spaCy" to analyze this text data. During the analysis, it extracts important elements based on interview evaluation metrics and generates scores for them. A generative AI model is used for the analysis, leveraging patterns learned through past interview data.

[0702] The device presents the generated score to the user. At this time, it displays the score in a visually comparable format, such as a graph or list, to allow the user to intuitively understand the evaluation. It also has a function to set customized evaluation metrics based on the needs of the company or organization.

[0703] As a concrete example, when a user opens the app on their smartphone and presses the record button, the conversation with the candidate is automatically recorded and sent to the server. The results, analyzed by the server, are displayed as a score on the smartphone screen in real time.

[0704] An example of a prompt for a generative AI model is, "Design a prompt to identify the necessary evaluation metrics in an interview and generate a score based on the candidate's voice data." This allows users to create a fairer and more efficient interview process.

[0705] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0706] Step 1:

[0707] The user begins the interview and starts recording audio data using an application on the device. The device receives and records the interview audio data via the microphone. It takes the user's audio data as input and generates an audio file saved to local storage as output.

[0708] Step 2:

[0709] The terminal sends the recorded audio file to the server. The server receives this audio data and converts it into text data using speech recognition technology. Specifically, it analyzes the audio waveform using the "SpeechRecognition" library and generates the corresponding text. It takes an audio file as input and generates text data in string format as output.

[0710] Step 3:

[0711] The server analyzes the generated text data using a natural language processing library ("NLTK" or "spaCy"). Here, keywords and phrases related to the evaluation metrics are extracted. The generative AI model identifies features for score generation based on prompts derived from past data. The input is text data, and the output is data with identified evaluation metrics.

[0712] Step 4:

[0713] The server calculates a score based on the identified evaluation metrics. The calculation uses machine learning algorithms (e.g., scikit-learn or TensorFlow) to generate numerical scores for each metric. The input is the evaluation metric data, and the output is the numerical score.

[0714] Step 5:

[0715] The terminal receives scores sent from the server and visualizes them through a user interface. Graphs and lists are used to allow for visual comparison of scores among multiple candidates. The input is a numerical score, and the output is a visualized evaluation result.

[0716] Step 6:

[0717] Users compare and evaluate candidates based on the provided evaluation results and make hiring decisions. By comparing candidates, users can identify their relative strengths and weaknesses, enabling more rational decision-making. The input is the visualized evaluation results, and the output is the final hiring decision.

[0718] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0719] This invention provides a system that enables more precise talent evaluation by combining an emotion engine to recognize and incorporate the emotional state of candidates during job interviews into the evaluation process. The system includes processing to detect emotions in addition to analyzing voice data.

[0720] Data reception and conversion

[0721] The user records the conversation with the candidate as audio during the interview and sends the data to the system.

[0722] Recognition of voice and emotion

[0723] The server processes the received audio data using a speech recognition engine to convert it into text data. Furthermore, it uses an emotion engine to analyze the tone and tempo of the voice and recognize the candidate's emotional state.

[0724] Data analysis and evaluation

[0725] The server generates scores for each evaluation metric based on the content of the text data and the recognized sentiment information. Sentiment information is incorporated into the evaluation of communication skills and personal characteristics.

[0726] Presentation of results and feedback

[0727] The device displays emotional information, along with the generated evaluation score, as part of the analysis results on the screen and presents it to the user acting as the interviewer. This enables a comprehensive evaluation that includes emotional aspects.

[0728] Specific example

[0729] In a typical interview, a user collects audio data from five candidates. The server converts the audio data into text and uses an emotion engine to analyze the emotions each candidate expressed. For example, it evaluates the level of confidence and nervousness that candidate A displayed when answering questions, and reflects these factors influencing their communication skills in a score. The device then graphs this information, providing the interviewer with an intuitive understanding. The user can leverage this detailed feedback to make informed decisions about selecting the best candidates.

[0730] Thus, by integrating an emotion engine, the system of this invention grasps the latent characteristics of candidates that cannot be captured by conventional language analysis alone, and supports talent evaluation that is more suitable for companies.

[0731] The following describes the processing flow.

[0732] Step 1:

[0733] After the interview begins, the user uses a dedicated recording device to collect audio data of the conversation with the candidate. This data is transmitted to the system in real time via a secure connection.

[0734] Step 2:

[0735] The server inputs the audio data sent by the user into the speech recognition engine and converts it into text data sequentially. This conversion process includes pre-processing such as background noise reduction and speaker separation.

[0736] Step 3:

[0737] The server simultaneously inputs voice data into the emotion engine and identifies the candidate's emotional state by analyzing features such as tone, volume, and speed of voice. Emotional information is classified into categories such as joy, tension, and surprise.

[0738] Step 4:

[0739] The server integrates the generated text data and sentiment data, inputs it into the evaluation model, and calculates scores for each evaluation metric. Specifically, it scores metrics such as communication skills, cooperativeness, and adaptability.

[0740] Step 5:

[0741] The device displays the evaluation score and sentiment analysis results for each candidate as visualized data on the screen. The display uses graphs, charts, and numerical data, providing information in a visually easy-to-understand format.

[0742] Step 6:

[0743] Users deepen their thinking based on the information presented and make a comprehensive evaluation of candidates from a more objective perspective. This improves the transparency and rationality of the selection process, enabling the selection of the most suitable talent.

[0744] (Example 2)

[0745] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0746] In recruitment interviews, traditional evaluation methods struggle to adequately grasp candidates' nonverbal characteristics and emotions, resulting in difficulties in selecting appropriate personnel. Furthermore, reliance on the interviewer's subjective evaluation leads to a lack of consistency and objectivity.

[0747] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0748] In this invention, the server includes means for receiving interview information, means for converting the interview information into text information using speech analysis technology, and means for recognizing the emotional state from the text information and generating numerical values ​​for each evaluation criterion using the results. This makes it possible to reflect the candidate's emotional state during the interview in the evaluation, enabling a more objective and consistent personnel evaluation.

[0749] "Interview information" refers to audio data and related data obtained from conversations and interactions with candidates during job interviews.

[0750] "Speech analysis technology" is a technology that processes speech data and converts it into text information, and in particular, it uses a speech recognition engine for analysis.

[0751] "Textual information" refers to text data that represents the content of audio data, converted using speech analysis technology.

[0752] "Emotional state" refers to the emotional characteristics and reactions that a candidate exhibits during a conversation, and is analyzed from their voice and tone.

[0753] "Evaluation criteria" are standards or scales used to measure a candidate's abilities and aptitudes, and are set according to the organization's needs and objectives.

[0754] "Numerical values" refer to data that quantitatively represents a candidate's abilities and characteristics, calculated based on evaluation criteria.

[0755] A "decision-maker" refers to an individual or organization that has the authority to make decisions regarding personnel selection based on the results of job interviews.

[0756] This invention is a system for recognizing and evaluating the emotional state of candidates during job interviews, and is primarily composed of audio data processing. This system operates as follows:

[0757] Data collection and transmission

[0758] The user records the conversation with the candidate during the interview as audio data using a recording device. This data is then sent to a server.

[0759] Analysis of audio data

[0760] The server analyzes the received audio data. Specifically, it uses a speech recognition engine (e.g., a common cloud-based speech recognition API) as an audio analysis technology to convert the audio data into text information. It also uses an emotion engine (e.g., commercially available emotion recognition software) to analyze the tone and speed of the voice and recognize the candidate's emotional state.

[0761] Generation and display of evaluation scores

[0762] The server uses a generative AI model to generate numerical values ​​for each evaluation criterion, based on the converted text information and recognized emotional states.

[0763] The terminal visualizes numerical data and sentiment information generated on the server, displaying it intuitively through graphs and charts. This information is provided to the user, who is the decision-maker, to assist in the overall evaluation of candidates.

[0764] As a concrete example, suppose a user collects audio data from five candidates during an interview session. The server processes this audio, performing speech recognition and sentiment analysis. For instance, it evaluates candidate A's confidence and nervousness, reflecting these in a numerical value corresponding to their communication skills. The terminal then intuitively displays these evaluation results, allowing the user to easily compare the characteristics of each candidate. This detailed information is extremely useful in making hiring decisions.

[0765] An example of a prompt might be: "Please describe a method for analyzing candidate voice data during interviews to quantify the influence of emotions on the evaluation of communication skills."

[0766] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0767] Step 1:

[0768] The user records the conversation with the candidate during the job interview as audio data using a recording device. The recorded audio data is sent to the server in its original format. In this step, the input is the candidate's audio data, and the output is the transmission of the audio data to the server. Specifically, the user presses the record button and continues recording until the audio data is of sufficient length.

[0769] Step 2:

[0770] The server converts received audio data into text data through a speech recognition engine. A cloud-based speech recognition API is used for this process. The input is audio data, and the output is the corresponding text data. Specifically, the server batch processes the audio data for analysis and generates recognition results in real time.

[0771] Step 3:

[0772] The server uses an emotion engine to analyze the converted text data and the original audio data to recognize the candidate's emotional state. In this step, text and audio data are taken as input, and information about the emotional state is obtained as output. Specifically, the server analyzes the tone and speed of the voice and extracts emotional characteristics.

[0773] Step 4:

[0774] The server utilizes a generative AI model to quantify the obtained text data and sentiment information based on evaluation criteria. The input consists of text data and sentiment information, and the output is a numerical value generated according to the evaluation criteria. The server processes the data using algorithms and generates numerical values ​​for each evaluation metric.

[0775] Step 5:

[0776] The terminal visualizes numerical data and emotional states transmitted from the server, displaying them on the screen as graphs and charts. Inputs are numerical data and emotional state information, while output is a visual evaluation report. Specifically, the terminal updates data in real time and displays it in a dashboard format for intuitive user understanding.

[0777] Step 6:

[0778] The user makes an overall evaluation of the candidates based on the data displayed on the device. The input is the displayed evaluation results, and the output is the final hiring decision. Specifically, the user reviews the evaluation results on the device and compares the candidates to reach a conclusion.

[0779] (Application Example 2)

[0780] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0781] In collaborative work between humans and robots in factories, it is not easy to grasp the emotional state of workers in real time and for robots to take appropriate actions and provide feedback in response to those emotions. Therefore, there is a need for systems that can improve work efficiency and safety.

[0782] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0783] In this invention, the server includes means for receiving voice data, means for converting it into text data using speech recognition technology, and means for analyzing voice tone and tempo to recognize emotional state. This enables real-time understanding of the emotional state of workers and allows for appropriate feedback and decision-making regarding actions.

[0784] "Interview data" refers to all information collected during an interview, including audio and video.

[0785] "Speech recognition technology" is a technology that converts speech data into text data.

[0786] "Text data" refers to data that has been converted using speech recognition technology and is represented as textual information.

[0787] "Tone" refers to elements that describe characteristics such as pitch and volume of sound.

[0788] "Tempo" refers to the speed and rhythm of speech.

[0789] "Emotional state" refers to the psychological state analyzed from audio data, and includes emotions such as joy, anger, sadness, and happiness.

[0790] "Evaluation indicators" refer to the criteria and perspectives used to evaluate individuals during interviews or work assignments.

[0791] A "score" is a rating or numerical value calculated based on evaluation metrics.

[0792] "User" is a general term referring to anyone who uses a system or receives a service.

[0793] This system is designed to facilitate collaborative work with factory workers. The server receives voice data from workers in real time and converts it to text using speech recognition technology. Here, speech recognition software such as Google Cloud Speech-to-Text API is used. Subsequently, an emotion engine, such as Microsoft Azure Emotion API, is used to analyze the tone and tempo of the voice and recognize the worker's emotional state.

[0794] Based on the information obtained through this emotion analysis, the server generates feedback and instructions tailored to the work situation and displays them on the terminal. The terminal refers to a smartphone or head-mounted display, providing information in a way that the worker can intuitively understand. For example, if the server determines that the worker is feeling anxious, the terminal will display a message such as, "Take your time, there's no need to rush."

[0795] In this system, the user refers to a worker, and by receiving feedback, improvements in work efficiency and safety can be expected. For example, if an incorrect operation is detected while handling specialized tools, the robot will suggest appropriate corrective steps.

[0796] An example of a prompt for a generative AI model is: "Analyze the worker's current emotions from this audio data and generate instructions for the robot to provide appropriate feedback."

[0797] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0798] Step 1:

[0799] Users record spoken audio in real time at the work site using a smartphone or head-mounted display and send the audio data to a server. The input is audio data, and the output is the raw audio data sent to the server. This process allows workers' voices to be stored as data.

[0800] Step 2:

[0801] The server converts the received audio data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose. The input is raw audio data, and the output is text data extracted from the audio. This process converts the audio content into a format that can be parsed as text information.

[0802] Step 3:

[0803] The server analyzes the converted text data and the tone and tempo of the speech, and uses the Microsoft Azure Emotion API to recognize the emotional state. The input is text data and speech feature information, and the output is the recognized emotional state. This process makes it possible to determine the user's psychological state.

[0804] Step 4:

[0805] The server determines appropriate feedback or action plans based on the worker's emotional state and work progress. Inputs are emotional state and individual work status data, while outputs are feedback content and action plans. This allows for appropriate advice and warnings to be given to the worker.

[0806] Step 5:

[0807] The device displays the determined feedback or action plan to the user. The input is the feedback content, and the output is visual or audio information presented to the user. This operation allows the user to intuitively receive guidance for improvement or correction while working.

[0808] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0809] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0810] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0811] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0812] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0813] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0814] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0815] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0816] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0817] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0818] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0819] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0820] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0821] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0822] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0823] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0824] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0825] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0826] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0827] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0828] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0829] The following is further disclosed regarding the embodiments described above.

[0830] (Claim 1)

[0831] Means for receiving interview data,

[0832] The means for converting the aforementioned interview data into text data using speech recognition technology,

[0833] A means for analyzing the aforementioned text data and generating scores for each interview evaluation metric,

[0834] A means for presenting the generated score to the interviewer,

[0835] A system that includes this.

[0836] (Claim 2)

[0837] The system according to claim 1, comprising means for training interview evaluation criteria using past interview record data, thereby improving the accuracy of the analysis.

[0838] (Claim 3)

[0839] The system according to claim 1, which includes means for setting customized evaluation indicators based on the needs of a company, and which supports the securing of suitable personnel for each company.

[0840] "Example 1"

[0841] (Claim 1)

[0842] Means of receiving interview information,

[0843] The means for converting the aforementioned interview information into text information using a speech recognition method,

[0844] A means for analyzing the aforementioned textual information and generating evaluation values ​​for each interview evaluation criterion,

[0845] A means for presenting the generated evaluation value to the interviewer,

[0846] A means of generating a prompt sentence and inputting it to a generative AI model,

[0847] A system that includes this.

[0848] (Claim 2)

[0849] The system according to claim 1, comprising means for training interview evaluation criteria using past interview record information, thereby improving the accuracy of the analysis.

[0850] (Claim 3)

[0851] The system according to claim 1, comprising means for setting evaluation criteria adjusted based on the requirements of the organization, and supporting the securing of suitable personnel for each organization.

[0852] "Application Example 1"

[0853] (Claim 1)

[0854] Means for receiving interview data,

[0855] The means for converting the aforementioned interview data into text data using speech recognition technology,

[0856] A means for analyzing the aforementioned text data and generating scores for each interview evaluation metric,

[0857] A means for displaying the generated score on the terminal,

[0858] A means to make the evaluation results of candidates visually comparable,

[0859] A system that includes this.

[0860] (Claim 2)

[0861] The system according to claim 1, which uses past interview record data to train interview evaluation criteria and improve analysis accuracy.

[0862] (Claim 3)

[0863] The system according to claim 1, which sets customized evaluation metrics based on the needs of the organization and supports the recruitment of suitable personnel for each organization.

[0864] "Example 2 of combining an emotion engine"

[0865] (Claim 1)

[0866] Means of receiving interview information,

[0867] A means for converting the aforementioned interview information into text information using speech analysis technology,

[0868] A means for recognizing emotional states from the aforementioned textual information and generating numerical values ​​for each evaluation criterion using the results,

[0869] A means of visualizing the aforementioned numerical values ​​and perceived emotional states and presenting them to decision-makers,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] The system according to claim 1, which includes a means for improving interview evaluation criteria using machine learning based on past interview information, thereby improving the accuracy of the analysis.

[0873] (Claim 3)

[0874] The system according to claim 1, which includes means for setting evaluation criteria adjusted according to the organization's needs, and assists in selecting personnel suitable for each organization.

[0875] "Application example 2 of combining emotional engines"

[0876] (Claim 1)

[0877] Means for receiving interview data,

[0878] The means for converting the aforementioned interview data into text data using speech recognition technology,

[0879] A means for analyzing the aforementioned text data and recognizing the emotional state using the tone and tempo of the voice,

[0880] A means for incorporating the recognized emotional state into an evaluation index and generating a score,

[0881] A means for presenting the generated score to the user,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, comprising means for learning evaluation criteria using past data to improve analysis accuracy, and determining actions based on the user's emotional state.

[0885] (Claim 3)

[0886] The system according to claim 1, comprising means for setting customized evaluation metrics based on the needs of the organization, and supporting the selection of suitable personnel. [Explanation of Symbols]

[0887] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for receiving interview data, The means for converting the aforementioned interview data into text data using speech recognition technology, A means for analyzing the aforementioned text data and generating scores for each interview evaluation metric, A means for presenting the generated score to the interviewer, A system that includes this.

2. The system according to claim 1, comprising means for training interview evaluation criteria using past interview record data, thereby improving the accuracy of analysis.

3. The system according to claim 1, which includes means for setting customized evaluation indicators based on the needs of a company, and which supports the securing of suitable personnel for each company.