System
The system objectively evaluates interviews by recording and analyzing video/audio data, using natural language processing and a learning model to provide fair and efficient feedback, addressing the subjectivity and inefficiency of conventional methods.
Patent Information
- Application Number
- JP2024119072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional interview evaluations are subjective and inefficient, prone to human bias and require significant time and effort, making it difficult to assess diversity and potential objectively.
A system that records video and audio data of interviews, applies noise reduction, text conversion, natural language processing, and facial/body language analysis, uses a learning model trained on existing employee data to evaluate objectively, and provides feedback.
Enables fair and efficient interview evaluations by eliminating human subjectivity, reducing the burden on interviewers, and providing real-time feedback.
Smart Images

Figure 2026018011000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional interview evaluations often involve the subjectivity and bias of human interviewers, making it difficult to achieve a fair evaluation. Furthermore, evaluations require a lot of time and effort, making it difficult to properly assess diversity and potential. There is a need for a system that can solve these issues and achieve more objective and efficient evaluations. [Means for solving the problem]
[0005] This invention provides a series of means for recording video and audio data of interviews and analyzing this data. First, the video and audio data are recorded and sent to a server. The server performs noise reduction and text conversion on the received data. Natural language processing technology is used to analyze the wording and content of the conversation for the text data, and speaking style and tone of voice are analyzed for the audio data. Furthermore, facial expressions and body language are analyzed from the video data, and these features are quantified and scored. Furthermore, a learning model is trained using interview data from existing employees, and this model is used to evaluate new interview data. Finally, the evaluation results are displayed and feedback is provided to the user. This realizes a system that can perform interview evaluations objectively and fairly.
[0006] An "interview" is a question and answer process designed to assess a candidate's aptitude and capabilities.
[0007] "Video and audio data" refers to digitally recorded data of the interview process.
[0008] A "terminal" is an electronic device that has the function of recording an interview and transmitting the data to a server.
[0009] A "server" is a computer system for analyzing video and audio data and training learning models.
[0010] "Noise reduction" is the process of removing unwanted background sounds and noise from recorded audio data.
[0011] "Text conversion" is the process of converting audio data into written information.
[0012] "Natural language processing technology" is an AI technology for analyzing and understanding human language.
[0013] "Language" refers to the choice of language and the manner of expression used by the interviewer.
[0014] "Conversation content" refers to the content of questions and answers and discussions that take place during the interview.
[0015] "Speaking style" refers to the speed, rhythm, and pronunciation characteristics of the interviewer's speech.
[0016] "Voice tone" refers to variations in tone, pitch, and volume of speech.
[0017] "Facial expressions" indicate emotions and reactions through the interviewer's facial movements and expressions.
[0018] "Body language" refers to gestures, posture, and other forms of non-verbal communication.
[0019] "Feature quantification" is the process of scoring extracted features based on standardized evaluation criteria.
[0020] "Scoring" is the process of assigning points to numerical characteristics and evaluating them.
[0021] A "learning model" is a machine learning algorithm that is trained using existing data.
[0022] A "rule of thumb" is a set of indicators or criteria for evaluating an interviewer's performance.
[0023] "Feedback" is information provided to the interviewer or interviewers based on the results of the evaluation. [Brief explanation of the drawings]
[0024] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0025] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0026] First, the terms used in the following description will be explained.
[0027] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0028] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0029] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0030] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0032] [First embodiment]
[0033] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0034] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0036] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0037] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0039] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0040] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0041] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0043] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0045] System Overview
[0046] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes that record the video and audio data of the interview, analyze it, and provide the results as feedback. This eliminates human subjectivity and bias, enabling fair evaluation.
[0047] System configuration
[0048] The system consists of the following main components:
[0049] 1. Recording device: A device that records the video and audio data of the interview.
[0050] 2. Server: A central computer that receives and analyzes the recorded data.
[0051] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[0052] 4. Feedback terminal: A device for displaying the analysis results to the user.
[0053] System Operation
[0054] 1. Data Collection and Transfer
[0055] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[0056] 2. Data Preprocessing
[0057] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[0058] 3. Feature extraction and quantification
[0059] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0060] 4. Training and evaluation of the learning model
[0061] The server trains a machine learning model using interview data from existing employees. This trained model is then used to evaluate newly acquired interview data and calculate various scores. For example, the server evaluates overall performance based on factors such as the consistency of language, conversational content, and facial expressions.
[0062] 5. Presentation of results
[0063] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) can check the evaluation results and receive feedback through this terminal. The evaluation results include an overall evaluation and detailed scores for each evaluation item.
[0064] Specific examples
[0065] For example, when an interview is conducted with interviewer A, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[0066] Noise is removed from video and audio data to extract high-quality data.
[0067] Converts audio data into text.
[0068] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[0069] Analyzes voice tone, rhythm, and tone from audio data.
[0070] Extracts changes in facial expressions and body language from video data.
[0071] As a result, an overall score of 85 / 100 is calculated for Interviewer A, and detailed evaluations are given, with 90 / 100 for language and 80 / 100 for tone of voice. These results are presented to the user via the feedback terminal.
[0072] This system will enable interview evaluation to be conducted objectively and efficiently, and is expected to be a significant improvement over conventional subjective evaluation methods.
[0073] The processing flow will be explained below.
[0074] Step 1:
[0075] The user (interviewer) starts the interview, and the device activates the camera and microphone to record the video and audio data of the interview.
[0076] Step 2:
[0077] The device sends the recorded data to the server in real time or after the interview is over. The sent data is in the form of video and audio files.
[0078] Step 3:
[0079] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[0080] Step 4:
[0081] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the speaker's voice is transcribed sequentially as text.
[0082] Step 5:
[0083] The server then analyzes the converted text data using natural language processing (NLP) technology, specifically evaluating the language used (for example, frequency of honorific use, use of technical terms), the grammatical structure of the conversation, and logic.
[0084] Step 6:
[0085] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[0086] Step 7:
[0087] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[0088] Step 8:
[0089] The server quantifies each extracted feature and assigns a score according to a standardized evaluation scale, such as 90 / 100 for appropriate language and 85 / 100 for friendly facial expressions.
[0090] Step 9:
[0091] The server uses interview data from existing employees to train the learning model, which uses machine learning algorithms to learn evaluation criteria that fit the company culture.
[0092] Step 10:
[0093] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[0094] Step 11:
[0095] The server sums up the scores for each evaluation item to generate an overall evaluation, which is then sent to the device.
[0096] Step 12:
[0097] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the overall evaluation might be 85 / 100, language 90 / 100, and facial expression 85 / 100.
[0098] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias and helping to select the right candidates.
[0099] Example 1
[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0101] Conventional interview evaluation systems have the problem that they are prone to human subjectivity and bias, making it difficult to achieve consistent and fair evaluations. Furthermore, the evaluation process is inefficient, placing a heavy burden on interviewers. Furthermore, conventional systems have difficulty providing real-time feedback, making it impossible to immediately identify areas for improvement.
[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0103] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to a central processing unit, means for removing noise from the transmitted video and audio data and converting the audio into text information, means for analyzing the phrasing and conversation content of the audio data using language analysis technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and sign language from the video data, means for quantifying and evaluating each analyzed feature, means for training a machine learning model using interview data of existing employees, means for evaluating new interview data using the machine learning model, and means for displaying the evaluation results and providing feedback to the user. This improves the objectivity and fairness of interview evaluations, reduces the burden on interviewers, and enables feedback to be provided in real time.
[0104] A "recording terminal" is a device for recording video and audio data of an interview.
[0105] A "central processing unit" is a computer or server that receives and analyzes the recorded data.
[0106] "Noise reduction" is the process of removing unnecessary noise from recorded audio and video data.
[0107] "Speech-to-text" refers to the process of converting recorded voice data into text data.
[0108] "Language analysis technology" is a technology that performs natural language processing on text data and analyzes wording and conversation content.
[0109] "Analyzing speaking style and tone of voice" refers to analyzing characteristics such as speaking rate, pitch, and tone based on the audio data.
[0110] "Analyzing facial expressions and sign language" means analyzing the interviewer's facial expressions and body language based on video data.
[0111] "Quantification" means expressing the analyzed features as numerical data.
[0112] A "machine learning model" is an algorithm that learns patterns based on past data and makes predictions and classifications.
[0113] "Providing feedback" means presenting the analysis and evaluation results to the user and informing them of areas for improvement and strengths.
[0114] "Setting evaluation criteria that fit the organizational culture" means setting evaluation criteria that are appropriate for the characteristics and values of the organization based on interview data from existing employees.
[0115] "Immediate analysis and display" means that interview data is processed in real time and the results are displayed immediately.
[0116] MODE FOR CARRYING OUT THE INVENTION
[0117] The present invention provides a system for objectively and efficiently evaluating interviews. This system includes a process for recording video and audio data of the interview, analyzing the data with a central processing unit, and providing the results as feedback to the user.
[0118] The system consists of the following main components:
[0119] 1. Recording device: This is a device that records the video and audio data of the interview. In this case, a standard webcam and microphone are used.
[0120] 2. Central Processing Unit (Server): A computer that receives and analyzes recorded data. A general cloud server or on-premise server can be used for high-performance computing.
[0121] 3. Analysis module: This is software that runs inside the central processing unit and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specifically, it uses software such as FFmpeg, Google Speech-to-Text API, NLTK, Spacy, Praat, OpenCV, Dlib, Scikit-learn, and TensorFlow.
[0122] 4. Feedback terminal: A device that displays analysis results to users and provides feedback. In this case, this applies to general PCs, tablets, smartphones, etc.
[0123] (Data Collection and Transfer)
[0124] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When the recording is finished, the recording device sends this data to the central processing unit. The transmission process can use storage services such as Google Drive API or AWS S3.
[0125] (Data preprocessing)
[0126] The server uses FFmpeg to filter the received recording data to remove noise, and then uses the Google Speech-to-Text API to convert the audio data into text.
[0127] (Feature extraction and quantification)
[0128] When the server receives the speech-recognized text data, it uses natural language processing technology to analyze the vocabulary and content of the conversation. This analysis is performed using NLTK and Spacy. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points.
[0129] (Training and evaluating learning models)
[0130] The server uses the feature data to train a machine learning model. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. The trained model is then used to evaluate new interview data and calculate an overall score and detailed scores for each evaluation item.
[0131] (Presentation of results)
[0132] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript as a dashboard operated by the user. The user (interviewer or interviewee) uses this feedback terminal to check the evaluation results and identify areas for improvement and strengths.
[0133] Specific examples
[0134] For example, when an interview with Interviewer A is conducted, the entire process is recorded and the recorded data is sent to the server. The server removes noise using FFmpeg and converts the audio data into text using the Google Speech-to-Text API. Next, it analyzes the text data using NLTK or Spacy, analyzes the audio data using Praat, and extracts facial expressions from the video data using OpenCV and Dlib. These results are quantified, and each evaluation item is scored using a model trained with Scikit-learn or TensorFlow. Finally, these results are provided to the user in a form that can be viewed on a feedback terminal. The evaluation results include detailed information such as Interviewer A's overall score of 85 / 100, language score of 90 / 100, and tone of voice score of 80 / 100.
[0135] Prompt Sentence Examples
[0136] Below are some example prompts to input to a generative AI model:
[0137] "Please explain in detail, step by step, the process of the program that analyzes the interview recording data and conducts a comprehensive evaluation based on the interviewer's language, tone of voice, facial expressions, and sign language."
[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0139] Step 1: Data collection and transfer
[0140] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When recording is finished, the recording device sends the recorded video and audio data to the server. This sending process uses storage services such as Google Drive API and AWS S3. The input is the video and audio data of the interview, and the output is the recorded data sent to the server.
[0141] Step 2: Preprocessing the data
[0142] The recording data received by the server is first denoised. Specifically, FFmpeg is used to filter the audio data and remove static noise. Next, the Google Speech-to-Text API is used to convert the audio data into text. This process converts the audio data into text information. The input is the recording data sent to the server, and the output is the noise-denoised audio data and text data.
[0143] Step 3: Feature extraction and quantification
[0144] The server uses natural language processing technology to analyze the phrasing and conversation content of the speech-recognized text data. NLTK and Spacy are used for this analysis. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points. The input is noise-removed audio data and text data, and the output is quantified feature data.
[0145] Step 4: Training and evaluating the learning model
[0146] The server trains a machine learning model using the feature data it has acquired so far. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. Next, it uses the trained model to evaluate newly acquired interview data and calculates an overall score and detailed scores for each evaluation item. The input is quantified feature data, and the output is an evaluation score.
[0147] Step 5: Presenting the results
[0148] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript to build a dashboard. The user (interviewer or interviewee) can check the evaluation results using this feedback terminal. The evaluation results include an overall score and detailed evaluation information for each evaluation item. The input is the evaluation score, and the output is the visualized evaluation results.
[0149] (Application example 1)
[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0151] In conventional interview evaluation systems, evaluations tend to be subjective and often result in bias. Furthermore, in factories and other workplaces, it is difficult to objectively evaluate worker performance and provide efficient feedback. Therefore, a new system is needed that can accurately and quickly evaluate worker performance in workplaces where improvements in work efficiency and safety are required.
[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0153] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using existing data, means for evaluating new data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for evaluating work efficiency, accuracy, safety, etc. based on the analysis results, and means for feeding back the scoring results to the worker and suggesting areas for improvement. This makes it possible to objectively and efficiently evaluate worker performance and suggest areas for improvement even in factories and on-site workplaces.
[0154] Definitions of important words
[0155] "Video Data" means visual information captured by a camera or other recording device.
[0156] "Audio Data" means audio information captured by a microphone or other recording device.
[0157] "Noise reduction" is the process of removing unwanted background noise and artifacts from video and audio recordings.
[0158] "Text conversion" is the process of converting audio data into written information.
[0159] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0160] "Language" refers to the choice of words and expressions used in speech.
[0161] "Conversation content" refers to specific topics and information contained in the voice data.
[0162] "Speaking style" refers to characteristics of the audio data such as pronunciation, intonation, and rhythm.
[0163] "Voice tone" refers to voice characteristics including the pitch, strength, rhythm, etc. of the voice in the voice data.
[0164] "Facial expressions" refer to facial changes and emotional expressions in video data.
[0165] "Body language" refers to the body movements and postures in video data.
[0166] "Scoring" is the process of quantifying and evaluating the analyzed features.
[0167] A "learning model" is an algorithm that learns from data to identify and predict specific patterns.
[0168] "Feedback" is the process of providing the evaluation results to the user and indicating areas for improvement and evaluation.
[0169] "Work efficiency" is an index that indicates how efficiently work is performed within a certain period of time.
[0170] "Accuracy" is an indicator of how accurately a task is performed.
[0171] "Safety" is an indicator of how safely the work was carried out.
[0172] MODE FOR CARRYING OUT THE INVENTION
[0173] 1. System Overview
[0174] This invention provides a system for objectively evaluating the performance of workers in a factory or other workplace. The system mainly consists of the following components:
[0175] Recording device
[0176] server
[0177] Analysis Module
[0178] Feedback Terminal
[0179] 2. Hardware and Software Configuration
[0180] Recording device
[0181] The recording terminal is equipped with a camera and microphone and is a device that records the worker's video and audio data in high quality. In this example, a general web camera and microphone are used.
[0182] server
[0183] The server acts as a central computer that receives and processes the recorded data. The following software is installed on the server:
[0184] OpenCV (video data analysis)
[0185] speech_recognition (converts speech data to text)
[0186] scikit-learn (machine learning model training and evaluation)
[0187] 3. Program Processing
[0188] Data Collection and Transfer
[0189] The recording terminal records the video and audio data of the worker's work and sends the data to the server. The recorded data is pre-processed to remove noise, and the audio data is converted into text.
[0190] Data analysis
[0191] The server applies natural language processing technology to the audio data to analyze the language used and the content of the conversation. It analyzes facial expressions and body language from the video data, and extracts speaking style and tone of voice from the audio data. These features are then quantified and scored.
[0192] Feature extraction and quantification
[0193] The server quantifies the characteristics obtained from the recorded data and scores it based on specific criteria, including work efficiency, accuracy, and safety.
[0194] Training and evaluating learning models
[0195] The server trains a learning model using existing data and evaluates newly acquired data. This process uses a machine learning model (e.g., SVM) using the scikit-learn library.
[0196] Feedback of results
[0197] The analysis results and the evaluated scoring results are sent to a feedback terminal where users can check the results, allowing workers to receive an evaluation of their performance and learn areas for improvement.
[0198] 4. Examples of concrete examples and prompts
[0199] Specific examples
[0200] For example, Worker A's work process for one day is recorded and an evaluation is made based on that data. The recorded data is converted into text using voice recognition, and the work movements are evaluated through video analysis. Finally, an overall evaluation score is obtained, and this score is fed back to the worker to indicate areas for improvement.
[0201] Prompt Sentence Examples
[0202] "Please create prompts for a scoring system that will record Worker A's work for one day, analyze the video and audio, and evaluate work efficiency, accuracy, safety, etc. Please include specific steps from data preprocessing to evaluation."
[0203] This invention makes it possible to make objective evaluations that were difficult to achieve with conventional methods, and is expected to improve work efficiency and safety.
[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0205] Program processing steps
[0206] Step 1:
[0207] The recording terminal records the video and audio data of the worker's work in real time using a camera and microphone, and saves this data as a file.
[0208] Input: Video and audio data of the worker
[0209] Output: Recording files (video files, audio files)
[0210] Step 2:
[0211] The recording device sends the recorded video and audio data to a server, where the data is centrally managed and analyzed.
[0212] Input: Recording file
[0213] Output: Data transferred to the server
[0214] Step 3:
[0215] The server performs noise reduction on the received video and audio data using filtering technology to improve the quality of the data.
[0216] Input: Transferred recording data (video files, audio files)
[0217] Output: High-quality data after noise removal
[0218] Step 4:
[0219] The server uses speech recognition technology to convert the audio data into text, specifically the speech_recognition library.
[0220] Input: High-quality audio data
[0221] Output: Text data
[0222] Step 5:
[0223] The server uses natural language processing technology to analyze the vocabulary and content of the conversation, thereby understanding the meaning and context of the text data.
[0224] Input: Text data
[0225] Output: Analysis results (characteristics of language and conversation content)
[0226] Step 6:
[0227] The server analyzes the speech data to determine speaking style and tone, and extracts features such as pitch, rhythm, and tone from the speech waveform data.
[0228] Input: High-quality audio data
[0229] Output: Audio feature data
[0230] Step 7:
[0231] The server analyzes facial expressions and body language from the video data, detecting facial expressions and body movements using OpenCV.
[0232] Input: High-quality video data
[0233] Output: Video feature data
[0234] Step 8:
[0235] The server quantifies and scores each analyzed feature, calculates the score for each feature, and performs an overall performance evaluation.
[0236] Input: Language feature data, audio feature data, video feature data
[0237] Output: Scoring results
[0238] Step 9:
[0239] The server trains a learning model using existing data, using scikit-learn for this process and optimizing the model based on existing evaluation data.
[0240] Input: Existing evaluation data
[0241] Output: A trained model
[0242] Step 10:
[0243] The server evaluates newly acquired data using the learning model, inputting new data to the trained model and obtaining the evaluation results.
[0244] Input: Newly acquired feature data
[0245] Output: Evaluation results
[0246] Step 11:
[0247] The server transmits the evaluation results to the feedback terminal, which displays the evaluation results so that the user can check them.
[0248] Input: Evaluation result (scoring result)
[0249] Output: Feedback to terminal
[0250] Step 12:
[0251] Users receive the evaluation results through a feedback terminal and identify areas for improvement, which helps workers improve their own performance.
[0252] Input: Data to be displayed on the feedback terminal
[0253] Output: Feedback information to the user
[0254] These steps allow for objective and efficient evaluation and feedback of worker performance.
[0255] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0256] System Overview
[0257] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to extract emotional changes during the interview and incorporate them into a comprehensive evaluation.
[0258] System configuration
[0259] The system consists of the following main components:
[0260] 1. Recording device: A device that records the video and audio data of the interview.
[0261] 2. Server: A central computer that receives and analyzes the recorded data.
[0262] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[0263] 4. Emotion engine: Software that runs inside the server and analyzes the user's emotions from video and audio data.
[0264] 5. Feedback terminal: A device for displaying the analysis results to the user.
[0265] System Operation
[0266] 1. Data Collection and Transfer
[0267] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[0268] 2. Data Preprocessing
[0269] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[0270] 3. Feature extraction and quantification
[0271] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0272] 4. Analysis by Emotion Engine
[0273] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[0274] 5. Quantifying and Scoring Emotional Data
[0275] The server quantifies the emotional change data obtained from the emotion engine and scores it in the same way as other evaluation items. Evaluation criteria include emotional stability and appropriate emotional expression.
[0276] 6. Training and Evaluation of the Learning Model
[0277] The server trains a machine learning model using interview data from existing employees. The model uses various features, including emotional data, to learn evaluation criteria that fit the company culture. When new interview data is input, the trained model is used to perform an overall evaluation.
[0278] 7. Presentation of results
[0279] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) checks the evaluation results through this terminal and receives feedback. The evaluation results include an overall evaluation including emotional evaluation and detailed scores for each evaluation item.
[0280] Specific examples
[0281] For example, when an interview with Interviewer B is conducted, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[0282] Noise is removed from video and audio data to extract high-quality data.
[0283] Converts audio data into text.
[0284] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[0285] Analyzes voice tone, rhythm, and tone from audio data.
[0286] Extracts changes in facial expressions and body language from video data.
[0287] An emotion engine is used to detect various emotions (happiness, surprise, confusion, etc.).
[0288] Each feature and emotion data is quantified and scored.
[0289] As a result, an overall score of 90 / 100 is calculated for Interviewer B, and detailed evaluations are given, including 85 / 100 for language, 82 / 100 for tone of voice, and 88 / 100 for emotional stability. These results are presented to the user via a feedback terminal.
[0290] This system allows for objective and efficient evaluation of interviews, and is expected to be a significant improvement over conventional subjective evaluation methods. In addition, by combining it with an emotion engine, changes in the interviewer's emotions can be reflected in the evaluation, enabling a more comprehensive evaluation.
[0291] The processing flow will be explained below.
[0292] Step 1:
[0293] The user (interviewer) starts the interview, and the recording terminal activates the camera and microphone to record the video and audio data of the interview.
[0294] Step 2:
[0295] The recording device sends the recorded data to the server in real time or after the interview ends. The sent data is in the form of video and audio files.
[0296] Step 3:
[0297] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[0298] Step 4:
[0299] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the voice is transcribed sequentially.
[0300] Step 5:
[0301] The server then analyzes the converted text data using natural language processing (NLP) techniques, specifically evaluating the language used (e.g., frequency of honorific use and use of technical terms), the grammatical structure of the conversation, coherence, and logic.
[0302] Step 6:
[0303] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[0304] Step 7:
[0305] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[0306] Step 8:
[0307] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[0308] Step 9:
[0309] The server quantifies the emotional change data obtained from the emotion engine and scores it according to standardized evaluation criteria, such as emotional stability and appropriate emotional expression.
[0310] Step 10:
[0311] The server quantifies and scores each extracted feature (language, speaking style, facial expressions, tone of voice, body language, emotional changes, etc.). For example, appropriate language is given a score of 90 / 100, and friendly facial expressions are given a score of 85 / 100.
[0312] Step 11:
[0313] The server uses interview data from existing employees to train a machine learning model, which uses a machine learning algorithm to learn evaluation criteria that fit the company culture.
[0314] Step 12:
[0315] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[0316] Step 13:
[0317] The server sums up the scores for each evaluation item to generate an overall evaluation, and sends the evaluation result to the terminal.
[0318] Step 14:
[0319] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the user can see detailed scores of 90 / 100 for overall evaluation, 85 / 100 for language, 82 / 100 for facial expressions, and 88 / 100 for emotional stability.
[0320] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias to help select the right candidates. Furthermore, by incorporating an emotion engine, changes in the interviewer's emotions are reflected in the evaluation, enabling a more comprehensive evaluation.
[0321] Example 2
[0322] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0323] Conventional interview evaluation systems often rely on the subjective judgment of the interviewer, resulting in problems such as a lack of fairness and consistency. Furthermore, it is difficult to perform a comprehensive evaluation that includes the interviewer's emotional changes, making it impossible to reflect emotional stability or appropriate emotional expression in the evaluation. Furthermore, it is difficult to perform these analyses in real time and provide feedback immediately after the interview.
[0324] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0325] In this invention, the server includes a means for recording video and audio data of the interview, a means for transmitting the recorded video and audio data to a central control device, a means for performing noise reduction and speech recognition on the transmitted video and audio data, a means for analyzing the phrasing and conversation content of the speech recognition results using natural language analysis technology, a means for analyzing speaking style and tone of voice from the audio data, a means for analyzing facial expressions and body movements from the video data, a means for quantifying and scoring each analyzed feature, a means for analyzing emotions from the recorded video and audio data, a means for quantifying and scoring the emotion data, a means for training a machine learning model using interview data of existing employees, a means for evaluating new interview data using the learning model, and a means for displaying the evaluation results and providing feedback to the user. This enables objective and efficient interview evaluation, and also allows the interviewer's emotional changes to be reflected in the evaluation. Analysis results can also be provided in real time, allowing feedback to be received immediately after the interview.
[0326] "Video and audio data of the interview" refers to the visual and audio information recorded during the interview.
[0327] "Recording means" refers to technical devices and methods for recording video and audio data.
[0328] "Central control unit" refers to the main computer system that centrally manages and processes multiple data.
[0329] "Transmission means" refers to the technical devices and methods for transferring data from one point to another.
[0330] "Noise reduction" refers to the process of removing unwanted noise and interference from video and audio data.
[0331] "Speech recognition" refers to the technology of analyzing voice data and converting it into text information.
[0332] "Natural language analysis technology" refers to analysis technology that enables computers to understand and interpret human language.
[0333] "Means for analyzing speaking style and tone of voice" refers to technology that analyzes a speaker's speech patterns and voice quality from audio data.
[0334] "Means for analyzing facial expressions and body movements" refers to technology that detects and analyzes changes in facial expressions and gestures of subjects from video data.
[0335] "Quantification" refers to the act of expressing analyzed information as quantitative data.
[0336] "Scoring" refers to the process of assigning a score based on quantified data according to evaluation criteria.
[0337] "Means for analyzing emotions" refers to technology that estimates emotional states from video and audio data.
[0338] A "machine learning model" refers to an algorithm that learns patterns from large amounts of data and makes inferences and predictions about new data.
[0339] "Means for displaying the evaluation results and providing feedback to the user" refers to technical devices and methods for visually presenting the analysis results to the user and facilitating their understanding.
[0340] This invention is a system for conducting interview evaluations objectively and efficiently, and includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. The system incorporates specific hardware and software, with the user, terminal, and server playing their respective roles.
[0341] Key Components of the System
[0342] 1. Recording device: A device for recording the video and audio data of the interview. Generally, a device equipped with a camera and microphone is used.
[0343] 2. Server: A central control unit that receives and analyzes recorded data. This device is a high-performance computer system that can process and analyze data at high speed. For example, a commonly used corporate server system can be used.
[0344] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. It uses deep learning libraries such as TensorFlow and PyTorch.
[0345] 4. Emotion engine: Software that analyzes emotions from video and audio data. Typically, a cloud-based API (e.g., Microsoft Azure's Emotion Recognition API) is used.
[0346] 5. Feedback terminal: A device used to display analysis results to users. This is typically a PC or mobile device.
[0347] Data collection and analysis flow
[0348] When the user (interviewer) starts the interview, the recording device records the entire process. After recording is finished, the device sends the recorded video and audio data to the server. The server performs noise reduction and speech recognition, and converts the audio data into text. For example, noise reduction is performed using FFmpeg, and the audio is converted to text information using the Google Cloud Speech-to-Text API.
[0349] For text data, natural language analysis techniques are used to analyze the phrasing and content of conversations. Natural language processing tools such as the BERT model are used. For audio data, speech analysis tools such as Praat are used to analyze speaking style and tone of voice. In addition, the server uses image analysis libraries such as OpenCV to analyze facial expressions and body language from video data.
[0350] Sentiment Analysis and Scoring
[0351] The emotion engine analyzes emotions from video and audio data to extract emotions such as joy, anger, surprise, and sadness, using Microsoft Azure's emotion recognition API.
[0352] The analyzed emotion data is quantified and scored by the server, using evaluation indices based on the intensity and frequency of each emotion.
[0353] The server uses interview data from existing employees to train a machine learning model, which uses tools like Scikit-learn to learn evaluation criteria that fit the company culture.
[0354] Presentation of evaluation results
[0355] Finally, when new interview data is input, the server uses the trained model to perform a comprehensive evaluation, and the generated evaluation results are sent to the feedback terminal, where users can check the evaluation results.
[0356] Examples and prompts
[0357] For example, if an interview with Interviewer B is conducted, the analysis will proceed as follows:
[0358] The recording device captures video and audio data.
[0359] The device sends data to the server.
[0360] The server uses FFmpeg to remove noise.
[0361] The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[0362] The server analyzes the text data using the BERT model.
[0363] The server uses Praat to extract audio features.
[0364] The server analyzes the video data using OpenCV.
[0365] The emotion engine analyzes emotions using Microsoft Azure's emotion recognition API.
[0366] The server quantifies the emotional data and scores it.
[0367] The server trains the machine learning model using Scikit-learn.
[0368] When new interview data is input, the server evaluates it using the trained model.
[0369] Evaluation results are presented through a feedback terminal.
[0370] Prompt Sentence Examples
[0371] Please provide a detailed description of each processing step in your interview evaluation system, including the specific actions and techniques used from data collection to providing results.
[0372] This system allows for objective and efficient interview evaluation. Furthermore, by using an emotion engine, the interviewer's emotional changes are incorporated into the evaluation, enabling a more comprehensive evaluation. Evaluation results are presented in real time, allowing users to receive immediate feedback.
[0373] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0374] Step 1:
[0375] The user (interviewer) starts the interview. Recording of audio and video data begins. The input is the video and audio during the interview, and the output is the video and audio data recorded on the recording terminal.
[0376] Step 2:
[0377] The device sends the recorded data to the server. After recording is complete, the recorded video and audio data is transferred to the server. The input is the video and audio data stored on the recording device, and the output is the same data sent to the server. Specifically, a secure data transfer protocol (such as SFTP) is used.
[0378] Step 3:
[0379] The server performs noise reduction. The server removes background noise from the received video and audio data. The input is the raw video and audio data sent to the server, and the output is the denoised video and audio data. FFmpeg noise reduction filters are used.
[0380] Step 4:
[0381] The server performs speech recognition. The server converts noise-removed speech data into text. The input is noise-removed speech data, and the output is text data. The Google Cloud Speech-to-Text API is used for speech recognition.
[0382] Step 5:
[0383] The server analyzes the text data. The server uses natural language analysis technology to analyze the transcribed text data. The input is text data generated by speech recognition, and the output is feature data of the analyzed phrasing and conversation content. The BERT model, for example, is used here.
[0384] Step 6:
[0385] The server extracts voice features. The server analyzes speaking style and tone of voice from the voice data. The input is the voice data after noise removal, and the output is feature data related to speaking style and tone of voice. Specifically, the voice analysis is performed using Praat.
[0386] Step 7:
[0387] The server analyzes the video data. The server analyzes facial expressions and body language from the video data. The input is the video data after noise removal, and the output is feature data related to facial expressions and body language. Image analysis libraries such as OpenCV are used.
[0388] Step 8:
[0389] The emotion engine analyzes emotions. The emotion engine extracts multiple emotional states from video and audio data. The input is noise-removed video and audio data, and the output is quantified data for each emotion. Microsoft Azure's emotion recognition API is used.
[0390] Step 9:
[0391] The server quantifies the analysis data and scores it. The server quantifies the acquired feature data and emotion data and displays it as a score. The input is feature data such as speaking style, tone of voice, facial expressions, body language, and emotions, and the output is a quantified evaluation score.
[0392] Step 10:
[0393] The server trains a machine learning model. The server trains the model using interview data of existing employees. The input is the interview data of existing employees, and the output is the trained machine learning model. Libraries such as Scikit-learn are used.
[0394] Step 11:
[0395] The server evaluates new interview data. The server evaluates new interview data using the trained model. The input is the newly collected interview data and the trained model, and the output is an overall evaluation score.
[0396] Step 12:
[0397] The server sends the evaluation result to the feedback terminal. The server sends the generated evaluation result to the feedback terminal. The input is the calculated evaluation score, and the output is the evaluation result displayed on the feedback terminal.
[0398] Step 13:
[0399] The user checks the evaluation results. The user checks the analysis results and the scores for each evaluation item through the feedback terminal. The input is the evaluation results displayed on the feedback terminal, and the output is the user's understanding of the evaluation results and feedback.
[0400] This series of processes enables interview evaluation to be conducted objectively and efficiently, and also enables comprehensive evaluation including emotional stability and appropriate expression.
[0401] (Application example 2)
[0402] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0403] In recent years, there has been a demand for strengthened security in offices and homes, but current security systems have difficulty grasping changes in visitors' emotions and behavior in real time. In particular, there is a lack of an objective evaluation system for quickly responding to suspicious behavior or abnormal emotional fluctuations. This can lead to potential risks being overlooked, making it difficult to implement effective security measures.
[0404] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using interview data of existing employees, means for evaluating new interview data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for detecting abnormal emotional fluctuations and suspicious behavior and issuing an alert, and means for generating a report based on the analysis results. This makes it possible to analyze the emotions and behavior of visitors in real time, quickly detect suspicious behavior, and issue an alert.
[0405] "Video data" refers to image information captured using a camera.
[0406] "Audio data" refers to audio information recorded using a microphone.
[0407] "Recording" refers to the act of saving video and audio data.
[0408] "Server" refers to a central computer that collects, stores, and analyzes data over a network.
[0409] "Noise reduction" refers to the method of removing unwanted information or interference from video and audio data.
[0410] "Text conversion" refers to the technology of converting voice data into text information.
[0411] "Natural language processing technology" refers to algorithms and methods for analyzing text data.
[0412] "Language" refers to the way a character speaks and the words they choose to use.
[0413] "Vocal tone" refers to characteristics such as the pitch, rhythm, and tone of a speaker's voice.
[0414] "Facial expression" refers to changes in emotions shown by the movement of facial muscles.
[0415] "Body language" refers to non-verbal communication expressed through bodily movements and posture.
[0416] "Quantification" refers to the act of expressing analyzed features as numerical data.
[0417] "Scoring" refers to the method of calculating an evaluation score based on numerical characteristics.
[0418] A "learning model" refers to a computational model that learns patterns from large amounts of data and performs predictions and classifications.
[0419] "Issuing an alarm" refers to the act of sending a notification to draw attention when the system determines that there is an abnormality.
[0420] "Report generation" refers to the act of creating a report summarizing the analysis results.
[0421] System Overview
[0422] This invention is a system that can be applied to a security system to analyze visitors' emotions and behavior in real time and detect suspicious behavior or abnormal emotional fluctuations. This system aims to improve security by analyzing video and audio data. It also issues an alarm and generates a report based on the analysis results.
[0423] System configuration
[0424] The system consists of the following main components:
[0425] 1. Recording device: A device that records video and audio data of visitors. Specifically, it uses a camera (e.g., Logitech C922 Pro Stream) and a microphone (e.g., Rode NT-USB).
[0426] 2. Server: A central computer that receives and analyzes the recorded data.
[0427] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specific examples of this software include OpenCV and the emotion_recognition library.
[0428] 4. Feedback terminal: A device for displaying the analysis results to the user.
[0429] 5. Alarm system: A device that issues an alarm if it detects abnormal emotional fluctuations or suspicious behavior.
[0430] 6. Report generation module: Software for generating reports based on the analysis results.
[0431] System Operation
[0432] 1. Data Collection and Transfer
[0433] The user starts the system, and the recording device records the video and audio of the visitor in real time. The recorded data is immediately sent to the server.
[0434] 2. Data Preprocessing
[0435] The server removes noise from the received data and converts the audio data into text. The OpenCV library is used for noise removal, and speech recognition technology is used for converting the audio data into text.
[0436] 3. Feature extraction and quantification
[0437] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0438] 4. Analysis by Emotion Engine
[0439] The server uses the emotion_recognition library to analyze emotions from video and audio data, detecting emotional changes in real time and identifying abnormal emotional fluctuations and suspicious behavior.
[0440] 5. Alert generation and report generation
[0441] If suspicious behavior or abnormal emotional changes are detected, the server will issue an alert through the alarm system. Furthermore, it will periodically generate reports based on the analysis results and provide them to the user.
[0442] Specific examples
[0443] For example, if a visitor shows signs of anger or fear at an office reception, the system can detect this in real time and immediately issue an alert, while a report summarizing the day's events is sent to the manager.
[0444] Example prompt sentence:
[0445] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[0446] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0447] Step 1:
[0448] Data Collection and Transfer
[0449] The user starts the system, and the recording terminal records the visitor's video and audio data in real time. The input is video data from the camera and audio data from the microphone. The output is to transfer this data to the server. Specifically, the camera captures the visitor's video and the microphone records the audio. The video data is then compressed into H.264 format or similar, and the audio data is converted into WAV format or similar before being sent to the server using the TCP / IP protocol.
[0450] Step 2:
[0451] Data Preprocessing
[0452] The server removes noise from the received video and audio data and converts the audio data to text. The input is the received video and audio data. The output is noise-removed video data, audio data, and text data. Specifically, the OpenCV library is used to remove noise from the video data, and filtering technology is used to clean the audio data. A speech recognition API (e.g., Google Cloud Speech-to-Text API) is used to convert the audio to text.
[0453] Step 3:
[0454] Feature extraction and quantification
[0455] The server uses natural language processing technology to analyze the phrasing and content of the conversation from the speech-recognized text data, analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. The inputs include cleaned video data, audio data, and text data. The output is quantified feature data. Specifically, NLTK or spaCy is used for natural language processing, and the LibROSA library is used to analyze speaking style and tone of voice from the audio data. Facial expression recognition and body language analysis are performed from the video data using OpenCV and Dlib libraries.
[0456] Step 4:
[0457] Analysis by emotion engine
[0458] The server uses the emotion_recognition library to analyze emotions from video and audio data. The input is quantified feature data. The output is detected emotion data. Specific operations include inferring emotions from facial expressions in the video, tone of voice, and linguistic expressions in the text. For example, by combining OpenCV and TensorFlow, emotions such as joy, anger, and surprise can be identified in real time based on facial expressions.
[0459] Step 5:
[0460] Issuance of an alert
[0461] If the server detects suspicious behavior or abnormal emotional fluctuations, it will issue an alarm through the alarm system. The inputs are detected emotional data and behavioral data. The output is an alarm message. Specifically, if a certain emotional score exceeds a threshold, the alarm system will be activated and an audio alert or text notification will be sent to security staff.
[0462] Step 6:
[0463] Report Generation
[0464] The server generates a report based on the analysis results and provides it to the user. The input is quantified and analyzed feature data. The output is a daily or monthly report. Specifically, it aggregates the analysis results stored in the database and automatically generates a report according to a template. The generated report is saved in PDF or CSV format and provided to the user via email or dashboard.
[0465] Example prompt sentence:
[0466] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[0467] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0468] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0469] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0470] [Second embodiment]
[0471] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0472] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0473] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0474] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0475] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0477] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0478] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0479] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0480] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0481] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0482] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0483] System Overview
[0484] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes that record the video and audio data of the interview, analyze it, and provide the results as feedback. This eliminates human subjectivity and bias, enabling fair evaluation.
[0485] System configuration
[0486] The system consists of the following main components:
[0487] 1. Recording device: A device that records the video and audio data of the interview.
[0488] 2. Server: A central computer that receives and analyzes the recorded data.
[0489] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[0490] 4. Feedback terminal: A device for displaying the analysis results to the user.
[0491] System Operation
[0492] 1. Data Collection and Transfer
[0493] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[0494] 2. Data Preprocessing
[0495] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[0496] 3. Feature extraction and quantification
[0497] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0498] 4. Training and evaluation of the learning model
[0499] The server trains a machine learning model using interview data from existing employees. This trained model is then used to evaluate newly acquired interview data and calculate various scores. For example, the server evaluates overall performance based on factors such as the consistency of language, conversational content, and facial expressions.
[0500] 5. Presentation of results
[0501] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) can check the evaluation results and receive feedback through this terminal. The evaluation results include an overall evaluation and detailed scores for each evaluation item.
[0502] Specific examples
[0503] For example, when an interview is conducted with interviewer A, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[0504] Noise is removed from video and audio data to extract high-quality data.
[0505] Converts audio data into text.
[0506] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[0507] Analyzes voice tone, rhythm, and tone from audio data.
[0508] Extracts changes in facial expressions and body language from video data.
[0509] As a result, an overall score of 85 / 100 is calculated for Interviewer A, and detailed evaluations are given, with 90 / 100 for language and 80 / 100 for tone of voice. These results are presented to the user via the feedback terminal.
[0510] This system will enable interview evaluation to be conducted objectively and efficiently, and is expected to be a significant improvement over conventional subjective evaluation methods.
[0511] The processing flow will be explained below.
[0512] Step 1:
[0513] The user (interviewer) starts the interview, and the device activates the camera and microphone to record the video and audio data of the interview.
[0514] Step 2:
[0515] The device sends the recorded data to the server in real time or after the interview is over. The sent data is in the form of video and audio files.
[0516] Step 3:
[0517] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[0518] Step 4:
[0519] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the speaker's voice is transcribed sequentially as text.
[0520] Step 5:
[0521] The server then analyzes the converted text data using natural language processing (NLP) technology, specifically evaluating the language used (for example, frequency of honorific use, use of technical terms), the grammatical structure of the conversation, and logic.
[0522] Step 6:
[0523] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[0524] Step 7:
[0525] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[0526] Step 8:
[0527] The server quantifies each extracted feature and assigns a score according to a standardized evaluation scale, such as 90 / 100 for appropriate language and 85 / 100 for friendly facial expressions.
[0528] Step 9:
[0529] The server uses interview data from existing employees to train the learning model, which uses machine learning algorithms to learn evaluation criteria that fit the company culture.
[0530] Step 10:
[0531] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[0532] Step 11:
[0533] The server sums up the scores for each evaluation item to generate an overall evaluation, which is then sent to the device.
[0534] Step 12:
[0535] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the overall evaluation might be 85 / 100, language 90 / 100, and facial expression 85 / 100.
[0536] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias and helping to select the right candidates.
[0537] Example 1
[0538] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0539] Conventional interview evaluation systems have the problem that they are prone to human subjectivity and bias, making it difficult to achieve consistent and fair evaluations. Furthermore, the evaluation process is inefficient, placing a heavy burden on interviewers. Furthermore, conventional systems have difficulty providing real-time feedback, making it impossible to immediately identify areas for improvement.
[0540] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0541] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to a central processing unit, means for removing noise from the transmitted video and audio data and converting the audio into text information, means for analyzing the phrasing and conversation content of the audio data using language analysis technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and sign language from the video data, means for quantifying and evaluating each analyzed feature, means for training a machine learning model using interview data of existing employees, means for evaluating new interview data using the machine learning model, and means for displaying the evaluation results and providing feedback to the user. This improves the objectivity and fairness of interview evaluations, reduces the burden on interviewers, and enables feedback to be provided in real time.
[0542] A "recording terminal" is a device for recording video and audio data of an interview.
[0543] A "central processing unit" is a computer or server that receives and analyzes the recorded data.
[0544] "Noise reduction" is the process of removing unnecessary noise from recorded audio and video data.
[0545] "Speech-to-text" refers to the process of converting recorded voice data into text data.
[0546] "Language analysis technology" is a technology that performs natural language processing on text data and analyzes wording and conversation content.
[0547] "Analyzing speaking style and tone of voice" refers to analyzing characteristics such as speaking rate, pitch, and tone based on the audio data.
[0548] "Analyzing facial expressions and sign language" means analyzing the interviewer's facial expressions and body language based on video data.
[0549] "Quantification" means expressing the analyzed features as numerical data.
[0550] A "machine learning model" is an algorithm that learns patterns based on past data and makes predictions and classifications.
[0551] "Providing feedback" means presenting the analysis and evaluation results to the user and informing them of areas for improvement and strengths.
[0552] "Setting evaluation criteria that fit the organizational culture" means setting evaluation criteria that are appropriate for the characteristics and values of the organization based on interview data from existing employees.
[0553] "Immediate analysis and display" means that interview data is processed in real time and the results are displayed immediately.
[0554] MODE FOR CARRYING OUT THE INVENTION
[0555] The present invention provides a system for objectively and efficiently evaluating interviews. This system includes a process for recording video and audio data of the interview, analyzing the data with a central processing unit, and providing the results as feedback to the user.
[0556] The system consists of the following main components:
[0557] 1. Recording device: This is a device that records the video and audio data of the interview. In this case, a standard webcam and microphone are used.
[0558] 2. Central Processing Unit (Server): A computer that receives and analyzes recorded data. A general cloud server or on-premise server can be used for high-performance computing.
[0559] 3. Analysis module: This is software that runs inside the central processing unit and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specifically, it uses software such as FFmpeg, Google Speech-to-Text API, NLTK, Spacy, Praat, OpenCV, Dlib, Scikit-learn, and TensorFlow.
[0560] 4. Feedback terminal: A device that displays analysis results to users and provides feedback. In this case, this applies to general PCs, tablets, smartphones, etc.
[0561] (Data Collection and Transfer)
[0562] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When the recording is finished, the recording device sends this data to the central processing unit. The transmission process can use storage services such as Google Drive API or AWS S3.
[0563] (Data preprocessing)
[0564] The server uses FFmpeg to filter the received recording data to remove noise, and then uses the Google Speech-to-Text API to convert the audio data into text.
[0565] (Feature extraction and quantification)
[0566] When the server receives the speech-recognized text data, it uses natural language processing technology to analyze the vocabulary and content of the conversation. This analysis is performed using NLTK and Spacy. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points.
[0567] (Training and evaluating learning models)
[0568] The server uses the feature data to train a machine learning model. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. The trained model is then used to evaluate new interview data and calculate an overall score and detailed scores for each evaluation item.
[0569] (Presentation of results)
[0570] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript as a dashboard operated by the user. The user (interviewer or interviewee) uses this feedback terminal to check the evaluation results and identify areas for improvement and strengths.
[0571] Specific examples
[0572] For example, when an interview with Interviewer A is conducted, the entire process is recorded and the recorded data is sent to the server. The server removes noise using FFmpeg and converts the audio data into text using the Google Speech-to-Text API. Next, it analyzes the text data using NLTK or Spacy, analyzes the audio data using Praat, and extracts facial expressions from the video data using OpenCV and Dlib. These results are quantified, and each evaluation item is scored using a model trained with Scikit-learn or TensorFlow. Finally, these results are provided to the user in a form that can be viewed on a feedback terminal. The evaluation results include detailed information such as Interviewer A's overall score of 85 / 100, language score of 90 / 100, and tone of voice score of 80 / 100.
[0573] Prompt Sentence Examples
[0574] Below are some example prompts to input to a generative AI model:
[0575] "Please explain in detail, step by step, the process of the program that analyzes the interview recording data and conducts a comprehensive evaluation based on the interviewer's language, tone of voice, facial expressions, and sign language."
[0576] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0577] Step 1: Data collection and transfer
[0578] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When recording is finished, the recording device sends the recorded video and audio data to the server. This sending process uses storage services such as Google Drive API and AWS S3. The input is the video and audio data of the interview, and the output is the recorded data sent to the server.
[0579] Step 2: Preprocessing the data
[0580] The recording data received by the server is first denoised. Specifically, FFmpeg is used to filter the audio data and remove static noise. Next, the Google Speech-to-Text API is used to convert the audio data into text. This process converts the audio data into text information. The input is the recording data sent to the server, and the output is the noise-denoised audio data and text data.
[0581] Step 3: Feature extraction and quantification
[0582] The server uses natural language processing technology to analyze the phrasing and conversation content of the speech-recognized text data. NLTK and Spacy are used for this analysis. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points. The input is noise-removed audio data and text data, and the output is quantified feature data.
[0583] Step 4: Training and evaluating the learning model
[0584] The server trains a machine learning model using the feature data it has acquired so far. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. Next, it uses the trained model to evaluate newly acquired interview data and calculates an overall score and detailed scores for each evaluation item. The input is quantified feature data, and the output is an evaluation score.
[0585] Step 5: Presenting the results
[0586] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript to build a dashboard. The user (interviewer or interviewee) can check the evaluation results using this feedback terminal. The evaluation results include an overall score and detailed evaluation information for each evaluation item. The input is the evaluation score, and the output is the visualized evaluation results.
[0587] (Application example 1)
[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] In conventional interview evaluation systems, evaluations tend to be subjective and often result in bias. Furthermore, in factories and other workplaces, it is difficult to objectively evaluate worker performance and provide efficient feedback. Therefore, a new system is needed that can accurately and quickly evaluate worker performance in workplaces where improvements in work efficiency and safety are required.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0591] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using existing data, means for evaluating new data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for evaluating work efficiency, accuracy, safety, etc. based on the analysis results, and means for feeding back the scoring results to the worker and suggesting areas for improvement. This makes it possible to objectively and efficiently evaluate worker performance and suggest areas for improvement even in factories and on-site workplaces.
[0592] Definitions of important words
[0593] "Video Data" means visual information captured by a camera or other recording device.
[0594] "Audio Data" means audio information captured by a microphone or other recording device.
[0595] "Noise reduction" is the process of removing unwanted background noise and artifacts from video and audio recordings.
[0596] "Text conversion" is the process of converting audio data into written information.
[0597] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[0598] "Language" refers to the choice of words and expressions used in speech.
[0599] "Conversation content" refers to specific topics and information contained in the voice data.
[0600] "Speaking style" refers to characteristics of the audio data such as pronunciation, intonation, and rhythm.
[0601] "Voice tone" refers to voice characteristics including the pitch, strength, rhythm, etc. of the voice in the voice data.
[0602] "Facial expressions" refer to facial changes and emotional expressions in video data.
[0603] "Body language" refers to the body movements and postures in video data.
[0604] "Scoring" is the process of quantifying and evaluating the analyzed features.
[0605] A "learning model" is an algorithm that learns from data to identify and predict specific patterns.
[0606] "Feedback" is the process of providing the evaluation results to the user and indicating areas for improvement and evaluation.
[0607] "Work efficiency" is an index that indicates how efficiently work is performed within a certain period of time.
[0608] "Accuracy" is an indicator of how accurately a task is performed.
[0609] "Safety" is an indicator of how safely the work was carried out.
[0610] MODE FOR CARRYING OUT THE INVENTION
[0611] 1. System Overview
[0612] This invention provides a system for objectively evaluating the performance of workers in a factory or other workplace. The system mainly consists of the following components:
[0613] Recording device
[0614] server
[0615] Analysis Module
[0616] Feedback Terminal
[0617] 2. Hardware and Software Configuration
[0618] Recording device
[0619] The recording terminal is equipped with a camera and microphone and is a device that records the worker's video and audio data in high quality. In this example, a general web camera and microphone are used.
[0620] server
[0621] The server acts as a central computer that receives and processes the recorded data. The following software is installed on the server:
[0622] OpenCV (video data analysis)
[0623] speech_recognition (converts speech data to text)
[0624] scikit-learn (machine learning model training and evaluation)
[0625] 3. Program Processing
[0626] Data Collection and Transfer
[0627] The recording terminal records the video and audio data of the worker's work and sends the data to the server. The recorded data is pre-processed to remove noise, and the audio data is converted into text.
[0628] Data analysis
[0629] The server applies natural language processing technology to the audio data to analyze the language used and the content of the conversation. It analyzes facial expressions and body language from the video data, and extracts speaking style and tone of voice from the audio data. These features are then quantified and scored.
[0630] Feature extraction and quantification
[0631] The server quantifies the characteristics obtained from the recorded data and scores it based on specific criteria, including work efficiency, accuracy, and safety.
[0632] Training and evaluating learning models
[0633] The server trains a learning model using existing data and evaluates newly acquired data. This process uses a machine learning model (e.g., SVM) using the scikit-learn library.
[0634] Feedback of results
[0635] The analysis results and the evaluated scoring results are sent to a feedback terminal where users can check the results, allowing workers to receive an evaluation of their performance and learn areas for improvement.
[0636] 4. Examples of concrete examples and prompts
[0637] Specific examples
[0638] For example, Worker A's work process for one day is recorded and an evaluation is made based on that data. The recorded data is converted into text using voice recognition, and the work movements are evaluated through video analysis. Finally, an overall evaluation score is obtained, and this score is fed back to the worker to indicate areas for improvement.
[0639] Prompt Sentence Examples
[0640] "Please create prompts for a scoring system that will record Worker A's work for one day, analyze the video and audio, and evaluate work efficiency, accuracy, safety, etc. Please include specific steps from data preprocessing to evaluation."
[0641] This invention makes it possible to make objective evaluations that were difficult to achieve with conventional methods, and is expected to improve work efficiency and safety.
[0642] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0643] Program processing steps
[0644] Step 1:
[0645] The recording terminal records the video and audio data of the worker's work in real time using a camera and microphone, and saves this data as a file.
[0646] Input: Video and audio data of the worker
[0647] Output: Recording files (video files, audio files)
[0648] Step 2:
[0649] The recording device sends the recorded video and audio data to a server, where the data is centrally managed and analyzed.
[0650] Input: Recording file
[0651] Output: Data transferred to the server
[0652] Step 3:
[0653] The server performs noise reduction on the received video and audio data using filtering technology to improve the quality of the data.
[0654] Input: Transferred recording data (video files, audio files)
[0655] Output: High-quality data after noise removal
[0656] Step 4:
[0657] The server uses speech recognition technology to convert the audio data into text, specifically the speech_recognition library.
[0658] Input: High-quality audio data
[0659] Output: Text data
[0660] Step 5:
[0661] The server uses natural language processing technology to analyze the vocabulary and content of the conversation, thereby understanding the meaning and context of the text data.
[0662] Input: Text data
[0663] Output: Analysis results (characteristics of language and conversation content)
[0664] Step 6:
[0665] The server analyzes the speech data to determine speaking style and tone, and extracts features such as pitch, rhythm, and tone from the speech waveform data.
[0666] Input: High-quality audio data
[0667] Output: Audio feature data
[0668] Step 7:
[0669] The server analyzes facial expressions and body language from the video data, detecting facial expressions and body movements using OpenCV.
[0670] Input: High-quality video data
[0671] Output: Video feature data
[0672] Step 8:
[0673] The server quantifies and scores each analyzed feature, calculates the score for each feature, and performs an overall performance evaluation.
[0674] Input: Language feature data, audio feature data, video feature data
[0675] Output: Scoring results
[0676] Step 9:
[0677] The server trains a learning model using existing data, using scikit-learn for this process and optimizing the model based on existing evaluation data.
[0678] Input: Existing evaluation data
[0679] Output: A trained model
[0680] Step 10:
[0681] The server evaluates newly acquired data using the learning model, inputting new data to the trained model and obtaining the evaluation results.
[0682] Input: Newly acquired feature data
[0683] Output: Evaluation results
[0684] Step 11:
[0685] The server transmits the evaluation results to the feedback terminal, which displays the evaluation results so that the user can check them.
[0686] Input: Evaluation result (scoring result)
[0687] Output: Feedback to terminal
[0688] Step 12:
[0689] Users receive the evaluation results through a feedback terminal and identify areas for improvement, which helps workers improve their own performance.
[0690] Input: Data to be displayed on the feedback terminal
[0691] Output: Feedback information to the user
[0692] These steps allow for objective and efficient evaluation and feedback of worker performance.
[0693] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0694] System Overview
[0695] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to extract emotional changes during the interview and incorporate them into a comprehensive evaluation.
[0696] System configuration
[0697] The system consists of the following main components:
[0698] 1. Recording device: A device that records the video and audio data of the interview.
[0699] 2. Server: A central computer that receives and analyzes the recorded data.
[0700] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[0701] 4. Emotion engine: Software that runs inside the server and analyzes the user's emotions from video and audio data.
[0702] 5. Feedback terminal: A device for displaying the analysis results to the user.
[0703] System Operation
[0704] 1. Data Collection and Transfer
[0705] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[0706] 2. Data Preprocessing
[0707] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[0708] 3. Feature extraction and quantification
[0709] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0710] 4. Analysis by Emotion Engine
[0711] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[0712] 5. Quantifying and Scoring Emotional Data
[0713] The server quantifies the emotional change data obtained from the emotion engine and scores it in the same way as other evaluation items. Evaluation criteria include emotional stability and appropriate emotional expression.
[0714] 6. Training and Evaluation of the Learning Model
[0715] The server trains a machine learning model using interview data from existing employees. The model uses various features, including emotional data, to learn evaluation criteria that fit the company culture. When new interview data is input, the trained model is used to perform an overall evaluation.
[0716] 7. Presentation of results
[0717] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) checks the evaluation results through this terminal and receives feedback. The evaluation results include an overall evaluation including an emotional evaluation and detailed scores for each evaluation item.
[0718] Specific examples
[0719] For example, when an interview with Interviewer B is conducted, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[0720] Noise is removed from video and audio data to extract high-quality data.
[0721] Converts audio data into text.
[0722] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[0723] Analyzes voice tone, rhythm, and tone from audio data.
[0724] Extracts changes in facial expressions and body language from video data.
[0725] An emotion engine is used to detect various emotions (happiness, surprise, confusion, etc.).
[0726] Each feature and emotion data is quantified and scored.
[0727] As a result, an overall score of 90 / 100 is calculated for Interviewer B, and detailed evaluations are given, including 85 / 100 for language, 82 / 100 for tone of voice, and 88 / 100 for emotional stability. These results are presented to the user via a feedback terminal.
[0728] This system allows for objective and efficient evaluation of interviews, and is expected to be a significant improvement over conventional subjective evaluation methods. In addition, by combining it with an emotion engine, changes in the interviewer's emotions can be reflected in the evaluation, enabling a more comprehensive evaluation.
[0729] The processing flow will be explained below.
[0730] Step 1:
[0731] The user (interviewer) starts the interview, and the recording terminal activates the camera and microphone to record the video and audio data of the interview.
[0732] Step 2:
[0733] The recording device sends the recorded data to the server in real time or after the interview ends. The sent data is in the form of video and audio files.
[0734] Step 3:
[0735] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[0736] Step 4:
[0737] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the voice is transcribed sequentially.
[0738] Step 5:
[0739] The server then analyzes the converted text data using natural language processing (NLP) techniques, specifically evaluating the language used (e.g., frequency of honorific use and use of technical terms), the grammatical structure of the conversation, coherence, and logic.
[0740] Step 6:
[0741] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[0742] Step 7:
[0743] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[0744] Step 8:
[0745] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[0746] Step 9:
[0747] The server quantifies the emotional change data obtained from the emotion engine and scores it according to standardized evaluation criteria, such as emotional stability and appropriate emotional expression.
[0748] Step 10:
[0749] The server quantifies and scores each extracted feature (language, speaking style, facial expressions, tone of voice, body language, emotional changes, etc.). For example, appropriate language is given a score of 90 / 100, and friendly facial expressions are given a score of 85 / 100.
[0750] Step 11:
[0751] The server uses interview data from existing employees to train a machine learning model, which uses a machine learning algorithm to learn evaluation criteria that fit the company culture.
[0752] Step 12:
[0753] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[0754] Step 13:
[0755] The server sums up the scores for each evaluation item to generate an overall evaluation, and sends the evaluation result to the terminal.
[0756] Step 14:
[0757] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the user can see detailed scores of 90 / 100 for overall evaluation, 85 / 100 for language, 82 / 100 for facial expressions, and 88 / 100 for emotional stability.
[0758] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias to help select the right candidates. Furthermore, by incorporating an emotion engine, changes in the interviewer's emotions are reflected in the evaluation, enabling a more comprehensive evaluation.
[0759] Example 2
[0760] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0761] Conventional interview evaluation systems often rely on the subjective judgment of the interviewer, resulting in problems such as a lack of fairness and consistency. Furthermore, it is difficult to perform a comprehensive evaluation that includes the interviewer's emotional changes, making it impossible to reflect emotional stability or appropriate emotional expression in the evaluation. Furthermore, it is difficult to perform these analyses in real time and provide feedback immediately after the interview.
[0762] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0763] In this invention, the server includes a means for recording video and audio data of the interview, a means for transmitting the recorded video and audio data to a central control device, a means for performing noise reduction and speech recognition on the transmitted video and audio data, a means for analyzing the phrasing and conversation content of the speech recognition results using natural language analysis technology, a means for analyzing speaking style and tone of voice from the audio data, a means for analyzing facial expressions and body movements from the video data, a means for quantifying and scoring each analyzed feature, a means for analyzing emotions from the recorded video and audio data, a means for quantifying and scoring the emotion data, a means for training a machine learning model using interview data of existing employees, a means for evaluating new interview data using the learning model, and a means for displaying the evaluation results and providing feedback to the user. This enables objective and efficient interview evaluation, and also allows the interviewer's emotional changes to be reflected in the evaluation. Analysis results can also be provided in real time, allowing feedback to be received immediately after the interview.
[0764] "Video and audio data of the interview" refers to the visual and audio information recorded during the interview.
[0765] "Recording means" refers to technical devices and methods for recording video and audio data.
[0766] "Central control unit" refers to the main computer system that centrally manages and processes multiple data.
[0767] "Transmission means" refers to the technical devices and methods for transferring data from one point to another.
[0768] "Noise reduction" refers to the process of removing unwanted noise and interference from video and audio data.
[0769] "Speech recognition" refers to the technology of analyzing voice data and converting it into text information.
[0770] "Natural language analysis technology" refers to analysis technology that enables computers to understand and interpret human language.
[0771] "Means for analyzing speaking style and tone of voice" refers to technology that analyzes a speaker's speech patterns and voice quality from audio data.
[0772] "Means for analyzing facial expressions and body movements" refers to technology that detects and analyzes changes in facial expressions and gestures of subjects from video data.
[0773] "Quantification" refers to the act of expressing analyzed information as quantitative data.
[0774] "Scoring" refers to the process of assigning a score based on quantified data according to evaluation criteria.
[0775] "Means for analyzing emotions" refers to technology that estimates emotional states from video and audio data.
[0776] A "machine learning model" refers to an algorithm that learns patterns from large amounts of data and makes inferences and predictions about new data.
[0777] "Means for displaying the evaluation results and providing feedback to the user" refers to technical devices and methods for visually presenting the analysis results to the user and facilitating their understanding.
[0778] This invention is a system for conducting interview evaluations objectively and efficiently, and includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. The system incorporates specific hardware and software, with the user, terminal, and server playing their respective roles.
[0779] Key Components of the System
[0780] 1. Recording device: A device for recording the video and audio data of the interview. Generally, a device equipped with a camera and microphone is used.
[0781] 2. Server: A central control unit that receives and analyzes recorded data. This device is a high-performance computer system that can process and analyze data at high speed. For example, a commonly used corporate server system can be used.
[0782] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. It uses deep learning libraries such as TensorFlow and PyTorch.
[0783] 4. Emotion engine: Software that analyzes emotions from video and audio data. Typically, a cloud-based API (e.g., Microsoft Azure's Emotion Recognition API) is used.
[0784] 5. Feedback terminal: A device used to display analysis results to users. This is typically a PC or mobile device.
[0785] Data collection and analysis flow
[0786] When the user (interviewer) starts the interview, the recording device records the entire process. After recording is finished, the device sends the recorded video and audio data to the server. The server performs noise reduction and speech recognition, and converts the audio data into text. For example, noise reduction is performed using FFmpeg, and the audio is converted to text information using the Google Cloud Speech-to-Text API.
[0787] For text data, natural language analysis techniques are used to analyze the phrasing and content of conversations. Natural language processing tools such as the BERT model are used. For audio data, speech analysis tools such as Praat are used to analyze speaking style and tone of voice. In addition, the server uses image analysis libraries such as OpenCV to analyze facial expressions and body language from video data.
[0788] Sentiment Analysis and Scoring
[0789] The emotion engine analyzes emotions from video and audio data to extract emotions such as joy, anger, surprise, and sadness, using Microsoft Azure's emotion recognition API.
[0790] The analyzed emotion data is quantified and scored by the server, using evaluation indices based on the intensity and frequency of each emotion.
[0791] The server uses interview data from existing employees to train a machine learning model, which uses tools like Scikit-learn to learn evaluation criteria that fit the company culture.
[0792] Presentation of evaluation results
[0793] Finally, when new interview data is input, the server uses the trained model to perform a comprehensive evaluation, and the generated evaluation results are sent to the feedback terminal, where users can check the evaluation results.
[0794] Examples and prompts
[0795] For example, if an interview with Interviewer B is conducted, the analysis will proceed as follows:
[0796] The recording device captures video and audio data.
[0797] The device sends data to the server.
[0798] The server uses FFmpeg to remove noise.
[0799] The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[0800] The server analyzes the text data using the BERT model.
[0801] The server uses Praat to extract audio features.
[0802] The server analyzes the video data using OpenCV.
[0803] The emotion engine analyzes emotions using Microsoft Azure's emotion recognition API.
[0804] The server quantifies the emotional data and scores it.
[0805] The server trains the machine learning model using Scikit-learn.
[0806] When new interview data is input, the server evaluates it using the trained model.
[0807] Evaluation results are presented through a feedback terminal.
[0808] Prompt Sentence Examples
[0809] Please provide a detailed description of each processing step in your interview evaluation system, including the specific actions and techniques used from data collection to providing results.
[0810] This system allows for objective and efficient interview evaluation. Furthermore, by using an emotion engine, the interviewer's emotional changes are incorporated into the evaluation, enabling a more comprehensive evaluation. Evaluation results are presented in real time, allowing users to receive immediate feedback.
[0811] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0812] Step 1:
[0813] The user (interviewer) starts the interview. Recording of audio and video data begins. The input is the video and audio during the interview, and the output is the video and audio data recorded on the recording terminal.
[0814] Step 2:
[0815] The device sends the recorded data to the server. After recording is complete, the recorded video and audio data is transferred to the server. The input is the video and audio data stored on the recording device, and the output is the same data sent to the server. Specifically, a secure data transfer protocol (such as SFTP) is used.
[0816] Step 3:
[0817] The server performs noise reduction. The server removes background noise from the received video and audio data. The input is the raw video and audio data sent to the server, and the output is the denoised video and audio data. FFmpeg noise reduction filters are used.
[0818] Step 4:
[0819] The server performs speech recognition. The server converts noise-removed speech data into text. The input is noise-removed speech data, and the output is text data. The Google Cloud Speech-to-Text API is used for speech recognition.
[0820] Step 5:
[0821] The server analyzes the text data. The server uses natural language analysis technology to analyze the transcribed text data. The input is text data generated by speech recognition, and the output is feature data of the analyzed phrasing and conversation content. The BERT model, for example, is used here.
[0822] Step 6:
[0823] The server extracts voice features. The server analyzes speaking style and tone of voice from the voice data. The input is the voice data after noise removal, and the output is feature data related to speaking style and tone of voice. Specifically, the voice analysis is performed using Praat.
[0824] Step 7:
[0825] The server analyzes the video data. The server analyzes facial expressions and body language from the video data. The input is the video data after noise removal, and the output is feature data related to facial expressions and body language. Image analysis libraries such as OpenCV are used.
[0826] Step 8:
[0827] The emotion engine analyzes emotions. The emotion engine extracts multiple emotional states from video and audio data. The input is noise-removed video and audio data, and the output is quantified data for each emotion. Microsoft Azure's emotion recognition API is used.
[0828] Step 9:
[0829] The server quantifies the analysis data and scores it. The server quantifies the acquired feature data and emotion data and displays it as a score. The input is feature data such as speaking style, tone of voice, facial expressions, body language, and emotions, and the output is a quantified evaluation score.
[0830] Step 10:
[0831] The server trains a machine learning model. The server trains the model using interview data of existing employees. The input is the interview data of existing employees, and the output is the trained machine learning model. Libraries such as Scikit-learn are used.
[0832] Step 11:
[0833] The server evaluates new interview data. The server evaluates new interview data using the trained model. The input is the newly collected interview data and the trained model, and the output is an overall evaluation score.
[0834] Step 12:
[0835] The server sends the evaluation result to the feedback terminal. The server sends the generated evaluation result to the feedback terminal. The input is the calculated evaluation score, and the output is the evaluation result displayed on the feedback terminal.
[0836] Step 13:
[0837] The user checks the evaluation results. The user checks the analysis results and the scores for each evaluation item through the feedback terminal. The input is the evaluation results displayed on the feedback terminal, and the output is the user's understanding of the evaluation results and feedback.
[0838] This series of processes enables interview evaluation to be conducted objectively and efficiently, and also enables comprehensive evaluation including emotional stability and appropriate expression.
[0839] (Application example 2)
[0840] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0841] In recent years, there has been a demand for strengthened security in offices and homes, but current security systems have difficulty grasping changes in visitors' emotions and behavior in real time. In particular, there is a lack of an objective evaluation system for quickly responding to suspicious behavior or abnormal emotional fluctuations. This can lead to potential risks being overlooked, making it difficult to implement effective security measures.
[0842] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using interview data of existing employees, means for evaluating new interview data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for detecting abnormal emotional fluctuations and suspicious behavior and issuing an alert, and means for generating a report based on the analysis results. This makes it possible to analyze the emotions and behavior of visitors in real time, quickly detect suspicious behavior, and issue an alert.
[0843] "Video data" refers to image information captured using a camera.
[0844] "Audio data" refers to audio information recorded using a microphone.
[0845] "Recording" refers to the act of saving video and audio data.
[0846] "Server" refers to a central computer that collects, stores, and analyzes data over a network.
[0847] "Noise reduction" refers to the method of removing unwanted information or interference from video and audio data.
[0848] "Text conversion" refers to the technology of converting voice data into text information.
[0849] "Natural language processing technology" refers to algorithms and methods for analyzing text data.
[0850] "Language" refers to the way a character speaks and the words they choose to use.
[0851] "Vocal tone" refers to characteristics such as the pitch, rhythm, and tone of a speaker's voice.
[0852] "Facial expression" refers to changes in emotions shown by the movement of facial muscles.
[0853] "Body language" refers to non-verbal communication expressed through bodily movements and posture.
[0854] "Quantification" refers to the act of expressing analyzed features as numerical data.
[0855] "Scoring" refers to the method of calculating an evaluation score based on numerical characteristics.
[0856] A "learning model" refers to a computational model that learns patterns from large amounts of data and performs predictions and classifications.
[0857] "Issuing an alarm" refers to the act of sending a notification to draw attention when the system determines that there is an abnormality.
[0858] "Report generation" refers to the act of creating a report summarizing the analysis results.
[0859] System Overview
[0860] This invention is a system that can be applied to a security system to analyze visitors' emotions and behavior in real time and detect suspicious behavior or abnormal emotional fluctuations. This system aims to improve security by analyzing video and audio data. It also issues an alarm and generates a report based on the analysis results.
[0861] System configuration
[0862] The system consists of the following main components:
[0863] 1. Recording device: A device that records video and audio data of visitors. Specifically, it uses a camera (e.g., Logitech C922 Pro Stream) and a microphone (e.g., Rode NT-USB).
[0864] 2. Server: A central computer that receives and analyzes the recorded data.
[0865] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specific examples of this software include OpenCV and the emotion_recognition library.
[0866] 4. Feedback terminal: A device for displaying the analysis results to the user.
[0867] 5. Alarm system: A device that issues an alarm if it detects abnormal emotional fluctuations or suspicious behavior.
[0868] 6. Report generation module: Software for generating reports based on the analysis results.
[0869] System Operation
[0870] 1. Data Collection and Transfer
[0871] The user starts the system, and the recording device records the video and audio of the visitor in real time. The recorded data is immediately sent to the server.
[0872] 2. Data Preprocessing
[0873] The server removes noise from the received data and converts the audio data into text. The OpenCV library is used for noise removal, and speech recognition technology is used for converting the audio data into text.
[0874] 3. Feature extraction and quantification
[0875] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0876] 4. Analysis by Emotion Engine
[0877] The server uses the emotion_recognition library to analyze emotions from video and audio data, detecting emotional changes in real time and identifying abnormal emotional fluctuations and suspicious behavior.
[0878] 5. Alert generation and report generation
[0879] If suspicious behavior or abnormal emotional changes are detected, the server will issue an alert through the alarm system. Furthermore, it will periodically generate reports based on the analysis results and provide them to the user.
[0880] Specific examples
[0881] For example, if a visitor shows signs of anger or fear at an office reception, the system can detect this in real time and immediately issue an alert, while a report summarizing the day's events is sent to the manager.
[0882] Example prompt sentence:
[0883] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[0884] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0885] Step 1:
[0886] Data Collection and Transfer
[0887] The user starts the system, and the recording terminal records the visitor's video and audio data in real time. The input is video data from the camera and audio data from the microphone. The output is to transfer this data to the server. Specifically, the camera captures the visitor's video and the microphone records the audio. The video data is then compressed into H.264 format or similar, and the audio data is converted into WAV format or similar before being sent to the server using the TCP / IP protocol.
[0888] Step 2:
[0889] Data Preprocessing
[0890] The server removes noise from the received video and audio data and converts the audio data to text. The input is the received video and audio data. The output is noise-removed video data, audio data, and text data. Specifically, the OpenCV library is used to remove noise from the video data, and filtering technology is used to clean the audio data. A speech recognition API (e.g., Google Cloud Speech-to-Text API) is used to convert the audio to text.
[0891] Step 3:
[0892] Feature extraction and quantification
[0893] The server uses natural language processing technology to analyze the phrasing and content of the conversation from the speech-recognized text data, analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. The inputs include cleaned video data, audio data, and text data. The output is quantified feature data. Specifically, NLTK or spaCy is used for natural language processing, and the LibROSA library is used to analyze speaking style and tone of voice from the audio data. Facial expression recognition and body language analysis are performed from the video data using OpenCV and Dlib libraries.
[0894] Step 4:
[0895] Analysis by emotion engine
[0896] The server uses the emotion_recognition library to analyze emotions from video and audio data. The input is quantified feature data. The output is detected emotion data. Specific operations include inferring emotions from facial expressions in the video, tone of voice, and linguistic expressions in the text. For example, by combining OpenCV and TensorFlow, emotions such as joy, anger, and surprise can be identified in real time based on facial expressions.
[0897] Step 5:
[0898] Issuance of an alert
[0899] If the server detects suspicious behavior or abnormal emotional fluctuations, it will issue an alarm through the alarm system. The inputs are detected emotional data and behavioral data. The output is an alarm message. Specifically, if a certain emotional score exceeds a threshold, the alarm system will be activated and an audio alert or text notification will be sent to security staff.
[0900] Step 6:
[0901] Report Generation
[0902] The server generates a report based on the analysis results and provides it to the user. The input is quantified and analyzed feature data. The output is a daily or monthly report. Specifically, it aggregates the analysis results stored in the database and automatically generates a report according to a template. The generated report is saved in PDF or CSV format and provided to the user via email or dashboard.
[0903] Example prompt sentence:
[0904] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[0905] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0906] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0907] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0908] [Third embodiment]
[0909] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0910] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0911] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0912] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0913] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0914] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0915] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0916] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0917] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0918] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0919] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0920] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0921] System Overview
[0922] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes that record the video and audio data of the interview, analyze it, and provide the results as feedback. This eliminates human subjectivity and bias, enabling fair evaluation.
[0923] System configuration
[0924] The system consists of the following main components:
[0925] 1. Recording device: A device that records the video and audio data of the interview.
[0926] 2. Server: A central computer that receives and analyzes the recorded data.
[0927] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[0928] 4. Feedback terminal: A device for displaying the analysis results to the user.
[0929] System Operation
[0930] 1. Data Collection and Transfer
[0931] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[0932] 2. Data Preprocessing
[0933] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[0934] 3. Feature extraction and quantification
[0935] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[0936] 4. Training and evaluation of the learning model
[0937] The server trains a machine learning model using interview data from existing employees. This trained model is then used to evaluate newly acquired interview data and calculate various scores. For example, the server evaluates overall performance based on factors such as the consistency of language, conversational content, and facial expressions.
[0938] 5. Presentation of results
[0939] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) can check the evaluation results and receive feedback through this terminal. The evaluation results include an overall evaluation and detailed scores for each evaluation item.
[0940] Specific examples
[0941] For example, when an interview is conducted with interviewer A, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[0942] Noise is removed from video and audio data to extract high-quality data.
[0943] Converts audio data into text.
[0944] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[0945] Analyzes voice tone, rhythm, and tone from audio data.
[0946] Extracts changes in facial expressions and body language from video data.
[0947] As a result, an overall score of 85 / 100 is calculated for Interviewer A, and detailed evaluations are given, with 90 / 100 for language and 80 / 100 for tone of voice. These results are presented to the user via the feedback terminal.
[0948] This system will enable interview evaluation to be conducted objectively and efficiently, and is expected to be a significant improvement over conventional subjective evaluation methods.
[0949] The processing flow will be explained below.
[0950] Step 1:
[0951] The user (interviewer) starts the interview, and the device activates the camera and microphone to record the video and audio data of the interview.
[0952] Step 2:
[0953] The device sends the recorded data to the server in real time or after the interview is over. The sent data is in the form of video and audio files.
[0954] Step 3:
[0955] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[0956] Step 4:
[0957] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the speaker's voice is transcribed sequentially as text.
[0958] Step 5:
[0959] The server then analyzes the converted text data using natural language processing (NLP) technology, specifically evaluating the language used (for example, frequency of honorific use, use of technical terms), the grammatical structure of the conversation, and logic.
[0960] Step 6:
[0961] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[0962] Step 7:
[0963] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[0964] Step 8:
[0965] The server quantifies each extracted feature and assigns a score according to a standardized evaluation scale, such as 90 / 100 for appropriate language and 85 / 100 for friendly facial expressions.
[0966] Step 9:
[0967] The server uses interview data from existing employees to train the learning model, which uses machine learning algorithms to learn evaluation criteria that fit the company culture.
[0968] Step 10:
[0969] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[0970] Step 11:
[0971] The server sums up the scores for each evaluation item to generate an overall evaluation, which is then sent to the device.
[0972] Step 12:
[0973] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the overall evaluation might be 85 / 100, language 90 / 100, and facial expression 85 / 100.
[0974] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias and helping to select the right candidates.
[0975] Example 1
[0976] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0977] Conventional interview evaluation systems have the problem that they are prone to human subjectivity and bias, making it difficult to achieve consistent and fair evaluations. Furthermore, the evaluation process is inefficient, placing a heavy burden on interviewers. Furthermore, conventional systems have difficulty providing real-time feedback, making it impossible to immediately identify areas for improvement.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0979] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to a central processing unit, means for removing noise from the transmitted video and audio data and converting the audio into text information, means for analyzing the phrasing and conversation content of the audio data using language analysis technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and sign language from the video data, means for quantifying and evaluating each analyzed feature, means for training a machine learning model using interview data of existing employees, means for evaluating new interview data using the machine learning model, and means for displaying the evaluation results and providing feedback to the user. This improves the objectivity and fairness of interview evaluations, reduces the burden on interviewers, and enables feedback to be provided in real time.
[0980] A "recording terminal" is a device for recording video and audio data of an interview.
[0981] A "central processing unit" is a computer or server that receives and analyzes the recorded data.
[0982] "Noise reduction" is the process of removing unnecessary noise from recorded audio and video data.
[0983] "Speech-to-text" refers to the process of converting recorded voice data into text data.
[0984] "Language analysis technology" is a technology that performs natural language processing on text data and analyzes wording and conversation content.
[0985] "Analyzing speaking style and tone of voice" refers to analyzing characteristics such as speaking rate, pitch, and tone based on the audio data.
[0986] "Analyzing facial expressions and sign language" means analyzing the interviewer's facial expressions and body language based on video data.
[0987] "Quantification" means expressing the analyzed features as numerical data.
[0988] A "machine learning model" is an algorithm that learns patterns based on past data and makes predictions and classifications.
[0989] "Providing feedback" means presenting the analysis and evaluation results to the user and informing them of areas for improvement and strengths.
[0990] "Setting evaluation criteria that fit the organizational culture" means setting evaluation criteria that are appropriate for the characteristics and values of the organization based on interview data from existing employees.
[0991] "Immediate analysis and display" means that interview data is processed in real time and the results are displayed immediately.
[0992] MODE FOR CARRYING OUT THE INVENTION
[0993] The present invention provides a system for objectively and efficiently evaluating interviews. This system includes a process for recording video and audio data of the interview, analyzing the data with a central processing unit, and providing the results as feedback to the user.
[0994] The system consists of the following main components:
[0995] 1. Recording device: This is a device that records the video and audio data of the interview. In this case, a standard webcam and microphone are used.
[0996] 2. Central Processing Unit (Server): A computer that receives and analyzes recorded data. A general cloud server or on-premise server can be used for high-performance computing.
[0997] 3. Analysis module: This is software that runs inside the central processing unit and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specifically, it uses software such as FFmpeg, Google Speech-to-Text API, NLTK, Spacy, Praat, OpenCV, Dlib, Scikit-learn, and TensorFlow.
[0998] 4. Feedback terminal: A device that displays analysis results to users and provides feedback. In this case, this applies to general PCs, tablets, smartphones, etc.
[0999] (Data Collection and Transfer)
[1000] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When the recording is finished, the recording device sends this data to the central processing unit. The transmission process can use storage services such as Google Drive API or AWS S3.
[1001] (Data preprocessing)
[1002] The server uses FFmpeg to filter the received recording data to remove noise, and then uses the Google Speech-to-Text API to convert the audio data into text.
[1003] (Feature extraction and quantification)
[1004] When the server receives the speech-recognized text data, it uses natural language processing technology to analyze the vocabulary and content of the conversation. This analysis is performed using NLTK and Spacy. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points.
[1005] (Training and evaluating learning models)
[1006] The server uses the feature data to train a machine learning model. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. The trained model is then used to evaluate new interview data and calculate an overall score and detailed scores for each evaluation item.
[1007] (Presentation of results)
[1008] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript as a dashboard operated by the user. The user (interviewer or interviewee) uses this feedback terminal to check the evaluation results and identify areas for improvement and strengths.
[1009] Specific examples
[1010] For example, when an interview with Interviewer A is conducted, the entire process is recorded and the recorded data is sent to the server. The server removes noise using FFmpeg and converts the audio data into text using the Google Speech-to-Text API. Next, it analyzes the text data using NLTK or Spacy, analyzes the audio data using Praat, and extracts facial expressions from the video data using OpenCV and Dlib. These results are quantified, and each evaluation item is scored using a model trained with Scikit-learn or TensorFlow. Finally, these results are provided to the user in a form that can be viewed on a feedback terminal. The evaluation results include detailed information such as Interviewer A's overall score of 85 / 100, language score of 90 / 100, and tone of voice score of 80 / 100.
[1011] Prompt Sentence Examples
[1012] Below are some example prompts to input to a generative AI model:
[1013] "Please explain in detail, step by step, the process of the program that analyzes the interview recording data and conducts a comprehensive evaluation based on the interviewer's language, tone of voice, facial expressions, and sign language."
[1014] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1015] Step 1: Data collection and transfer
[1016] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When recording is finished, the recording device sends the recorded video and audio data to the server. This sending process uses storage services such as Google Drive API and AWS S3. The input is the video and audio data of the interview, and the output is the recorded data sent to the server.
[1017] Step 2: Preprocessing the data
[1018] The recording data received by the server is first denoised. Specifically, FFmpeg is used to filter the audio data and remove static noise. Next, the Google Speech-to-Text API is used to convert the audio data into text. This process converts the audio data into text information. The input is the recording data sent to the server, and the output is the noise-denoised audio data and text data.
[1019] Step 3: Feature extraction and quantification
[1020] The server uses natural language processing technology to analyze the phrasing and conversation content of the speech-recognized text data. NLTK and Spacy are used for this analysis. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points. The input is noise-removed audio data and text data, and the output is quantified feature data.
[1021] Step 4: Training and evaluating the learning model
[1022] The server trains a machine learning model using the feature data it has acquired so far. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. Next, it uses the trained model to evaluate newly acquired interview data and calculates an overall score and detailed scores for each evaluation item. The input is quantified feature data, and the output is an evaluation score.
[1023] Step 5: Presenting the results
[1024] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript to build a dashboard. The user (interviewer or interviewee) can check the evaluation results using this feedback terminal. The evaluation results include an overall score and detailed evaluation information for each evaluation item. The input is the evaluation score, and the output is the visualized evaluation results.
[1025] (Application example 1)
[1026] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1027] In conventional interview evaluation systems, evaluations tend to be subjective and often result in bias. Furthermore, in factories and other workplaces, it is difficult to objectively evaluate worker performance and provide efficient feedback. Therefore, a new system is needed that can accurately and quickly evaluate worker performance in workplaces where improvements in work efficiency and safety are required.
[1028] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1029] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using existing data, means for evaluating new data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for evaluating work efficiency, accuracy, safety, etc. based on the analysis results, and means for feeding back the scoring results to the worker and suggesting areas for improvement. This makes it possible to objectively and efficiently evaluate worker performance and suggest areas for improvement even in factories and on-site workplaces.
[1030] Definitions of important words
[1031] "Video Data" means visual information captured by a camera or other recording device.
[1032] "Audio Data" means audio information captured by a microphone or other recording device.
[1033] "Noise reduction" is the process of removing unwanted background noise and artifacts from video and audio recordings.
[1034] "Text conversion" is the process of converting audio data into written information.
[1035] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[1036] "Language" refers to the choice of words and expressions used in speech.
[1037] "Conversation content" refers to specific topics and information contained in the voice data.
[1038] "Speaking style" refers to characteristics of the audio data such as pronunciation, intonation, and rhythm.
[1039] "Voice tone" refers to voice characteristics including the pitch, strength, rhythm, etc. of the voice in the voice data.
[1040] "Facial expressions" refer to facial changes and emotional expressions in video data.
[1041] "Body language" refers to the body movements and postures in video data.
[1042] "Scoring" is the process of quantifying and evaluating the analyzed features.
[1043] A "learning model" is an algorithm that learns from data to identify and predict specific patterns.
[1044] "Feedback" is the process of providing the evaluation results to the user and indicating areas for improvement and evaluation.
[1045] "Work efficiency" is an index that indicates how efficiently work is performed within a certain period of time.
[1046] "Accuracy" is an indicator of how accurately a task is performed.
[1047] "Safety" is an indicator of how safely the work was carried out.
[1048] MODE FOR CARRYING OUT THE INVENTION
[1049] 1. System Overview
[1050] This invention provides a system for objectively evaluating the performance of workers in a factory or other workplace. The system mainly consists of the following components:
[1051] Recording device
[1052] server
[1053] Analysis Module
[1054] Feedback Terminal
[1055] 2. Hardware and Software Configuration
[1056] Recording device
[1057] The recording terminal is equipped with a camera and microphone and is a device that records the worker's video and audio data in high quality. In this example, a general web camera and microphone are used.
[1058] server
[1059] The server acts as a central computer that receives and processes the recorded data. The following software is installed on the server:
[1060] OpenCV (video data analysis)
[1061] speech_recognition (converts speech data to text)
[1062] scikit-learn (machine learning model training and evaluation)
[1063] 3. Program Processing
[1064] Data Collection and Transfer
[1065] The recording terminal records the video and audio data of the worker's work and sends the data to the server. The recorded data is pre-processed to remove noise, and the audio data is converted to text.
[1066] Data analysis
[1067] The server applies natural language processing technology to the audio data to analyze the vocabulary and content of the conversation. It analyzes facial expressions and body language from the video data, and extracts speaking style and tone of voice from the audio data. These features are then quantified and scored.
[1068] Feature extraction and quantification
[1069] The server quantifies the characteristics obtained from the recorded data and scores it based on specific criteria, including work efficiency, accuracy, and safety.
[1070] Training and evaluating learning models
[1071] The server trains a learning model using existing data and evaluates newly acquired data. This process uses a machine learning model (e.g., SVM) using the scikit-learn library.
[1072] Feedback of results
[1073] The analysis results and the evaluated scoring results are sent to a feedback terminal where users can check the results, allowing workers to receive an evaluation of their performance and learn areas for improvement.
[1074] 4. Examples of concrete examples and prompts
[1075] Specific examples
[1076] For example, Worker A's work process for one day is recorded and an evaluation is made based on that data. The recorded data is converted into text using voice recognition, and the work movements are evaluated through video analysis. Finally, an overall evaluation score is obtained, and this score is fed back to the worker to indicate areas for improvement.
[1077] Prompt Sentence Examples
[1078] "Please create prompts for a scoring system that will record Worker A's work for one day, analyze the video and audio, and evaluate work efficiency, accuracy, safety, etc. Please include specific steps from data preprocessing to evaluation."
[1079] This invention makes it possible to make objective evaluations that were difficult to achieve with conventional methods, and is expected to improve work efficiency and safety.
[1080] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1081] Program processing steps
[1082] Step 1:
[1083] The recording terminal records the video and audio data of the worker's work in real time using a camera and microphone, and saves this data as a file.
[1084] Input: Video and audio data of the worker
[1085] Output: Recording files (video files, audio files)
[1086] Step 2:
[1087] The recording device sends the recorded video and audio data to a server, where the data is centrally managed and analyzed.
[1088] Input: Recording file
[1089] Output: Data transferred to the server
[1090] Step 3:
[1091] The server performs noise reduction on the received video and audio data using filtering technology to improve the quality of the data.
[1092] Input: Transferred recording data (video files, audio files)
[1093] Output: High-quality data after noise removal
[1094] Step 4:
[1095] The server uses speech recognition technology to convert the audio data into text, specifically the speech_recognition library.
[1096] Input: High-quality audio data
[1097] Output: Text data
[1098] Step 5:
[1099] The server uses natural language processing technology to analyze the vocabulary and content of the conversation, thereby understanding the meaning and context of the text data.
[1100] Input: Text data
[1101] Output: Analysis results (characteristics of language and conversation content)
[1102] Step 6:
[1103] The server analyzes the speech data to determine speaking style and tone, and extracts features such as pitch, rhythm, and tone from the speech waveform data.
[1104] Input: High-quality audio data
[1105] Output: Audio feature data
[1106] Step 7:
[1107] The server analyzes facial expressions and body language from the video data, detecting facial expressions and body movements using OpenCV.
[1108] Input: High-quality video data
[1109] Output: Video feature data
[1110] Step 8:
[1111] The server quantifies and scores each analyzed feature, calculates the score for each feature, and performs an overall performance evaluation.
[1112] Input: Language feature data, audio feature data, video feature data
[1113] Output: Scoring results
[1114] Step 9:
[1115] The server trains a learning model using existing data, using scikit-learn for this process and optimizing the model based on existing evaluation data.
[1116] Input: Existing evaluation data
[1117] Output: A trained model
[1118] Step 10:
[1119] The server evaluates newly acquired data using the learning model, inputting new data to the trained model and obtaining the evaluation results.
[1120] Input: Newly acquired feature data
[1121] Output: Evaluation results
[1122] Step 11:
[1123] The server transmits the evaluation results to the feedback terminal, which displays the evaluation results so that the user can check them.
[1124] Input: Evaluation result (scoring result)
[1125] Output: Feedback to terminal
[1126] Step 12:
[1127] Users receive the evaluation results through a feedback terminal and identify areas for improvement, which helps workers improve their own performance.
[1128] Input: Data to be displayed on the feedback terminal
[1129] Output: Feedback information to the user
[1130] These steps allow for objective and efficient evaluation and feedback of worker performance.
[1131] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1132] System Overview
[1133] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to extract emotional changes during the interview and incorporate them into a comprehensive evaluation.
[1134] System configuration
[1135] The system consists of the following main components:
[1136] 1. Recording device: A device that records the video and audio data of the interview.
[1137] 2. Server: A central computer that receives and analyzes the recorded data.
[1138] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[1139] 4. Emotion engine: Software that runs inside the server and analyzes the user's emotions from video and audio data.
[1140] 5. Feedback terminal: A device for displaying the analysis results to the user.
[1141] System Operation
[1142] 1. Data Collection and Transfer
[1143] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[1144] 2. Data Preprocessing
[1145] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[1146] 3. Feature extraction and quantification
[1147] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[1148] 4. Analysis by Emotion Engine
[1149] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[1150] 5. Quantifying and Scoring Emotional Data
[1151] The server quantifies the emotional change data obtained from the emotion engine and scores it in the same way as other evaluation items. Evaluation criteria include emotional stability and appropriate emotional expression.
[1152] 6. Training and Evaluation of the Learning Model
[1153] The server trains a machine learning model using interview data from existing employees. The model uses various features, including emotional data, to learn evaluation criteria that fit the company culture. When new interview data is input, the trained model is used to perform an overall evaluation.
[1154] 7. Presentation of results
[1155] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) checks the evaluation results through this terminal and receives feedback. The evaluation results include an overall evaluation including emotional evaluation and detailed scores for each evaluation item.
[1156] Specific examples
[1157] For example, when an interview with Interviewer B is conducted, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[1158] Noise is removed from video and audio data to extract high-quality data.
[1159] Converts audio data into text.
[1160] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[1161] Analyzes voice tone, rhythm, and tone from audio data.
[1162] Extracts changes in facial expressions and body language from video data.
[1163] An emotion engine is used to detect various emotions (happiness, surprise, confusion, etc.).
[1164] Each feature and emotion data is quantified and scored.
[1165] As a result, an overall score of 90 / 100 is calculated for Interviewer B, and detailed evaluations are given, including 85 / 100 for language, 82 / 100 for tone of voice, and 88 / 100 for emotional stability. These results are presented to the user via a feedback terminal.
[1166] This system allows for objective and efficient evaluation of interviews, and is expected to be a significant improvement over conventional subjective evaluation methods. In addition, by combining it with an emotion engine, changes in the interviewer's emotions can be reflected in the evaluation, enabling a more comprehensive evaluation.
[1167] The processing flow will be explained below.
[1168] Step 1:
[1169] The user (interviewer) starts the interview, and the recording terminal activates the camera and microphone to record the video and audio data of the interview.
[1170] Step 2:
[1171] The recording device sends the recorded data to the server in real time or after the interview ends. The sent data is in the form of video and audio files.
[1172] Step 3:
[1173] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[1174] Step 4:
[1175] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the voice is transcribed sequentially.
[1176] Step 5:
[1177] The server then analyzes the converted text data using natural language processing (NLP) techniques, specifically evaluating the language used (e.g., frequency of honorific use and use of technical terms), the grammatical structure of the conversation, coherence, and logic.
[1178] Step 6:
[1179] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[1180] Step 7:
[1181] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[1182] Step 8:
[1183] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[1184] Step 9:
[1185] The server quantifies the emotional change data obtained from the emotion engine and scores it according to standardized evaluation criteria, such as emotional stability and appropriate emotional expression.
[1186] Step 10:
[1187] The server quantifies and scores each extracted feature (language, speaking style, facial expressions, tone of voice, body language, emotional changes, etc.). For example, appropriate language is given a score of 90 / 100, and friendly facial expressions are given a score of 85 / 100.
[1188] Step 11:
[1189] The server uses interview data from existing employees to train a machine learning model, which uses a machine learning algorithm to learn evaluation criteria that fit the company culture.
[1190] Step 12:
[1191] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[1192] Step 13:
[1193] The server sums up the scores for each evaluation item to generate an overall evaluation, and sends the evaluation result to the terminal.
[1194] Step 14:
[1195] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the user can see detailed scores of 90 / 100 for overall evaluation, 85 / 100 for language, 82 / 100 for facial expressions, and 88 / 100 for emotional stability.
[1196] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias to help select the right candidates. Furthermore, by incorporating an emotion engine, changes in the interviewer's emotions are reflected in the evaluation, enabling a more comprehensive evaluation.
[1197] Example 2
[1198] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1199] Conventional interview evaluation systems often rely on the subjective judgment of the interviewer, resulting in problems such as a lack of fairness and consistency. Furthermore, it is difficult to perform a comprehensive evaluation that includes the interviewer's emotional changes, making it impossible to reflect emotional stability or appropriate emotional expression in the evaluation. Furthermore, it is difficult to perform these analyses in real time and provide feedback immediately after the interview.
[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1201] In this invention, the server includes a means for recording video and audio data of the interview, a means for transmitting the recorded video and audio data to a central control device, a means for performing noise reduction and speech recognition on the transmitted video and audio data, a means for analyzing the phrasing and conversation content of the speech recognition results using natural language analysis technology, a means for analyzing speaking style and tone of voice from the audio data, a means for analyzing facial expressions and body movements from the video data, a means for quantifying and scoring each analyzed feature, a means for analyzing emotions from the recorded video and audio data, a means for quantifying and scoring the emotion data, a means for training a machine learning model using interview data of existing employees, a means for evaluating new interview data using the learning model, and a means for displaying the evaluation results and providing feedback to the user. This enables objective and efficient interview evaluation, and also allows the interviewer's emotional changes to be reflected in the evaluation. Analysis results can also be provided in real time, allowing feedback to be received immediately after the interview.
[1202] "Video and audio data of the interview" refers to the visual and audio information recorded during the interview.
[1203] "Recording means" refers to technical devices and methods for recording video and audio data.
[1204] "Central control unit" refers to the main computer system that centrally manages and processes multiple data.
[1205] "Transmission means" refers to the technical devices and methods for transferring data from one point to another.
[1206] "Noise reduction" refers to the process of removing unwanted noise and interference from video and audio data.
[1207] "Speech recognition" refers to the technology of analyzing voice data and converting it into text information.
[1208] "Natural language analysis technology" refers to analysis technology that enables computers to understand and interpret human language.
[1209] "Means for analyzing speaking style and tone of voice" refers to technology that analyzes a speaker's speech patterns and voice quality from audio data.
[1210] "Means for analyzing facial expressions and body movements" refers to technology that detects and analyzes changes in facial expressions and gestures of subjects from video data.
[1211] "Quantification" refers to the act of expressing analyzed information as quantitative data.
[1212] "Scoring" refers to the process of assigning a score based on quantified data according to evaluation criteria.
[1213] "Means for analyzing emotions" refers to technology that estimates emotional states from video and audio data.
[1214] A "machine learning model" refers to an algorithm that learns patterns from large amounts of data and makes inferences and predictions about new data.
[1215] "Means for displaying the evaluation results and providing feedback to the user" refers to technical devices and methods for visually presenting the analysis results to the user and facilitating their understanding.
[1216] This invention is a system for conducting interview evaluations objectively and efficiently, and includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. The system incorporates specific hardware and software, with the user, terminal, and server playing their respective roles.
[1217] Key Components of the System
[1218] 1. Recording device: A device for recording the video and audio data of the interview. Generally, a device equipped with a camera and microphone is used.
[1219] 2. Server: A central control unit that receives and analyzes recorded data. This device is a high-performance computer system that can process and analyze data at high speed. For example, a commonly used corporate server system can be used.
[1220] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. It uses deep learning libraries such as TensorFlow and PyTorch.
[1221] 4. Emotion engine: Software that analyzes emotions from video and audio data. Typically, a cloud-based API (e.g., Microsoft Azure's Emotion Recognition API) is used.
[1222] 5. Feedback terminal: A device used to display analysis results to users. This is typically a PC or mobile device.
[1223] Data collection and analysis flow
[1224] When the user (interviewer) starts the interview, the recording device records the entire process. After recording is finished, the device sends the recorded video and audio data to the server. The server performs noise reduction and speech recognition, and converts the audio data into text. For example, noise reduction is performed using FFmpeg, and the audio is converted to text information using the Google Cloud Speech-to-Text API.
[1225] For text data, natural language analysis techniques are used to analyze the phrasing and content of conversations. Natural language processing tools such as the BERT model are used. For audio data, speech analysis tools such as Praat are used to analyze speaking style and tone of voice. In addition, the server uses image analysis libraries such as OpenCV to analyze facial expressions and body language from video data.
[1226] Sentiment Analysis and Scoring
[1227] The emotion engine analyzes emotions from video and audio data to extract emotions such as joy, anger, surprise, and sadness, using Microsoft Azure's emotion recognition API.
[1228] The analyzed emotion data is quantified and scored by the server, using evaluation indices based on the intensity and frequency of each emotion.
[1229] The server uses interview data from existing employees to train a machine learning model, which uses tools like Scikit-learn to learn evaluation criteria that fit the company culture.
[1230] Presentation of evaluation results
[1231] Finally, when new interview data is input, the server uses the trained model to perform a comprehensive evaluation, and the generated evaluation results are sent to the feedback terminal, where users can check the evaluation results.
[1232] Examples and prompts
[1233] For example, if an interview with Interviewer B is conducted, the analysis will proceed as follows:
[1234] The recording device captures video and audio data.
[1235] The device sends data to the server.
[1236] The server uses FFmpeg to remove noise.
[1237] The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[1238] The server analyzes the text data using the BERT model.
[1239] The server uses Praat to extract audio features.
[1240] The server analyzes the video data using OpenCV.
[1241] The emotion engine analyzes emotions using Microsoft Azure's emotion recognition API.
[1242] The server quantifies the emotional data and scores it.
[1243] The server trains the machine learning model using Scikit-learn.
[1244] When new interview data is input, the server evaluates it using the trained model.
[1245] Evaluation results are presented through a feedback terminal.
[1246] Prompt Sentence Examples
[1247] Please provide a detailed description of each processing step in your interview evaluation system, including the specific actions and techniques used from data collection to providing results.
[1248] This system allows for objective and efficient interview evaluation. Furthermore, by using an emotion engine, the interviewer's emotional changes are incorporated into the evaluation, enabling a more comprehensive evaluation. Evaluation results are presented in real time, allowing users to receive immediate feedback.
[1249] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1250] Step 1:
[1251] The user (interviewer) starts the interview. Recording of audio and video data begins. The input is the video and audio during the interview, and the output is the video and audio data recorded on the recording terminal.
[1252] Step 2:
[1253] The device sends the recorded data to the server. After recording is complete, the recorded video and audio data is transferred to the server. The input is the video and audio data stored on the recording device, and the output is the same data sent to the server. Specifically, a secure data transfer protocol (such as SFTP) is used.
[1254] Step 3:
[1255] The server performs noise reduction. The server removes background noise from the received video and audio data. The input is the raw video and audio data sent to the server, and the output is the denoised video and audio data. FFmpeg noise reduction filters are used.
[1256] Step 4:
[1257] The server performs speech recognition. The server converts noise-removed speech data into text. The input is noise-removed speech data, and the output is text data. The Google Cloud Speech-to-Text API is used for speech recognition.
[1258] Step 5:
[1259] The server analyzes the text data. The server uses natural language analysis technology to analyze the transcribed text data. The input is text data generated by speech recognition, and the output is feature data of the analyzed phrasing and conversation content. The BERT model, for example, is used here.
[1260] Step 6:
[1261] The server extracts voice features. The server analyzes speaking style and tone of voice from the voice data. The input is the voice data after noise removal, and the output is feature data related to speaking style and tone of voice. Specifically, the voice analysis is performed using Praat.
[1262] Step 7:
[1263] The server analyzes the video data. The server analyzes facial expressions and body language from the video data. The input is the video data after noise removal, and the output is feature data related to facial expressions and body language. Image analysis libraries such as OpenCV are used.
[1264] Step 8:
[1265] The emotion engine analyzes emotions. The emotion engine extracts multiple emotional states from video and audio data. The input is noise-removed video and audio data, and the output is quantified data for each emotion. Microsoft Azure's emotion recognition API is used.
[1266] Step 9:
[1267] The server quantifies the analysis data and scores it. The server quantifies the acquired feature data and emotion data and displays it as a score. The input is feature data such as speaking style, tone of voice, facial expressions, body language, and emotions, and the output is a quantified evaluation score.
[1268] Step 10:
[1269] The server trains a machine learning model. The server trains the model using interview data of existing employees. The input is the interview data of existing employees, and the output is the trained machine learning model. Libraries such as Scikit-learn are used.
[1270] Step 11:
[1271] The server evaluates new interview data. The server evaluates new interview data using the trained model. The input is the newly collected interview data and the trained model, and the output is an overall evaluation score.
[1272] Step 12:
[1273] The server sends the evaluation result to the feedback terminal. The server sends the generated evaluation result to the feedback terminal. The input is the calculated evaluation score, and the output is the evaluation result displayed on the feedback terminal.
[1274] Step 13:
[1275] The user checks the evaluation results. The user checks the analysis results and the scores for each evaluation item through the feedback terminal. The input is the evaluation results displayed on the feedback terminal, and the output is the user's understanding of the evaluation results and feedback.
[1276] This series of processes enables interview evaluation to be conducted objectively and efficiently, and also enables comprehensive evaluation including emotional stability and appropriate expression.
[1277] (Application example 2)
[1278] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1279] In recent years, there has been a demand for strengthened security in offices and homes, but current security systems have difficulty grasping changes in visitors' emotions and behavior in real time. In particular, there is a lack of an objective evaluation system for quickly responding to suspicious behavior or abnormal emotional fluctuations. This can lead to potential risks being overlooked, making it difficult to implement effective security measures.
[1280] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using interview data of existing employees, means for evaluating new interview data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for detecting abnormal emotional fluctuations and suspicious behavior and issuing an alert, and means for generating a report based on the analysis results. This makes it possible to analyze the emotions and behavior of visitors in real time, quickly detect suspicious behavior, and issue an alert.
[1281] "Video data" refers to image information captured using a camera.
[1282] "Audio data" refers to audio information recorded using a microphone.
[1283] "Recording" refers to the act of saving video and audio data.
[1284] "Server" refers to a central computer that collects, stores, and analyzes data over a network.
[1285] "Noise reduction" refers to the method of removing unwanted information or interference from video and audio data.
[1286] "Text conversion" refers to the technology of converting voice data into text information.
[1287] "Natural language processing technology" refers to algorithms and methods for analyzing text data.
[1288] "Language" refers to the way a character speaks and the words they choose to use.
[1289] "Vocal tone" refers to characteristics such as the pitch, rhythm, and tone of a speaker's voice.
[1290] "Facial expression" refers to changes in emotions shown by the movement of facial muscles.
[1291] "Body language" refers to non-verbal communication expressed through bodily movements and posture.
[1292] "Quantification" refers to the act of expressing analyzed features as numerical data.
[1293] "Scoring" refers to the method of calculating an evaluation score based on numerical characteristics.
[1294] A "learning model" refers to a computational model that learns patterns from large amounts of data and performs predictions and classifications.
[1295] "Issuing an alarm" refers to the act of sending a notification to draw attention when the system determines that there is an abnormality.
[1296] "Report generation" refers to the act of creating a report summarizing the analysis results.
[1297] System Overview
[1298] This invention is a system that can be applied to a security system to analyze visitors' emotions and behavior in real time and detect suspicious behavior or abnormal emotional fluctuations. This system aims to improve security by analyzing video and audio data. It also issues an alarm and generates a report based on the analysis results.
[1299] System configuration
[1300] The system consists of the following main components:
[1301] 1. Recording device: A device that records video and audio data of visitors. Specifically, it uses a camera (e.g., Logitech C922 Pro Stream) and a microphone (e.g., Rode NT-USB).
[1302] 2. Server: A central computer that receives and analyzes the recorded data.
[1303] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specific examples of this software include OpenCV and the emotion_recognition library.
[1304] 4. Feedback terminal: A device for displaying the analysis results to the user.
[1305] 5. Alarm system: A device that issues an alarm if it detects abnormal emotional fluctuations or suspicious behavior.
[1306] 6. Report generation module: Software for generating reports based on the analysis results.
[1307] System Operation
[1308] 1. Data Collection and Transfer
[1309] The user starts the system, and the recording device records the video and audio of the visitor in real time. The recorded data is immediately sent to the server.
[1310] 2. Data Preprocessing
[1311] The server removes noise from the received data and converts the audio data into text. The OpenCV library is used for noise removal, and speech recognition technology is used for converting the audio data into text.
[1312] 3. Feature extraction and quantification
[1313] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. These characteristics are then quantified and scored.
[1314] 4. Analysis by Emotion Engine
[1315] The server uses the emotion_recognition library to analyze emotions from video and audio data, detecting emotional changes in real time and identifying abnormal emotional fluctuations and suspicious behavior.
[1316] 5. Alert generation and report generation
[1317] If suspicious behavior or abnormal emotional changes are detected, the server will issue an alert through the alarm system. Furthermore, it will periodically generate reports based on the analysis results and provide them to the user.
[1318] Specific examples
[1319] For example, if a visitor shows signs of anger or fear at an office reception, the system can detect this in real time and immediately issue an alert, while a report summarizing the day's events is sent to the manager.
[1320] Example prompt sentence:
[1321] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[1322] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1323] Step 1:
[1324] Data Collection and Transfer
[1325] The user starts the system, and the recording terminal records the visitor's video and audio data in real time. The input is video data from the camera and audio data from the microphone. The output is to transfer this data to the server. Specifically, the camera captures the visitor's video and the microphone records the audio. The video data is then compressed into H.264 format or similar, and the audio data is converted into WAV format or similar before being sent to the server using the TCP / IP protocol.
[1326] Step 2:
[1327] Data Preprocessing
[1328] The server removes noise from the received video and audio data and converts the audio data to text. The input is the received video and audio data. The output is noise-removed video data, audio data, and text data. Specifically, the OpenCV library is used to remove noise from the video data, and filtering technology is used to clean the audio data. A speech recognition API (e.g., Google Cloud Speech-to-Text API) is used to convert the audio to text.
[1329] Step 3:
[1330] Feature extraction and quantification
[1331] The server uses natural language processing technology to analyze the phrasing and content of the conversation from the speech-recognized text data, analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. The inputs include cleaned video data, audio data, and text data. The output is quantified feature data. Specifically, NLTK or spaCy is used for natural language processing, and the LibROSA library is used to analyze speaking style and tone of voice from the audio data. Facial expression recognition and body language analysis are performed from the video data using OpenCV and Dlib libraries.
[1332] Step 4:
[1333] Analysis by emotion engine
[1334] The server uses the emotion_recognition library to analyze emotions from video and audio data. The input is quantified feature data. The output is detected emotion data. Specific operations include inferring emotions from facial expressions in the video, tone of voice, and linguistic expressions in the text. For example, by combining OpenCV and TensorFlow, emotions such as joy, anger, and surprise can be identified in real time based on facial expressions.
[1335] Step 5:
[1336] Issuance of an alert
[1337] If the server detects suspicious behavior or abnormal emotional fluctuations, it will issue an alarm through the alarm system. The inputs are detected emotional data and behavioral data. The output is an alarm message. Specifically, if a certain emotional score exceeds a threshold, the alarm system will be activated and an audio alert or text notification will be sent to security staff.
[1338] Step 6:
[1339] Report Generation
[1340] The server generates a report based on the analysis results and provides it to the user. The input is quantified and analyzed feature data. The output is a daily or monthly report. Specifically, it aggregates the analysis results stored in the database and automatically generates a report according to a template. The generated report is saved in PDF or CSV format and provided to the user via email or dashboard.
[1341] Example prompt sentence:
[1342] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[1343] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1344] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1345] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1346] [Fourth embodiment]
[1347] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1348] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1349] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1350] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1351] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1352] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1353] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1354] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1355] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1356] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1357] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1358] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1359] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1360] System Overview
[1361] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes that record the video and audio data of the interview, analyze it, and provide the results as feedback. This eliminates human subjectivity and bias, enabling fair evaluation.
[1362] System configuration
[1363] The system consists of the following main components:
[1364] 1. Recording device: A device that records the video and audio data of the interview.
[1365] 2. Server: A central computer that receives and analyzes the recorded data.
[1366] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[1367] 4. Feedback terminal: A device for displaying the analysis results to the user.
[1368] System Operation
[1369] 1. Data Collection and Transfer
[1370] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[1371] 2. Data Preprocessing
[1372] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[1373] 3. Feature extraction and quantification
[1374] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[1375] 4. Training and evaluation of the learning model
[1376] The server trains a machine learning model using interview data from existing employees. This trained model is then used to evaluate newly acquired interview data and calculate various scores. For example, the server evaluates overall performance based on factors such as the consistency of language, conversational content, and facial expressions.
[1377] 5. Presentation of results
[1378] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) can check the evaluation results and receive feedback through this terminal. The evaluation results include an overall evaluation and detailed scores for each evaluation item.
[1379] Specific examples
[1380] For example, when an interview is conducted with interviewer A, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[1381] Noise is removed from video and audio data to extract high-quality data.
[1382] Converts audio data into text.
[1383] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[1384] Analyzes voice tone, rhythm, and tone from audio data.
[1385] Extracts changes in facial expressions and body language from video data.
[1386] As a result, an overall score of 85 / 100 is calculated for Interviewer A, and detailed evaluations are given, with 90 / 100 for language and 80 / 100 for tone of voice. These results are presented to the user via the feedback terminal.
[1387] This system will enable interview evaluation to be conducted objectively and efficiently, and is expected to be a significant improvement over conventional subjective evaluation methods.
[1388] The processing flow will be explained below.
[1389] Step 1:
[1390] The user (interviewer) starts the interview, and the device activates the camera and microphone to record the video and audio data of the interview.
[1391] Step 2:
[1392] The device sends the recorded data to the server in real time or after the interview is over. The sent data is in the form of video and audio files.
[1393] Step 3:
[1394] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[1395] Step 4:
[1396] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the speaker's voice is transcribed sequentially as text.
[1397] Step 5:
[1398] The server then analyzes the converted text data using natural language processing (NLP) technology, specifically evaluating the language used (for example, frequency of honorific use, use of technical terms), the grammatical structure of the conversation, and logic.
[1399] Step 6:
[1400] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[1401] Step 7:
[1402] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[1403] Step 8:
[1404] The server quantifies each extracted feature and assigns a score according to a standardized evaluation scale, such as 90 / 100 for appropriate language and 85 / 100 for friendly facial expressions.
[1405] Step 9:
[1406] The server uses interview data from existing employees to train the learning model, which uses machine learning algorithms to learn evaluation criteria that fit the company culture.
[1407] Step 10:
[1408] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[1409] Step 11:
[1410] The server sums up the scores for each evaluation item to generate an overall evaluation, which is then sent to the device.
[1411] Step 12:
[1412] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the overall evaluation might be 85 / 100, language 90 / 100, and facial expression 85 / 100.
[1413] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias and helping to select the right candidates.
[1414] Example 1
[1415] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1416] Conventional interview evaluation systems have the problem that they are prone to human subjectivity and bias, making it difficult to achieve consistent and fair evaluations. Furthermore, the evaluation process is inefficient, placing a heavy burden on interviewers. Furthermore, conventional systems have difficulty providing real-time feedback, making it impossible to immediately identify areas for improvement.
[1417] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1418] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to a central processing unit, means for removing noise from the transmitted video and audio data and converting the audio into text information, means for analyzing the phrasing and conversation content of the audio data using language analysis technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and sign language from the video data, means for quantifying and evaluating each analyzed feature, means for training a machine learning model using interview data of existing employees, means for evaluating new interview data using the machine learning model, and means for displaying the evaluation results and providing feedback to the user. This improves the objectivity and fairness of interview evaluations, reduces the burden on interviewers, and enables feedback to be provided in real time.
[1419] A "recording terminal" is a device for recording video and audio data of an interview.
[1420] A "central processing unit" is a computer or server that receives and analyzes the recorded data.
[1421] "Noise reduction" is the process of removing unnecessary noise from recorded audio and video data.
[1422] "Speech-to-text" refers to the process of converting recorded voice data into text data.
[1423] "Language analysis technology" is a technology that performs natural language processing on text data and analyzes wording and conversation content.
[1424] "Analyzing speaking style and tone of voice" refers to analyzing characteristics such as speaking rate, pitch, and tone based on the audio data.
[1425] "Analyzing facial expressions and sign language" means analyzing the interviewer's facial expressions and body language based on video data.
[1426] "Quantification" means expressing the analyzed features as numerical data.
[1427] A "machine learning model" is an algorithm that learns patterns based on past data and makes predictions and classifications.
[1428] "Providing feedback" means presenting the analysis and evaluation results to the user and informing them of areas for improvement and strengths.
[1429] "Setting evaluation criteria that fit the organizational culture" means setting evaluation criteria that are appropriate for the characteristics and values of the organization based on interview data from existing employees.
[1430] "Immediate analysis and display" means that interview data is processed in real time and the results are displayed immediately.
[1431] MODE FOR CARRYING OUT THE INVENTION
[1432] The present invention provides a system for objectively and efficiently evaluating interviews. This system includes a process for recording video and audio data of the interview, analyzing the data with a central processing unit, and providing the results as feedback to the user.
[1433] The system consists of the following main components:
[1434] 1. Recording device: This is a device that records the video and audio data of the interview. In this case, a standard webcam and microphone are used.
[1435] 2. Central Processing Unit (Server): A computer that receives and analyzes recorded data. A general cloud server or on-premise server can be used for high-performance computing.
[1436] 3. Analysis module: This is software that runs inside the central processing unit and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specifically, it uses software such as FFmpeg, Google Speech-to-Text API, NLTK, Spacy, Praat, OpenCV, Dlib, Scikit-learn, and TensorFlow.
[1437] 4. Feedback terminal: A device that displays analysis results to users and provides feedback. In this case, this applies to general PCs, tablets, smartphones, etc.
[1438] (Data Collection and Transfer)
[1439] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When the recording is finished, the recording device sends this data to the central processing unit. The transmission process can use storage services such as Google Drive API or AWS S3.
[1440] (Data preprocessing)
[1441] The server uses FFmpeg to filter the received recording data to remove noise, and then uses the Google Speech-to-Text API to convert the audio data into text.
[1442] (Feature extraction and quantification)
[1443] When the server receives the speech-recognized text data, it uses natural language processing technology to analyze the vocabulary and content of the conversation. This analysis is performed using NLTK and Spacy. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points.
[1444] (Training and evaluating learning models)
[1445] The server uses the feature data to train a machine learning model. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. The trained model is then used to evaluate new interview data and calculate an overall score and detailed scores for each evaluation item.
[1446] (Presentation of results)
[1447] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript as a dashboard operated by the user. The user (interviewer or interviewee) uses this feedback terminal to check the evaluation results and identify areas for improvement and strengths.
[1448] Specific examples
[1449] For example, when an interview with Interviewer A is conducted, the entire process is recorded and the recorded data is sent to the server. The server removes noise using FFmpeg and converts the audio data into text using the Google Speech-to-Text API. Next, it analyzes the text data using NLTK or Spacy, analyzes the audio data using Praat, and extracts facial expressions from the video data using OpenCV and Dlib. These results are quantified, and each evaluation item is scored using a model trained with Scikit-learn or TensorFlow. Finally, these results are provided to the user in a form that can be viewed on a feedback terminal. The evaluation results include detailed information such as Interviewer A's overall score of 85 / 100, language score of 90 / 100, and tone of voice score of 80 / 100.
[1450] Prompt Sentence Examples
[1451] Below are some example prompts to input to a generative AI model:
[1452] "Please explain in detail, step by step, the process of the program that analyzes the interview recording data and conducts a comprehensive evaluation based on the interviewer's language, tone of voice, facial expressions, and sign language."
[1453] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1454] Step 1: Data collection and transfer
[1455] When the user (interviewer) starts the interview, the recording device records the entire interview as video and audio. When recording is finished, the recording device sends the recorded video and audio data to the server. This sending process uses storage services such as Google Drive API and AWS S3. The input is the video and audio data of the interview, and the output is the recorded data sent to the server.
[1456] Step 2: Preprocessing the data
[1457] The recording data received by the server is first denoised. Specifically, FFmpeg is used to filter the audio data and remove static noise. Next, the Google Speech-to-Text API is used to convert the audio data into text. This process converts the audio data into text information. The input is the recording data sent to the server, and the output is the noise-denoised audio data and text data.
[1458] Step 3: Feature extraction and quantification
[1459] The server uses natural language processing technology to analyze the phrasing and conversation content of the speech-recognized text data. NLTK and Spacy are used for this analysis. Praat is also used to analyze speaking style (e.g., speaking rate, pitch, tone, etc.) from the audio data. OpenCV and Dlib are used to analyze facial expressions and sign language from the video data. These features are quantified and scored as evaluation points. The input is noise-removed audio data and text data, and the output is quantified feature data.
[1460] Step 4: Training and evaluating the learning model
[1461] The server trains a machine learning model using the feature data it has acquired so far. Specifically, it uses Scikit-learn and TensorFlow to train a random forest or neural network based on interview data from existing employees. Next, it uses the trained model to evaluate newly acquired interview data and calculates an overall score and detailed scores for each evaluation item. The input is quantified feature data, and the output is an evaluation score.
[1462] Step 5: Presenting the results
[1463] The server sends the generated evaluation results to a feedback terminal. For example, the results are visualized using HTML / CSS and JavaScript to build a dashboard. The user (interviewer or interviewee) can check the evaluation results using this feedback terminal. The evaluation results include an overall score and detailed evaluation information for each evaluation item. The input is the evaluation score, and the output is the visualized evaluation results.
[1464] (Application example 1)
[1465] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1466] In conventional interview evaluation systems, evaluations tend to be subjective and often result in bias. Furthermore, in factories and other workplaces, it is difficult to objectively evaluate worker performance and provide efficient feedback. Therefore, a new system is needed that can accurately and quickly evaluate worker performance in workplaces where improvements in work efficiency and safety are required.
[1467] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1468] In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using existing data, means for evaluating new data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for evaluating work efficiency, accuracy, safety, etc. based on the analysis results, and means for feeding back the scoring results to the worker and suggesting areas for improvement. This makes it possible to objectively and efficiently evaluate worker performance and suggest areas for improvement even in factories and on-site workplaces.
[1469] Definitions of important words
[1470] "Video Data" means visual information captured by a camera or other recording device.
[1471] "Audio Data" means audio information captured by a microphone or other recording device.
[1472] "Noise reduction" is the process of removing unwanted background noise and artifacts from video and audio recordings.
[1473] "Text conversion" is the process of converting audio data into written information.
[1474] "Natural language processing technology" is a technology that enables computers to understand and analyze human language.
[1475] "Language" refers to the choice of words and expressions used in speech.
[1476] "Conversation content" refers to specific topics and information contained in the voice data.
[1477] "Speaking style" refers to characteristics of the audio data such as pronunciation, intonation, and rhythm.
[1478] "Voice tone" refers to voice characteristics including the pitch, strength, rhythm, etc. of the voice in the voice data.
[1479] "Facial expressions" refer to facial changes and emotional expressions in video data.
[1480] "Body language" refers to the body movements and postures in video data.
[1481] "Scoring" is the process of quantifying and evaluating the analyzed features.
[1482] A "learning model" is an algorithm that learns from data to identify and predict specific patterns.
[1483] "Feedback" is the process of providing the evaluation results to the user and indicating areas for improvement and evaluation.
[1484] "Work efficiency" is an index that indicates how efficiently work is performed within a certain period of time.
[1485] "Accuracy" is an indicator of how accurately a task is performed.
[1486] "Safety" is an indicator of how safely the work was carried out.
[1487] MODE FOR CARRYING OUT THE INVENTION
[1488] 1. System Overview
[1489] This invention provides a system for objectively evaluating the performance of workers in a factory or other workplace. The system mainly consists of the following components:
[1490] Recording device
[1491] server
[1492] Analysis Module
[1493] Feedback Terminal
[1494] 2. Hardware and Software Configuration
[1495] Recording device
[1496] The recording terminal is equipped with a camera and microphone and is a device that records the worker's video and audio data in high quality. In this example, a general web camera and microphone are used.
[1497] server
[1498] The server acts as a central computer that receives and processes the recorded data. The following software is installed on the server:
[1499] OpenCV (video data analysis)
[1500] speech_recognition (converts speech data to text)
[1501] scikit-learn (machine learning model training and evaluation)
[1502] 3. Program Processing
[1503] Data Collection and Transfer
[1504] The recording terminal records the video and audio data of the worker's work and sends the data to the server. The recorded data is pre-processed to remove noise, and the audio data is converted to text.
[1505] Data analysis
[1506] The server applies natural language processing technology to the audio data to analyze the vocabulary and content of the conversation. It analyzes facial expressions and body language from the video data, and extracts speaking style and tone of voice from the audio data. These features are then quantified and scored.
[1507] Feature extraction and quantification
[1508] The server quantifies the characteristics obtained from the recorded data and scores it based on specific criteria, including work efficiency, accuracy, and safety.
[1509] Training and evaluating learning models
[1510] The server trains a learning model using existing data and evaluates newly acquired data. This process uses a machine learning model (e.g., SVM) using the scikit-learn library.
[1511] Feedback of results
[1512] The analysis results and the evaluated scoring results are sent to a feedback terminal where users can check the results, allowing workers to receive an evaluation of their performance and learn areas for improvement.
[1513] 4. Examples of concrete examples and prompts
[1514] Specific examples
[1515] For example, Worker A's work process for one day is recorded and an evaluation is made based on that data. The recorded data is converted into text using voice recognition, and the work movements are evaluated through video analysis. Finally, an overall evaluation score is obtained, and this score is fed back to the worker to indicate areas for improvement.
[1516] Prompt Sentence Examples
[1517] "Please create prompts for a scoring system that will record Worker A's work for one day, analyze the video and audio, and evaluate work efficiency, accuracy, safety, etc. Please include specific steps from data preprocessing to evaluation."
[1518] This invention makes it possible to make objective evaluations that were difficult to achieve with conventional methods, and is expected to improve work efficiency and safety.
[1519] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1520] Program processing steps
[1521] Step 1:
[1522] The recording terminal records the video and audio data of the worker's work in real time using a camera and microphone, and saves this data as a file.
[1523] Input: Video and audio data of the worker
[1524] Output: Recording files (video files, audio files)
[1525] Step 2:
[1526] The recording device sends the recorded video and audio data to a server, where the data is centrally managed and analyzed.
[1527] Input: Recording file
[1528] Output: Data transferred to the server
[1529] Step 3:
[1530] The server performs noise reduction on the received video and audio data using filtering technology to improve the quality of the data.
[1531] Input: Transferred recording data (video files, audio files)
[1532] Output: High-quality data after noise removal
[1533] Step 4:
[1534] The server uses speech recognition technology to convert the audio data into text, specifically the speech_recognition library.
[1535] Input: High-quality audio data
[1536] Output: Text data
[1537] Step 5:
[1538] The server uses natural language processing technology to analyze the vocabulary and content of the conversation, thereby understanding the meaning and context of the text data.
[1539] Input: Text data
[1540] Output: Analysis results (characteristics of language and conversation content)
[1541] Step 6:
[1542] The server analyzes the speech data to determine speaking style and tone, and extracts features such as pitch, rhythm, and tone from the speech waveform data.
[1543] Input: High-quality audio data
[1544] Output: Audio feature data
[1545] Step 7:
[1546] The server analyzes facial expressions and body language from the video data, detecting facial expressions and body movements using OpenCV.
[1547] Input: High-quality video data
[1548] Output: Video feature data
[1549] Step 8:
[1550] The server quantifies and scores each analyzed feature, calculates the score for each feature, and performs an overall performance evaluation.
[1551] Input: Language feature data, audio feature data, video feature data
[1552] Output: Scoring results
[1553] Step 9:
[1554] The server trains a learning model using existing data, using scikit-learn for this process and optimizing the model based on existing evaluation data.
[1555] Input: Existing evaluation data
[1556] Output: A trained model
[1557] Step 10:
[1558] The server evaluates newly acquired data using the learning model, inputting new data to the trained model and obtaining the evaluation results.
[1559] Input: Newly acquired feature data
[1560] Output: Evaluation results
[1561] Step 11:
[1562] The server transmits the evaluation results to the feedback terminal, which displays the evaluation results so that the user can check them.
[1563] Input: Evaluation result (scoring result)
[1564] Output: Feedback to terminal
[1565] Step 12:
[1566] Users receive the evaluation results through a feedback terminal and identify areas for improvement, which helps workers improve their own performance.
[1567] Input: Data to be displayed on the feedback terminal
[1568] Output: Feedback information to the user
[1569] These steps allow for objective and efficient evaluation and feedback of worker performance.
[1570] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1571] System Overview
[1572] This invention is a system for objectively and efficiently evaluating interviews. The system includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. In addition, by combining it with an emotion engine that recognizes the user's emotions, it is possible to extract emotional changes during the interview and incorporate them into a comprehensive evaluation.
[1573] System configuration
[1574] The system consists of the following main components:
[1575] 1. Recording device: A device that records the video and audio data of the interview.
[1576] 2. Server: A central computer that receives and analyzes the recorded data.
[1577] 3. Analysis module: Software that runs inside the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of evaluation models.
[1578] 4. Emotion engine: Software that runs inside the server and analyzes the user's emotions from video and audio data.
[1579] 5. Feedback terminal: A device for displaying the analysis results to the user.
[1580] System Operation
[1581] 1. Data Collection and Transfer
[1582] The user (interviewer) starts the interview, and the recording device records the entire interview as video and audio. When the recording is finished, the device sends this data to the server.
[1583] 2. Data Preprocessing
[1584] The server first removes noise from the data received to improve the quality of the recorded data, then uses voice recognition technology to convert the audio data into text.
[1585] 3. Feature extraction and quantification
[1586] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also extracts speaking style and tone from the audio data, and analyzes facial expressions and body language from the video data. These characteristics are then quantified and scored.
[1587] 4. Analysis by Emotion Engine
[1588] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[1589] 5. Quantifying and Scoring Emotional Data
[1590] The server quantifies the emotional change data obtained from the emotion engine and scores it in the same way as other evaluation items. Evaluation criteria include emotional stability and appropriate emotional expression.
[1591] 6. Training and Evaluation of the Learning Model
[1592] The server trains a machine learning model using interview data from existing employees. The model uses various features, including emotional data, to learn evaluation criteria that fit the company culture. When new interview data is input, the trained model is used to perform an overall evaluation.
[1593] 7. Presentation of results
[1594] The server sends the generated evaluation results to a feedback terminal. The user (interviewer or interviewer) checks the evaluation results through this terminal and receives feedback. The evaluation results include an overall evaluation including emotional evaluation and detailed scores for each evaluation item.
[1595] Specific examples
[1596] For example, when an interview with Interviewer B is conducted, the entire process is recorded and the data is sent to the server. The server then performs the following analysis:
[1597] Noise is removed from video and audio data to extract high-quality data.
[1598] Converts audio data into text.
[1599] Text data is analyzed using natural language processing technology to evaluate the wording and logic of the conversation.
[1600] Analyzes voice tone, rhythm, and tone from audio data.
[1601] Extracts changes in facial expressions and body language from video data.
[1602] An emotion engine is used to detect various emotions (happiness, surprise, confusion, etc.).
[1603] Each feature and emotion data is quantified and scored.
[1604] As a result, an overall score of 90 / 100 is calculated for Interviewer B, and detailed evaluations are given, including 85 / 100 for language, 82 / 100 for tone of voice, and 88 / 100 for emotional stability. These results are presented to the user via a feedback terminal.
[1605] This system allows for objective and efficient evaluation of interviews, and is expected to be a significant improvement over conventional subjective evaluation methods. In addition, by combining it with an emotion engine, changes in the interviewer's emotions can be reflected in the evaluation, enabling a more comprehensive evaluation.
[1606] The processing flow will be explained below.
[1607] Step 1:
[1608] The user (interviewer) starts the interview, and the recording terminal activates the camera and microphone to record the video and audio data of the interview.
[1609] Step 2:
[1610] The recording device sends the recorded data to the server in real time or after the interview ends. The sent data is in the form of video and audio files.
[1611] Step 3:
[1612] The server applies a noise reduction filter to the video and audio data received to improve the quality of the recording data, using an algorithm that effectively removes background noise.
[1613] Step 4:
[1614] The server analyzes the voice data using a speech recognition engine and converts it into text data. During this process, the voice is transcribed sequentially.
[1615] Step 5:
[1616] The server then analyzes the converted text data using natural language processing (NLP) techniques, specifically evaluating the language used (e.g., frequency of honorific use and use of technical terms), the grammatical structure of the conversation, coherence, and logic.
[1617] Step 6:
[1618] The server analyzes the waveform and spectrum of the voice data and extracts speaking characteristics (e.g., speaking speed, pauses, and pitch).
[1619] Step 7:
[1620] The server analyzes the video data frame by frame to extract facial expressions and body language. Using a facial expression recognition algorithm, it captures changes in the interviewer's emotions (smile, surprise, confusion, etc.). Body language is evaluated based on posture, hand movements, and gaze direction.
[1621] Step 8:
[1622] The emotion engine analyzes the user's emotions from video and audio data. The emotion engine analyzes the interviewer's facial expressions, tone of voice, and choice of words to extract emotional changes in real time. For example, emotions such as joy, anger, surprise, and sadness can be detected.
[1623] Step 9:
[1624] The server quantifies the emotional change data obtained from the emotion engine and scores it according to standardized evaluation criteria, such as emotional stability and appropriate emotional expression.
[1625] Step 10:
[1626] The server quantifies and scores each extracted feature (language, speaking style, facial expressions, tone of voice, body language, emotional changes, etc.). For example, appropriate language is given a score of 90 / 100, and friendly facial expressions are given a score of 85 / 100.
[1627] Step 11:
[1628] The server uses interview data from existing employees to train a machine learning model, which uses a machine learning algorithm to learn evaluation criteria that fit the company culture.
[1629] Step 12:
[1630] The server evaluates newly acquired interview data using the trained model and calculates a score for each evaluation item. For example, to evaluate logical thinking ability, the server evaluates whether the candidate was able to present a coherent and persuasive argument.
[1631] Step 13:
[1632] The server sums up the scores for each evaluation item to generate an overall evaluation, and sends the evaluation result to the terminal.
[1633] Step 14:
[1634] The terminal displays the evaluation results to the user (interviewer or interviewer). The display includes detailed scores for each item and an overall evaluation. For example, the user can see detailed scores of 90 / 100 for overall evaluation, 85 / 100 for language, 82 / 100 for facial expressions, and 88 / 100 for emotional stability.
[1635] In this way, the system evaluates interviews objectively and fairly, eliminating human subjectivity and bias to help select the right candidates. Furthermore, by incorporating an emotion engine, changes in the interviewer's emotions are reflected in the evaluation, enabling a more comprehensive evaluation.
[1636] Example 2
[1637] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1638] Conventional interview evaluation systems often rely on the subjective judgment of the interviewer, resulting in problems such as a lack of fairness and consistency. Furthermore, it is difficult to perform a comprehensive evaluation that includes the interviewer's emotional changes, making it impossible to reflect emotional stability or appropriate emotional expression in the evaluation. Furthermore, it is difficult to perform these analyses in real time and provide feedback immediately after the interview.
[1639] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1640] In this invention, the server includes a means for recording video and audio data of the interview, a means for transmitting the recorded video and audio data to a central control device, a means for performing noise reduction and speech recognition on the transmitted video and audio data, a means for analyzing the phrasing and conversation content of the speech recognition results using natural language analysis technology, a means for analyzing speaking style and tone of voice from the audio data, a means for analyzing facial expressions and body movements from the video data, a means for quantifying and scoring each analyzed feature, a means for analyzing emotions from the recorded video and audio data, a means for quantifying and scoring the emotion data, a means for training a machine learning model using interview data of existing employees, a means for evaluating new interview data using the learning model, and a means for displaying the evaluation results and providing feedback to the user. This enables objective and efficient interview evaluation, and also allows the interviewer's emotional changes to be reflected in the evaluation. Analysis results can also be provided in real time, allowing feedback to be received immediately after the interview.
[1641] "Video and audio data of the interview" refers to the visual and audio information recorded during the interview.
[1642] "Recording means" refers to technical devices and methods for recording video and audio data.
[1643] "Central control unit" refers to the main computer system that centrally manages and processes multiple data.
[1644] "Transmission means" refers to the technical devices and methods for transferring data from one point to another.
[1645] "Noise reduction" refers to the process of removing unwanted noise and interference from video and audio data.
[1646] "Speech recognition" refers to the technology of analyzing voice data and converting it into text information.
[1647] "Natural language analysis technology" refers to analysis technology that enables computers to understand and interpret human language.
[1648] "Means for analyzing speaking style and tone of voice" refers to technology that analyzes a speaker's speech patterns and voice quality from audio data.
[1649] "Means for analyzing facial expressions and body movements" refers to technology that detects and analyzes changes in facial expressions and gestures of subjects from video data.
[1650] "Quantification" refers to the act of expressing analyzed information as quantitative data.
[1651] "Scoring" refers to the process of assigning a score based on quantified data according to evaluation criteria.
[1652] "Means for analyzing emotions" refers to technology that estimates emotional states from video and audio data.
[1653] A "machine learning model" refers to an algorithm that learns patterns from large amounts of data and makes inferences and predictions about new data.
[1654] "Means for displaying the evaluation results and providing feedback to the user" refers to technical devices and methods for visually presenting the analysis results to the user and facilitating their understanding.
[1655] This invention is a system for conducting interview evaluations objectively and efficiently, and includes a series of processes for recording video and audio data of the interview, analyzing the data, and providing the evaluation results. The system incorporates specific hardware and software, with the user, terminal, and server playing their respective roles.
[1656] Key Components of the System
[1657] 1. Recording device: A device for recording the video and audio data of the interview. Generally, a device equipped with a camera and microphone is used.
[1658] 2. Server: A central control unit that receives and analyzes recorded data. This device is a high-performance computer system that can process and analyze data at high speed. For example, a commonly used corporate server system can be used.
[1659] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. It uses deep learning libraries such as TensorFlow and PyTorch.
[1660] 4. Emotion engine: Software that analyzes emotions from video and audio data. Typically, a cloud-based API (e.g., Microsoft Azure's Emotion Recognition API) is used.
[1661] 5. Feedback terminal: A device used to display analysis results to users. This is typically a PC or mobile device.
[1662] Data collection and analysis flow
[1663] When the user (interviewer) starts the interview, the recording device records the entire process. After recording is finished, the device sends the recorded video and audio data to the server. The server performs noise reduction and speech recognition, and converts the audio data into text. For example, noise reduction is performed using FFmpeg, and the audio is converted to text information using the Google Cloud Speech-to-Text API.
[1664] For text data, natural language analysis techniques are used to analyze the phrasing and content of conversations. Natural language processing tools such as the BERT model are used. For audio data, speech analysis tools such as Praat are used to analyze speaking style and tone of voice. In addition, the server uses image analysis libraries such as OpenCV to analyze facial expressions and body language from video data.
[1665] Sentiment Analysis and Scoring
[1666] The emotion engine analyzes emotions from video and audio data to extract emotions such as joy, anger, surprise, and sadness, using Microsoft Azure's emotion recognition API.
[1667] The analyzed emotion data is quantified and scored by the server, using evaluation indices based on the intensity and frequency of each emotion.
[1668] The server uses interview data from existing employees to train a machine learning model, which uses tools like Scikit-learn to learn evaluation criteria that fit the company culture.
[1669] Presentation of evaluation results
[1670] Finally, when new interview data is input, the server uses the trained model to perform a comprehensive evaluation, and the generated evaluation results are sent to the feedback terminal, where users can check the evaluation results.
[1671] Examples and prompts
[1672] For example, if an interview with Interviewer B is conducted, the analysis will proceed as follows:
[1673] The recording device captures video and audio data.
[1674] The device sends data to the server.
[1675] The server uses FFmpeg to remove noise.
[1676] The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[1677] The server analyzes the text data using the BERT model.
[1678] The server uses Praat to extract audio features.
[1679] The server analyzes the video data using OpenCV.
[1680] The emotion engine analyzes emotions using Microsoft Azure's emotion recognition API.
[1681] The server quantifies the emotional data and scores it.
[1682] The server trains the machine learning model using Scikit-learn.
[1683] When new interview data is input, the server evaluates it using the trained model.
[1684] Evaluation results are presented through a feedback terminal.
[1685] Prompt Sentence Examples
[1686] Please provide a detailed description of each processing step in your interview evaluation system, including the specific actions and techniques used from data collection to providing results.
[1687] This system allows for objective and efficient interview evaluation. Furthermore, by using an emotion engine, the interviewer's emotional changes are incorporated into the evaluation, enabling a more comprehensive evaluation. Evaluation results are presented in real time, allowing users to receive immediate feedback.
[1688] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1689] Step 1:
[1690] The user (interviewer) starts the interview. Recording of audio and video data begins. The input is the video and audio during the interview, and the output is the video and audio data recorded on the recording terminal.
[1691] Step 2:
[1692] The device sends the recorded data to the server. After recording is complete, the recorded video and audio data is transferred to the server. The input is the video and audio data stored on the recording device, and the output is the same data sent to the server. Specifically, a secure data transfer protocol (such as SFTP) is used.
[1693] Step 3:
[1694] The server performs noise reduction. The server removes background noise from the received video and audio data. The input is the raw video and audio data sent to the server, and the output is the denoised video and audio data. FFmpeg noise reduction filters are used.
[1695] Step 4:
[1696] The server performs speech recognition. The server converts noise-removed speech data into text. The input is noise-removed speech data, and the output is text data. The Google Cloud Speech-to-Text API is used for speech recognition.
[1697] Step 5:
[1698] The server analyzes the text data. The server uses natural language analysis technology to analyze the transcribed text data. The input is text data generated by speech recognition, and the output is feature data of the analyzed phrasing and conversation content. The BERT model, for example, is used here.
[1699] Step 6:
[1700] The server extracts voice features. The server analyzes speaking style and tone of voice from the voice data. The input is the voice data after noise removal, and the output is feature data related to speaking style and tone of voice. Specifically, the voice analysis is performed using Praat.
[1701] Step 7:
[1702] The server analyzes the video data. The server analyzes facial expressions and body language from the video data. The input is the video data after noise removal, and the output is feature data related to facial expressions and body language. Image analysis libraries such as OpenCV are used.
[1703] Step 8:
[1704] The emotion engine analyzes emotions. The emotion engine extracts multiple emotional states from video and audio data. The input is noise-removed video and audio data, and the output is quantified data for each emotion. Microsoft Azure's emotion recognition API is used.
[1705] Step 9:
[1706] The server quantifies the analysis data and scores it. The server quantifies the acquired feature data and emotion data and displays it as a score. The input is feature data such as speaking style, tone of voice, facial expressions, body language, and emotions, and the output is a quantified evaluation score.
[1707] Step 10:
[1708] The server trains a machine learning model. The server trains the model using interview data of existing employees. The input is the interview data of existing employees, and the output is the trained machine learning model. Libraries such as Scikit-learn are used.
[1709] Step 11:
[1710] The server evaluates new interview data. The server evaluates new interview data using the trained model. The input is the newly collected interview data and the trained model, and the output is an overall evaluation score.
[1711] Step 12:
[1712] The server sends the evaluation result to the feedback terminal. The server sends the generated evaluation result to the feedback terminal. The input is the calculated evaluation score, and the output is the evaluation result displayed on the feedback terminal.
[1713] Step 13:
[1714] The user checks the evaluation results. The user checks the analysis results and the scores for each evaluation item through the feedback terminal. The input is the evaluation results displayed on the feedback terminal, and the output is the user's understanding of the evaluation results and feedback.
[1715] This series of processes enables interview evaluation to be conducted objectively and efficiently, and also enables comprehensive evaluation including emotional stability and appropriate expression.
[1716] (Application example 2)
[1717] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1718] In recent years, there has been a demand for strengthened security in offices and homes, but current security systems have difficulty grasping changes in visitors' emotions and behavior in real time. In particular, there is a lack of an objective evaluation system for quickly responding to suspicious behavior or abnormal emotional fluctuations. This can lead to potential risks being overlooked, making it difficult to implement effective security measures.
[1719] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording video and audio data of the interview, means for transmitting the recorded video and audio data to the server, means for performing noise reduction and text conversion on the transmitted video and audio data, means for analyzing the wording and conversation content of the text data using natural language processing technology, means for analyzing speaking style and tone of voice from the audio data, means for analyzing facial expressions and body language from the video data, means for quantifying and scoring each analyzed feature, means for training a learning model using interview data of existing employees, means for evaluating new interview data using the learning model, means for displaying the evaluation results and providing feedback to the user, means for detecting abnormal emotional fluctuations and suspicious behavior and issuing an alert, and means for generating a report based on the analysis results. This makes it possible to analyze the emotions and behavior of visitors in real time, quickly detect suspicious behavior, and issue an alert.
[1720] "Video data" refers to image information captured using a camera.
[1721] "Audio data" refers to audio information recorded using a microphone.
[1722] "Recording" refers to the act of saving video and audio data.
[1723] "Server" refers to a central computer that collects, stores, and analyzes data over a network.
[1724] "Noise reduction" refers to the method of removing unwanted information or interference from video and audio data.
[1725] "Text conversion" refers to the technology of converting voice data into text information.
[1726] "Natural language processing technology" refers to algorithms and methods for analyzing text data.
[1727] "Language" refers to the way a character speaks and the words they choose to use.
[1728] "Vocal tone" refers to characteristics such as the pitch, rhythm, and tone of a speaker's voice.
[1729] "Facial expression" refers to changes in emotions shown by the movement of facial muscles.
[1730] "Body language" refers to non-verbal communication expressed through bodily movements and posture.
[1731] "Quantification" refers to the act of expressing analyzed features as numerical data.
[1732] "Scoring" refers to the method of calculating an evaluation score based on numerical characteristics.
[1733] A "learning model" refers to a computational model that learns patterns from large amounts of data and performs predictions and classifications.
[1734] "Issuing an alarm" refers to the act of sending a notification to draw attention when the system determines that there is an abnormality.
[1735] "Report generation" refers to the act of creating a report summarizing the analysis results.
[1736] System Overview
[1737] This invention is a system that can be applied to a security system to analyze visitors' emotions and behavior in real time and detect suspicious behavior or abnormal emotional fluctuations. This system aims to improve security by analyzing video and audio data. It also issues an alarm and generates a report based on the analysis results.
[1738] System configuration
[1739] The system consists of the following main components:
[1740] 1. Recording device: A device that records video and audio data of visitors. Specifically, it uses a camera (e.g., Logitech C922 Pro Stream) and a microphone (e.g., Rode NT-USB).
[1741] 2. Server: A central computer that receives and analyzes the recorded data.
[1742] 3. Analysis module: This software runs on the server and performs data preprocessing, feature extraction, quantification, and training and evaluation of the evaluation model. Specific examples of this software include OpenCV and the emotion_recognition library.
[1743] 4. Feedback terminal: A device for displaying the analysis results to the user.
[1744] 5. Alarm system: A device that issues an alarm if it detects abnormal emotional fluctuations or suspicious behavior.
[1745] 6. Report generation module: Software for generating reports based on the analysis results.
[1746] System Operation
[1747] 1. Data Collection and Transfer
[1748] The user starts the system, and the recording device records the video and audio of the visitor in real time. The recorded data is immediately sent to the server.
[1749] 2. Data Preprocessing
[1750] The server removes noise from the received data and converts the audio data into text. The OpenCV library is used for noise removal, and speech recognition technology is used for converting the audio data into text.
[1751] 3. Feature extraction and quantification
[1752] The server uses natural language processing technology to analyze the vocabulary and content of the conversation based on the speech recognition data. It also analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. These characteristics are then quantified and scored.
[1753] 4. Analysis by Emotion Engine
[1754] The server uses the emotion_recognition library to analyze emotions from video and audio data, detecting emotional changes in real time and identifying abnormal emotional fluctuations and suspicious behavior.
[1755] 5. Alert generation and report generation
[1756] If suspicious behavior or abnormal emotional changes are detected, the server will issue an alert through the alarm system. Furthermore, it will periodically generate reports based on the analysis results and provide them to the user.
[1757] Specific examples
[1758] For example, if a visitor shows signs of anger or fear at an office reception, the system can detect this in real time and immediately issue an alert, while a report summarizing the day's events is sent to the manager.
[1759] Example prompt sentence:
[1760] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[1761] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1762] Step 1:
[1763] Data Collection and Transfer
[1764] The user starts the system, and the recording terminal records the visitor's video and audio data in real time. The input is video data from the camera and audio data from the microphone. The output is to transfer this data to the server. Specifically, the camera captures the visitor's video and the microphone records the audio. The video data is then compressed into H.264 format or similar, and the audio data is converted into WAV format or similar before being sent to the server using the TCP / IP protocol.
[1765] Step 2:
[1766] Data Preprocessing
[1767] The server removes noise from the received video and audio data and converts the audio data to text. The input is the received video and audio data. The output is noise-removed video data, audio data, and text data. Specifically, the OpenCV library is used to remove noise from the video data, and filtering technology is used to clean the audio data. A speech recognition API (e.g., Google Cloud Speech-to-Text API) is used to convert the audio to text.
[1768] Step 3:
[1769] Feature extraction and quantification
[1770] The server uses natural language processing technology to analyze the phrasing and content of the conversation from the speech-recognized text data, analyzes speaking style and tone of voice from the audio data, and extracts facial expressions and body language from the video data. The inputs include cleaned video data, audio data, and text data. The output is quantified feature data. Specifically, NLTK or spaCy is used for natural language processing, and the LibROSA library is used to analyze speaking style and tone of voice from the audio data. Facial expression recognition and body language analysis are performed from the video data using OpenCV and Dlib libraries.
[1771] Step 4:
[1772] Analysis by emotion engine
[1773] The server uses the emotion_recognition library to analyze emotions from video and audio data. The input is quantified feature data. The output is detected emotion data. Specific operations include inferring emotions from facial expressions in the video, tone of voice, and linguistic expressions in the text. For example, by combining OpenCV and TensorFlow, emotions such as joy, anger, and surprise can be identified in real time based on facial expressions.
[1774] Step 5:
[1775] Issuance of an alert
[1776] If the server detects suspicious behavior or abnormal emotional fluctuations, it will issue an alarm through the alarm system. The inputs are detected emotional data and behavioral data. The output is an alarm message. Specifically, if a certain emotional score exceeds a threshold, the alarm system will be activated and an audio alert or text notification will be sent to security staff.
[1777] Step 6:
[1778] Report Generation
[1779] The server generates a report based on the analysis results and provides it to the user. The input is quantified and analyzed feature data. The output is a daily or monthly report. Specifically, it aggregates the analysis results stored in the database and automatically generates a report according to a template. The generated report is saved in PDF or CSV format and provided to the user via email or dashboard.
[1780] Example prompt sentence:
[1781] "Can you give me an example of a program that analyzes emotions from audio and video data and detects abnormal emotional fluctuations?"
[1782] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1783] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1784] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1785] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1786] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1787] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1788] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1789] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1790] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1791] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1792] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1793] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1794] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1795] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1796] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1797] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1798] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1799] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1800] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1801] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1802] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1803] The following is further disclosed regarding the above embodiment.
[1804] (Claim 1)
[1805] a means for recording video and audio data of the interview;
[1806] means for transmitting the recorded video and audio data to a server;
[1807] means for performing noise reduction and text conversion on the transmitted video and audio data;
[1808] A means of analyzing the wording and content of conversations using natural language processing technology for text data;
[1809] means for analyzing speaking style and tone of voice from the audio data;
[1810] A means for analyzing facial expressions and body language from video data;
[1811] A means for quantifying and scoring each analyzed feature;
[1812] A means of training the learning model using interview data from existing employees;
[1813] a means of evaluating new interview data using the learning model;
[1814] A means of displaying the evaluation results and providing feedback to the user
[1815] A system including:
[1816] (Claim 2)
[1817] 10. The system of claim 1, further comprising means for setting evaluation criteria that match the company culture based on interview data of existing employees.
[1818] (Claim 3)
[1819] 10. The system of claim 1, further comprising means for analyzing the video and audio data of the interview in real time and displaying the results in real time.
[1820] "Example 1"
[1821] (Claim 1)
[1822] a means for recording video and audio data of the interview;
[1823] means for transmitting the recorded video and audio data to a central processing unit;
[1824] means for removing noise from the transmitted video and audio data and converting the audio into text information;
[1825] A means for analyzing the wording and content of conversation using language analysis technology for the voice data;
[1826] means for analyzing speaking style and tone of voice from the audio data;
[1827] A means for analyzing facial expressions and sign language from video data;
[1828] A means for quantifying and evaluating each analyzed feature;
[1829] A means of training a machine learning model using interview data from existing employees; and
[1830] a means of evaluating new interview data using machine learning models; and
[1831] a means for displaying the evaluation results and providing feedback to the user;
[1832] A system including:
[1833] (Claim 2)
[1834] 10. The system of claim 1, further comprising means for setting evaluation criteria that fit with organizational culture based on interview data of existing employees.
[1835] (Claim 3)
[1836] 10. The system of claim 1, further comprising means for instantly analyzing the video and audio data of the interview and instantly displaying the results.
[1837] "Application Example 1"
[1838] Rewritten claims
[1839] (Claim 1)
[1840] a means for recording video and audio data of the interview;
[1841] means for transmitting the recorded video and audio data to a server;
[1842] means for performing noise reduction and text conversion on the transmitted video and audio data;
[1843] A means of analyzing the wording and content of conversations using natural language processing technology for text data;
[1844] means for analyzing speaking style and tone of voice from the audio data;
[1845] A means for analyzing facial expressions and body language from video data;
[1846] A means for quantifying and scoring each analyzed feature;
[1847] a means for training a learning model using existing data;
[1848] a means for evaluating new data using the learned model;
[1849] means for displaying the evaluation results and providing feedback to the user;
[1850] A means of evaluating work efficiency, accuracy, safety, etc. based on the analysis results, and
[1851] A means of providing feedback on scoring results to workers and suggesting areas for improvement
[1852] A system including:
[1853] (Claim 2)
[1854] 10. The system of claim 1, further comprising means for establishing standardized evaluation criteria based on existing data.
[1855] (Claim 3)
[1856] 10. The system of claim 1, further comprising means for analyzing the video and audio data in real time and displaying the results in real time.
[1857] "Example 2: Combining Emotion Engines"
[1858] (Claim 1)
[1859] a means for recording video and audio data of the interview;
[1860] means for transmitting the recorded video and audio data to a central control unit;
[1861] means for performing noise removal and voice recognition on the transmitted video and audio data;
[1862] A means for analyzing the wording and conversation content of the speech recognition results using natural language analysis technology;
[1863] means for analyzing speaking style and tone of voice from the audio data;
[1864] A means for analyzing facial expressions and body movements from video data;
[1865] A means for quantifying and scoring each analyzed feature;
[1866] means for analyzing emotions from the recorded video and audio data;
[1867] A means of quantifying and scoring emotional data;
[1868] A means to train machine learning models using interview data from existing employees; and
[1869] a means of evaluating new interview data using the learning model;
[1870] A means of displaying the evaluation results and providing feedback to the user
[1871] A system including:
[1872] (Claim 2)
[1873] 10. The system of claim 1, further comprising means for setting evaluation criteria that match the company culture based on interview data of existing employees.
[1874] (Claim 3)
[1875] 10. The system of claim 1, further comprising means for analyzing the video and audio data and the emotional data of the interview in real time and displaying the results in real time.
[1876] "Application example 2 when combining emotion engines"
[1877] (Claim 1)
[1878] a means for recording video and audio data of the interview;
[1879] means for transmitting the recorded video and audio data to a server;
[1880] means for performing noise reduction and text conversion on the transmitted video and audio data;
[1881] A means of analyzing the wording and content of conversations using natural language processing technology for text data;
[1882] means for analyzing speaking style and tone of voice from the audio data;
[1883] A means for analyzing facial expressions and body language from video data;
[1884] A means for quantifying and scoring each analyzed feature;
[1885] A means of training the learning model using interview data from existing employees;
[1886] a means of evaluating new interview data using the learning model;
[1887] means for displaying the evaluation results and providing feedback to the user;
[1888] means for detecting and alerting to abnormal emotional fluctuations and suspicious behavior;
[1889] A means of generating reports based on analysis results
[1890] A system incl...
Claims
1. a means for recording video and audio data of the interview; means for transmitting the recorded video and audio data to a server; means for performing noise reduction and text conversion on the transmitted video and audio data; A means of analyzing the wording and content of conversations using natural language processing technology for text data; means for analyzing speaking style and tone of voice from the audio data; A means for analyzing facial expressions and body language from video data; A means for quantifying and scoring each analyzed feature; A means of training the learning model using interview data from existing employees; a means of evaluating new interview data using the learning model; A means of displaying the evaluation results and providing feedback to the user A system including:
2. The system according to claim 1, further comprising means for setting evaluation criteria suited to the company culture based on interview data of existing employees.
3. 2. The system of claim 1, further comprising means for analyzing the video and audio data of the interview in real time and displaying the results in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A