AI video interview evaluation method and device, electronic equipment and storage medium

By obtaining and analyzing interviewers' multimodal information in AI video interviews, the problem of ignoring psychological state in the existing technology is solved, and more accurate evaluation results are achieved, providing enterprises with objective evaluation data.

CN120220019APending Publication Date: 2025-06-27BEISEN CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279625.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In AI video interviews, the existing technology only tests based on the interviewer's final reply content, ignoring the interviewer's psychological state, resulting in inaccurate evaluation results.

Method used

By obtaining the multimodal information of the interviewer when answering the current interview questions, including facial images, eye tracking data, audio data and answering text, after preprocessing, this information is input into the preconstructed video interview evaluation model to obtain the evaluation results of the current interview questions, and based on this, select the next interview question and generate an evaluation report.

Benefits of technology

Through multi-dimensional comprehensive assessment, the accuracy of the assessment results is significantly improved, and the company provides objective and comprehensive real assessment data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220019A_ABST
    Figure CN120220019A_ABST
Patent Text Reader

Abstract

The invention provides an AI video interview evaluation method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining multi-modal information when an interviewer answers a current test question; respectively preprocessing the face image, the eye tracking data, the audio data and the answer text; jointly inputting the preprocessed face image, the eye tracking data, the audio data and the answer text into a pre-constructed video interview evaluation model to obtain an evaluation result of the current interview test question; selecting a next face test question in a face test question database based on the evaluation result of the current face test question; and after the interview is finished, generating an evaluation report based on the evaluation result of each test question and the actual reply result. According to the method, comprehensive evaluation is carried out from multiple dimensions, and the accuracy of the evaluation result is greatly improved. Therefore, objective and comprehensive real evaluation data are provided for enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an AI video interview evaluation method, device, electronic device and storage medium. Background Art

[0002] During an AI video interview, the AI interviewer often gives the next interview question based on the current actual answer of the interviewee, or directly conducts a question-and-answer interview according to a preset answer sequence. In this way, the interviewee is only tested from one dimension, such as the final text reply content of the interviewee, ignoring the psychological state of the interviewee. For example, in order to get a high score or meet the management requirements, the answerer often chooses some answers that he is not good at or does not like, so as to meet the enterprise's requirements for talents. As a result, the obtained interview evaluation results are often inaccurate. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an AI video interview evaluation method, device, electronic device and storage medium to improve the evaluation accuracy of AI video interviews.

[0004] In a first aspect, an AI video interview evaluation method is provided, which is applied to the server side. The method includes:

[0005] Obtain multi-modal information of the interviewee when answering the current interview question. The multi-modal information at least includes a facial image, eye tracking data, audio data and answer text;

[0006] Preprocess the facial image, eye tracking data, audio data and answer text respectively;

[0007] Input the preprocessed facial image, eye tracking data, audio data and answer text into a pre-constructed video interview evaluation model together to obtain the evaluation result of the current interview question;

[0008] Select the next interview question from the interview question database based on the evaluation result of the current interview question;

[0009] After the interview ends, generate an evaluation report based on the evaluation results and actual answer results of each interview question.

[0010] Optionally, preprocessing the facial image, eye tracking data, audio data and answer text respectively includes:

[0011] Perform one or more of the following processing operations on the facial image: cropping, rotating and enhancing contrast;

[0012] And perform denoising and frame preprocessing on the audio data;

[0013] And convert the answer text into a word feature vector through a word vector model;

[0014] and performing noise removal, missing value filling, and data standardization on the eye tracking data.

[0015] Optionally, the video interview assessment model includes an expression recognition unit, a voice emotion recognition unit, an eye tracking unit, an eye recognition unit, a text recognition unit, and a fusion output unit; the preprocessed facial images, eye images, audio data, and answer texts are jointly input into the pre-constructed video interview assessment model, and the assessment results of the current interview question include:

[0016] Input the preprocessed facial image into the expression recognition unit for recognition to obtain an expression category label; the expression category label includes at least nervous, hesitant, confused, and relaxed;

[0017] Input the preprocessed audio data into the voice emotion recognition unit for recognition to obtain a voice emotion category label; the voice emotion category label includes at least: nervous, hesitant, and confident;

[0018] Input the preprocessed answer text into the text recognition unit to obtain a text classification label, and the text classification label includes at least: a language expression fluency label and a confidence label;

[0019] Input the preprocessed eye tracking data into the eye emotion recognition unit for recognition to obtain an eye emotion category label, and the eye emotion category label includes at least nervous, hesitant, and confident;

[0020] Input the expression category label, the voice emotion category label, the text classification label, and the emotion category label into the fusion output unit for weighted fusion to output the assessment result of the current interview question.

[0021] Optionally, the eye emotion recognition unit includes a first eye emotion recognition subunit and a second eye emotion recognition subunit, and the eye tracking data includes the line of sight direction of the pupil relative to the video interview display, the dwell time of each line of sight direction, and the blink frequency; input the preprocessed eye tracking data into the eye emotion recognition unit for recognition to obtain an eye emotion category label including:

[0022] When the type of the current interview question is a multiple-choice question, input the line of sight direction and the dwell time of each line of sight direction into the first eye emotion recognition subunit, so that the first eye emotion recognition subunit determines the speculated option of the current interview question corresponding to the line of sight direction based on the line of sight direction with the longest dwell time, and outputs an eye emotion category label based on the speculated option of the interviewee and the actually selected option;

[0023] When the type of the current interview question is a short-answer question, input the blink frequency into the second eye emotion recognition subunit, so that the second eye emotion recognition subunit outputs an eye emotion category label based on the blink frequency.

[0024] Optionally, the first eye emotion recognition subunit determines that the speculated options for the current interview question corresponding to the line-of-sight direction with the longest stay duration include:

[0025] Map the line-of-sight direction with the longest stay duration to the coordinates on the video interview display through coordinate system conversion;

[0026] Based on the coordinates of the line-of-sight direction on the video interview display and the coordinate information of each option of the current interview question on the video interview display, determine the speculated option for the current interview question corresponding to the line-of-sight direction with the longest stay duration.

[0027] Optionally, the evaluation result includes a positive evaluation result and a negative evaluation result; selecting the next interview question from the interview question database based on the evaluation result of the current interview question includes:

[0028] If the evaluation result of the current interview question is a negative evaluation result, select an interview question related to the current interview question from the interview question database; the negative evaluation result is at least confusion or hesitation;

[0029] If the evaluation result of the current interview question is a positive evaluation result, select an interview question not related to the current interview question from the interview question database; the positive evaluation result is at least confidence or relaxation.

[0030] Optionally, generating an evaluation report based on the evaluation results and actual response results of each interview question includes:

[0031] Generate a speculated evaluation report based on the evaluation results of each interview question;

[0032] Generate an actual evaluation report based on the actual response results of each interview question.

[0033] On the second aspect, provide an AI video interview evaluation device, which is applied to the server side. The device includes:

[0034] An acquisition unit, configured to acquire multimodal information of the interviewee when answering the current interview question. The multimodal information at least includes a facial image, eye tracking data, audio data, and answer text;

[0035] A preprocessing unit, configured to preprocess the facial image, eye tracking data, audio data, and answer text respectively;

[0036] An evaluation unit, configured to jointly input the preprocessed facial image, eye tracking data, audio data, and answer text into a pre-constructed video interview evaluation model to obtain the evaluation result of the current interview question;

[0037] A selection unit for selecting the next interview question in the interview question database based on the evaluation result of the current interview question;

[0038] A generation unit for generating an evaluation report based on the evaluation results and actual response results of each interview question after the interview ends.

[0039] In a third aspect, there is provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0040] The memory is used for storing a computer program;

[0041] The processor, when executing the program stored on the memory, implements the method steps described in any one of the first aspect.

[0042] In a fourth aspect, there is provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps described in any one of the first aspect are implemented.

[0043] An AI video interview evaluation method, device, electronic device, and storage medium provided by the present invention obtain multi-modal information of an interviewee when answering the current interview question; preprocess the facial image, eye tracking data, audio data, and answer text respectively; jointly input the preprocessed facial image, eye tracking data, audio data, and answer text into a pre-constructed video interview evaluation model to obtain the evaluation result of the current interview question; select the next interview question in the interview question database based on the evaluation result of the current interview question; and generate an evaluation report based on the evaluation results and actual response results of each interview question after the interview ends. The present invention conducts comprehensive evaluation from multiple dimensions, greatly improving the accuracy of the evaluation result. Thus, objective and comprehensive real evaluation data are provided for enterprises.

[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1Shows a flowchart of an AI video interview assessment method provided by an embodiment of the present invention;

[0047] Figure 2 Shows a schematic structural diagram of an AI video interview assessment device provided by an embodiment of the present invention;

[0048] Figure 3 Shows a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Considering that during an AI video interview, the AI interviewer often gives the next interview question based on the current actual answer of the interviewee, or directly conducts a question-and-answer interview according to a preset answer sequence. In this way, the interviewee is only tested from one dimension, such as the final reply content of the interviewee, ignoring the psychological state of the interviewee. For example, in order to get a high score or meet the management requirements, the answerer often chooses some answers that he is not good at or does not like, so as to conform to the enterprise's requirements for talents. As a result, the obtained assessment results are often inaccurate.

[0051] Based on this, the embodiments of the present invention provide an AI video interview assessment method and device, which will be described below through embodiments.

[0052] An embodiment of the present invention provides an AI video interview assessment method, which is applied to an AI video interview system. The system includes an AI video interview terminal and a server side, and the AI video interview terminal and the server side are connected by a wired or wireless communication method. The AI video interview terminal includes a display, a camera embedded in the display, an eye tracker (such as an eye movement instrument), and a microphone. After the interviewee enters the interview room and stands in front of the display, the camera is used to collect the facial image of the interviewee, the eye tracker is used to track the eye data of the interviewee, the microphone is used to collect the voice data of the interviewee, and the display is used to display the interview questions and the virtual characters of the AI interviewer. A video interview assessment model is deployed in the server, which is used to assess the interview process of the interviewee.

[0053] An embodiment of the present invention provides an AI video interview assessment method, which is applied to the server side. As Figure 1 shown, the method includes the following steps:

[0054] Step S101: Obtain the multimodal information of the interviewee when answering the current interview question. The multimodal information includes at least facial images, eye tracking data, audio data, and answer texts.

[0055] In this step, the facial image can be obtained through the camera; the eye tracking data can be obtained through an eye tracker such as an eye movement instrument; the audio data of the interviewee can be picked up through the microphone; and the answer text finally presented on the display by the interviewee through oral narration, handwriting, etc.

[0056] Step S102: Preprocess the facial image, eye tracking data, audio data, and answer text respectively.

[0057] Before inputting the collected data into the video interview assessment model, preprocessing it is a very crucial step. Data preprocessing not only helps to improve the quality of the data, but also directly or indirectly affects the effectiveness and reliability of the final analysis results.

[0058] The specific preprocessing method will be described in detail in the following embodiments and will not be elaborated here.

[0059] Step S103: Input the preprocessed facial image, eye tracking data, audio data, and answer text into the pre-constructed video interview assessment model together to obtain the assessment result of the current interview question.

[0060] In this step, the video interview assessment model is a fusion model that can process multiple modalities of information simultaneously and outputs a more accurate assessment result by fusing the relationships between multiple modalities of information.

[0061] The network structure and specific evaluation process of the video interview evaluation model will be described in detail in the following embodiments and will not be elaborated here.

[0062] Step S104: Select the next interview question from the interview question database based on the evaluation result of the current interview question.

[0063] In a feasible implementation manner, the evaluation results include positive evaluation results and negative evaluation results; selecting the next interview question from the interview question database based on the evaluation result of the current interview question includes:

[0064] If the evaluation result of the current interview question is a negative evaluation result, select an interview question related to the current interview question from the interview question database; the negative evaluation result is at least confusion or hesitation.

[0065] If the evaluation result shows that the candidate is confused or hesitant, this may indicate that they have difficulties or knowledge blind spots in dealing with specific types of questions, or in order to meet the enterprise's requirements, they choose some answers that they are not good at or do not like, so as to get closer to the enterprise's requirements for talents.

[0066] To solve these problems, the system will select other questions related to the current interview question. These related questions may re-examine the same knowledge points or skills from different angles or difficulty levels, thereby helping to improve the accuracy of the evaluation results.

[0067] If the evaluation result of the current interview question is a positive evaluation result, select an interview question unrelated to the current interview question from the interview question database; the positive evaluation result is at least confidence or relaxation.

[0068] When the interviewee shows a high degree of confidence and relaxation, it means that they may have mastered the knowledge or skills related to the current question and can perform well in similar scenarios. In this case, the system will select other questions unrelated to the current interview question from the interview question database. Doing so can further explore the performance of the candidate in different fields or different types of challenges and ensure a comprehensive evaluation of their comprehensive qualities.

[0069] Step S105: After the interview, generate an evaluation report based on the evaluation results and actual response results of each interview question.

[0070] In the embodiment of the present invention, generating an evaluation report based on the evaluation results and actual response results of each interview question includes:

[0071] Generate a speculative evaluation report based on the evaluation results of each interview question.

[0072] In the embodiments of the present invention, the speculative evaluation report is a report obtained by speculation based on data from multiple dimensions of the interviewee. For example, when the interviewee shows hesitation and their eyes are fixed on option A while doing a certain interview question but finally selects option B, then by combining these two types of data, the original intention of the interviewee is speculated to generate a speculative evaluation report.

[0073] Generate an actual evaluation report based on the actual response results of each interview question.

[0074] For the actual evaluation report, it is based on the actual response results of the interviewee. For example, in the previous example, if the interviewee selects option B, then the evaluation report is generated according to the final selected result.

[0075] Based on the differences between the two evaluation reports, the recruiter can further communicate deeply with the candidate to understand their true coping ideas and potential abilities.

[0076] For the recruitment party, this two-track evaluation method helps to make a more scientific and reasonable employment decision, and at the same time can improve the professionalism and transparency of the entire recruitment process.

[0077] The present invention comprehensively evaluates by collecting various modal information of the interviewee, greatly improving the accuracy of the evaluation results. Thus, it provides objective and comprehensive real evaluation data for the enterprise.

[0078] Based on the above embodiments, preprocessing is respectively performed on the facial image, eye tracking data, audio data, and answer text, including:

[0079] Step S102A: Perform one or more of the following processing operations on the facial image: cropping, rotating, and enhancing contrast.

[0080] In this step, by cropping, the irrelevant parts in the image are removed, and only the facial area is retained, which helps to reduce the influence of background noise on the analysis. Adjust the image angle to keep the face horizontal, which helps to standardize the input and is convenient for using pre-trained models or algorithms. By adjusting the contrast of the image, the facial features can be made more clear, which helps to improve the effect of tasks such as face recognition or expression analysis.

[0081] Step S102B: Perform denoising and frame segmentation preprocessing on the audio data.

[0082] In this step, by eliminating environmental noise or other unnecessary sound interferences in the recording, the quality of speech recognition or emotion analysis is improved.

[0083] In addition, the continuous audio signal is segmented into shorter time periods (frames) for subsequent spectral analysis or feature extraction.

[0084] Step S102C: Convert the answer text into word feature vectors through a word vector model.

[0085] In one example, use word vector models such as Word2Vec, GloVe, or BERT to convert the text into a numerical form (word feature vectors) that can be understood by a computer. This method not only considers the vocabulary itself but also captures the semantic relationships between words, helping to improve the performance of tasks such as text classification and sentiment analysis.

[0086] Step S102D: Perform noise removal, missing value filling, and data standardization on the eye tracking data.

[0087] In this step, ensure the data quality by filtering out outliers or errors that may be introduced during the data collection process. For missing data points, they can be estimated and supplemented through interpolation or other methods to ensure the integrity and continuity of the time series. Convert data of different scales to the same scale, such as a standard normal distribution with a mean of 0 and a standard deviation of 1, which can avoid inappropriate effects on the results due to scale differences of certain features.

[0088] Based on the above embodiments, the video interview assessment model includes a facial expression recognition unit, a voice emotion recognition unit, an eye tracking unit, an eye recognition unit, a text recognition unit, and a fusion output unit.

[0089] Each unit is an independent small model that can directly output predicted emotion categories, etc. through pre-training.

[0090] Input the preprocessed facial images, eye images, audio data, and answer text into the pre-constructed video interview assessment model to obtain the assessment results for the current interview question, including:

[0091] Step S103A: Input the preprocessed facial image into the facial expression recognition unit for recognition to obtain a facial expression category label.

[0092] In the embodiments of the present invention, the facial expression recognition unit uses a convolutional neural network CNN, which includes an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer.

[0093] Among them, the input layer is for receiving the preprocessed facial image.

[0094] The convolutional layer uses multiple convolutional kernels to extract the features of the facial image, and the expression of the convolutional kernel is:

[0095]

[0096] Among them, Z ij is the convolutional output; x is the input facial image; w mnis a convolution kernel; M is the width of the convolution kernel; N is the height of the convolution kernel; m is the index of the width, ranging from 0 to M - 1; n is the index of the height, ranging from 0 to N - 1.

[0097] The pooling layer then performs downsampling to reduce the data dimension of the image, thereby reducing the data calculation amount of the subsequent layers. In an example, the expression of this pooling layer is:

[0098] y ij = max m,n input(R ij ,x m,n )(2);

[0099] where, R ij is the pooling region; y ij is the pooling output; x m,n is the facial feature image output by one of the convolution kernels.

[0100] The fully connected layer maps the features extracted by the convolutional layer and the pooling layer to the final expression categories, and finally the output layer outputs the probability distribution of each expression category through the Softmax function.

[0101] In this step, the expression category labels at least include nervous, hesitant, confused, and relaxed.

[0102] When pre-training the convolutional neural network model CNN, it is necessary to collect a large amount of face image data of users showing different expressions (such as hesitant, painful, happy, relaxed, etc.) when answering questions, and perform accurate expression annotation. For example, 1 represents hesitant, 2 represents painful, 3 represents happy, 4 represents relaxed, etc. Then, the annotated data and the collected image data are input into the CNN for iterative training.

[0103] Step S103B: Input the preprocessed audio data into the voice emotion recognition unit for recognition to obtain the voice emotion category label.

[0104] In this step, since the audio data belongs to a kind of sequential data, the voice emotion recognition unit uses a recurrent neural network RNN or a long short-term memory network LSTM to recognize the voice emotion. In a feasible implementation manner, first, the preprocessed audio data is used to extract voice features through mel-frequency cepstral coefficients, such as pitch features, volume features, etc., and then the LSTM network is used to recognize the voice emotion category. The expression of this LSTM network is:

[0105] h t = σ(W xh x t + W hh h t-1 + b h )(3);

[0106] Among them, h t is the hidden state at time point t; x t is the input vector at time point t; W xh , W hh are the weight matrices from the input to the hidden layer and from the hidden layer at the previous moment to the hidden layer at the current moment respectively; h t-1 is the previous hidden state; b h is the bias; σ is the activation function.

[0107] The voice emotion category labels at least include: nervous, hesitant, and confident.

[0108] Step S103C: Input the preprocessed answer text into the text recognition unit to obtain text classification labels.

[0109] The text classification labels at least include: language expression fluency labels, confidence labels.

[0110] In this step, the text recognition unit can also adopt a convolutional neural network CNN. Details are not elaborated here.

[0111] In one example, the language expression fluency labels can be divided into low, medium, and high, represented by 1, 2, and 3 respectively (1 - low, 2 - medium, 3 - high); similarly, the confidence labels can be divided into low, medium, and high, represented by 1, 2, and 3 respectively (1 - low, 2 - medium, 3 - high).

[0112] Step S103D: Input the preprocessed eye tracking data into the eye emotion recognition unit for recognition to obtain eye emotion category labels, and the eye emotion category labels at least include nervous, hesitant, and confident.

[0113] In this step, by combining the eye tracking data with the facial image data, the mental state of the interviewee can be more accurately recognized, so as to evaluate the interview state and result of the interviewee in combination with the objective answer text, improving the accuracy of the evaluation.

[0114] Step S103E: Input the expression category labels, voice emotion category labels, text classification labels, and emotion category labels into the fusion output unit for weighted fusion to output the evaluation result of the current interview question.

[0115] The embodiments of the present invention adopt a late fusion method, that is, after each unit outputs the recognition category results respectively, the results are weighted and fused to comprehensively judge the answering state and emotion of the interviewee. The specific calculation method of late fusion is that first, weights need to be set for each input label. Multiply the classification results in each modality by their corresponding weights, and add these weighted results to obtain the final probability distribution. Based on the calculated final probability distribution, the system can make a final judgment on the current state and emotion of the interviewee.

[0116] Suppose in an interview scenario, the system analyzes the interviewee's facial expressions, intonation changes, and answering content simultaneously. Facial expression analysis may show that the candidate is nervous; voice analysis reveals some signs of uncertainty; while text analysis shows a logically clear answer. By weighted fusing this information, the system can give a more comprehensive and accurate assessment, reflecting the true state and emotional response of the candidate during the interview.

[0117] Late fusion is a powerful and flexible method, suitable for situations where multiple different types of data need to be integrated to gain in-depth insights. For complex tasks such as interview assessment, it can effectively improve the quality and reliability of decision-making.

[0118] Based on the above embodiments, the eye emotion recognition unit includes a first eye emotion recognition subunit and a second eye emotion recognition subunit. The eye tracking data includes the line of sight direction of the pupil relative to the video interview display, the dwell time of each line of sight direction, and the blink frequency. The preprocessed eye tracking data is input into the eye emotion recognition unit for recognition to obtain eye emotion category labels including:

[0119] Step S103D1: When the type of the current interview question is a multiple-choice question, input the line of sight direction and the dwell time of each line of sight direction into the first eye emotion recognition subunit, so that the first eye emotion recognition subunit determines the speculated option of the current interview question corresponding to the line of sight direction based on the line of sight direction with the longest dwell time, and outputs the eye emotion category label based on the interviewee's speculated option and the actually selected option.

[0120] In the embodiments of the present invention, based on the direction with the longest dwell time, the system infers the answer that the interviewee may tend to choose. Compare the inferred answer with the answer actually selected by the interviewee, and output the eye emotion category label according to the degree of difference. For example, if the interviewee's line of sight is concentrated on option A for a long time but finally selects option B, this may indicate that the interviewee has hesitation or uncertainty; there may be a situation where the interviewee chooses an answer he doesn't like in order to meet the enterprise's requirements; vice versa.

[0121] Step S103D2: When the type of the current interview question is a short-answer question, input the blink frequency into the second eye emotion recognition subunit, so that the second eye emotion recognition subunit outputs an eye emotion category label based on the blink frequency.

[0122] For short-answer questions, since there are no clear options for the interviewee to choose, the blink frequency is used as one of the indicators to measure the emotional state. High-frequency blinking may indicate emotions such as tension and anxiety; while a stable blink frequency may be a manifestation of relaxation or confidence. According to different patterns of the blink frequency, the second eye emotion recognition subunit will output corresponding eye emotion category labels, such as "tense", "relaxed", etc.

[0123] Through more detailed eye emotion recognition in the embodiments of the present invention, finer emotional fluctuations can be captured, thereby providing more accurate and reliable data support for the evaluation of the entire interview result.

[0124] Based on the above embodiments, the first eye emotion recognition subunit determines the speculated options for the current interview question corresponding to the line-of-sight direction with the longest dwell time, including:

[0125] Step A: Map the line-of-sight direction with the longest dwell time to the coordinates on the video interview display through coordinate transformation.

[0126] The line-of-sight direction is the line-of-sight coordinate data in the coordinate system with the eye tracker as the origin, so it is converted into the coordinate data in the coordinate system with the video display as the origin.

[0127] It should be noted that if the eye tracker is embedded in the display, coordinate transformation may not be required, and the coordinates of the line-of-sight direction on the display can be directly obtained.

[0128] Step B: Determine the speculated options for the current interview question corresponding to the line-of-sight direction with the longest dwell time based on the coordinates of the line-of-sight direction on the video interview display and the coordinate information of each option of the current interview question on the video interview display.

[0129] In a specific example, assume that the resolution of the video interview display is 1920x1080 pixels.

[0130] The layout of the interview questions and options on the display is as follows:

[0131] Question area: Located in the center of the top of the screen (specific coordinates are not considered);

[0132] The four options are respectively located in the middle and lower part of the screen, arranged in a rectangle, and each option occupies 1 / 4 of the screen width, with a fixed height.

[0133] Option A: Upper left corner coordinates (0,700), lower right corner coordinates (480,850);

[0134] Option B: The upper left corner coordinates are (480, 700), and the lower right corner coordinates are (960, 850);

[0135] Option C: The upper left corner coordinates are (960, 700), and the lower right corner coordinates are (1440, 850);

[0136] Option D: The upper left corner coordinates are (1440, 700), and the lower right corner coordinates are (1920, 850);

[0137] Assume that the line-of-sight direction provided by the eye tracker corresponds to the coordinates (600, 800) on the display. This coordinate is within the area of Option B. Therefore, it can be inferred that the speculated option for the current interview question corresponding to the line-of-sight direction with the longest dwell time is Option B.

[0138] Based on the same inventive concept, an AI video interview assessment device is provided, which is applied to the server side. As Figure 2 shown, the device includes:

[0139] An acquisition unit for acquiring multi-modal information of the interviewee when answering the current interview question. The multi-modal information at least includes a facial image, eye tracking data, audio data, and answer text.

[0140] A preprocessing unit for preprocessing the facial image, eye tracking data, audio data, and answer text respectively.

[0141] An assessment unit for jointly inputting the preprocessed facial image, eye tracking data, audio data, and answer text into a pre-constructed video interview assessment model to obtain the assessment result of the current interview question.

[0142] A selection unit for selecting the next interview question in the interview question database based on the assessment result of the current interview question.

[0143] A generation unit for generating an assessment report based on the assessment results and actual answer results of each interview question after the interview ends.

[0144] Based on the same technical concept, an embodiment of the present invention also provides an electronic device. As Figure 3 shown, it includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communication interface 302, and the memory 303 complete mutual communication through the communication bus 304.

[0145] The memory 303 is used to store a computer program;

[0146] The processor 301 is used to implement the steps of the AI video interview assessment method when executing the program stored on the memory 303.

[0147] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0148] The communication interface is used for communication between the above electronic device and other devices.

[0149] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0150] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0151] The computer program product for performing AI video interview assessment provided by the embodiments of the present invention includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.

[0152] The device for AI video interview assessment provided by the embodiments of the present invention can be specific hardware on a device, or software or firmware installed on the device, etc. The implementation principle and the technical effects generated by the device provided by the embodiments of the present invention are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference can be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing-described systems, devices, and units can all refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0153] In the embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0154] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, each functional unit in the embodiments provided by the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0156] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0157] It should be noted that like reference numerals and letters indicate like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.

[0158] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An AI video interview evaluation method, characterized in that: Applied to the server side, the method includes: Acquire multimodal information of the interviewee when answering the current interview question, wherein the multimodal information includes at least facial image, eye tracking data, audio data and answer text; Preprocessing the facial image, eye tracking data, audio data and answer text respectively; The pre-processed facial image, eye tracking data, audio data and answer text are input into a pre-built video interview evaluation model to obtain an evaluation result of the current interview question; Selecting a next interview question from an interview question database based on the evaluation result of the current interview question; After the interview, an evaluation report is generated based on the evaluation results of each interview question and the actual answer results.

2. The method according to claim 1, characterized in that The preprocessing of the facial image, the eye tracking data, the audio data and the answer text respectively includes: Performing one or more of the following processing operations on the facial image: cropping, rotating, and contrast enhancement; and performing denoising and frame preprocessing on the audio data; and converting the answer text into a word feature vector through a word vector model; The eye tracking data is subjected to noise removal, missing value filling and data standardization processing.

3. The method according to claim 1, characterized in that The video interview evaluation model includes an expression recognition unit, a voice emotion recognition unit, an eye tracking unit, an eye recognition unit, a text recognition unit and a fusion output unit; the pre-processed facial image, eye image, audio data and answer text are input into the pre-built video interview evaluation model, and the evaluation results of the current interview question are obtained, including: Inputting the preprocessed facial image into the expression recognition unit for recognition to obtain expression category labels; the expression category labels at least include tension, hesitation, confusion and relaxation; Inputting the pre-processed audio data into the voice emotion recognition unit for recognition to obtain a voice emotion category label; the voice emotion category label at least includes: nervous, hesitant and confident; Inputting the preprocessed answer text into the text recognition unit to obtain a text classification label, wherein the text classification label at least includes: a language expression fluency label and a confidence label; Inputting the pre-processed eye tracking data into an eye emotion recognition unit for recognition to obtain eye emotion category labels, wherein the eye emotion category labels at least include nervousness, hesitation and confidence; The expression category label, the sound emotion category label, the text classification label and the emotion category label are input into the fusion output unit for weighted fusion to output the evaluation result of the current interview question.

4. The method according to claim 3, characterized in that The eye emotion recognition unit includes a first eye emotion recognition subunit and a second eye emotion recognition subunit, and the eye tracking data includes the sight direction of the pupil relative to the video interview display, the dwell time of each sight direction, and the blinking frequency; The step of inputting the pre-processed eye tracking data into an eye emotion recognition unit for recognition to obtain an eye emotion category label comprises: When the type of the current interview question is a multiple-choice question, the gaze directions and the duration of each gaze direction are input into the first eye emotion recognition subunit, so that the first eye emotion recognition subunit determines the inferred option of the current interview question corresponding to the gaze direction based on the gaze direction with the longest duration, and outputs the eye emotion category label based on the interviewee's inferred option and the option actually selected; When the type of the current interview question is a question-and-answer question, the blinking frequency is input into the second eye emotion recognition subunit, so that the second eye emotion recognition subunit outputs an eye emotion category label based on the blinking frequency.

5. The method according to claim 4, characterized in that The first eye emotion recognition sub-unit determines, based on the gaze direction with the longest stay time, the guessing options for the current interview question corresponding to the gaze direction, including: The direction of the gaze with the longest dwell time is mapped to the coordinates on the video interview monitor through coordinate system conversion; Based on the coordinates of the gaze direction on the video interview display and the coordinate information of each option of the current interview question on the video interview display, the inferred option of the current interview question corresponding to the gaze direction with the longest stay time is determined.

6. The method according to claim 1, characterized in that The evaluation result includes a positive evaluation result and a negative evaluation result; and the selecting the next interview question in the interview question database based on the evaluation result of the current interview question includes: If the evaluation result of the current interview question is a negative evaluation result, then selecting an interview question related to the current interview question from the interview question database; the negative evaluation result is at least confusion and hesitation; If the evaluation result of the current interview question is a positive evaluation result, an interview question that is not related to the current interview question is selected from the interview question database; the positive evaluation result is at least confidence and ease.

7. The method according to claim 1, characterized in that The evaluation report generated based on the evaluation results of each interview question and the actual answer results includes: Generate an inferred evaluation report based on the evaluation results of each interview question; Generate an actual assessment report based on the actual response results of each interview question.

8. An AI video interview assessment device, characterized in that: Applied to the server side, the device comprises: An acquisition unit, used to acquire multimodal information of the interviewee when answering the current interview question, wherein the multimodal information at least includes facial image, eye tracking data, audio data and answer text; A preprocessing unit, used for preprocessing the facial image, eye tracking data, audio data and answer text respectively; An evaluation unit, used to input the pre-processed facial image, eye tracking data, audio data and answer text into a pre-built video interview evaluation model to obtain an evaluation result of the current interview question; A selection unit, configured to select a next interview question from an interview question database based on the evaluation result of the current interview question; The generation unit is used to generate an evaluation report based on the evaluation results of each interview question and the actual answer results after the interview is completed.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.