A recruitment interview monitoring system based on computer vision

Through a computer vision-based recruitment interview monitoring system, video streams, audio streams and text data are used to extract interviewee behavior characteristics and conduct real-time analysis and feedback, the subjectivity of interview evaluation is solved, and the objectivity and intelligence of the interview process are improved.

CN119991063BActive Publication Date: 2025-07-18KUNYUAN MOMENTARY CALCULATION DATA (HUBEI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510479025.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

During the recruitment interview process, the interview results are easily affected by the manager's subjective judgment, resulting in poor objectivity and consistency of evaluations. The existing multi-person selection methods are inefficient and cannot meet the needs of modern enterprises.

Method used

The computer vision-based recruitment interview monitoring system is adopted, and the video stream, audio stream and text data of the interview video are extracted through the pre-processing unit. The feature extraction unit integrates behavioral characteristics. The performance analysis unit uses the trained instantaneous performance analysis model to perform interviewer performance analysis, generates comprehensive interview performance analysis results, and displays them in real time through the feedback unit.

Benefits of technology

The objectivity and intelligence level of interview evaluation have been improved, and managers can quickly, comprehensively and in real time to grasp the performance of candidates, and assist in the objectivity and efficiency of the recruitment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991063B_ABST
    Figure CN119991063B_ABST
Patent Text Reader

Abstract

The present invention provides a recruitment interview monitoring system based on computer vision, which includes a preprocessing unit for preprocessing the acquired interview record video data to obtain video stream data, audio stream data, and corresponding text data; a feature extraction unit for extracting behavioral features based on the obtained video stream data, audio stream data, and text data, and further performing feature fusion processing on the obtained behavioral features to obtain a behavioral fusion feature vector of the interviewee; a performance analysis unit for inputting the behavioral fusion feature vector of the interviewee into a trained instantaneous performance analysis model to obtain an instantaneous performance analysis result output by the instantaneous performance analysis model; and performing continuous performance analysis on the interviewee according to the instantaneous performance analysis results within a period of time to obtain a continuous performance analysis result and generate a comprehensive interview performance analysis result. The present invention helps to improve the objectivity and intelligent level of recruitment interview monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and recruitment interview, and particularly to a recruitment interview monitoring system based on computer vision. Background Art

[0002] As human resources are increasingly valued, enterprises are facing more and more challenges in talent recruitment.

[0003] Currently, when an enterprise conducts an interview for a candidate for a position, the interview process is mostly completed by the enterprise's managers. However, when interviewing candidates, the interview results are easily affected by the managers' attention and subjective judgment, etc. (for example, the evaluation of the interviewee's interview performance is highly subjective, or the managers miss the key performance of the interviewee), which easily leads to a large difference in the interview results of the same candidate by different managers. To address the above problems, when current relatively mature enterprises conduct candidate interviews, they mostly record the interview videos and evaluate the comprehensive performance of the candidates through multi-person evaluation or secondary evaluation based on the recorded videos to improve the accuracy and objectivity of the candidate interview results. However, the above methods still have a relatively high subjectivity and low efficiency, and cannot meet the needs of modern enterprises for job recruitment. Summary of the Invention

[0004] To address the above problems, the present invention aims to provide a recruitment interview monitoring system based on computer vision.

[0005] The object of the present invention is achieved by the following technical solutions:

[0006] The present invention provides a recruitment interview monitoring system based on computer vision, including a preprocessing unit, a feature extraction unit, and a performance analysis unit. Among them,

[0007] The preprocessing unit is used to preprocess the acquired interview record video data, including extracting the video frame and sound from the interview video record data to obtain video stream data and audio stream data, and further extracting text according to the audio stream data to obtain corresponding text data;

[0008] The feature extraction unit is used to extract behavioral features based on the obtained video stream data, audio stream data, and text data, and further perform feature fusion processing on the obtained behavioral features to obtain the behavioral fusion feature vector of the interviewee;

[0009] The performance analysis unit is used to input the behavior fusion feature vector of the interviewee into the trained instantaneous performance analysis model to obtain the instantaneous performance analysis result output by the instantaneous performance analysis model; further, continuous performance analysis is performed on the interviewee according to the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result; and the comprehensive interview performance analysis result is generated according to the obtained instantaneous performance analysis result and continuous performance analysis result.

[0010] Preferably, the system further includes a data acquisition unit;

[0011] The data acquisition unit is used to acquire the interview video record data collected by the camera.

[0012] Preferably, the system further includes a feedback unit;

[0013] The feedback unit is used to perform real-time display according to the obtained comprehensive interview performance analysis result.

[0014] Preferably, the preprocessing unit includes a video extraction unit, an audio extraction unit, and a text conversion unit; where

[0015] The video extraction unit is used to extract video stream data according to the acquired interview record video data;

[0016] The audio extraction unit is used to extract audio stream data according to the acquired interview record video data;

[0017] The text conversion unit is used to extract the corresponding text data according to the acquired audio stream data;

[0018] Among them, the obtained video stream data, audio stream data, and text data are aligned to the same time axis.

[0019] Preferably, the feature extraction unit includes an expression feature extraction unit, an emotion feature extraction unit, an attention extraction unit, and a feature fusion unit; where

[0020] The expression feature extraction unit is used to perform expression analysis and processing according to the obtained video stream data to obtain the expression features of the interviewee C face (t) ;

[0021] The emotion feature extraction unit is used to perform emotion analysis and processing according to the obtained audio stream data to obtain the emotion features of the interviewee C Emo (t) ;

[0022] The attention extraction unit is used to perform keyword density analysis and processing according to the obtained text data to obtain the attention features of the interviewee C word (t);

[0023] The feature fusion unit is used to perform timeline alignment and feature fusion processing according to the obtained facial expression features C face (t) , emotional features C Emo (t) and attention features C word (t) to obtain a behavior fusion feature vector CS(t) = {C face (t), C Emo (t), C word (t)} .

[0024] Preferably, the performance analysis unit includes an instantaneous analysis unit, a continuous analysis unit, and a generation unit;

[0025] The instantaneous analysis unit is used to input the behavior fusion feature vector at t time into the trained instantaneous performance analysis model to obtain the instantaneous performance analysis result output by the instantaneous performance analysis model, where the instantaneous performance analysis result includes the probability features of the interviewee for different performance types;

[0026] The continuous analysis unit is used to perform continuous performance analysis on the interviewee according to the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result, where the continuous performance analysis result includes the performance stability score of the interviewee;

[0027] The generation unit is used to generate a real-time comprehensive interview performance analysis result according to the obtained instantaneous performance analysis result and continuous performance analysis result of the interviewee.

[0028] Preferably, in the instantaneous analysis unit, the adopted instantaneous performance analysis model is built based on a CNN neural network, where the instantaneous performance analysis model includes an input layer, a first convolutional layer, a second convolutional layer, a feature compression layer, a fully connected layer, a classifier layer, and an output layer connected in sequence; among them, the input layer is used to input 128 dimensional behavior fusion feature vector; the first convolutional layer includes 64 convolutional filters, the kernel size is 5×5 , the stride is 1 , and the activation function is ReLU ; a max pooling layer is also set after the first convolutional layer, where the pooling kernel size is 3×3 , the stride is 2 ; the second convolutional layer includes 128 convolutional filters, the kernel size is

[0029] 3×3 , with a step size of 1 , and the activation function is ReLU ; After the second convolutional layer, there is still a max pooling layer, where the pooling kernel size is 2×2 , with a step size of 2 ; The feature compression layer performs global average pooling on the output of the second convolutional layer to obtain 128 -dimensional intermediate feature quantity; The fully connected layer contains 512 neurons, and the activation function is ReLU , and a Dropout layer is set after the fully connected layer, with a dropout rate of 50% , to obtain the output neuron features; The classifier layer classifies the interview performance based on the neuron features, where the classifier used is Softmax, The obtained feature output corresponds to the corresponding interview performance classification result; The output layer outputs the final instantaneous interview performance classification result according to the feature output of the classifier layer.

[0030] Preferably, the continuous analysis unit specifically includes:

[0031] Obtain t The instantaneous interview performance analysis results within a time period before the Y(t - K + 1), …, Y(t - k), …, Y(t) ; Among them, the instantaneous interview analysis results at each moment include N -dimensional features; Calculate the current stability evaluation factor according to the instantaneous interview performance analysis results at each moment: , where in the formula, KP(t) represents t The stability evaluation factor at the moment, P n represents the average value of the corresponding n -dimensional features in each interview performance analysis result within the time period, and the variable n = 1, 2, …, N, N represents the total number of feature dimensions; φ represents the set balance coefficient, JS(Y(c) || Y(c - 1)) represents the instantaneous interview analysis result at time c and c-1 The JS divergence value between the instantaneous interview analysis results at the moment; The variable c = t - K +1,…,t ;

[0032] Generate the current continuous performance analysis result according to the obtained stability evaluation factor V(t) .

[0033] Preferably, the generation unit specifically includes:

[0034] According to the obtained instantaneous performance analysis results of the interviewee Y(t) and the continuous performance analysis results V(t)Integrate them to obtain the comprehensive performance analysis result corresponding to time t U(t).

[0035] The beneficial effects of the present invention are as follows: The present invention proposes a recruitment interview monitoring system based on computer vision. Based on the interview record video data collected during the interview process of the interviewee, first, the preprocessing unit separates the image and sound from the interview record video data, and further performs text recognition according to the audio data, so as to obtain the recorded data in three dimensions: image, sound, and text. Based on the recorded data in the three dimensions obtained, the feature extraction unit extracts the behavioral features of the candidate during the interview respectively, and analyzes the real-time performance and continuous performance of the interviewee within a continuous period based on the obtained behavioral features, so as to objectively evaluate the real-time performance and the stability of the performance of the interviewee, assisting the manager to quickly master the performance of the candidate during the interview process and improving the intelligent level of interview monitoring.

[0036] Monitoring and feedback on the interview performance of the interviewee in real time based on computer vision can assist the manager to comprehensively, real-time and objectively monitor the performance of the interviewee during the interview process, as a reference for the manager's evaluation of the interviewee's performance, and improve the effect of recruitment interview monitoring. Description of the Drawings

[0037] The present invention is further described with the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to the following drawings without creative efforts.

[0038] Figure 1 It is the framework structure diagram of a recruitment interview monitoring system based on computer vision shown in the embodiment of the present invention;

[0039] Figure 2 It is the schematic diagram of the exemplary functional module setting of the present invention. Detailed Embodiments

[0040] The present invention is further described in combination with the following application scenarios.

[0041] See Figure 1 As shown in the embodiment, it shows a recruitment interview monitoring system based on computer vision, including a preprocessing unit, a feature extraction unit and a performance analysis unit; among them,

[0042] The preprocessing unit is used to preprocess the obtained interview record video data, including extracting the picture and sound from the interview video record data to obtain the video stream data and the audio stream data, and further performing text extraction according to the audio stream data to obtain the corresponding text data;

[0043] The feature extraction unit is used to extract behavior features based on the obtained video stream data, audio stream data, and text data, and further perform feature fusion processing on the obtained behavior features to obtain the behavior fusion feature vector of the interviewee.

[0044] The performance analysis unit is used to input the behavior fusion feature vector of the interviewee into the trained instantaneous performance analysis model to obtain the instantaneous performance analysis result output by the instantaneous performance analysis model; further perform continuous performance analysis on the interviewee based on the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result; generate a comprehensive interview performance analysis result based on the obtained instantaneous performance analysis result and continuous performance analysis result.

[0045] In the above embodiment of the present invention, a recruitment interview monitoring system based on computer vision is proposed. Based on the interview record video data collected during the interview of the interviewee, the preprocessing unit is first used to separate the image and sound of the interview record video data, and further perform text recognition based on the audio data, so as to obtain the recorded data in three dimensions: image, sound, and text. Based on the obtained recorded data in three dimensions, the feature extraction unit is used to extract the behavior features of the candidate during the interview respectively. Based on the obtained behavior features, the real-time performance and continuous performance of the interviewee within a continuous period of time are analyzed, so as to objectively evaluate the real-time performance and performance stability of the interviewee, and assist the manager to quickly master the performance of the candidate during the interview, improving the intelligent level of interview monitoring.

[0046] Monitoring and feedback on the interview performance of the interviewee in real time based on computer vision can assist the manager to comprehensively, real-time, and objectively monitor the performance of the interviewee during the interview, as a reference for the manager's evaluation of the interviewee's performance, and improve the effect of recruitment interview monitoring.

[0047] Optionally, the recruitment interview monitoring system based on computer vision proposed by the present invention can be built based on an enterprise private cloud server or a local server, etc. By collecting the video record data during the interview of the interviewee (including on-site interview or remote interview), and analyzing and processing the collected interview record video data based on the server, it can quickly respond to the real-time performance of the interviewee, and thus objectively and comprehensively feedback the performance of the interviewee based on the facial expression features, emotional features, and attention features of the interviewee, assisting the manager to quickly master the current performance of the interviewee during the interview.

[0048] Preferably, the system further includes a data acquisition unit;

[0049] The data acquisition unit is used to acquire the interview video record data collected by the camera.

[0050] For different scenarios such as on-site enterprise interviews or remote interviews, by setting up a dedicated data acquisition unit to acquire the interview video record data collected for the interviewees, it can adapt to the use of different recruitment interview scenarios.

[0051] Preferably, referring to Figure 2 , the preprocessing unit includes a video frame extraction unit, an audio extraction unit, and a text conversion unit; among them,

[0052] The video frame extraction unit is used to extract video stream data according to the acquired interview record video data;

[0053] The audio extraction unit is used to extract audio stream data according to the acquired interview record video data;

[0054] The text conversion unit is used to extract corresponding text data according to the acquired audio stream data;

[0055] Among them, the obtained video stream data, audio stream data, and text data are aligned to the same timeline.

[0056] For the real-time acquired interview video record data, the preprocessing unit is used to extract the video stream data and audio stream data from the interview video record data. At the same time, based on the extracted audio stream data, the corresponding text data is further obtained through text extraction, and the three-dimensional behavior data obtained is aligned on the timeline, laying a foundation for further extracting behavior characteristics and performance analysis according to the real-time behavior data of the interviewees.

[0057] Among them, for the video frame extraction and audio extraction of the interview record video data, different or the same video editing applications can be used to complete, such as Adobe Premiere, Final Cut Pro, etc. For the extraction of text data, audio-to-text models such as Wav2Vec 2.0, Whisper, etc. can be used.

[0058] Among them, considering that in the acquired video stream data, since the picture of the interview video recording data is relatively stable, when extracting the facial expression features of the interviewee based on the video data subsequently, it is often affected by the background of the interview scene (for example, the front brightness of the interviewee is insufficient due to the interviewee being backlit, or the complexity of the background area (such as interference information, reflection information, etc.)), resulting in interference with the accuracy of the subsequent extraction of the facial expression features of the interviewee. Therefore, in the following embodiments, a picture enhancement unit is further provided in the picture extraction unit, and the picture enhancement unit is used to adaptively enhance the clarity of the extracted video stream data picture, so as to improve the quality of the image, and indirectly improve the accuracy and reliability of the subsequent extraction and performance analysis of the facial expression features of the interviewee. It also indirectly improves the robustness of the present invention to complex interview scenes.

[0059] Preferably, the picture extraction unit further includes a picture enhancement unit; wherein,

[0060] The picture enhancement unit is used to perform adaptive enhancement processing on each frame of video picture image in the video stream data in sequence to obtain the enhanced video stream data, specifically including:

[0061] For t the video picture image at the moment X t , respectively obtain t the average brightness feature of the video picture image at the moment Lh t and the average brightness feature of the foreground target Lm t and the background complexity feature Cb t , construct t the environmental feature vector of the video picture image at the moment S t ={Lh t , Lm t , Cb t } ;

[0062] Input the constructed environmental feature vector into the trained Q-learning model to obtain the optimal brightness value Lg t and the optimal contrast enhancement coefficient Ag t of the current image output by the Q-learning model;

[0063] According to the obtained optimal brightness value Lg t for t the video picture image at the momentX t Perform adaptive brightness adjustment, where the brightness value adjustment function used is: , where in the formula, and L t respectively represent the brightness values of the video frame image after and before adaptive brightness adjustment t at time t; ω1 and ω2 represent the set weight factors, where ω1 + ω2 = 1 , ω1 the value range of 0.3-0.5 is

[0064] Further, according to the obtained optimal contrast enhancement coefficient Ag t for t the video frame image at time t X t perform adaptive contrast adjustment, where the contrast adjustment function used is: , where in the formula, and C t respectively represent the contrasts of the video frame image after and before adaptive contrast adjustment at time t;

[0065] After sequentially completing the adaptive brightness adjustment and the adaptive contrast adjustment, output the enhanced t video frame image at time t;

[0066] Obtain the enhanced video stream data based on the enhanced video frame images at each time.

[0067] Transmit the enhanced video stream data to the feature extraction unit for further behavioral feature extraction processing.

[0068] Preferably, in the picture enhancement unit, it includes: where the reward function set in the Q-learning model is R t = PSNR(X t-1 ,X t ) , where PSNR(X t-1 ,X t ) represents t the video frame image at time X t and t-1 the video frame image at time X t-1The peak signal-to-noise ratio between is used to measure X t For X t-1 the degree of distortion; according to the obtained Q value, update the environmental factor, where the update function used is: , in the formula, Y t and Y t-1 respectively represent t the moment and t-1 the moment of the environmental factor, η represents the set learning rate, γ represents the set discount factor, R t-1 represents t-1 the reward function value at the moment; represents according to the environmental feature vector S t , brightness value Lg and contrast enhancement coefficient Ag obtained Q the maximum value of the value; Q(S t-1 , Lg t-1 , Ag t-1 ) represents according to the environmental feature vector S t-1 , t-1 the optimal brightness value at the moment Lg t-1 and the optimal contrast enhancement coefficient Ag t-1 obtained Q the value;

[0069] The Q-learning model outputs the optimal brightness value of the current image according to the obtained environmental factor , in the formula, Lb represents the set standard brightness value, and the value is 75%-80% of the maximum brightness value, Lc is the set brightness adjustment value, and the value is 10%-15% of the maximum brightness value, Y max represents the maximum value of the environmental factors at each current moment; based on the Q value corresponding to the optimal brightness value, the Q-learning model further obtains and outputs the optimal contrast enhancement coefficient , where argmax() represents the maximum value index function; that is Ag t the value of makesQ(S t , Lg t , Ag t ) of Q has the maximum value.

[0070] Optionally, set the range of the learning rate to 0.8 - 0.9, and set the range of the discount factor to 0.1 - 0.2.

[0071] Among them, the brightness value of the video frame image can adopt the brightness information obtained based on the Lab color space, HSI color space, RGB color space, etc. according to the actual situation, and the maximum brightness value is the maximum value within the value range of the brightness information in different color spaces.

[0072] Optionally, for the initial value of the environmental factor, a reasonable empirical value can be set according to historical data.

[0073] Optionally, the average brightness feature Lh of the video frame image t is obtained by calculating the average value of the brightness values of each pixel point in the image. The average brightness feature Lm of the foreground target t and the background complexity feature Cb t are obtained by using a foreground extraction model (such as a foreground extraction method based on edge detection, or a foreground extraction model built by a deep learning model based on GMM, CNN, etc.) to obtain the foreground area in the video frame image, and obtaining the average brightness feature of the foreground target according to the average brightness value of the obtained foreground area. And further perform edge pixel point detection on the background area except the foreground area (such as Canny, Sobel edge pixel detection), and obtain the background complexity feature Cb according to the ratio of the number of edge pixel points to the total number of pixel points in the background area. t .

[0074] The above-described embodiments of the present invention propose a technical solution for real-time image enhancement of acquired video stream data. Among them, first, based on the video frame image processed at the current moment, the average brightness feature of the frame, the average brightness feature of the foreground object, and the background complexity feature are respectively extracted as reference features, and then combined to reflect the brightness level of the foreground part and the complexity level of the background part in the current frame. Through the proposed Q-learning model, based on the current reference features of the image and the peak signal-to-noise ratio of the image frame change as influencing factors, and based on the Q value of the model, a comprehensive feedback on the factors of the foreground object and background complexity of the image is performed to obtain the optimal brightness value of the image at the current moment, and the brightness level of the current image is adaptively adjusted based on the optimal brightness value. Based on the image frame after brightness adjustment, again based on the idea of the maximum Q value, the contrast of the image is further adaptively adjusted, thereby improving the overall clarity of the image and the performance level of the foreground detail features, which helps to improve the accuracy and reliability of subsequent expression feature extraction and further performance analysis of the interviewee based on the commodity image frame.

[0075] Among them, in the process of performing adaptive brightness enhancement adjustment, based on the above-mentioned proposed environmental factor, the optimal brightness value of the image is updated, which can not only improve the overall brightness effect of the frame image, but also improve the smoothness of the brightness adjustment of the frame image, thereby enhancing the overall effect of the video stream data enhancement adjustment.

[0076] Preferably, the feature extraction unit includes an expression feature extraction unit, an emotion feature extraction unit, an attention degree extraction unit, and a feature fusion unit; among them,

[0077] The expression feature extraction unit is used to perform expression analysis processing on the obtained video stream data (enhanced video stream data) to obtain the expression features of the interviewee C face (t) ;

[0078] The emotion feature extraction unit is used to perform emotion analysis processing on the obtained audio stream data to obtain the emotion features of the interviewee C Emo (t) ;

[0079] The attention degree extraction unit is used to perform keyword density analysis processing on the obtained text data to obtain the attention degree features of the interviewee C word (t) ;

[0080] The feature fusion unit is used to according to the obtained expression features C face (t) 、emotion featuresC Emo (t) and attention characteristics C word (t) Perform timeline alignment and feature fusion processing to obtain a behavior fusion feature vector CS(t) = {C face (t), C Emo (t), C word (t)} 。

[0081] Based on the obtained video stream data, audio stream data, and text data as the basis, extract the facial expression features, emotional features, and attention characteristics of the interviewee respectively, and fuse the extracted features to obtain the three-dimensional behavior fusion features of the interviewee, so as to provide feedback on the performance of the interviewee during the interview, which helps to improve the reliability and objectivity of the subsequent analysis of the interviewee's performance based on this in turn.

[0082] Preferably, the facial expression feature extraction unit includes:

[0083] Based on the obtained video stream data (enhanced video stream data), use a micro-expression recognition model to obtain the facial expression features of the interviewee at each moment C face (t) ,where the facial expression features C face (t) include anger, happiness, sadness, surprise, fear, disgust, contempt, and calmness.

[0084] Optionally, the micro-expression recognition model can be implemented using a micro-expression recognition model based on 3D-ResNet or RNN / LSTM, or an image processing model that can implement facial expression recognition in the prior art. The present invention does not specifically limit this here.

[0085] Preferably, the emotional feature extraction unit includes:

[0086] Based on the obtained audio data, use a speech emotional polarity analysis model to obtain the emotional feature C Emo (t) of the interviewee, where the emotional feature C Emo (t) includes positive, neutral, and negative.

[0087] Among them, the speech emotion polarity analysis model can be implemented by using a rain sound emotion polarity analysis model built based on Wav2Vec2.0 or CNN, etc. The result output by the model is usually an emotion polarity feature quantity ranging from -1 to 1. Based on the obtained emotion polarity feature quantity and the set classification criteria, the emotion polarity feature quantity can be divided into positive, neutral, and negative. The larger the emotion polarity feature quantity, the more positive the emotion feature.

[0088] Preferably, the attention extraction unit includes:

[0089] Detect the number N of interview keywords included in the text data obtained within the current time period key , according to the number of interview keywords N key and the total number of words within this time period N t to obtain the attention feature of the interviewee C word (t) = N key / N t .

[0090] Among them, the interview keywords are keywords set according to the position for which the current interviewee is interviewed, including technical terms, technical nouns, high-frequency words of the position, etc., and can be reasonably set according to the position content or interview questions, etc.

[0091] In one scenario, the behavior fusion feature vector CS(t) is a 128-dimensional behavior fusion feature vector, and the behavior fusion feature vector is arranged in the order of " Facial expression features | Emotional features | Attention features ". Among them, the expression feature is a 64-dimensional feature vector, which contains the feature probabilities corresponding to different expressions. For example [0.1 (angry), 0.8 (happy), 0.2 (sad), 0.1 (surprised), 0 (fear), 0 (disgust), 0.1 (contemptuous), 0.3 (calm)] ; the emotion feature is a 48-dimensional feature vector, which contains the feature probabilities corresponding to different emotion feature classifications. For example [0.9 (positive), 0.1 (neutral), 0 (negative)] ; the attention feature is a 16-dimensional feature vector, which contains the obtained attention feature [0.3 (keyword occurrence frequency)] .

[0092] Through the above-mentioned method for obtaining behavioral characteristics, it is possible to provide feedback based on the key characteristics demonstrated by the interviewee during the interview. Based on the changes in the interviewee's facial expressions, emotions, and attention, the real-time behavioral characteristics of the interviewee can be captured and analyzed, serving as the basis for subsequent analysis of the interviewee's performance. Among them, through facial expression characteristics and emotional characteristics, the true situation of the interviewee in the interview scenario can be accurately represented, thus serving as the basis for performance evaluation. Furthermore, the attention characteristic is added as one of the interviewee's behavioral characteristics. Through the pertinence and relevance of the interviewee's answers to questions, the response situation of the interviewee during the process of answering interview questions can be further reflected (for example, in order to conceal nervousness or negative emotions, the interviewee delays answering questions through meaningless descriptions, etc.), further reflecting the true response of the interviewee to interview questions.

[0093] Based on the obtained behavioral fusion feature vector, the real-time interview performance of the interviewee is further objectively and accurately evaluated.

[0094] Preferably, the performance analysis unit includes an instantaneous analysis unit, a continuous analysis unit, and a generation unit;

[0095] The instantaneous analysis unit is used to input the t behavioral fusion feature vector at a certain moment into the trained instantaneous performance analysis model to obtain the instantaneous performance analysis result output by the instantaneous performance analysis model, where the instantaneous performance analysis result includes the probability characteristics of the interviewee for different performance types;

[0096] The continuous analysis unit is used to conduct continuous performance analysis on the interviewee based on the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result, where the continuous performance analysis result includes the performance stability score of the interviewee;

[0097] The generation unit is used to generate a real-time comprehensive interview performance analysis result based on the obtained instantaneous performance analysis result and continuous performance analysis result of the interviewee.

[0098] Based on the obtained behavioral fusion characteristics, first, the instantaneous performance analysis method is adopted to accurately analyze the current performance of the interviewee. Based on the obtained instantaneous performance analysis results within a period of time, further stability analysis can be carried out, thereby characterizing the stability of the interviewee's performance, assisting the manager to capture the positivity or negativity of the interviewee's instantaneous performance, and evaluating the stability of the interviewee's positive performance, so as to reflect the current performance evaluation situation of the interviewee.

[0099] Preferably, in the instantaneous analysis unit, the instantaneous performance analysis model adopted is built based on a CNN neural network. The instantaneous performance analysis model includes an input layer, a first convolutional layer, a second convolutional layer, a feature compression layer, a fully connected layer, a classifier layer, and an output layer connected in sequence. Among them, the input layer is used to input 128 -dimensional behavior fusion feature vectors; the first convolutional layer includes 64 convolutional filters, with a kernel size of 5×5 , a stride of 1 , and an activation function of ReLU ; a max pooling layer is also set after the first convolutional layer, with a pooling kernel size of 3×3 and a stride of 2 ; the second convolutional layer includes 128 convolutional filters, with a kernel size of 3×3 , a stride of 1 , and an activation function of ReLU ; a max pooling layer is also set after the second convolutional layer, with a pooling kernel size of 2×2 and a stride of 2 ; the feature compression layer performs global average pooling on the output of the second convolutional layer to obtain 128 -dimensional intermediate feature quantities; the fully connected layer includes 512 neurons, with an activation function of ReLU , and a Dropout layer is set after the fully connected layer, with a dropout rate of 50% to obtain the output neuron features; the classifier layer classifies the interview performance based on the neuron features, and the classifier adopted is Softmax, to obtain the corresponding interview performance classification results for the feature outputs; the output layer outputs the final instantaneous interview performance classification results based on the feature outputs of the classifier layer.

[0100] Optionally, in the classifier layer, the 6-dimensional feature vectors set as outputs respectively correspond to the feature probabilities of "nervous", "relaxed", "confident", "hesitant", "chaotic", and "clear" in the interview analysis results; based on the obtained feature probabilities, the output layer can be set to select the top 1-3 features with the highest probabilities as the instantaneous interview performance analysis results. Additionally, a comprehensive performance value can also be calculated based on the 6 analysis feature quantities as the instantaneous interview performance analysis result.

[0101] Among them, for the instantaneous performance analysis of the interviewee, based on the real-time behavior fusion feature vectors of the interviewee as the basis, the trained deep learning model is used to obtain the probability weights of the interviewee for different interview performances, so as to obtain the corresponding interview performance analysis results. The intelligent acquisition of the interviewee's instantaneous performance analysis results based on the deep learning method can meet the real-time requirements of interview monitoring, and help improve the accuracy, objectivity, and intelligence level of real-time interview performance analysis.

[0102] Among them, the training of the instantaneous performance analysis model can be completed based on a calibrated training set. After training the instantaneous performance analysis model based on the training set and passing the verification test of the test set, a trained instantaneous performance analysis model can be obtained.

[0103] Preferably, the continuous analysis unit specifically includes:

[0104] Obtain t The instantaneous interview performance analysis results within a time period before the moment Y(t - K + 1), …, Y(t - k), …, Y(t) ; Among them, the instantaneous interview analysis results at each moment include N d-dimensional features; calculate the current stability evaluation factor based on the instantaneous interview performance analysis results at each moment: , where in the formula, KP(t) represents t the stability evaluation factor at the moment P n represents the average value of the corresponding d-dimensional features in each interview performance analysis result within the time period, and the variable n represents the total number of feature dimensions; n = 1, 2, …, N, N represents a set balance coefficient, φ represents JS(Y(c) || Y(c - 1)) the Jensen-Shannon divergence value between the instantaneous interview analysis result at time c and c-1 the instantaneous interview analysis result at the moment JS ; the variable c = t - K +1,…,t ;

[0105] Generate the current continuous performance analysis result according to the obtained stability evaluation factor V(t) .

[0106] Among them, JS the divergence value is the Jensen-Shannon divergence value, which is used to measure the difference between two instantaneous interview performance analysis results.

[0107] Optionally, according to the obtained stability evaluation factor KP(t) , the stability evaluation factor can be compared with a set stability threshold. When the stability evaluation factor is greater than the set threshold, the current continuous performance analysis result V(t) is output as unstable; on the contrary, when the stability evaluation factor is less than the set threshold, the current continuous performance analysis result V(t) is output as stable;

[0108] In addition, the obtained stability evaluation factor KP(t) can also be standardized, and the stability evaluation factor KP(t)Quantify it into a stable value between 0 and 1 as the result of continuous performance analysis V(t) , which is used to intuitively measure the stability of the interviewee's performance. For example, the smaller the stable value, the more stable the interviewee's performance.

[0109] Based on the instantaneous performance analysis results within a certain time period, further analyze the stability of the interviewee's performance. The proposed stability evaluation factor can accurately evaluate the volatility of the interviewee's performance based on the idea of characteristic behavior entropy, and accurately reflect the consistency and difference of the interviewee's performance over a continuous period of time, so as to evaluate the performance stability of the interviewee during the interview process, which helps to improve the objectivity and accuracy of the evaluation of the interviewee's performance.

[0110] Preferably, the generation unit specifically includes:

[0111] According to the obtained instantaneous performance analysis results of the interviewee Y(t) and the continuous performance analysis results V(t) Integrate them to obtain the comprehensive performance analysis results corresponding to the t moment U(t) .

[0112] Optionally, the displayed comprehensive performance analysis results can directly display the corresponding values or classification features of the instantaneous performance analysis results and continuous performance analysis results at the current moment. At the same time, it can also further monitor the instantaneous performance analysis results and continuous performance analysis results. When the instantaneous performance analysis results or continuous performance analysis results fall into the set abnormal range, corresponding prompts for the abnormal situation are further made.

[0113] Preferably, the system further includes a feedback unit;

[0114] The feedback unit is used to perform real-time display according to the obtained comprehensive interview performance analysis results.

[0115] According to the obtained real-time comprehensive performance interview analysis results, real-time display can be performed through the feedback unit, which helps managers access the feedback unit through intelligent devices and obtain the corresponding comprehensive interview performance analysis results, assisting managers in making corresponding adjustments to the interview process or behavior, and improving the pertinence and effectiveness of the interview.

[0116] It should be noted that in each embodiment of the present invention, each functional unit / module can be integrated in a processing unit / module, or each unit / module exists physically alone, or two or more units / module are integrated in one unit / module. The above integrated unit / module can be implemented in the form of hardware or in the form of software functional unit / module.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any appropriate combination thereof. For hardware implementation, the processor can be implemented in one or more of the following units: application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, other electronic units designed to implement the functions described herein, or a combination thereof. For software implementation, part or all of the processes of the embodiments can be completed by instructing the relevant hardware through a computer program. When implemented, the above program can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a computer. The computer-readable medium can include, but is not limited to, RAM, ROM, EEPROM, CD-ROM, or other optical disc storage, magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A recruitment interview monitoring system based on computer vision, characterized in that, It includes a preprocessing unit, a feature extraction unit, and a performance analysis unit; among them, The preprocessing unit is used to preprocess the acquired interview record video data, including extracting the video stream data and audio stream data from the interview video record data, and further extracting the corresponding text data according to the audio stream data; The feature extraction unit is used to extract behavioral features based on the obtained video stream data, audio stream data, and text data, and further perform feature fusion processing based on the obtained behavioral features to obtain the behavioral fusion feature vector of the interviewee; The performance analysis unit is used to input the behavioral fusion feature vector of the interviewee into the trained instantaneous performance analysis model to obtain the instantaneous performance analysis result output by the instantaneous performance analysis model; further perform continuous performance analysis on the interviewee according to the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result; generate a comprehensive interview performance analysis result according to the obtained instantaneous performance analysis result and continuous performance analysis result; Among them, the continuous performance analysis of the interviewee according to the instantaneous performance analysis results within a period of time includes: Obtain t The instantaneous interview performance analysis results within a time period before a certain moment Y(t - K + 1), …, Y(t - k), …, Y(t) ; among which the instantaneous interview analysis results at each moment include N d-dimensional features; calculate the current stability evaluation factor according to the instantaneous interview performance analysis results at each moment: , in the formula, KP(t) represents t the stability evaluation factor at moment P n represents the average value of the corresponding d-dimensional features in each interview performance analysis result within the time period, and the variable n represents the total number of feature dimensions; n = 1, 2, …, N, N represents the set balance coefficient, φ represents JS(Y(c)||Y (c-1)) represents the divergence value between the instantaneous interview analysis result at moment c and c-1 the instantaneous interview analysis result at moment JS ; the variable c = t - K + 1, …, t ; Generate the current continuous performance analysis result according to the obtained stable evaluation factor.

2. The recruitment interview monitoring system based on computer vision according to claim 1, characterized in that, It also includes a data acquisition unit; The data acquisition unit is used to acquire the interview video record data collected by the camera.

3. The recruitment interview monitoring system based on computer vision according to claim 1, characterized in that, It also includes a feedback unit; The feedback unit is used to perform real-time display according to the obtained comprehensive interview performance analysis result.

4. A recruitment interview monitoring system based on computer vision according to claim 1, characterized in that, The preprocessing unit includes a video extraction unit, an audio extraction unit, and a text conversion unit; among them, The video extraction unit is used to extract video stream data according to the acquired interview record video data; The audio extraction unit is used to extract audio stream data according to the acquired interview record video data; The text conversion unit is used to extract the corresponding text data according to the acquired audio stream data; Among them, the obtained video stream data, audio stream data, and text data are aligned to the same time axis.

5. The recruitment interview monitoring system based on computer vision according to claim 4, characterized in that, The feature extraction unit includes an expression feature extraction unit, an emotion feature extraction unit, an attention extraction unit, and a feature fusion unit; Among them, The facial expression feature extraction unit is used to perform facial expression analysis and processing based on the obtained video stream data to obtain the facial expression features of the interviewee C face (t) ; The emotional feature extraction unit is used to perform emotional analysis processing on the obtained audio stream data to obtain the emotional features of the interviewee C Emo (t) ; The attention extraction unit is used to perform keyword density analysis processing on the obtained text data to obtain the attention characteristics of the interviewee C word (t) ; The feature fusion unit is used to perform timeline alignment and feature fusion processing based on the obtained expression features C face (t) , emotional features C Emo (t) and attention features C word (t) to obtain a behavior fusion feature vector CS(t) = {C face (t), C Emo (t), C word (t)} .

6. The recruitment interview monitoring system based on computer vision according to claim 5, characterized in that, The performance analysis unit includes an instantaneous analysis unit, a continuous analysis unit, and a generation unit; The instantaneous analysis unit is used to input the behavior fusion feature vector at t into the trained instantaneous performance analysis model, and obtain the instantaneous performance analysis result output by the instantaneous performance analysis model, where the instantaneous performance analysis result includes the probability characteristics of the interviewee for different performance types; The continuous analysis unit is used to perform continuous performance analysis on the interviewee according to the instantaneous performance analysis results within a period of time to obtain the continuous performance analysis result, where the continuous performance analysis result includes the performance stability score of the interviewee; The generation unit is used to generate a real-time comprehensive interview performance analysis result according to the obtained instantaneous performance analysis result and continuous performance analysis result of the interviewee.

7. The recruitment interview monitoring system based on computer vision according to claim 6, characterized in that instantaneous In the analysis unit, the instantaneous performance analysis model adopted is built based on a CNN neural network. The instantaneous performance analysis model includes an input layer, a first convolutional layer, a second convolutional layer, a feature compression layer, a fully connected layer, a classifier layer, and an output layer connected in sequence. Among them, the input layer is used to input 128 -dimensional behavior fusion feature vectors; the first convolutional layer contains 64 convolutional filters, with a kernel size of 5×5 , a stride of 1 , and an activation function of ReLU ; after the first convolutional layer, a max pooling layer is also set, with a pooling kernel size of 3×3 , and a stride of 2 ; the second convolutional layer contains 128 convolutional filters, with a kernel size of 3×3 , a stride of 1 , and an activation function of ReLU ; after the second convolutional layer, a max pooling layer is still set, with a pooling kernel size of 2×2 , and a stride of 2 ; the feature compression layer performs global average pooling on the output of the second convolutional layer to obtain a 128 -dimensional intermediate feature quantity; the fully connected layer contains 512 neurons, with an activation function of ReLU , and a Dropout layer is set after the fully connected layer, with a dropout rate of 50% to obtain the output neuron features; the classifier layer classifies the interview performance based on the neuron features, and the classifier adopted is Softmax, to obtain the corresponding interview performance classification results for the feature output; the output layer outputs the final instantaneous interview performance classification results according to the feature output of the classifier layer.

8. The recruitment interview monitoring system based on computer vision according to claim 6, characterized in that The generation unit specifically includes: According to the obtained instantaneous performance analysis results of the interviewees Y(t) and the continuous performance analysis results V(t) integrate them to obtain the comprehensive performance analysis results at the corresponding time t U(t).

Citation Information

Patent Citations

  • Intelligent integrated recruitment system capable of interviewing online

    CN118154143A

  • Personnel stability evaluation and model training method and device

    CN119721845A