Risk assessment method and device for video face review, equipment and medium

By asking pre-set questions and performing micro-expression recognition and anomaly detection in AI video interviews, the problem of users concealing or deceiving in their answers has been solved, achieving more accurate risk assessment.

CN120954110APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511058167.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing AI video interview technology cannot effectively assess whether users' answers contain concealment or deception, leading to inaccurate risk assessment.

Method used

By asking pre-set questions to the target interviewee, collecting video clips and performing micro-expression recognition, a baseline expression feature is generated. Based on the baseline feature and the answer feature, anomaly detection and question transformation are performed to generate pressure questions. Finally, expression anomaly detection is performed to assess interview anomalies.

Benefits of technology

It improves the accuracy and reliability of risk assessment in video interviews, and can effectively identify deceptive or concealed behavior in user responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954110A_ABST
    Figure CN120954110A_ABST
Patent Text Reader

Abstract

The invention provides a risk assessment method and device for video review, equipment and a medium, relates to the technical field of artificial intelligence, and is suitable for the field of financial science and technology. The method comprises the following steps: carrying out image acquisition on a target surface examination object answering a first preset question to obtain a baseline expression feature; performing image acquisition on the target face examination object answering the second preset question to obtain a first answer expression feature; performing anomaly detection on a second preset question according to the baseline expression feature and the first answer expression feature to obtain an abnormal question, and then performing conversion to obtain a pressure exerting question; carrying out image acquisition on the target surface examination object for answering the pressure exerting problem to obtain target answer expression features; and performing expression anomaly detection according to the baseline expression features and the target answer expression features to obtain face examination anomaly information. According to the method, whether the content answered by the user is hidden or cheated can be effectively evaluated, and the risk evaluation accuracy and reliability of video face examination are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applicable to the financial technology field, particularly to a risk assessment method, apparatus, device, and medium for video interviews. Background Technology

[0002] AI video interviews refer to interview and review processes conducted via video technology, with an AI digital android acting as the interviewer. For example, in the fintech field, customers can use video interviews to apply for inclusive lending, thereby automating the approval process and significantly shortening the workflow.

[0003] Currently, AI video interviews have applied some visual risk control methods, such as facial recognition, liveness detection, and video face-swapping recognition, which can achieve risk control to a certain extent. However, to determine whether a user is intentionally deceiving or concealing information, current large language models can only ask customers for routine information at the conversational level, and cannot effectively assess whether the user's answers contain concealment or deception.

[0004] Therefore, there are issues of concealment or deception in video interview technology. Summary of the Invention

[0005] The main objective of this application is to provide a risk assessment method, apparatus, device, and medium for video interviews, which can effectively assess whether the content of the user's answers contains concealment or deception, thereby improving the accuracy and reliability of risk assessment in video interviews.

[0006] To achieve the above objectives, a first aspect of this application proposes a risk assessment method for video interviews, the method comprising:

[0007] Ask the target interviewee a first preset question, and capture an image of the target interviewee who answers the first preset question to obtain a first video clip of the target interviewee;

[0008] Micro-expression recognition is performed on the video segment of the first object to obtain baseline expression features;

[0009] Ask the target interviewee a second preset question, and capture an image of the target interviewee who answers the second preset question to obtain a second video clip of the target interviewee;

[0010] Micro-expression recognition is performed on the video clip of the second object to obtain the facial expression features of the first response;

[0011] Based on the baseline facial expression features and the first response facial expression features, anomaly detection is performed on the second preset question to obtain an abnormal question, and the abnormal question is transformed into a pressure question;

[0012] Ask the target interviewee the pressure question, and capture images of the target interviewee answering the pressure question to obtain a video clip of the target interviewee;

[0013] Micro-expression recognition is performed on the video clip of the target object to obtain the target's response expression features;

[0014] Anomaly detection is performed based on the baseline facial expression features and the target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

[0015] Optionally, before transforming the anomalous problem into a stress problem, the method further includes:

[0016] Obtain the answer content of the target interviewee to the second preset question;

[0017] Based on the content of the answer, further questions are derived.

[0018] Ask the target interviewee the derivative questions, and capture images of the target interviewee who answers the derivative questions to obtain a video clip of the third object;

[0019] Micro-expression recognition is performed on the video clip of the third object to obtain the second response expression features;

[0020] Anomaly detection is performed on the derived question based on the baseline facial expression features and the second response facial expression features to obtain the abnormal question.

[0021] Optionally, the step of mining questions based on the answer content to obtain derivative questions includes:

[0022] Obtain the material information submitted by the target interviewee, the material information including the entity attribute information of the candidate entity;

[0023] Key entities are extracted from the answer content, and attributes are extracted from the entity attribute information based on the matching of the key entities and the candidate entities to obtain the target interest entity attributes.

[0024] The derived questions are obtained by constructing questions based on the attributes of the key entities and the target interest entities.

[0025] Optionally, the first object video segment includes multiple video sub-segments that are time-incrementing, and two adjacent video sub-segments overlap.

[0026] The step of performing micro-expression recognition on the first object video segment to obtain baseline expression features includes:

[0027] For each of the video segments, facial expression features are extracted to obtain initial facial expression and posture features; wherein, the initial facial expression and posture features include at least one of the following: facial expression motion unit features, emotion features, head posture features, and eye gaze point features.

[0028] The initial facial expression and pose features are extracted using a pre-defined temporal convolutional network to obtain the facial expression dynamic features of each video sub-segment.

[0029] The baseline facial expression features are obtained by generating baseline features for the facial expression dynamic features of each video sub-segment using a preset personalized Gaussian mixture model.

[0030] Optionally, before generating baseline features from the facial expression dynamic features of each of the video sub-segments using a preset personalized Gaussian mixture model, the method further includes:

[0031] Based on the facial expression dynamic features of each video sub-segment, facial expression posture is evaluated to obtain neutral facial expression posture features;

[0032] The distance between the neutral facial expression posture features and the facial expression dynamic features is calculated to obtain the facial expression posture distance;

[0033] The facial expression dynamic features of the video sub-segment are filtered based on the facial expression pose distance.

[0034] Optionally, the step of detecting facial expression anomalies based on the baseline facial expression features and the target response facial expression features to obtain abnormal interview information includes:

[0035] The baseline facial expression features are modeled using a graph neural network to obtain the target baseline dynamic facial expression features;

[0036] The target's response facial expression features are modeled using a graph neural network to obtain the target's real-time dynamic facial expression features.

[0037] The deviation is calculated based on the target baseline dynamic facial expression features and the target real-time dynamic facial expression features to obtain the facial expression deviation.

[0038] The abnormal information of the face review is generated based on the degree of facial expression deviation.

[0039] Optionally, the target baseline dynamic facial expression feature includes at least two baseline dynamic sub-facial expression features, and the target real-time dynamic facial expression feature includes at least two real-time dynamic sub-facial expression features;

[0040] The step of calculating the deviation based on the target baseline dynamic facial expression features and the target real-time dynamic facial expression features to obtain the facial expression deviation includes:

[0041] The distance between feature points is calculated based on the real-time dynamic sub-expression features and each baseline dynamic sub-expression feature.

[0042] Based on the distance between the feature points, K neighboring features are selected from the baseline dynamic sub-expression features; where K is a positive integer greater than 1.

[0043] Local reachability density is calculated based on the real-time dynamic sub-expression features and the K neighboring features to obtain the local anomaly factor score;

[0044] The expression deviation is obtained by fusing the scores of the local anomaly factors of each of the real-time dynamic sub-expression features.

[0045] To achieve the above objectives, a second aspect of this application provides a risk assessment device for video interviews, the device comprising:

[0046] The first acquisition module is used to ask the target interviewee a first preset question, and to acquire an image of the target interviewee who answers the first preset question to obtain a first video clip of the target;

[0047] The first recognition module is used to perform micro-expression recognition on the video clip of the first object to obtain baseline expression features;

[0048] The second acquisition module is used to ask the target interviewee a second preset question, and to acquire an image of the target interviewee who answers the second preset question to obtain a second video clip of the target interviewee.

[0049] The second recognition module is used to perform micro-expression recognition on the video clip of the second object to obtain the first response expression features;

[0050] The problem anomaly detection module is used to perform anomaly detection on the second preset problem based on the baseline facial expression features and the first response facial expression features, to obtain an abnormal problem, and to transform the abnormal problem into a pressure problem;

[0051] The third acquisition module is used to ask the target interviewee the pressure question and to acquire images of the target interviewee who answers the pressure question to obtain a video clip of the target interviewee;

[0052] The third recognition module is used to perform micro-expression recognition on the video clip of the target object to obtain the target's response expression features;

[0053] The facial expression anomaly detection module is used to detect facial expression anomalies based on the baseline facial expression features and the target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

[0054] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the risk assessment method for video interviews described in the first aspect.

[0055] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the risk assessment method for video interviews described in the first aspect.

[0056] The risk assessment method, apparatus, electronic device, and storage medium for video interviews proposed in this application first ask the target interviewee a first preset question, then perform micro-expression recognition on the captured video clips of the first interviewee to obtain baseline expression features. These baseline expression features characterize the target interviewee's facial expression habits when answering questions, and are used to assess whether the target interviewee deviates from these habits when answering subsequent questions. Further, the target interviewee is asked a second preset question, and then micro-expression recognition is performed on the captured video clips of the second interviewee to obtain first answer expression features. If the first answer expression features deviate significantly from the baseline expression features, it indicates that the target interviewee may be deceiving or concealing information in their answer to the second preset question; in this case, the second preset question is considered an abnormal question. Further, the abnormal question is converted into a pressure question for further verification when deception or concealment is suspected. Furthermore, the target interviewee is asked pressure questions, and then micro-expression recognition is performed on the collected video clips of the target interviewee to obtain the target's answer expression features. If the target's answer expression features deviate significantly from the baseline expression features, it indicates that the target interviewee is deceiving or concealing information in their answer to the pressure questions. In this case, the generated interview abnormality information indicates that the target interviewee's interview is abnormal; otherwise, it indicates that the target interviewee's interview is normal. In summary, this application, by identifying the micro-expression features in video clips of the target interviewee answering different questions, can effectively assess whether the user's answers contain concealment or deception, thereby improving the accuracy and reliability of risk assessment in video interviews.

[0057] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0058] Figure 1 This is a flowchart of a risk assessment method for video interviews provided in an embodiment of this application;

[0059] Figure 2 yes Figure 1 The flowchart for step 102 in the document;

[0060] Figure 3 This is a flowchart of a risk assessment method for video interviews provided in another embodiment of this application;

[0061] Figure 4 yes Figure 3 The flowchart for step 302 in the document;

[0062] Figure 5 yes Figure 1 The flowchart for step 108 in the document;

[0063] Figure 6 yes Figure 5 The flowchart for step 503 in the document;

[0064] Figure 7 This is a block diagram of the module structure of the risk assessment device for video interviews provided in the embodiments of this application;

[0065] Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0067] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0069] First, let's analyze some of the terms used in this application:

[0070] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0071] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information and image processing, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0072] The risk assessment method for video interviews provided in this application can be applied to both terminals and servers, or it can be software running on the server. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The software can be an application that implements the risk assessment method for video interviews, but it is not limited to the above forms.

[0073] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0074] This application provides a risk assessment method, a risk assessment device, an electronic device, and a computer-readable storage medium for video interviews. The specific details are illustrated in the following embodiments. First, the risk assessment method for video interviews in this application is described.

[0075] It should be noted that in each specific embodiment of this application, when it is necessary to process data related to the user's identity or characteristics, such as the user's image data, the user's permission or consent will be obtained first. Moreover, the collection, use and processing of this data will comply with relevant laws, regulations and standards.

[0076] Reference Figure 1 , Figure 1 This is an optional flowchart of a risk assessment method for video interviews provided in the embodiments of this application. The method may include, but is not limited to, steps 101 to 108.

[0077] Step 101: Ask the target interviewee a first preset question, and capture an image of the target interviewee who answers the first preset question to obtain a video clip of the first subject;

[0078] Step 102: Perform micro-expression recognition on the first object video segment to obtain baseline expression features;

[0079] Step 103: Ask the target interviewee a second preset question, and capture an image of the target interviewee who answers the second preset question to obtain a video clip of the second subject;

[0080] Step 104: Perform micro-expression recognition on the second object video segment to obtain the first response expression features;

[0081] Step 105: Based on the baseline facial expression features and the facial expression features of the first response, perform anomaly detection on the second preset question to obtain the abnormal question, and transform the abnormal question into a pressure question;

[0082] Step 106: Ask the target interviewee questions about applying pressure, and capture images of the target interviewee answering the questions about applying pressure to obtain video clips of the target interviewee;

[0083] Step 107: Perform micro-expression recognition on the video clip of the target object to obtain the target's response expression features;

[0084] Step 108: Perform facial expression anomaly detection based on baseline facial expression features and target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

[0085] Steps 101 to 108, as illustrated in this embodiment, involve first asking the target interviewee a first preset question, then performing micro-expression recognition on a captured video clip of the first interviewee to obtain baseline expression features. These baseline expression features characterize the target interviewee's facial expression habits when answering questions and are used to assess whether the target interviewee deviates from these habits when answering subsequent questions. Further, a second preset question is asked the target interviewee, and then micro-expression recognition is performed on a captured video clip of the second interviewee to obtain first answer expression features. If these first answer expression features deviate significantly from the baseline expression features, it indicates that the target interviewee may be deceiving or concealing information in their answer to the second preset question, and in this case, the second preset question is considered an abnormal question. Further, the abnormal question is converted into a pressure question for further verification when deception or concealment is suspected. Furthermore, the target interviewee is asked pressure questions, and then micro-expression recognition is performed on the collected video clips of the target interviewee to obtain the target's answer expression features. If the target's answer expression features deviate significantly from the baseline expression features, it indicates that the target interviewee is deceiving or concealing information in their answer to the pressure questions. In this case, the generated interview abnormality information indicates that the target interviewee's interview is abnormal; otherwise, it indicates that the target interviewee's interview is normal. In summary, this application, by identifying the micro-expression features in video clips of the target interviewee answering different questions, can effectively assess whether the user's answers contain concealment or deception, thereby improving the accuracy and reliability of risk assessment in video interviews.

[0086] In one example, within the fintech sector's inclusive lending scenario, a user might want to apply for a loan from Bank A. The user can initiate a video interview via a mobile app or web platform connected to Bank A's linked terminal or server. During the video interview, an AI-powered digital assistant will ask the user multiple questions, including inquiries about identity information, asset status, proof of income, and the purpose of the loan. If the AI ​​determines that the user has lied or concealed information in their answers, the loan application will be rejected. If the AI ​​determines that the user has not lied or concealed information, the loan application will be approved, and the loan can be disbursed. This automates the approval process and significantly shortens the workflow.

[0087] In another example, within the fintech sector's insurance account opening scenario, a user might wish to register an insurance account with Insurance Company B. In this case, the user can initiate a video interview via a mobile application or web platform to the terminal or server linked to Insurance Company B. During the video interview, an AI-powered digital assistant will ask the user multiple questions, including inquiries about identity information, health status, proof of income, intention to open an account, and desired insurance policies. If the AI ​​determines that the user has lied or concealed information in their answers, the user's insurance account opening application will be rejected. If the AI ​​determines that the user has not lied or concealed information in their answers, the user's insurance account opening application will be approved, and an insurance account can be assigned to the user.

[0088] In step 101 of some embodiments, a first preset question is asked to the target interviewee, and images of the target interviewee answering the first preset question are captured to obtain a first video segment of the target interviewee. The first preset question refers to a question that is weakly related to the business the target interviewee is conducting. When the business the target interviewee is conducting is a loan business, the first preset question refers to a question that is weakly related to whether the repayment will be overdue. For example, the first preset question includes some everyday questions such as name and home address. During the video interview, the camera continuously captures images of the target interviewee, and multiple images of the target interviewee answering the first preset question are determined as the first video segment of the target interviewee.

[0089] It should be noted that multiple pre-set questions can be asked to the target interviewee. In this case, the first target video segment includes video segments of the target interviewee answering multiple pre-set questions.

[0090] In step 102 of some embodiments, micro-expression recognition is performed on the first object video segment to obtain baseline expression features. Micro-expression recognition can be performed using optical flow.

[0091] In one embodiment, step 102 may include: performing eyebrow detection on the first object image in the first object video segment to obtain eyebrow upward displacement features; performing mouth corner detection on the first object image to obtain mouth corner drooping rate features; performing blink detection on the first object image to obtain blink interval time features; performing feature fusion based on the eyebrow upward displacement features, mouth corner drooping rate features, and blink interval time features to obtain micro-expression features of the first object image; and performing feature fusion based on the micro-expression features of each first object feature to obtain baseline expression features.

[0092] In one embodiment, the eyebrow center detection is performed on the first object image to obtain the eyebrow center upward displacement feature, including: using a face detection algorithm to locate the position of the face in the first object image; extracting facial key points, especially feature points around the eyebrows, using a facial landmark extractor; obtaining the eyebrow center position coordinates by calculating the average of the two left eyebrow feature points and the right eyebrow feature points; and calculating the vertical displacement by comparing the eyebrow center position coordinates in two frames of images. If the y-value of the eyebrow center coordinates decreases significantly in images at different time points (i.e., moves upward on the y-axis), it can be determined that the eyebrow center is upward.

[0093] In one embodiment, mouth corner detection is performed on a first object image to obtain a mouth corner drooping rate feature, including: using a face detection algorithm to locate the position of a face in the first object image; extracting facial key points, especially feature points around the mouth corners, using a facial landmark extractor; obtaining the feature point coordinates of the left and right mouth corners. Typically, the left mouth corner corresponds to the 48th feature point, and the right mouth corner corresponds to the 54th feature point; comparing the mouth corner coordinates in two frames and calculating their change on the y-axis. If the y-value of the mouth corner coordinates increases significantly (i.e., moves downward) in consecutive images, mouth corner drooping can be inferred; combining the detected drooping displacement with the time interval yields the mouth corner drooping rate. For example, the frame rate is used to calculate the drooping change per unit time.

[0094] In one embodiment, blink detection is performed on a first object image to obtain blink interval time features, including: using a face detection algorithm to locate the position of a face in the first object image; extracting facial key points, especially feature points around the eyes, using a facial landmark extractor; determining whether the eyes are closed by analyzing the position of eye feature points, for example, based on the ratio of the vertical height to the horizontal length of the eyes (the distance between the upper and lower points of the eyebrows and the upper and lower points of the eyes); and recording a timestamp when the eyes are detected to go from open to closed and back to open. This can be continuously monitored using a sliding window method; if eye closure is detected within a certain time window, it can be determined as one blink.

[0095] In another embodiment, the first video segment comprises multiple video sub-segments in ascending order of time, with adjacent video sub-segments overlapping. For example, the first video sub-segment has a length of 0 to 5 seconds, the second video sub-segment has a length of 1 to 6 seconds, the third video sub-segment has a length of 2 to 7 seconds, and so on. (See also...) Figure 2 Step 102 may include:

[0096] Step 201: Extract facial expression features from each video segment to obtain initial facial expression and pose features;

[0097] Step 202: Extract features from the initial facial expression and pose features using a pre-defined temporal convolutional network to obtain the facial expression dynamic features of each video sub-segment;

[0098] Step 203: Baseline facial expression features are generated by using a preset personalized Gaussian mixture model to generate baseline facial expression features for the facial expression dynamic features of each video sub-segment.

[0099] In step 201, the initial facial expression and pose features include at least one of the following: facial expression motion unit features, emotion features, head pose features, and eye gaze point features. These features can be extracted from images in video sub-segments using deep learning models (such as convolutional models).

[0100] In step 202, Temporal Convolutional Network (TCN) is an algorithm used to solve time series prediction.

[0101] In step 203, the personalized Gaussian Mixture Model (GMM) is an unsupervised learning algorithm based on probability density functions, widely used in cluster analysis and probability density estimation. It models the data distribution through a linear combination of multiple Gaussian distributions, thereby revealing the underlying structure in the data.

[0102] The advantage of the embodiments of steps 201 to 203 above is that, by obtaining facial expression dynamic features through TCN and utilizing the distribution prediction function of GMM on feature sequences, incremental updates can be supported to adapt to the emotional fluctuations of the interviewee.

[0103] In one embodiment, prior to step 203, the risk assessment method for video interviews may further include:

[0104] Based on the facial expression dynamic features of each video segment, facial expression and posture are evaluated to obtain neutral facial expression and posture features;

[0105] The distance between neutral facial expression posture features and facial expression dynamic features is calculated to obtain the facial expression posture distance;

[0106] The facial expression dynamics features of video segments are filtered based on the distance between facial expressions and postures.

[0107] Specifically, a neutral facial expression pose feature can be the facial expression dynamic feature whose pose evaluation result ranks in the middle among multiple facial expression dynamic features. Alternatively, a neutral facial expression pose feature can be the average facial expression dynamic feature obtained by fusing multiple facial expression dynamic features. The purpose of this embodiment is to select relatively neutral facial expression pose data from the multiple facial expression dynamic features generated by TCN, that is, to select facial expression dynamic features whose facial expression pose distance is less than or equal to a preset distance threshold, while facial expression dynamic features whose facial expression pose distance is greater than the preset distance threshold are deleted.

[0108] The advantage of the above embodiments is that they can filter out some overly average or overly excited facial expression dynamic features, thereby improving the stability of the subsequently generated facial expression feature baseline.

[0109] In step 103 of some embodiments, a second preset question is asked to the target interviewee, and images of the target interviewee answering the second preset question are captured to obtain a second video segment of the target interviewee. The second preset question refers to a question that is highly relevant to the business the target interviewee is conducting. When the business the target interviewee is conducting is a loan business, the second preset question refers to a question that is highly relevant to whether the repayment will be overdue. For example, the second preset question includes key questions such as asset status, debt status, and loan purpose. During the video interview, the camera continuously captures images of the target interviewee, and multiple images of the target interviewee answering the second preset question are determined as the second video segment of the target interviewee.

[0110] In step 104 of some embodiments, micro-expression recognition is performed on the second object video segment to obtain the first response expression features. The specific process of micro-expression recognition is basically the same as step 102, and will not be described again here.

[0111] In step 105 of some embodiments, anomaly detection is performed on the second preset question based on baseline facial expression features and first answer facial expression features to obtain an abnormal question. This abnormal question is then transformed into a pressure question. An abnormal question refers to the second preset question corresponding to a first answer facial expression feature that deviates significantly from the baseline facial expression features. A pressure question is a follow-up question asked when there is doubt about the answer to the abnormal question. For example, if the second preset question is "Do you have any debt?", and the target interviewee says they have no debt, but the background check materials show a loan from Bank A, the generated pressure question could be "Do you have any bank loans?" If the target interviewee still answers that they have no loan debt, the generated pressure question could also be "Do you have any loans from Bank A?" Generally, once it is confirmed that the customer has engaged in deception or concealment regarding the second preset question (such as the target interviewee exhibiting unnatural frowning, fake smiling, or pursing lips), the loan application process for the target interviewee will be terminated.

[0112] In some embodiments, the process of performing anomaly detection on the second preset question based on the baseline facial expression features and the first response facial expression features in step 105 to obtain an abnormal question may include: calculating the deviation based on the baseline facial expression features and the first response facial expression features to obtain a first facial expression deviation; if the first facial expression deviation is greater than the deviation threshold, then the second preset question is determined to be an abnormal question.

[0113] It should be noted that if the deviation of the first expression is less than or equal to the deviation threshold, the second preset question is determined to be a normal question. The deviation can be calculated using vector distance formulas, such as cosine similarity and Euclidean distance, etc., and this embodiment does not specifically limit this method.

[0114] In some embodiments, refer to Figure 3 Before transforming unusual issues into pressure-related issues, risk assessment methods used for video interviews may also include:

[0115] Step 301: Obtain the answer content of the target interviewee to the second preset question;

[0116] Step 302: Based on the answer content, conduct question mining to obtain derivative questions;

[0117] Step 303: Ask the target interviewee derivative questions and capture images of the target interviewee who answers the derivative questions to obtain a video clip of the third object;

[0118] Step 304: Perform micro-expression recognition on the video clip of the third object to obtain the facial expression features of the second response;

[0119] Step 305: Perform anomaly detection on the derived question based on the baseline facial expression features and the facial expression features of the second response to obtain the abnormal question.

[0120] In steps 301 and 302, considering that some target interviewees may conceal or deceive while answering the second preset question without showing it, the AI ​​digital human will tentatively ask some derivative questions of the second preset question and determine whether the other party is concealing or deceiving based on the user's answers. Specifically, assuming the second preset question is "Please answer your asset status," after asking the target interviewee this question, the target interviewee's answer to the question is obtained. This answer may contain concealment or deception, but sometimes it is not possible to determine whether the second preset question is an abnormal question based on the facial expression characteristics of the first answer and the baseline facial expression characteristics. Therefore, in this embodiment, derivative questions can be generated based on the answer content to ask the target interviewee some detailed follow-up questions.

[0121] In step 303, during the video interview, the camera continuously captures images of the target interviewee, identifying multiple images of the target interviewee answering derivative questions as third-party video clips. Specifically, the target interviewee will provide further answers to the interviewer's questions. However, if the user's asset information is falsified, their answers to these specific questions may contradict the facts (e.g., unable to answer the specific name of the residential complex, or the stated purchase price differing significantly from the average market price at the time, or the stated apartment type and size not existing in that complex). Most people, after being exposed for lying, tend to exhibit noticeable panic, silence, and tension. This is equivalent to the questions triggering an internal conflict within the user, which is then expressed through facial expressions or behavior, specifically reflected in the third-party video clips.

[0122] The specific process of step 304 is basically the same as that of step 104 above, and will not be repeated here.

[0123] In step 305, the deviation is calculated based on the baseline facial expression features and the second response facial expression features to obtain the second facial expression deviation; if the second facial expression deviation is greater than the deviation threshold, the derived question is identified as an abnormal question.

[0124] It should be noted that if the deviation of the second expression is less than or equal to the deviation threshold, the derived problem is determined to be a normal problem, and no further pressure problems need to be generated based on this derived problem. The deviation can be calculated using vector distance formulas, such as cosine similarity and Euclidean distance, etc., and this embodiment does not specifically limit it.

[0125] The advantage of the embodiments of steps 301 to 305 above is that converting the second preset question into a derived question and asking the subject can effectively capture subtle changes in the subject's facial expressions when answering detailed questions, thereby improving the accuracy of risk assessment.

[0126] In one embodiment, reference is made to Figure 4 Step 302 may include:

[0127] Step 401: Obtain the material information submitted by the target interviewee. The material information includes the entity attribute information of the candidate entity.

[0128] Step 402: Extract key entities from the answer content, and extract attributes from the entity attribute information based on the matching of key entities and candidate entities to obtain the target interest entity attributes.

[0129] Step 403: Construct questions based on the attributes of key entities and target interest entities to obtain derived questions.

[0130] In step 401, the material information may include the following: personal identification information, income verification, asset verification (property ownership certificate or real estate certificate, vehicle registration certificate and ownership certificate, bank statement, bank statement), debt information (existing loan contracts, repayment records and bills), and housing information (purchase contract, purchase invoice or receipt). A candidate entity refers to content or elements in the material information that, after extraction or identification, may represent a specific object, object category, or entity. For example, "ID number" or "name" in personal identification information is a candidate entity. "Property address" or "property certificate number" in housing information is also a candidate entity. Entity attribute information describes the specific attributes or characteristics of a candidate entity, reflecting the entity's specific information content. For example, for the candidate entity "name," its attribute information might be "ABC." For "number of properties," its attribute information might be "2 units." For "property address," its attribute information might be "No. D, C Road, B District, City A."

[0131] In step 402, the key entity refers to the entity extracted from the answer content. For example, if the answer content is "has savings of xxx million yuan and 2 properties", the key entities could be "savings" and "property". Entities can be matched using rules based on dictionaries, regular expressions, etc. If the key entity matches a candidate entity, it indicates a match; if the key entity does not match a candidate entity, it indicates a mismatch. For example, if the key entity is "property", and the candidate entities include "number of properties" and "property address", then the key entity matches these two candidate entities. The entity attribute information corresponding to the candidate entities "number of properties" and "property address" can be determined as the target interest entity attribute, meaning the target interest entity attribute could include "2 properties" and "City A, District B, Road C, No. D".

[0132] In step 403, questions can be constructed based on the attributes of key entities and target interest entities using a large language model to obtain derived questions. For example, a derived question could be "In which region are your properties located, are they currently owner-occupied or rented out, and what is the rental income?" If candidate entities also include "Year of property purchase," "Purchase price of property," "Property type," and "Property area," then the derived question could also be "In which year did you purchase your property, what was the purchase price, what is the property type, and what is the area?"

[0133] The advantage of the embodiments of steps 401 to 403 described above is that it is possible to dynamically generate derivative questions based on the subject's answers during the video interview process, thereby improving the flexibility and applicability of risk assessment.

[0134] In some embodiments, after obtaining an unusual question, it can be transformed into a pressure question using a large language model. For example, pressure questions may include "Why did you purchase it at a price much lower than the average price at the time?" or "Are you sure you purchased a house of xx type and xx size in xx community?" etc. The AI ​​digital human will determine whether the target interviewee is deceiving or concealing information based on their answers to the pressure questions.

[0135] In step 106 of some embodiments, the target interviewee is asked a pressure question, and images of the target interviewee answering the pressure question are captured to obtain a video clip of the target interviewee. During the video interview, the camera continuously captures images of the target interviewee, and multiple images of the target interviewee answering the pressure question are identified as video clips of the target interviewee.

[0136] In step 107 of some embodiments, micro-expression recognition is performed on the target object video clip to obtain the target's response expression features. The specific process of step 107 is basically the same as that of step 102, and will not be described again here.

[0137] In step 108 of some embodiments, facial expression anomaly detection is performed based on baseline facial expression features and target response facial expression features to obtain interview anomaly information. This interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

[0138] In one embodiment, reference is made to Figure 5 Step 108 may include:

[0139] Step 501: Model the baseline facial expression features using a graph neural network to obtain the target baseline dynamic facial expression features;

[0140] Step 502: Model the target's response facial expression features using a graph neural network to obtain the target's real-time dynamic facial expression features;

[0141] Step 503: Calculate the deviation based on the target baseline dynamic facial expression features and the target real-time dynamic facial expression features to obtain the facial expression deviation.

[0142] Step 504: Generate abnormal information for the face-to-face interview based on the degree of facial expression deviation.

[0143] In step 501, macroscopic actions (such as frequently touching the nose and pursing the lips) and microscopic signals (slight frowning, pouting, etc.) are fused together, and the action correlation is modeled by graph neural networks (GNN) to obtain the target baseline dynamic expression features.

[0144] Step 502 is similar to step 501, and will not be described again here.

[0145] In step 503, the expression deviation degree is used to characterize the degree to which the target's real-time dynamic expression features deviate from the target's baseline dynamic expression features. A higher expression deviation degree indicates a greater degree of deviation, and a higher probability that the target candidate's interview will be abnormal. A lower expression deviation degree indicates a lesser degree of deviation, and a higher probability that the target candidate's interview will be normal. A normal interview means that the target candidate has not engaged in deception or concealment.

[0146] In step 504, if the facial expression deviation is greater than or equal to the deviation threshold, then facial expression abnormality information is generated to characterize the abnormality of the target interviewee; if the facial expression deviation is less than the deviation threshold, then facial expression abnormality information is generated to characterize the normality of the target interviewee.

[0147] The advantage of the embodiments of steps 501 to 504 above is that by modeling actions through graph neural networks, the denoising and enhancement effects of GNN on TCN feature sequences can be utilized to make the entire prediction more stable and improve the stability and accuracy of risk assessment.

[0148] In some embodiments, the target baseline dynamic facial expression feature includes at least two baseline dynamic sub-facial expression features, and the target real-time dynamic facial expression feature includes at least two real-time dynamic sub-facial expression features. (Refer to...) Figure 6 Step 503 may include:

[0149] Step 601: Calculate the distance between feature points based on the real-time dynamic sub-expression features and the dynamic sub-expression features of each baseline;

[0150] Step 602: Select K neighboring features from the baseline dynamic sub-expression features based on the feature point distance; where K is a positive integer greater than 1.

[0151] Step 603: Calculate the local reachability density based on the real-time dynamic sub-expression features and K neighboring features to obtain the local anomaly factor score;

[0152] Step 604: The local anomaly factor scores of each real-time dynamic sub-expression feature are fused to obtain the expression deviation.

[0153] In step 601, the distance calculation can use either the cosine distance formula or the Euclidean distance formula. Feature point distance is used to characterize the feature distance between real-time dynamic sub-expression features and baseline dynamic sub-expression features. For example, assume the target baseline dynamic expression feature is divided into three baseline dynamic sub-expression features, denoted as j1, j2, and j3; and the target real-time dynamic expression feature is divided into three real-time dynamic sub-expression features, denoted as i1, i2, and i3. The total number of baseline dynamic sub-expression features is the same as the total number of real-time dynamic sub-expression features, and the dimension of any baseline dynamic sub-expression feature is the same as the dimension of any real-time dynamic sub-expression feature. After distance calculation, the feature point distance S between i1 and j1 can be obtained. i1ji The feature point distance S between i1 and j2 i1j2 The distance S between feature points i1 and j3 i1j3 The feature point distance S between i2 and j1 i2ji The feature point distance S between i2 and j2 i2j2 The feature point distance S between i2 and j3 i2j3 The feature point distance S between i3 and j1 i3ji The feature point distance S between i3 and j2 i3j2 The feature point distance S between i3 and j3 i3j3 .

[0154] In step 602, the nearest neighbor feature refers to the baseline dynamic sub-expression feature whose feature point distance is sorted from smallest to largest and is ranked before the K position.

[0155] In step 603, Local Reachability Density (LRD) is used to measure the density of a data point within a local region. In this embodiment, Local Reachability Density specifically refers to the reciprocal of the average distance between neighboring features and real-time dynamic sub-expression features.

[0156] In one embodiment, step 603 may include: calculating the local reachability density based on the real-time dynamic sub-expression features and K neighboring features to obtain a first local reachability density; obtaining the local reachability density corresponding to each neighboring feature to obtain a second local reachability density; calculating the ratio between the second local reachability density and the first local reachability density to obtain an initial local anomaly factor score corresponding to each neighboring feature; and averaging the K initial local anomaly factor scores to obtain the local anomaly factor score of the real-time dynamic sub-expression features.

[0157] In step 604, the average score of the local anomaly factor of each real-time dynamic sub-expression feature can be calculated to obtain the expression deviation.

[0158] The advantage of the embodiments of steps 601 to 604 described above is that they can improve the accuracy of calculating facial expression deviation, making them suitable for video interview scenarios.

[0159] In some embodiments, when training the model used in the above embodiments, synthetic noise (such as sudden changes in illumination or partial facial occlusion) is added to the training data to improve the robustness of the model. Alternatively, based on video frame interpolation algorithms, videos of different frame rates can be aligned to the 30 FPS required for micro-expression capture, ensuring the consistency of the temporal model input.

[0160] In summary, the present application can achieve at least the following beneficial effects: (1) High accuracy: By constructing a personalized facial expression baseline and conducting detailed analysis of micro-expressions, it can more accurately identify deceptive or concealing behaviors. (2) Strong real-time performance: It can monitor changes in customers' micro-expressions in real time and promptly detect potential fraud risks. (3) Good adaptability: It can adapt to the facial expression characteristics of different customers, improving the universality and applicability of risk assessment. (4) Minimal intrusion: It does not require additional hardware modifications to the original AI interview process, and can directly use the interview script data. Only minor modifications to the process are needed, and the modifications are imperceptible to the customer.

[0161] Please see Figure 7 This application also provides a risk assessment device for video interviews, applied to an account opening server, which can implement the above-mentioned risk assessment method for video interviews. Figure 7 A module structure block diagram of a risk assessment device for video interviews provided in this application embodiment, the device comprising:

[0162] The first acquisition module 701 is used to ask the target interviewee a first preset question, acquire images of the target interviewee who answers the first preset question, and obtain a video clip of the first object.

[0163] The first recognition module 702 is used to perform micro-expression recognition on the first object video segment to obtain baseline expression features;

[0164] The second acquisition module 703 is used to ask the target interviewee a second preset question, acquire images of the target interviewee who answers the second preset question, and obtain a video clip of the second object.

[0165] The second recognition module 704 is used to perform micro-expression recognition on the second object video segment to obtain the first response expression features;

[0166] The problem anomaly detection module 705 is used to perform anomaly detection on the second preset problem based on the baseline facial expression features and the first response facial expression features, to obtain the abnormal problem, and to transform the abnormal problem into a pressure problem;

[0167] The third acquisition module 706 is used to ask the target interviewee questions about pressure and to acquire images of the target interviewee who answers the pressure questions, thereby obtaining video clips of the target interviewee.

[0168] The third recognition module 707 is used to perform micro-expression recognition on the video clip of the target object to obtain the target's response expression features;

[0169] The facial expression anomaly detection module 708 is used to detect facial expression anomalies based on baseline facial expression features and target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interviewee's interview is normal or abnormal.

[0170] In one embodiment, the risk assessment device for video interviews further includes a question generation module, configured to: obtain the answer content of the target interviewee to a second preset question; perform question mining based on the answer content to obtain derivative questions; ask the target interviewee the derivative questions and capture an image of the target interviewee answering the derivative questions to obtain a third object video segment; perform micro-expression recognition on the third object video segment to obtain a second answer expression feature; and perform anomaly detection on the derivative questions based on the baseline expression feature and the second answer expression feature to obtain abnormal questions.

[0171] In one embodiment, the risk assessment device for video interviews further includes a feature filtering module, which is used to: evaluate facial expression and posture based on the facial expression dynamic features of each video sub-segment to obtain neutral facial expression and posture features; calculate the distance between the neutral facial expression and posture features and the facial expression dynamic features to obtain facial expression and posture distance; and filter the facial expression dynamic features of the video sub-segments based on the facial expression and posture distance.

[0172] It should be noted that the specific implementation of the risk assessment device for video interviews is basically the same as the specific implementation of the risk assessment method for video interviews described above, and will not be repeated here.

[0173] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned risk assessment method for video interviews. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0174] Please see Figure 8 , Figure 8The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0175] The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0176] The memory 802 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 to execute the risk assessment method for video interviews in the embodiments of this application.

[0177] The 803 input / output interface is used to implement information input and output.

[0178] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0179] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);

[0180] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0181] This application embodiment also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described risk assessment method for video interviews.

[0182] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0183] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0184] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0185] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0186] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0187] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0188] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0191] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0193] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A risk assessment method for video interviews, characterized in that, The method includes: Ask the target interviewee a first preset question, and capture an image of the target interviewee who answers the first preset question to obtain a first video clip of the target interviewee; Micro-expression recognition is performed on the video segment of the first object to obtain baseline expression features; Ask the target interviewee a second preset question, and capture an image of the target interviewee who answers the second preset question to obtain a second video clip of the target interviewee; Micro-expression recognition is performed on the video clip of the second object to obtain the facial expression features of the first response; Based on the baseline facial expression features and the first response facial expression features, anomaly detection is performed on the second preset question to obtain an abnormal question, and the abnormal question is transformed into a pressure question; Ask the target interviewee the pressure question, and capture images of the target interviewee answering the pressure question to obtain a video clip of the target interviewee; Micro-expression recognition is performed on the video clip of the target object to obtain the target's response expression features; Anomaly detection is performed based on the baseline facial expression features and the target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

2. The method according to claim 1, characterized in that, Before transforming the anomalous problem into a stress problem, the method further includes: Obtain the answer content of the target interviewee to the second preset question; Based on the content of the answer, further questions are derived. Ask the target interviewee the derivative questions, and capture images of the target interviewee who answers the derivative questions to obtain a video clip of the third object; Micro-expression recognition is performed on the video clip of the third object to obtain the second response expression features; Anomaly detection is performed on the derived question based on the baseline facial expression features and the second response facial expression features to obtain the abnormal question.

3. The method according to claim 2, characterized in that, The process of mining questions based on the answer content to obtain derivative questions includes: Obtain the material information submitted by the target interviewee, the material information including the entity attribute information of the candidate entity; Key entities are extracted from the answer content, and attributes are extracted from the entity attribute information based on the matching of the key entities and the candidate entities to obtain the target interest entity attributes. The derived questions are obtained by constructing questions based on the attributes of the key entities and the target interest entities.

4. The method according to any one of claims 1 to 3, characterized in that, The first object video segment includes multiple video sub-segments in ascending order of time, and adjacent video sub-segments overlap; The step of performing micro-expression recognition on the first object video segment to obtain baseline expression features includes: For each of the video segments, facial expression features are extracted to obtain initial facial expression and posture features; wherein, the initial facial expression and posture features include at least one of the following: facial expression motion unit features, emotion features, head posture features, and eye gaze point features. The initial facial expression and pose features are extracted using a pre-defined temporal convolutional network to obtain the facial expression dynamic features of each video sub-segment. The baseline facial expression features are obtained by generating baseline features for the facial expression dynamic features of each video sub-segment using a preset personalized Gaussian mixture model.

5. The method according to claim 4, characterized in that, Before generating baseline facial expression features by performing baseline feature generation on the facial expression dynamic features of each of the video sub-segments using a preset personalized Gaussian mixture model, the method further includes: Based on the facial expression dynamic features of each video sub-segment, facial expression posture is evaluated to obtain neutral facial expression posture features; The distance between the neutral facial expression posture features and the facial expression dynamic features is calculated to obtain the facial expression posture distance; The facial expression dynamic features of the video sub-segment are filtered based on the facial expression pose distance.

6. The method according to any one of claims 1 to 3, characterized in that, The step of detecting facial expression anomalies based on the baseline facial expression features and the target response facial expression features to obtain abnormal information during the face-to-face interview includes: The baseline facial expression features are modeled using a graph neural network to obtain the target baseline dynamic facial expression features; The target's response facial expression features are modeled using a graph neural network to obtain the target's real-time dynamic facial expression features. The deviation is calculated based on the target baseline dynamic facial expression features and the target real-time dynamic facial expression features to obtain the facial expression deviation. The abnormal information of the face review is generated based on the degree of facial expression deviation.

7. The method according to claim 6, characterized in that, The target baseline dynamic facial expression feature includes at least two baseline dynamic sub-facial expression features, and the target real-time dynamic facial expression feature includes at least two real-time dynamic sub-facial expression features. The step of calculating the deviation based on the target baseline dynamic facial expression features and the target real-time dynamic facial expression features to obtain the facial expression deviation includes: The distance between feature points is calculated based on the real-time dynamic sub-expression features and each baseline dynamic sub-expression feature. Based on the distance between the feature points, K neighboring features are selected from the baseline dynamic sub-expression features; where K is a positive integer greater than 1. Local reachability density is calculated based on the real-time dynamic sub-expression features and the K neighboring features to obtain the local anomaly factor score; The expression deviation is obtained by fusing the scores of the local anomaly factors of each of the real-time dynamic sub-expression features.

8. A risk assessment device for video interviews, characterized in that, The device includes: The first acquisition module is used to ask the target interviewee a first preset question, and to acquire an image of the target interviewee who answers the first preset question to obtain a first video clip of the target; The first recognition module is used to perform micro-expression recognition on the video clip of the first object to obtain baseline expression features; The second acquisition module is used to ask the target interviewee a second preset question, and to acquire an image of the target interviewee who answers the second preset question to obtain a second video clip of the target interviewee. The second recognition module is used to perform micro-expression recognition on the video clip of the second object to obtain the first response expression features; The problem anomaly detection module is used to perform anomaly detection on the second preset problem based on the baseline facial expression features and the first response facial expression features, to obtain an abnormal problem, and to transform the abnormal problem into a pressure problem; The third acquisition module is used to ask the target interviewee the pressure question and to acquire images of the target interviewee who answers the pressure question to obtain a video clip of the target interviewee; The third recognition module is used to perform micro-expression recognition on the video clip of the target object to obtain the target's response expression features; The facial expression anomaly detection module is used to detect facial expression anomalies based on the baseline facial expression features and the target response facial expression features to obtain interview anomaly information; wherein, the interview anomaly information is used to characterize whether the interview of the target interviewee is normal or abnormal.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.