Face review monitoring method, device, equipment and storage medium
Through the combination of face detection and key point detection, the problem of inaccurate identification of fraud in facial review monitoring is solved, and the accuracy of facial review evaluation is improved.
Patent Information
- Application Number
- CN202210858890.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-21
AI Technical Summary
In the existing face-to-face review monitoring technology, the identification of fraud is insufficient, resulting in inaccurate face-to-face review evaluation results, especially when the head of the face-to-face review is not fully presented or interacts with a third person.
By calling face detection model and key point detection, we can judge whether the face is complete in the face review scene, and use neural network models to detect the head posture to identify whether the person being interviewed is the target person and whether there are abnormal behaviors.
It improves the accuracy of identifying fraud during the face-to-face review process and ensures the accuracy and reliability of face-to-face review evaluation.
Smart Images

Figure CN115424311B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a face review monitoring method, device, equipment and storage medium. Background Art
[0002] In the existing face review monitoring technology, generally, by asking questions to the user, collecting facial images or video data during the user's question-and-answer process, and performing user micro-expression recognition through facial recognition technology to determine whether the applicant's information is the applicant himself / herself and whether there is a possibility of fraud.
[0003] In the existing technology, for the fraudulent behavior of the interviewee cheating by deliberately disguising, the recognition accuracy is insufficient. For example, if the user's head is not fully presented in the camera or there is an abnormal angle deviation, and the user may interact with a third person to achieve cheating, resulting in inaccurate face review evaluation results. Summary of the Invention
[0004] The main purpose of the present invention is to solve the problem that the face review evaluation results are inaccurate due to insufficient recognition accuracy of fraudulent behaviors in the existing face review monitoring method.
[0005] The first aspect of the present invention provides a face review monitoring method, including:
[0006] When receiving a face review monitoring request, calling the interviewee's terminal to collect a pre-face review scene image;
[0007] Calling a pre-set face detection model to perform face detection on the pre-face review scene image to obtain a first face region image;
[0008] Performing key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0009] Judging whether the face in the first face region image is complete according to the set of face key points;
[0010] If it is determined that the face in the first face region image is incomplete, sending a face review specification prompt to the interviewee's terminal;
[0011] If it is determined that the face in the first face region image is complete, judging whether the interviewee in the first face region image is a pre-set target interviewee;
[0012] If it is determined that the interviewee in the first face region image is the target interviewee, sending a face review ready prompt to the interviewer's terminal to start the face review.
[0013] Optionally, in the first implementation manner of the first aspect of the present invention, after determining that the person being interviewed in the first face region image is the target person being interviewed and sending a ready-for-interview prompt to the interviewer terminal to start the interview, the following steps are further included:
[0014] Call the interviewed person's terminal to collect the first interview scene image based on a preset frequency;
[0015] Call the face detection model to perform face detection on the first interview scene image to obtain a second face region image;
[0016] Perform regression training on the preset neural network model for the deflection angle to obtain a head pose detection model;
[0017] Call the head pose detection model to perform pose detection on the second face region image to obtain the deflection angles of the head of the target person being interviewed in at least one preset direction;
[0018] Compare the deflection angles in the at least one preset direction with a preset standard deflection interval respectively. If the deflection angle in any direction is not within the standard deflection interval, it is determined that the behavior of the target person being interviewed is abnormal, and an abnormality prompt is sent to the interviewer terminal.
[0019] Optionally, in the second implementation manner of the first aspect of the present invention, the preset neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The performing regression training on the preset neural network model for the deflection angle to obtain a head pose detection model includes:
[0020] Call the pose feature extraction network to extract target face pose features from the target training images in the preset face pose training image set;
[0021] Call the deflection angle classification network to calculate the multi-classification deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and the preset deflection interval probability matrix;
[0022] Select the deflection interval with the largest probability value from the multi-classification deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtain the regression function corresponding to the target deflection interval;
[0023] Call the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose features and the regression function to obtain the deflection angles of the target training image in at least one preset direction;
[0024] Calculate the loss value corresponding to the deflection angle of the target training image in the at least one preset direction based on the preset loss function and the preset deflection angle annotation information of the target training image;
[0025] Adjust the network parameters of the preset neural network model according to the loss value until the loss value is less than the preset threshold, determine that the preset neural network model converges, and save the current network parameters of the preset neural network model to obtain a head pose detection model.
[0026] Optionally, in the third implementation manner of the first aspect of the present invention, after determining that the person being interviewed in the first face region image is the target person being interviewed and sending a face interview ready prompt to the interviewer terminal to start the face interview, it further includes:
[0027] Call the terminal of the person being interviewed to collect a second face interview scene image based on a preset frequency;
[0028] Perform face detection on the second face interview scene image to determine whether the second face interview scene image contains a face;
[0029] If it is determined that the second face interview scene image does not contain a face, determine that the current face interview scene is abnormal and send an abnormality prompt to the interviewer terminal;
[0030] If it is determined that the second face interview scene image contains a face, extract the target face parameters in the second face interview scene image;
[0031] Obtain the preset face parameters of the target person being interviewed from the preset database, and match the target face parameters with the preset face parameters. If the match fails, determine that the person being interviewed has been replaced and send a replacement prompt to the interviewer terminal.
[0032] Optionally, in the fourth implementation manner of the first aspect of the present invention, the calling the preset face detection model to perform face detection on the pre-face interview scene image to obtain a first face region image includes:
[0033] Input the pre-face interview scene image into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;
[0034] Call the face feature extraction network to extract face features of at least one feature size from the pre-face interview scene image;
[0035] Call the face feature recognition network to identify alternative face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size;
[0036] Invoke the face feature screening network to screen each of the alternative face regions to obtain a first face region image.
[0037] Optionally, in the fifth implementation manner of the first aspect of the present invention, the performing key point detection on the first face region image to obtain a set of face key points in the first face region image includes:
[0038] Perform a key point detection training task on a preset initial network model to obtain a face key point detection model;
[0039] Perform image preprocessing on the first face region image, and input the preprocessed first face region image into the face key point detection model for detection to obtain a set of face key points in the first face region image, where the set of face key points includes position coordinate information of each face key point.
[0040] Optionally, in the sixth implementation manner of the first aspect of the present invention, the performing a key point detection training task on a preset initial network model to obtain a face key point detection model includes:
[0041] Obtain a preset face image sample and a sample label corresponding to the face image sample, where the sample label includes face key point position data in the face image sample;
[0042] Perform prior calculation on the face key point position data to obtain a structural prior feature of the face image sample;
[0043] Invoke the initial network model to extract the self-attention feature of the face image sample, and perform image recognition based on the self-attention feature to obtain face key point detection data in the face image sample;
[0044] Based on a preset first loss function, calculate a first loss value between the self-attention feature and the structural prior feature, and based on a preset second loss function, calculate a second loss value between the face key point detection data and the face key point position data;
[0045] Perform weighted summation calculation on the first loss value and the second loss value according to a preset weighted algorithm to obtain the current global loss value of the initial network model;
[0046] Adjust the network parameters of the initial network model according to the global loss value to obtain a face key point detection model.
[0047] The second aspect of the present invention provides a face review monitoring device, including:
[0048] An image acquisition module, configured to, when receiving a face review monitoring request, call the interviewee's terminal to acquire pre-face review scenario images;
[0049] A face detection module, configured to call a pre-set face detection model to perform face detection on the pre-face review scenario images, and obtain a first face region image;
[0050] A key point detection module, configured to perform key point detection on the first face region image, and obtain a set of face key points in the first face region image;
[0051] A scenario correction module, configured to, according to the set of face key points, determine whether the face in the first face region image is complete; if it is determined that the face in the first face region image is incomplete, send a face review specification prompt to the interviewee's terminal;
[0052] An identity verification module, configured to, if it is determined that the face in the first face region image is complete, determine whether the interviewee in the first face region image is a pre-set target interviewee;
[0053] A face review start module, configured to, if it is determined that the interviewee in the first face region image is the target interviewee, send a face review ready prompt to the interviewer's terminal to start the face review.
[0054] Optionally, in the first implementation manner of the second aspect of the present invention, the pre-set neural network model includes a pose feature extraction network, a yaw angle classification network, and a yaw angle regression network, and the face review monitoring device further includes a yaw detection module, and the yaw detection module specifically includes:
[0055] An acquisition unit, configured to, based on a pre-set frequency, call the interviewee's terminal to acquire first face review scenario images;
[0056] A face detection unit, configured to call the face detection model to perform face detection on the first face review scenario images, and obtain a second face region image;
[0057] A model training unit, configured to perform regression training on the pre-set neural network model for the yaw angle, and obtain a head pose detection model;
[0058] A pose detection unit, configured to call the head pose detection model to perform pose detection on the second face region image, and obtain the yaw angle of the head of the target interviewee in at least one pre-set direction;
[0059] A comparison unit is configured to compare the deflection angles in the at least one preset direction with preset standard deflection intervals respectively. If the deflection angle in any direction is not within the standard deflection interval, it is determined that the behavior of the target interviewee is abnormal, and an abnormality prompt is sent to the interviewer terminal.
[0060] Optionally, in the second implementation manner of the second aspect of the present invention, the preset neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The model training unit is configured to:
[0061] Invoke the pose feature extraction network to extract target face pose features from a target training image in a preset face pose training image set;
[0062] Invoke the deflection angle classification network to calculate the multi-classification deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and a preset deflection interval probability matrix;
[0063] Select the deflection interval with the largest probability value from the multi-classification deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtain the regression function corresponding to the target deflection interval;
[0064] Invoke the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose features and the regression function, and obtain the deflection angles of the target training image in at least one preset direction;
[0065] Based on a preset loss function and the preset deflection angle annotation information of the target training image, calculate the loss value corresponding to the deflection angles of the target training image in the at least one preset direction;
[0066] Adjust the network parameters of the preset neural network model according to the loss value until the loss value is less than a preset threshold, determine that the preset neural network model converges, and save the current network parameters of the preset neural network model to obtain a head pose detection model.
[0067] Optionally, in the third implementation manner of the second aspect of the present invention, the interview monitoring device further includes a face recheck module, and the face recheck module includes:
[0068] An acquisition unit is configured to call the interviewee terminal to acquire a second interview scenario image based on a preset frequency;
[0069] A face detection unit is configured to perform face detection on the second interview scenario image to determine whether a face is included in the second interview scenario image;
[0070] Anomaly prompt unit, configured to determine that the current face review scenario is abnormal if it is determined that the second face review scenario image does not contain a human face, and send an anomaly prompt to the face reviewer terminal;
[0071] Parameter extraction unit, configured to extract target human face parameters in the second face review scenario image if it is determined that the second face review scenario image contains a human face;
[0072] Parameter matching unit, configured to obtain preset human face parameters of the target person to be face reviewed from a preset database, and match the target human face parameters with the preset human face parameters. If the match fails, it is determined that the person to be face reviewed has been replaced, and a replacement prompt is sent to the face reviewer terminal.
[0073] Optionally, in the fourth implementation manner of the second aspect of the present invention, the face detection module specifically includes:
[0074] Input unit, configured to input the pre-face review scenario image into the face detection model, where the face detection model includes a human face feature extraction network, a human face feature recognition network, and a human face feature screening network;
[0075] Feature extraction unit, configured to call the human face feature extraction network to extract human face features of at least one feature size from the pre-face review scenario image;
[0076] Recognition unit, configured to call the human face feature recognition network to recognize alternative human face regions in the human face features of each feature size according to a preset number of prior boxes corresponding to each feature size;
[0077] Screening unit, configured to call the human face feature screening network to screen each of the alternative human face regions to obtain a first human face region image.
[0078] Optionally, in the fifth implementation manner of the second aspect of the present invention, the key point detection module specifically includes:
[0079] Model construction unit, configured to perform a key point detection training task on a preset initial network model to obtain a human face key point detection model;
[0080] Detection unit, configured to perform image preprocessing on the first human face region image, and input the preprocessed first human face region image into the human face key point detection model for detection to obtain a set of human face key points in the first human face region image, where the set of human face key points includes position coordinate information of each human face key point.
[0081] Optionally, in the sixth implementation manner of the second aspect of the present invention, the model construction unit is specifically configured to:
[0082] Obtain a pre-set face image sample and the sample label corresponding to the face image sample, where the sample label includes the face key point position data in the face image sample;
[0083] Perform a priori calculation on the face key point position data to obtain the structural prior feature of the face image sample;
[0084] Call the initial network model to extract the self-attention feature of the face image sample, and perform image recognition based on the self-attention feature to obtain the face key point detection data in the face image sample;
[0085] Based on a pre-set first loss function, calculate the first loss value between the self-attention feature and the structural prior feature, and based on a pre-set second loss function, calculate the second loss value between the face key point detection data and the face key point position data;
[0086] Perform weighted summation calculation on the first loss value and the second loss value according to a pre-set weighting algorithm to obtain the current global loss value of the initial network model;
[0087] Adjust the network parameters of the initial network model according to the global loss value to obtain a face key point detection model.
[0088] The third aspect of the present invention provides a face review monitoring device, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the face review monitoring device to execute each step of the above face review monitoring method.
[0089] The fourth aspect of the present invention provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when it runs on a computer, it enables the computer to execute each step of the above face review monitoring method.
[0090] In the technical solution provided by the present invention, by detecting whether the head pose angle of the person being face-reviewed meets the standard, and by performing key point detection to determine whether the face in the captured face review scene is complete, to determine whether the person being face-reviewed interacts with a third person, so as to accurately identify fraud behavior, and further improve the accuracy of the face review evaluation result. Description of the Drawings
[0091] Figure 1 It is a schematic diagram of the first embodiment of the face review monitoring method in the embodiment of the present invention;
[0092] Figure 2 It is a schematic diagram of the second embodiment of the face review monitoring method in the embodiment of the present invention;
[0093] Figure 3 Schematic diagram of the third embodiment of the face review monitoring method in the embodiments of the present invention;
[0094] Figure 4 Schematic diagram of an embodiment of the face review monitoring device in the embodiments of the present invention;
[0095] Figure 5 Schematic diagram of another embodiment of the face review monitoring device in the embodiments of the present invention;
[0096] Figure 6 Schematic diagram of an embodiment of the face review monitoring device in the embodiments of the present invention;
[0097] Figure 7 Schematic diagram of an embodiment of the head pose detection in the embodiments of the present invention. Detailed implementation manners
[0098] The embodiments of the present invention provide a face review monitoring method, device, equipment and storage medium, which have higher recognition accuracy for fraud behaviors.
[0099] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "including" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0100] It can be understood that the execution subject of the present invention can be a face review monitoring device, or a terminal or a server, and specific limitations are not made here. The embodiments of the present invention are described by taking the server as the execution subject as an example.
[0101] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0102] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0103] For ease of understanding, the specific process of the embodiment of the present invention will be described below. Please refer to Figure 1 In the first embodiment of the in-person interview monitoring method in the embodiment of the present invention, it includes:
[0104] 101. When receiving an in-person interview monitoring request, call the interviewee's terminal to collect a pre-in-person interview scene image;
[0105] It can be understood that the interviewee's terminal refers to the terminal where the interviewee is located. This terminal includes but is not limited to devices such as mobile phones and monitors, all of which should be equipped with a camera device. Thus, when the server receives an in-person interview monitoring request, the camera device in this terminal is called to collect a pre-in-person interview scene image in real time. It should be noted that the pre-in-person interview scene image is the image captured in front of the camera device when the interviewee is ready for the in-person interview and the interviewer has not yet initiated an in-person interview question and answer to the interviewee. It is used to detect the in-person interview environment.
[0106] It can be understood that the camera device first needs to select a focusing object before collecting the pre-in-person interview scene image. Optionally, the server can schedule the camera device to adjust to the maximum aperture parameter and give priority to focusing on the human face, so as to obtain a better blurring effect, thereby reducing the effective pixels in the collected pre-in-person interview scene image, and further reducing the image calculation amount when detecting this pre-in-person interview scene image to improve the processing efficiency.
[0107] Optionally, when the camera device captures multiple people when selecting a focusing object, that is, when there are multiple people in the in-person interview scene, the server can obtain the person image of the target interviewee from the database, and then extract the person image parameters, so as to focus on the target interviewee.
[0108] 102. Call the pre-set face detection model to perform face detection on the pre-in-person interview scene image to obtain the first face region image;
[0109] It can be understood that the pre-set face detection model can be trained based on a convolutional neural network model in a public dataset. The convolutional neural network model can be a large network model such as LeNet, AlexNet, VGG, etc., so as to be deployed on a high-computing-power server for precise cloud computing, or it can be a lightweight network model such as LFFD, Mobile-Net, Slim-320, etc., so as to be deployed on a low-computing-power mobile device for high-efficiency edge computing. This embodiment does not limit it.
[0110] Optionally, the server first inputs the pre-audit scenario image into the face detection model obtained after training. The face detection model includes multiple layers of networks. The server first calls the face feature extraction network therein to extract face features of at least one feature size from the pre-audit scenario image, and then calls the face feature recognition network, so as to identify candidate face regions in the face features of each feature size according to the pre-set number of prior boxes corresponding to each feature size. Finally, the server calls the face feature screening network to screen each candidate face region to obtain the first face region image.
[0111] It can be understood that face features are used to describe the image features of a face. Exemplarily, the image features can include the color features, texture features, shape features, and spatial relationship features of the image, etc. The face detection model is used to extract features from the input image to obtain face features. The feature size is the size of the face feature. The face feature can be a feature map, and the feature size can include parameters such as the width and height of the map. Exemplarily, the face detection model can extract face features with a feature size of 32*32, face features with a feature size of 16*16, or face features with a feature size of 8*8. The face detection model can extract face features through a feature extraction layer. Exemplarily, the feature extraction layer includes at least one convolutional layer. Optionally, the convolutional layer includes a DW convolution (Depthwise separable convolution), where using the DW convolution can reduce the number of parameters and the computing cost, thus realizing lightweight operations.
[0112] An anchor can refer to boxes with different sizes and / or aspect ratios pre-set on an image in advance. The anchors are used for adjustment to continuously approach the true box containing the object to be detected. In fact, an anchor can be understood as pre-defining the width and height of the face region to be detected. During the model prediction process, the width and height are used to process the face features, and the face region is predicted. Among them, the prediction process is to adjust the (center) position and size of the anchor to obtain an alternative face region closest to the true face region. The sizes and / or aspect ratios of different anchors are different. The number of anchors with the same size and aspect ratio is one. Among them, the face detection model can identify the alternative face regions through a classification layer. Exemplarily, the classification layer can include a cascaded Feature Pyramid Networks (FPN).
[0113] It should be noted that the feature size is inversely proportional to the number of anchors. The fact that the feature size is inversely proportional to the number of anchors indicates that the feature size corresponds to the number of types of anchors. The number of types of anchors corresponding to different feature sizes is different. The sizes and / or aspect ratios of the anchors corresponding to different feature sizes can be partially the same or completely different. Exemplarily, for the 32*32 feature size, there is 1 anchor with an aspect ratio of 1:1. Another example is that for the 16*16 feature size, there are 2 anchors with aspect ratios of 1:1 and 2:1 respectively.
[0114] It can be understood that the screening of alternative face regions can be to reduce redundant alternative face regions and incorrect face regions, etc. The number of alternative face regions representing the same face can be multiple, and the multiple repeated alternative face regions representing the same face can be screened to reduce redundant alternative face regions. In addition, some alternative face regions are incorrect, and the incorrect alternative face regions can be deleted to improve the accuracy of the face region. At least one alternative face region obtained by screening is determined as the target face region. The face detection model can screen the alternative face regions through a post-processing layer. Exemplarily, the post-processing layer can use an algorithm such as Non-maximum suppression (NMS) or an algorithm based on Intersection over Union (IoU) to achieve screening.
[0115] 103. Perform key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0116] It can be understood that the process of key point detection is the process of deep semantic classification. For example, in the key point detection rule, the positions of 7 key points are defined during detection, including the positions of the left eye, right eye, nose, left corner of the mouth, right corner of the mouth, left eyebrow, and right eyebrow. If the server detects a face, these positions will be marked with dots in the corresponding face frame diagram. Optionally, the server can identify 68 key points for identifying the facial feature contours in the first face region image based on the 68-point detection method.
[0117] Before performing key point detection, it is necessary to construct a key point detection model. In this embodiment, the server performs a key point detection training task on a pre-set initial network model to obtain a face key point detection model. For example, the initial network model can be HRNet-v2, etc., and this embodiment does not limit it.
[0118] Optionally, the server first obtains a pre-set face image sample and the sample label corresponding to the face image sample. The sample label has pre-annotated the face key point position data in the face image sample. Secondly, prior calculation is performed on the face key point position data to obtain the structural prior feature of the face image sample. Then, the initial network model is called to extract the self-attention feature of the face image sample, and image recognition is performed based on the self-attention feature to obtain the face key point detection data in the face image sample. Finally, based on a pre-set first loss function, the first loss value between the self-attention feature and the structural prior feature is calculated, and based on a pre-set second loss function, the second loss value between the face key point detection data and the face key point position data is calculated. The first loss value and the second loss value are weighted and summed according to a pre-set weighted algorithm to obtain the current global loss value of the initial network model, and the network parameters of the initial network model are adjusted according to this global loss value to obtain the face key point detection model.
[0119] Optionally, to make the key point detection accurate and reduce the computational amount of image processing, the server also performs preprocessing methods on the first face region image, including but not limited to grayscale conversion, geometric transformation, smoothing, enhancement, restoration, noise reduction, etc., and this embodiment does not limit it. The server inputs the preprocessed first face region image into the face key point detection model for detection to obtain a set of face key points in the first face region image, and the set of face key points includes the position coordinate information of each face key point.
[0120] 104. According to this set of face key points, determine whether the face in the first face region image is complete. If it is determined that the face in the first face region image is incomplete, a face review specification prompt is sent to the face-reviewed person's terminal.
[0121] It can be understood that the server determines whether the face in the first face region image is complete based on the number of key points in the face key point set and the key point detection algorithm used in the face key point detection model. For example, if the 5-point detection algorithm is used in the face key point detection model, that is, 5 key points correspond to the facial features of a person. If the number of key points in the face key point set is 5, it is determined that the face is complete; otherwise, it is determined that the face is incomplete.
[0122] It can be understood that the face review specification prompt is sent to the terminal of the person being face-reviewed to prompt the person being face-reviewed to comply with the behavior specifications before the face review. For example, it checks whether the lens of the terminal of the person being face-reviewed is blocked by dirt, whether the lens is unclear, and whether the position of the person being face-reviewed is reasonable. Through this specification prompt, the person being face-reviewed is instructed to conduct a self-check to ensure that the face review process can proceed normally.
[0123] 105. If it is determined that the face in the first face region image is complete, then it is determined whether the person being face-reviewed in the first face region image is the preset target person being face-reviewed. If it is determined that the person being face-reviewed in the first face region image is the target person being face-reviewed, a face review ready prompt is sent to the face reviewer's terminal to start the face review.
[0124] It can be understood that the face review plan information is usually stored in a database in the server. The face review plan information should at least include data such as the identity information of the face reviewer, the identity information of the person being face-reviewed, and the face review plan time. Among them, the identity information of the person being face-reviewed contains the face image of the person being face-reviewed. Therefore, the server reads this face image from the database and then compares and matches it with the face image shown in the first face region image to determine whether the person being face-reviewed in the first face region image is the target person being face-reviewed. If the identity is confirmed correctly, the face reviewer is prompted to start the face review questions and answers through the terminal where the face reviewer is located.
[0125] In the embodiments of the present invention, by determining whether the face in the face review scene image is complete to determine whether the person being face-reviewed interacts with a third party, and by judging whether the current person being face-reviewed is the target person being face-reviewed, fraud behavior can be accurately identified, and the accuracy of the face review evaluation result can be improved.
[0126] Please refer to Figure 2 , the second embodiment of the face review monitoring method in the embodiments of the present invention includes:
[0127] 201. When a face review monitoring request is received, the terminal of the person being face-reviewed is called to collect a pre-face review scene image;
[0128] 202. The preset face detection model is called to perform face detection on the pre-face review scene image to obtain a first face region image;
[0129] 203. Perform key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0130] 204. According to the set of face key points, determine whether the face in the first face region image is complete. If it is determined that the face in the first face region image is incomplete, send a face review specification prompt to the interviewee terminal;
[0131] 205. If it is determined that the face in the first face region image is complete, then determine whether the interviewee in the first face region image is a preset target interviewee. If it is determined that the interviewee in the first face region image is the target interviewee, send a face review ready prompt to the interviewer terminal to start the face review;
[0132] Among them, the execution steps of steps 201-205 are similar to those of the above steps 101-105, and will not be elaborated here specifically.
[0133] 206. Based on a preset frequency, call the interviewee terminal to collect the first face review scenario image, and call the face detection model to perform face detection on the first face review scenario image to obtain a second face region image;
[0134] It can be understood that when starting the face review and entering the face review Q&A, the server also schedules the interviewee terminal to collect the current face review scenario image at a certain frequency, that is, the first face review scenario image, and then extracts the face image in the first face review image through the face detection model to obtain a second face region image. It should be noted that this frequency can be adjusted according to actual needs. A reasonable frequency can not only meet the monitoring of the face review scenario, but also save server energy consumption.
[0135] 207. Perform regression training on the preset neural network model for the deflection angle to obtain a head pose detection model;
[0136] It can be understood that in this embodiment, the type of the preset neural network model is not specifically limited. For example, networks such as Mobile-Net, ShuffleNet, and SqueezeNet. The preset neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network. The server first calls the pose feature extraction network to extract target face pose features from the target training images in the preset face pose training image set; secondly, calls the deflection angle classification network to calculate the multi-class deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and the preset deflection interval probability matrix; selects the deflection interval with the largest probability value from the multi-class deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtains the regression function corresponding to the target deflection interval;
[0137] Then, call the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose feature and the regression function, so as to obtain the deflection angles of the target training image in at least one preset direction;
[0138] Finally, based on the preset loss function and the preset deflection angle annotation information of the target training image, calculate the loss value corresponding to the deflection angle of the target training image in the above preset direction, and adjust the network parameters of the preset neural network model according to this loss value until the model converges, save the current network parameters, and obtain the head pose detection model. Specifically, the stochastic gradient descent algorithm can be used to adjust the network parameters of the model, so as to improve the fitting speed of the model. This embodiment does not limit it.
[0139] 208. Call the head pose detection model to perform pose detection on the second face region image, so as to obtain the deflection angles of the head of the target interviewee in at least one preset direction;
[0140] It can be understood that the server directly inputs the second face region image into the trained head pose detection model, and directly outputs the deflection angle of the head of the target interviewee in the preset direction in the output layer of the model. Optionally, before the server inputs the second face region image into the head pose detection model, it can also perform image preprocessing on the second face region image, so as to improve the model calculation speed.
[0141] Preferably, in head pose detection, there are three preset directions, namely pitch, roll, and yaw. As Figure 7 shown, pitch refers to the angle of rotation around the x-axis, specifically the angle of looking up and down in practice, and the angle range is [-99 degrees, 99 degrees]; yaw refers to the angle of rotation around the y-axis, that is, the deflection angle of the head to the left or right, and the angle range is [-99 degrees, 99 degrees]; roll refers to the angle of rotation around the z-axis, that is, the angle of the head turning to the left or right, and the angle range is [-180 degrees, 180 degrees].
[0142] 209. Compare the deflection angles in at least one preset direction with the preset standard deflection interval respectively. If the deflection angle in any one direction is not within the standard deflection interval, it is determined that the behavior of the target interviewee is abnormal, and an abnormal prompt is sent to the interviewer terminal.
[0143] It can be understood that during a normal in-person interview, the standard deflection angle of the interviewee's head in any preset direction should be 0 degrees, that is, no deflection occurs. In a specific in-person interview scenario, the standard deflection angle actually cannot reach the ideal state. Therefore, a certain tolerance angle is set for the standard deflection angle in any preset direction to obtain the corresponding standard deflection interval. For example, in one embodiment, the same tolerance angle, such as 10 degrees, is set for the deflection angle in each preset direction. Then the standard deflection interval is [-10 degrees, 10 degrees]. When the obtained deflection angle is not within this interval, it is determined that the behavior of the target interviewee is abnormal and may be interacting with a third party.
[0144] In an embodiment of the present invention, the process of detecting the deflection of the interviewee's head after the in-person interview starts is described in detail. The deflection angle of the interviewee's head in at least one direction is detected by a trained head pose detection model, and then it is compared with a reasonable standard deflection interval to accurately determine whether the current interviewee's behavior is abnormal.
[0145] Please refer to Figure 3 , the third embodiment of the in-person interview monitoring method in an embodiment of the present invention includes:
[0146] 301. When receiving an in-person interview monitoring request, call the interviewee's terminal to collect a pre-in-person interview scenario image;
[0147] 302. Call a preset face detection model to perform face detection on the pre-in-person interview scenario image to obtain a first face region image;
[0148] 303. Perform key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0149] 304. According to the set of face key points, determine whether the face in the first face region image is complete. If it is determined that the face in the first face region image is incomplete, send an in-person interview specification prompt to the interviewee's terminal;
[0150] 305. If it is determined that the face in the first face region image is complete, determine whether the interviewee in the first face region image is the preset target interviewee. If it is determined that the interviewee in the first face region image is the target interviewee, send an in-person interview ready prompt to the interviewer's terminal to start the in-person interview;
[0151] Among them, the execution steps of steps 301-305 are similar to those of the above steps 101-105, and will not be elaborated here specifically.
[0152] 306. Based on a preset frequency, call the interviewee's terminal to collect an image of the second face interview scenario, and perform face detection on the image of the second face interview scenario to determine whether the image of the second face interview scenario contains a face.
[0153] It can be understood that the preset frequency can be adjusted according to actual needs, and this embodiment does not limit it. In this embodiment, the specific face detection method is not specifically limited either. For example, it can be detected through a face detection model.
[0154] 307. If it is determined that the image of the second face interview scenario does not contain a face, it is determined that the current face interview scenario is abnormal, and an abnormality prompt is sent to the interviewer's terminal.
[0155] It can be understood that when it is determined that the image of the second face interview scenario does not contain a face, it is determined that there is no person in the current face interview scenario. At this time, the interviewer should be prompted through the terminal where the interviewer is located.
[0156] 308. If it is determined that the image of the second face interview scenario contains a face, extract the target face parameters in the image of the second face interview scenario.
[0157] It can be understood that the server can extract the target face parameters through a key point detection algorithm. The target face parameters are the face features of a person, such as interpupillary distance, nasal width, etc. This embodiment does not specifically limit them.
[0158] 309. Obtain the preset face parameters of the target interviewee from a preset database, and match the target face parameters with the preset face parameters. If the match fails, it is determined that the interviewee has been replaced, and a replacement prompt is sent to the interviewer's terminal.
[0159] It can be understood that the server usually pre-enters the information of the interviewee in the database, which includes the preset face parameters of the interviewee. The target face parameters are the face parameters to be detected. The server compares the two. If the two are consistent, it is determined that the interviewee in the current face interview scenario is the target interviewee; otherwise, it is determined that a person has been replaced.
[0160] Optionally, after the server determines that the interviewee has not been replaced, it also detects whether the mouth of the target interviewee is within the screen through a key point detection algorithm. When the mouth is not within the screen, it is determined that there is a risk of proxy answering in the face interview.
[0161] In the embodiment of the present invention, the process of detecting whether there is no person or the interviewee has been replaced in the face interview scenario after the face interview starts is described in detail. By face detection to determine whether there is a person in the face interview scenario, and by extracting face parameters and comparing them with the recorded preset face parameters to accurately identify whether the interviewee has been replaced, thereby improving the accuracy of identifying face interview fraud behavior.
[0162] The above describes the in-person review monitoring method in the embodiments of the present invention. Next, the in-person review monitoring device in the embodiments of the present invention will be described. Please refer to Figure 4 In one embodiment, the in-person review monitoring device in the embodiments of the present invention includes:
[0163] An image acquisition module 401, configured to call the terminal of the person being reviewed to acquire a pre-in-person review scenario image when receiving an in-person review monitoring request;
[0164] A face detection module 402, configured to call a pre-set face detection model to perform face detection on the pre-in-person review scenario image to obtain a first face region image;
[0165] A key point detection module 403, configured to perform key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0166] A scene correction module 404, configured to determine whether the face in the first face region image is complete according to the set of face key points; if it is determined that the face in the first face region image is incomplete, a face review specification prompt is sent to the terminal of the person being reviewed;
[0167] An identity verification module 405, configured to determine whether the person being reviewed in the first face region image is a pre-set target person being reviewed if it is determined that the face in the first face region image is complete;
[0168] An in-person review start module 406, configured to send an in-person review ready prompt to the in-person reviewer's terminal to start the in-person review if it is determined that the person being reviewed in the first face region image is the target person being reviewed.
[0169] In the embodiments of the present invention, by determining whether the face in the in-person review scenario image is complete to determine whether the person being reviewed interacts with a third party, and judging whether the current person being reviewed is the target person being reviewed, fraudulent behavior can be accurately identified, and the accuracy of the in-person review evaluation result can be improved.
[0170] Please refer to Figure 5 In another embodiment, the in-person review monitoring device in the embodiments of the present invention includes:
[0171] An image acquisition module 401, configured to call the terminal of the person being reviewed to acquire a pre-in-person review scenario image when receiving an in-person review monitoring request;
[0172] A face detection module 402, configured to call a pre-set face detection model to perform face detection on the pre-in-person review scenario image to obtain a first face region image;
[0173] The key point detection module 403 is used to perform key point detection on the first face region image to obtain a set of face key points in the first face region image;
[0174] The scene correction module 404 is used to determine whether the face in the first face region image is complete according to the set of face key points; if it is determined that the face in the first face region image is incomplete, a face review specification prompt is sent to the interviewee terminal;
[0175] The identity verification module 405 is used to determine whether the interviewee in the first face region image is a preset target interviewee if it is determined that the face in the first face region image is complete;
[0176] The face review start module 406 is used to send a face review ready prompt to the reviewer terminal to start the face review if it is determined that the interviewee in the first face region image is the target interviewee.
[0177] The deflection detection module 407 is used to call the interviewee terminal to collect the first face review scene image based on a preset frequency; call the face detection model to perform face detection on the first face review scene image to obtain a second face region image; perform regression training on the deflection angle of a preset neural network model to obtain a head pose detection model; call the head pose detection model to perform pose detection on the second face region image to obtain the deflection angle of the head of the target interviewee in at least one preset direction; compare the deflection angles in the at least one preset direction with a preset standard deflection interval respectively, and if the deflection angle in any direction is not within the standard deflection interval, it is determined that the behavior of the target interviewee is abnormal, and an abnormality prompt is sent to the reviewer terminal.
[0178] The face re-inspection module 408 is used to call the interviewee terminal to collect a second face review scene image based on a preset frequency, perform face detection on the second face review scene image to determine whether the second face review scene image contains a face, if it is determined that the second face review scene image does not contain a face, it is determined that the current face review scene is abnormal, and an abnormality prompt is sent to the reviewer terminal, if it is determined that the second face review scene image contains a face, the target face parameters in the second face review scene image are extracted, the preset face parameters of the target interviewee are obtained from a preset database, and the target face parameters are matched with the preset face parameters, if the matching fails, it is determined that the interviewee has been replaced, and a replacement prompt is sent to the reviewer terminal.
[0179] Among them, the face detection module 402 specifically includes:
[0180] An input unit 4021 for inputting the pre-in-person review scenario image into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network;
[0181] A feature extraction unit 4022 for calling the face feature extraction network to extract face features of at least one feature size from the pre-in-person review scenario image;
[0182] An identification unit 4023 for calling the face feature recognition network to identify candidate face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size;
[0183] A screening unit 4024 for calling the face feature screening network to screen each of the candidate face regions to obtain a first face region image.
[0184] Wherein, the key point detection module 403 specifically includes:
[0185] A model construction unit 4031 for performing a key point detection training task on a preset initial network model to obtain a face key point detection model;
[0186] A detection unit 4032 for performing image preprocessing on the first face region image and inputting the preprocessed first face region image into the face key point detection model for detection to obtain a set of face key points in the first face region image, where the set of face key points includes position coordinate information of each face key point.
[0187] Wherein, the model construction unit 4031 is specifically used for:
[0188] Obtaining a preset face image sample and a sample label corresponding to the face image sample, where the sample label includes face key point position data in the face image sample;
[0189] Performing prior calculation on the face key point position data to obtain a structural prior feature of the face image sample;
[0190] Calling the initial network model to extract the self-attention feature of the face image sample and performing image recognition based on the self-attention feature to obtain face key point detection data in the face image sample;
[0191] Calculating a first loss value between the self-attention feature and the structural prior feature based on a preset first loss function, and calculating a second loss value between the face key point detection data and the face key point position data based on a preset second loss function;
[0192] Perform a weighted summation calculation on the first loss value and the second loss value according to a preset weighted algorithm to obtain the current global loss value of the initial network model;
[0193] Adjust the network parameters of the initial network model according to the global loss value to obtain a face key point detection model.
[0194] In the embodiments of the present invention, the modular design enables the hardware of each part of the clinical pathway construction device to focus on the realization of a certain function, maximizing the performance of the hardware. At the same time, the modular design also reduces the coupling between the modules of the device, making it more convenient to maintain.
[0195] Above Figure 4 And Figure 5 The face review monitoring device in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the face review monitoring device in the embodiments of the present invention is described in detail from the perspective of hardware processing.
[0196] Figure 6 FIG. is a schematic structural diagram of a face review monitoring device provided by an embodiment of the present invention. The face review monitoring device 600 may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 610 (for example, one or more processors) and a memory 620, and one or more storage media 630 for storing application programs 633 or data 632 (for example, one or more mass storage devices). Among them, the memory 620 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the face review monitoring device 600. Further, the processor 610 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the face review monitoring device 600.
[0197] The face review monitoring device 600 may further include one or more power supplies 640, one or more wired or wireless network interfaces 650, one or more input / output interfaces 660, and / or one or more operating systems 631, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 6 The shown structure of the face review monitoring device does not constitute a limitation on the face review monitoring device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0198] The present invention also provides a face review monitoring device. The computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute each step of the face review monitoring method in the above-mentioned various embodiments.
[0199] The present invention also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute each step of the face review monitoring method.
[0200] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0202] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0203] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A face-to-face review monitoring method, characterized in that, The face review monitoring method includes: When a face review monitoring request is received, call the terminal of the person being interviewed to collect pre-face review scenario images; Call a pre-set face detection model to perform face detection on the pre-face review scenario images to obtain a first face region image; Perform key point detection on the first face region image to obtain a set of face key points in the first face region image; Based on the set of face key points, determine whether the face in the first face region image is complete; If it is determined that the face in the first face region image is incomplete, send a face review specification prompt to the terminal of the person being interviewed; If it is determined that the face in the first face region image is complete, determine whether the person being interviewed in the first face region image is a pre-set target person being interviewed; If it is determined that the person being interviewed in the first face region image is the target person being interviewed, send a face review ready prompt to the face reviewer's terminal to start the face review; Perform regression training on the deflection angle of a pre-set neural network model to obtain a head pose detection model, where the pre-set neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network; The performing regression training on the deflection angle of the pre-set neural network model to obtain a head pose detection model includes: calling the pose feature extraction network to extract target face pose features from a target training image in a pre-set face pose training image set; calling the deflection angle classification network to calculate a multi-classification deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and a pre-set deflection interval probability matrix; selecting the deflection interval with the largest probability value from the multi-classification deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtaining the regression function corresponding to the target deflection interval; calling the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose features and the regression function to obtain the deflection angle of the target training image in at least one pre-set direction; based on a pre-set loss function and the pre-set deflection angle annotation information of the target training image, calculate the loss value corresponding to the deflection angle of the target training image in the at least one pre-set direction; adjust the network parameters of the pre-set neural network model according to the loss value until the loss value is less than a pre-set threshold, determine that the pre-set neural network model converges, and save the current network parameters of the pre-set neural network model to obtain a head pose detection model.
2. The face review monitoring method according to claim 1, wherein After the step of if it is determined that the person being interviewed in the first face region image is the target person being interviewed, send a face review ready prompt to the face reviewer's terminal to start the face review, it further includes: Based on a pre-set frequency, call the terminal of the person being interviewed to collect first face review scenario images; Call the face detection model to perform face detection on the first face review scenario images to obtain a second face region image; Call the head pose detection model to perform pose detection on the second face region image to obtain the deflection angle of the head of the target person being interviewed in at least one pre-set direction; Compare the deflection angles in the at least one preset direction with a preset standard deflection range respectively. If the deflection angle in any direction is not within the standard deflection range, it is determined that the behavior of the target interviewee is abnormal, and an abnormality prompt is sent to the interviewer terminal.
3. The face-to-face review monitoring method according to claim 1, characterized in that, After it is determined that the interviewee in the first face region image is the target interviewee and an interview ready prompt is sent to the interviewer terminal to start the interview, the following steps are further included: Call the interviewee terminal to collect the second interview scene image based on a preset frequency. Perform face detection on the second interview scene image to determine whether a face is included in the second interview scene image. If it is determined that the second interview scene image does not include a face, it is determined that the current interview scene is abnormal, and an abnormality prompt is sent to the interviewer terminal. If it is determined that the second interview scene image includes a face, extract the target face parameters in the second interview scene image. Obtain the preset face parameters of the target interviewee from a preset database, and match the target face parameters with the preset face parameters. If the match fails, it is determined that the interviewee has been replaced, and a replacement prompt is sent to the interviewer terminal.
4. The face review monitoring method according to claim 1, characterized in that The step of performing face detection on the pre-interview scene image by calling a preset face detection model to obtain the first face region image includes: Input the pre-interview scene image into the face detection model, where the face detection model includes a face feature extraction network, a face feature recognition network, and a face feature screening network. Call the face feature extraction network to extract face features of at least one feature size from the pre-interview scene image. Call the face feature recognition network to identify candidate face regions in the face features of each feature size according to a preset number of prior boxes corresponding to each feature size. Call the face feature screening network to screen each of the candidate face regions to obtain the first face region image.
5. The face review monitoring method according to any one of claims 1-4, characterized in that The step of performing key point detection on the first face region image to obtain the set of face key points in the first face region image includes: Execute a key point detection training task on a preset initial network model to obtain a face key point detection model. Perform image preprocessing on the first face region image, and input the preprocessed first face region image into the face key point detection model for detection to obtain the set of face key points in the first face region image, where the set of face key points includes the position coordinate information of each face key point.
6. The face review monitoring method according to claim 5, wherein The step of executing a key point detection training task on a preset initial network model to obtain a face key point detection model includes: Obtain a preset face image sample and the sample label corresponding to the face image sample, where the sample label includes the face key point position data in the face image sample. Perform prior calculation on the face key point position data to obtain the structural prior feature of the face image sample. Call the initial network model to extract the self-attention features of the face image sample, and perform image recognition based on the self-attention features to obtain the face key point detection data in the face image sample; Calculate the first loss value between the self-attention features and the structural prior features based on a preset first loss function, and calculate the second loss value between the face key point detection data and the face key point position data based on a preset second loss function; Perform weighted summation calculation on the first loss value and the second loss value according to a preset weighting algorithm to obtain the current global loss value of the initial network model; Adjust the network parameters of the initial network model according to the global loss value to obtain a face key point detection model.
7. A face-to-face review monitoring device, characterized in that, The in-person review monitoring device includes: An image acquisition module, configured to call the terminal of the person being reviewed to acquire a pre-in-person review scene image when receiving an in-person review monitoring request; A face detection module, configured to call a preset face detection model to perform face detection on the pre-in-person review scene image to obtain a first face region image; A key point detection module, configured to perform key point detection on the first face region image to obtain a set of face key points in the first face region image; A scene correction module, configured to determine whether the face in the first face region image is complete according to the set of face key points; if it is determined that the face in the first face region image is incomplete, send an in-person review specification prompt to the terminal of the person being reviewed; An identity verification module, configured to determine whether the person being reviewed in the first face region image is a preset target person being reviewed if it is determined that the face in the first face region image is complete; An in-person review start module, configured to send an in-person review ready prompt to the in-person reviewer's terminal to start the in-person review if it is determined that the person being reviewed in the first face region image is the target person being reviewed; A deflection detection module, configured to perform regression training on the deflection angle of a preset neural network model to obtain a head pose detection model, where the preset neural network model includes a pose feature extraction network, a deflection angle classification network, and a deflection angle regression network; Performing regression training on the preset neural network model for the deflection angle to obtain a head pose detection model, including: calling the pose feature extraction network to extract target face pose features from the target training images in the preset face pose training image set; calling the deflection angle classification network to calculate the multi-classification deflection interval probability distribution corresponding to the target face pose features according to the target face pose features and the preset deflection interval probability matrix; selecting the deflection interval with the largest probability value from the multi-classification deflection interval probability distribution as the target deflection interval corresponding to the target face pose features, and obtaining the regression function corresponding to the target deflection interval; calling the deflection angle regression network to perform regression calculation of the deflection angle according to the target face pose features and the regression function to obtain the deflection angle of the target training image in at least one preset direction; calculating the loss value corresponding to the deflection angle of the target training image in the at least one preset direction based on the preset loss function and the preset deflection angle annotation information of the target training image; adjusting the network parameters of the preset neural network model according to the loss value until the loss value is less than the preset threshold, determining that the preset neural network model converges, and saving the current network parameters of the preset neural network model to obtain a head pose detection model.
8. A face review monitoring device, characterized in that, The face audit monitoring device includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the face audit monitoring device executes each step of the face audit monitoring method according to any one of claims 1-6.
9. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, each step of the face audit monitoring method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Head posture detection device and method
CN112307806A
Face false detection optimization method and system based on depth camera and face detection equipment
CN112580434A