Video interview assisting method and device, equipment and storage medium

By analyzing the target person's gaze, head, and hand movements through audio and video data, fraudulent behavior in video interviews can be identified, solving the problems of inconsistent standards and decreased attention among approval personnel, and improving the accuracy of fraud detection.

CN115273221BActive Publication Date: 2026-05-08PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2022-06-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Because different approvers have different standards for abnormal behavior, some abnormal behaviors are not given enough attention. Furthermore, due to the varying levels of expertise among approvers and the decreased attention due to long working hours, fraud risks are difficult to detect during video interviews.

Method used

By acquiring audio and video data for audio detection, the system identifies the target person's eye movements, head movements, and hand movements at the end of the approval personnel's audio recording. It then uses preset action templates to match and identify fraudulent behavior, generating and sending reminder messages.

Benefits of technology

It improves the accuracy of identifying fraudulent behavior during video interviews and reduces the risk of missed detections due to decreased attention of approvers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273221B_ABST
    Figure CN115273221B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses a video face examination auxiliary method, device and equipment and a storage medium, which are used for improving the identification accuracy of fraud in video face examination. The video face examination auxiliary method comprises the following steps: when an examination personnel and a target personnel are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is carried out based on the audio and video data; if the audio and video data contain the audio of the examination personnel, the target personnel is identified for a plurality of actions when the audio of the examination personnel ends, the plurality of actions including a gaze action, a head action and a hand action; if the gaze action meets a preset gaze action, the head action meets a preset head action or the hand action forms an obstacle to the face of the target personnel, it is determined that the target personnel has a face examination fraud; alert information corresponding to the face examination fraud is generated, and the alert information is sent to a face examination alert terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a video interview assistance method, apparatus, device, and storage medium. Background Technology

[0002] Video interviews refer to the review of users through video and human methods. Approval personnel need to combine their own interview experience with the materials provided by the user to ask interview questions and pay close attention to whether the user has any abnormal behavior when answering questions, so as to determine whether the user is lying or committing fraud.

[0003] Because different approvers have inconsistent standards for abnormal behavior, some abnormal behaviors are not given enough attention, which hides huge fraud risks. At the same time, due to the varying levels of competence among approvers and the fact that they are prone to lapse in attention when working long hours, abnormal user behavior is overlooked, making it difficult to detect fraud risks. Summary of the Invention

[0004] This invention provides a video interview assistance method, apparatus, device, and storage medium to improve the accuracy of fraud identification in video interviews.

[0005] The first aspect of this invention provides a video interview assistance method, comprising: when the approver and the target person are in a preset audio and video detection area, acquiring corresponding audio and video data, and performing audio detection based on the audio and video data; if the audio of the approver is present in the audio and video data, then when the audio of the approver ends, identifying multiple actions of the target person, including eye movements, head movements, and hand movements; if the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the face of the target person, then determining that the target person has engaged in interview fraud, wherein the preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right; generating a reminder message corresponding to the interview fraud, and sending the reminder message to an interview reminder terminal.

[0006] In one feasible implementation, if the audio of the approver is present in the audio-visual data, then when the audio of the approver ends, multiple actions of the target person are identified. These actions include eye contact, head movements, and hand movements. The steps include: if the audio of the approver is present in the audio-visual data, then when the audio of the approver ends, acquiring a facial video of the target person; performing eye contact recognition based on the facial video to obtain an eye contact recognition result; performing head movement recognition based on the facial video to obtain a head movement recognition result; and performing hand movement recognition based on the facial video to obtain a hand movement recognition result.

[0007] In one feasible implementation, the step of performing gaze action recognition based on the facial video to obtain gaze action recognition results includes: mapping the gaze angle values ​​of the target person in each frame of the facial video to a Cartesian coordinate system, and connecting the gaze coordinate points corresponding to the gaze angle values ​​in each frame to generate gaze action line segments of the target person, where the horizontal coordinate of the gaze coordinate points indicates the video frame and the vertical coordinate indicates the gaze angle value; calling a preset gaze point detection model to perform template matching on the gaze action line segments; if the matching distance between any line segment of the gaze action line segments and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, then it is determined that the gaze action of the target person conforms to the preset gaze action, which includes slow gaze, fast gaze, and gaze jitter; if the matching distance between each line segment of the gaze action line segments and the preset gaze action curve template is less than the preset gaze action matching distance, then it is determined that the gaze action of the target person does not conform to the preset gaze action.

[0008] In one feasible implementation, the step of performing head motion recognition based on the face video to obtain head motion recognition results includes: mapping the head posture angle values ​​of the target person in each frame of the face video to a Cartesian coordinate system, and connecting the head posture coordinate points corresponding to the head posture angle values ​​in each frame of the video to generate head motion line segments of the target person. The horizontal coordinate of the head posture coordinate points is used to indicate the video frame, and the vertical coordinate is used to indicate the head posture angle value. The head motion line segments are template matched using a preset head posture detection model. If the matching distance between any line segment of the head motion line segment and the preset head motion curve template is greater than or equal to the preset head motion matching distance, it is determined that the head motion of the target person conforms to the preset head motion, which includes rapid head rotation, head rotation to the left, and head rotation to the right. If the matching distance between each line segment of the head motion line segment and the preset head motion curve template is less than the preset head motion matching distance, it is determined that the head motion of the target person does not conform to the preset head motion.

[0009] In one feasible implementation, the step of performing hand gesture recognition based on the face video to obtain a hand gesture recognition result includes: generating a face region bounding box of the target person based on the face video; performing hand detection on the face video; if a hand is present in the face video, generating a hand bounding box corresponding to the hand; calculating the intersection value between the face region bounding box and the hand bounding box, the intersection value indicating the ratio of the area of ​​the overlapping area between the face region bounding box and the hand bounding box to the total area of ​​the face region bounding box and the hand bounding box; if the intersection value is greater than or equal to a preset value, determining that the target person's hand gesture obstructs the target person's face; if the intersection value is less than the preset value, determining that the target person's hand gesture does not obstruct the target person's face.

[0010] In one feasible implementation, the step of acquiring corresponding audio and video data when the approver and the target personnel are in a preset audio and video detection area, and performing audio detection based on the audio and video data, includes: acquiring corresponding audio and video data when the approver and the target personnel are in the preset audio and video detection area; extracting audio data from the audio and video data to obtain audio data; extracting voiceprint features from the audio data to obtain a voiceprint feature sequence; if the voiceprint feature sequence matches a preset approver voiceprint feature sequence, then it is determined that the audio of the approver exists in the audio and video data; if the voiceprint feature sequence does not match the preset approver voiceprint feature sequence, then it is determined that the audio of the approver does not exist in the audio and video data.

[0011] In one feasible implementation, after acquiring corresponding audio and video data when the approver and the target person are in a preset audio and video detection area, and performing audio detection based on the audio and video data, before generating a reminder message corresponding to the face-to-face fraud behavior and sending the reminder message to the face-to-face reminder terminal, the method further includes: if the audio of the approver is present in the audio and video data, then acquiring the face video of the target person when the approver's audio ends; performing color detection on the ear of the target person based on the face video; if the color of the ear matches a preset color, then determining that the target person has engaged in face-to-face fraud behavior.

[0012] A second aspect of the present invention provides a video interview assistance device, comprising: an audio detection module, configured to acquire corresponding audio and video data and perform audio detection based on the audio and video data when the approver and the target person are in a preset audio and video detection area; an action recognition module, configured to identify multiple actions of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data, the multiple actions including eye movements, head movements, and hand movements; a first determination module, configured to determine that the target person has engaged in interview fraud if the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the face of the target person, wherein the preset eye movements include slow eye glances, fast eye glances, and eye tremors, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right; and an information sending module, configured to generate reminder information corresponding to the interview fraud and send the reminder information to an interview reminder terminal.

[0013] In one feasible implementation, the action recognition module includes: an acquisition unit, configured to acquire a facial video of the target person when the audio of the approver ends if the audio of the approver is present in the audio-visual data; an eye movement recognition unit, configured to perform eye movement recognition based on the facial video to obtain an eye movement recognition result; a head movement recognition unit, configured to perform head movement recognition based on the facial video to obtain a head movement recognition result; and a hand movement recognition unit, configured to perform hand movement recognition based on the facial video to obtain a hand movement recognition result.

[0014] In one feasible implementation, the gaze action recognition unit is specifically used to: map the gaze angle value of the target person in each frame of the face video to a Cartesian coordinate system, and connect the gaze coordinate points corresponding to the gaze angle values ​​in each frame of the video to generate gaze action line segments of the target person, where the horizontal coordinate of the gaze coordinate points indicates the video frame and the vertical coordinate indicates the gaze angle value; call a preset gaze point detection model to perform template matching on the gaze action line segments; if the matching distance between any line segment of the gaze action line segments and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, then it is determined that the gaze action of the target person conforms to the preset gaze action, which includes slow gaze, fast gaze, and gaze jitter; if the matching distance between each line segment of the gaze action line segments and the preset gaze action curve template is less than the preset gaze action matching distance, then it is determined that the gaze action of the target person does not conform to the preset gaze action.

[0015] In one feasible implementation, the head movement recognition unit is specifically used to: map the head posture angle values ​​of the target person in each frame of the face video to a Cartesian coordinate system, and connect the head posture coordinate points corresponding to the head posture angle values ​​in each frame of the video to generate head movement line segments of the target person, wherein the horizontal coordinate of the head posture coordinate points is used to indicate the video frame, and the corresponding vertical coordinate is used to indicate the head posture angle value; perform template matching on the head movement line segments using a preset head posture detection model; if the matching distance between any line segment in the head movement line segments and the preset head movement curve template is greater than or equal to the preset head movement matching distance, then it is determined that the head movement of the target person conforms to the preset head movement, wherein the preset head movement includes rapid head rotation, head rotation to the left, and head rotation to the right; if the matching distance between each line segment in the head movement line segments and the preset head movement curve template is less than the preset head movement matching distance, then it is determined that the head movement of the target person does not conform to the preset head movement.

[0016] In one feasible implementation, the hand gesture recognition unit is specifically used to: generate a face region bounding box of the target person based on the face video; perform hand detection on the face video; if a hand is present in the face video, generate a hand bounding box corresponding to the hand; calculate the intersection value between the face region bounding box and the hand bounding box, the intersection value indicating the ratio of the area of ​​the overlapping area between the face region bounding box and the hand bounding box to the total area of ​​the face region bounding box and the hand bounding box; if the intersection value is greater than or equal to a preset value, determine that the target person's hand gesture obstructs the target person's face; if the intersection value is less than the preset value, determine that the target person's hand gesture does not obstruct the target person's face.

[0017] In one feasible implementation, the audio detection module is specifically used for: acquiring corresponding audio and video data when the approver and the target person are in a preset audio and video detection area; extracting audio data from the audio and video data to obtain audio data; extracting voiceprint features from the audio data to obtain a voiceprint feature sequence; if the voiceprint feature sequence matches a preset approver voiceprint feature sequence, then it is determined that the audio of the approver exists in the audio and video data; if the voiceprint feature sequence does not match the preset approver voiceprint feature sequence, then it is determined that the audio of the approver does not exist in the audio and video data.

[0018] In one feasible implementation, the video interview assistance device further includes: an acquisition module, configured to acquire the target person's face video when the audio of the approver ends if the audio of the approver is present in the audio and video data; a color detection module, configured to perform color detection on the target person's ears based on the face video; and a second determination module, configured to determine that the target person has engaged in interview fraud if the color of the ears matches a preset color.

[0019] A third aspect of the present invention provides a video interview assistance device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the video interview assistance device to perform the above-described video interview assistance method.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the video interview assistance method described above.

[0021] In the technical solution provided by this invention, when the approver and the target are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is performed based on the audio and video data. If the audio and video data contains the approver's audio, then when the approver's audio ends, multiple actions of the target are identified, including eye movements, head movements, and hand movements. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, then it is determined that the target has engaged in face-to-face verification fraud. The preset eye movements include slow glances, fast glances, and eye tremors, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. A reminder message corresponding to the face-to-face verification fraud is generated and sent to the face-to-face verification reminder terminal. In this embodiment of the invention, video communication is established between the approver and the target. When the approver and the target are in a preset audio and video detection area, audio and video data are acquired and audio detection is performed. If the approver's audio is present, the target's eye movements, head movements, and hand movements are identified when the approver's audio ends. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, it is determined that the target has engaged in fraudulent behavior during the face-to-face interview. A reminder message corresponding to the fraudulent behavior is generated and sent to the face-to-face interview reminder terminal, thereby improving the accuracy of fraud identification in video face-to-face interviews. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of one embodiment of the video interview assistance method in this invention;

[0023] Figure 2 This is a schematic diagram of another embodiment of the video interview assistance method in this invention;

[0024] Figure 3 This is a schematic diagram of one embodiment of the video interview assistance device in this invention;

[0025] Figure 4 This is a schematic diagram of another embodiment of the video interview assistance device in this invention;

[0026] Figure 5 This is a schematic diagram of one embodiment of the video interview auxiliary device in this invention. Detailed Implementation

[0027] This invention provides a video interview assistance method, apparatus, device, and storage medium to improve the accuracy of fraud identification in video interviews.

[0028] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0030] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0031] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the video interview assistance method in this invention includes:

[0032] 101. When the approver and the target personnel are in the preset audio and video detection area, acquire the corresponding audio and video data, and perform audio detection based on the audio and video data;

[0033] It is understood that the executing entity of this invention can be a video interview assistance device or a terminal, and no specific limitation is made here. This embodiment of the invention will be described using a terminal as an example.

[0034] The terminal establishes video communication between the approver and the target personnel. When the video connection is established, each personnel has their own video screen window, which is the audio and video detection area. Audio and video detection is performed on both personnel through these video screen windows. The audio in the audio and video data can be the approver's audio, the target personnel's audio, or the audio of both personnel.

[0035] 102. If the audio and video data contains the audio of the approver, then when the audio of the approver ends, the target person will be identified by multiple actions, including eye movements, head movements and hand movements.

[0036] Audio detection is used to detect whether the audio of the approver exists in the audio and video data. If the audio of the approver exists, the terminal will perform action recognition on the target person when the audio of the approver ends. This can not only more accurately identify the target person's actions when answering questions, but also reduce energy consumption.

[0037] 103. If the eye movements match the preset eye movements, the head movements match the preset head movements, or the hand movements obscure the face of the target person, then it is determined that the target person has committed fraud during the face-to-face interview. Among them, the preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head turning, turning the head to the left, and turning the head to the right.

[0038] For example, if the target person's eye movements are characterized by slow glances, head movements by rapid head turns, or hand movements that obscure the target person's face, then it is determined that the target person has engaged in fraudulent behavior during the face-to-face interview.

[0039] 104. Generate a reminder message corresponding to the face-to-face interview fraud behavior and send the reminder message to the face-to-face interview reminder terminal.

[0040] The reminder message can be either a voice message or a text message. For example, the terminal can generate a reminder message corresponding to fraudulent behavior during face-to-face review. The reminder message could be a voice message: "Please note that the target person is engaging in fraudulent behavior!" The terminal sends the voice reminder message to the face-to-face review reminder terminal, which then plays the voice message: "Please note that the target person is engaging in fraudulent behavior!" Alternatively, the reminder message could be a text message: "Please note that the target person is engaging in fraudulent behavior!" The terminal sends the text reminder message to the face-to-face review reminder terminal, which displays the text "Please note that the target person is engaging in fraudulent behavior!" in the target person's video screen window, thereby alerting the approval personnel to the potential fraudulent behavior of the target person.

[0041] In this embodiment of the invention, when the approver and the target are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is performed based on the audio and video data. If the audio and video data contains the approver's audio, then when the approver's audio ends, multiple actions of the target are identified, including eye movements, head movements, and hand movements. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, then it is determined that the target has engaged in fraudulent behavior during the face-to-face interview. The preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. A reminder message corresponding to the fraudulent behavior during the face-to-face interview is generated and sent to the face-to-face interview reminder terminal, thereby improving the accuracy of fraud identification in video interviews.

[0042] Please see Figure 2 Another embodiment of the video interview assistance method in this invention includes:

[0043] 201. When the approver and the target personnel are in the preset audio and video detection area, acquire the corresponding audio and video data, and perform audio detection based on the audio and video data;

[0044] Specifically, (1) when the approver and the target are in the preset audio and video detection area, the terminal acquires the corresponding audio and video data; (2) the terminal extracts the audio data from the audio and video data to obtain audio data; (3) the terminal extracts the voiceprint features from the audio data to obtain the voiceprint feature sequence; (4) if the voiceprint feature sequence matches the preset approver voiceprint feature sequence, the terminal determines that the audio of the approver exists in the audio and video data; (5) if the voiceprint feature sequence does not match the preset approver voiceprint feature sequence, the terminal determines that the audio of the approver does not exist in the audio and video data.

[0045] For example, when the approver and the target personnel are in a preset audio and video detection area, the terminal acquires the corresponding audio and video data; the terminal extracts the audio data from the audio and video data to obtain audio data; the terminal extracts voiceprint features from the audio data to obtain a voiceprint feature sequence, which is used to indicate the sound wave spectrum; if the voiceprint feature sequence matches the preset approver voiceprint feature sequence, the terminal determines that the approver's audio exists in the audio and video data; if the voiceprint feature sequence does not match the preset approver voiceprint feature sequence, the terminal determines that the approver's audio does not exist in the audio and video data.

[0046] 202. If the audio data contains the audio of the approver, then obtain the facial video of the target person when the audio of the approver ends.

[0047] The terminal uses Dynamic Time Warping (DTW) to recognize eye movements and head movements. DTW is used to calculate the similarity between two time series, and is especially suitable for time series of different lengths and rhythms. For example, if two people say the same word, they will get two different audio sequences. DTW will automatically warp the time series, that is, perform local scaling on the time axis to make the two audio sequences as consistent as possible and obtain the maximum possible similarity.

[0048] DTW employs dynamic programming (DP) for time warping calculations. For example, consider two action time series Q and C, with lengths n and m respectively. In an action matching scenario, one series serves as the reference template, and the other as the test template. Action time series Q has n frames, and the feature value (a number or a vector) of the i-th frame is qi, i.e., Q = q1, q2, ..., qi, ..., qn; C = c1, c2, ..., cj, ..., cm. An n*m matrix grid is constructed, where matrix elements (i, j) represent the distance d(qi, cj) between points qi and cj, representing the similarity between each point in action time series Q and each point in action time series C. A smaller distance indicates higher similarity. Euclidean distance is typically used, with the formula: d(qi, cj) = ... i ,c j )=(q i -c j ) 2 .

[0049] 203. Perform gaze and motion recognition based on facial video to obtain gaze and motion recognition results;

[0050] Based on the audio and video detection area corresponding to the target person, the angle formed by the plane corresponding to the audio and video detection area and the extension line of the target person's gaze direction is determined as the gaze landing angle value. With the center point of the audio and video detection area as the origin, four directional axes are established, namely: left horizontal axis, right horizontal axis, up vertical axis and down vertical axis, which are used to determine the gaze landing direction of the target person.

[0051] Specifically, (1) the terminal maps the angle value of the gaze landing point in each frame of the face video of the target person to a Cartesian coordinate system, and connects the gaze coordinate points corresponding to the angle value of the gaze landing point in each frame of the video to generate the gaze action line segment of the target person. The horizontal coordinate of the gaze coordinate point is used to indicate the video frame, and the vertical coordinate is used to indicate the angle value of the gaze landing point; (2) the terminal calls the preset gaze point detection model to perform template matching on the gaze action line segment; (3) if the matching distance between any line segment in the gaze action line segment and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, the terminal determines that the gaze action of the target person conforms to the preset gaze action. The preset gaze action includes slow gaze, fast gaze and gaze jitter; (4) if the matching distance between each line segment in the gaze action line segment and the preset gaze action curve template is less than the preset gaze action matching distance, the terminal determines that the gaze action of the target person does not conform to the preset gaze action.

[0052] For example, the terminal maps the angle values ​​of the gaze points of the target person in each frame of the face video to a Cartesian coordinate system. The multiple gaze angle values ​​in the left direction are: 10 degrees, 20 degrees, 30 degrees, 20 degrees, 10 degrees, and 0 degrees. The terminal then connects the gaze coordinates corresponding to the gaze angle values ​​in each frame to generate a line segment representing the target person's gaze action. The horizontal coordinate of each gaze coordinate indicates the video frame, and the vertical coordinate indicates the gaze angle value. The multiple gaze coordinates are: (1, 10), (2, 20), (3, 30), (4, 20), and (5, 6). 0), (5, 10) and (6, 0); The terminal calls the preset gaze point detection model to perform template matching on the gaze action line segments; If the matching distance between any line segment in the gaze action line segment and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, the terminal determines that the gaze action of the target person conforms to the preset gaze action, which includes slow gaze, fast gaze and gaze jitter; If the matching distance between each line segment in the gaze action line segment and the preset gaze action curve template is less than the preset gaze action matching distance, the terminal determines that the gaze action of the target person does not conform to the preset gaze action.

[0053] 204. Perform head motion recognition based on the facial video to obtain the head motion recognition results;

[0054] Based on the audio and video detection area corresponding to the target person, the angle formed by the plane corresponding to the audio and video detection area and the extension line of the center point of the target person's forehead is determined as the head posture angle value, and the direction axis of the four directions in step 203 is used to determine the head direction of the target person.

[0055] Specifically, (1) the terminal maps the head posture angle values ​​of the target person in each frame of the face video to a Cartesian coordinate system, and connects the head posture coordinate points corresponding to the head posture angle values ​​in each frame of the video to generate head action line segments of the target person. The horizontal coordinate of the head posture coordinate points is used to indicate the video frame, and the vertical coordinate is used to indicate the head posture angle value; (2) the terminal performs template matching on the head action line segments through a preset head posture detection model; (3) if the matching distance between any line segment in the head action line segment and the preset head action curve template is greater than or equal to the preset head action matching distance, the terminal determines that the head action of the target person conforms to the preset head action. The preset head action includes rapid head rotation, head turning to the left, and head turning to the right; (4) if the matching distance between each line segment in the head action line segment and the preset head action curve template is less than the preset head action matching distance, the terminal determines that the head action of the target person does not conform to the preset head action.

[0056] For example, the terminal maps the head pose angle values ​​of the target person in each frame of the face video to a Cartesian coordinate system. The multiple right-direction head pose angle values ​​are 5 degrees, 15 degrees, 25 degrees, 15 degrees, 5 degrees, and 0 degrees. The terminal then connects the head pose angle values ​​corresponding to the head pose coordinate points in each frame to generate a line segment representing the target person's head movement. The horizontal coordinate of each head pose coordinate point indicates the video frame, and the vertical coordinate indicates the head pose angle value. The multiple head pose coordinate points are: (1, 5), (2, 15), (3, 25), and (4, 15). (5, 5) and (6, 0); The terminal performs template matching on the head action line segments using a preset head posture detection model; If the matching distance between any line segment in the head action line segment and the preset head action curve template is greater than or equal to the preset head action matching distance, the terminal determines that the target person's head action conforms to the preset head action, which includes rapid head rotation, head turning to the left, and head turning to the right; If the matching distance between each line segment in the head action line segment and the preset head action curve template is less than the preset head action matching distance, the terminal determines that the target person's head action does not conform to the preset head action.

[0057] 205. Perform hand motion recognition based on facial video to obtain hand motion recognition results;

[0058] Specifically, (1) the terminal generates a bounding box of the target person's face region based on the face video; (2) the terminal performs hand detection on the face video; (3) if there is a hand in the face video, the terminal generates a bounding box of the hand corresponding to the hand; (4) the terminal calculates the intersection value between the bounding box of the face region and the bounding box of the hand, the intersection value is used to indicate the ratio of the area of ​​the overlapping area between the bounding box of the face region and the bounding box of the hand to the total area of ​​the bounding box of the face region and the bounding box of the hand; (5) if the intersection value is greater than or equal to a preset value, the terminal determines that the target person's hand movements occlude the target person's face; (6) if the intersection value is less than the preset value, the terminal determines that the target person's hand movements do not occlude the target person's face.

[0059] For example, based on the audio and video detection area corresponding to the target person, the lower left corner of the audio and video detection area is determined as the origin, and a Cartesian coordinate system is established. The terminal generates a bounding box of the target person's face based on the face video. Based on the Cartesian coordinate system, the coordinate information corresponding to the bounding box of the face is generated. The terminal performs hand detection on the face video. If a hand is present in the face video, the terminal generates a bounding box of the hand. Based on the Cartesian coordinate system, the coordinate information corresponding to the bounding box of the hand is generated. Based on the coordinate information corresponding to the bounding box of the face and the bounding box of the hand, the terminal calculates the intersection value between the bounding box of the face and the bounding box of the hand. The intersection value is used to indicate the ratio of the area of ​​the overlapping area between the bounding box of the face and the bounding box of the hand to the total area of ​​the bounding box of the face and the bounding box of the hand. If the intersection value is greater than or equal to a preset value, the terminal determines that the target person's hand movements occlude the target person's face. If the intersection value is less than the preset value, the terminal determines that the target person's hand movements do not occlude the target person's face.

[0060] It should be noted that steps 203, 204, and 205 are executed simultaneously.

[0061] 206. If the eye movements match the preset eye movements, the head movements match the preset head movements, or the hand movements obscure the face of the target person, then it is determined that the target person has committed fraud during the face-to-face interview. Among them, the preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head turning, turning the head to the left, and turning the head to the right.

[0062] For example, if the target person's eye movements are characterized by quick glances, head movements are characterized by turning their head to the left, or hand movements obscure the target person's face, then it is determined that the target person has engaged in fraudulent behavior during the face-to-face interview.

[0063] 207. Generate a reminder message corresponding to the face-to-face interview fraud behavior and send the reminder message to the face-to-face interview reminder terminal.

[0064] The reminder message can be a voice message, a text message, or a light message. For example, if the terminal generates a reminder message for fraudulent behavior during the face-to-face review, the reminder message can be a light message indicating "screen flashing". The terminal sends the light message to the face-to-face review reminder terminal, which then controls the screen to flash, thereby alerting the approval personnel to the possibility of fraudulent behavior by the target person.

[0065] Optionally, steps 202 to 206 can be replaced with the following steps:

[0066] (1) If the audio of the approver is present in the audio and video data, the terminal will obtain the face video of the target person when the audio of the approver ends; (2) The terminal will perform color detection on the ear of the target person based on the face video; (3) If the color of the ear matches the preset color, the terminal will determine that the target person has committed fraud in the face review.

[0067] It should be noted that the default color is red. For example, if the audio / video data contains the audio of the approver, the terminal will acquire the target person's facial video when the approver's audio ends. The terminal will then perform color detection on the target person's ears based on the facial video. If the ear color matches red, the terminal will determine that the target person has engaged in fraudulent face-to-face interview behavior.

[0068] Optionally, steps 202 to 206 can also be replaced with the following steps:

[0069] 1) If the audio of the approver is present in the audio and video data, the terminal acquires the face video of the target person when the audio of the approver ends; 2) The terminal identifies the behavior of touching the nose of the target person based on the face video; Step 2) includes: (1) The terminal generates the nose position box of the target person based on the face video; (2) The terminal performs hand detection on the face video; (3) If there is a hand in the face video, the terminal generates the hand position box corresponding to the hand; (4) The terminal determines whether there is an overlapping area between the nose position box and the hand position box; (5) If there is an overlapping area, the terminal determines that the target person has touched the nose; (6) If there is no overlapping area, the terminal determines that the target person has not touched the nose. 3) If the target person touches the nose, the terminal determines that the target person has committed fraud during the face review.

[0070] For example, based on the audio and video detection area corresponding to the target person, the lower left corner of the audio and video detection area is determined as the origin, and a Cartesian coordinate system is established. If the audio of the approver is present in the audio and video data, the terminal acquires the target person's facial video when the approver's audio ends. The terminal generates a nose bounding box for the target person based on the facial video and generates the coordinate information corresponding to the nose bounding box based on the Cartesian coordinate system. The terminal performs hand detection on the facial video. If a hand is present in the facial video, the terminal generates a hand bounding box and generates the coordinate information corresponding to the hand bounding box based on the Cartesian coordinate system. Based on the coordinate information corresponding to the nose bounding box and the hand bounding box, the terminal determines whether there is an overlapping area between the nose bounding box and the hand bounding box. If there is an overlapping area, the terminal determines that the target person has touched their nose. If there is no overlapping area, the terminal determines that the target person has not touched their nose. If the target person has touched their nose, the terminal determines that the target person has engaged in face-to-face verification fraud.

[0071] In this embodiment of the invention, when the approver and the target are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is performed based on the audio and video data. If the audio and video data contains the approver's audio, then when the approver's audio ends, multiple actions of the target are identified, including eye movements, head movements, and hand movements. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, then it is determined that the target has engaged in fraudulent behavior during the face-to-face interview. The preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. A reminder message corresponding to the fraudulent behavior during the face-to-face interview is generated and sent to the face-to-face interview reminder terminal, thereby improving the accuracy of fraud identification in video interviews.

[0072] The video interview assistance method in the embodiments of the present invention has been described above. The video interview assistance device in the embodiments of the present invention will be described below. Please refer to [link / reference]. Figure 3 One embodiment of the video interview assistance device in this invention includes:

[0073] The audio detection module 301 is used to acquire corresponding audio and video data and perform audio detection based on the audio and video data when the approver and the target personnel are in the preset audio and video detection area.

[0074] The action recognition module 302 is used to recognize multiple actions of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data. The multiple actions include eye movements, head movements and hand movements.

[0075] The first determining module 303 is used to determine that the target person has engaged in face-to-face fraud if the eye movement matches the preset eye movement, the head movement matches the preset head movement, or the hand movement obscures the face of the target person. The preset eye movements include slow eye glance, fast eye glance, and eye trembling, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right.

[0076] The information sending module 304 is used to generate reminder information corresponding to face-to-face audit fraud and send the reminder information to the face-to-face audit reminder terminal.

[0077] In this embodiment of the invention, when the approver and the target are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is performed based on the audio and video data. If the audio and video data contains the approver's audio, then when the approver's audio ends, multiple actions of the target are identified, including eye movements, head movements, and hand movements. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, then it is determined that the target has engaged in fraudulent behavior during the face-to-face interview. The preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. A reminder message corresponding to the fraudulent behavior during the face-to-face interview is generated and sent to the face-to-face interview reminder terminal, thereby improving the accuracy of fraud identification in video interviews.

[0078] Please see Figure 4 Another embodiment of the video interview assistance device in this invention includes:

[0079] The audio detection module 301 is used to acquire corresponding audio and video data and perform audio detection based on the audio and video data when the approver and the target personnel are in the preset audio and video detection area.

[0080] The action recognition module 302 is used to recognize multiple actions of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data. The multiple actions include eye movements, head movements and hand movements.

[0081] The first determining module 303 is used to determine that the target person has engaged in face-to-face fraud if the eye movement matches the preset eye movement, the head movement matches the preset head movement, or the hand movement obscures the face of the target person. The preset eye movements include slow eye glance, fast eye glance, and eye trembling, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right.

[0082] The information sending module 304 is used to generate reminder information corresponding to face-to-face audit fraud and send the reminder information to the face-to-face audit reminder terminal.

[0083] Optionally, the action recognition module 302 includes:

[0084] The acquisition unit 3021 is used to acquire the face video of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data.

[0085] The gaze action recognition unit 3022 is used to perform gaze action recognition based on face video and obtain gaze action recognition results.

[0086] The head action recognition unit 3023 is used to perform head action recognition based on the face video and obtain the head action recognition result.

[0087] The hand motion recognition unit 3024 is used to perform hand motion recognition based on face video to obtain hand motion recognition results.

[0088] Optionally, the gaze movement recognition unit 3022 is specifically used for:

[0089] Map the angle value of the gaze point of the target person in each frame of the face video to a Cartesian coordinate system, and connect the gaze coordinate points corresponding to the gaze angle value in each frame of the video to generate the gaze action line segment of the target person. The horizontal coordinate of the gaze coordinate point is used to indicate the video frame, and the vertical coordinate is used to indicate the gaze angle value.

[0090] The preset gaze point detection model is invoked to perform template matching on the gaze action line segment;

[0091] If the matching distance between any line segment in the eye movement line segment and the preset eye movement curve template is greater than or equal to the preset eye movement matching distance, then it is determined that the target person's eye movement conforms to the preset eye movement. The preset eye movement includes slow eye glance, fast eye glance, and eye trembling.

[0092] If the matching distance between each line segment in the eye movement line segment and the preset eye movement curve template is less than the preset eye movement matching distance, then it is determined that the target person's eye movement does not conform to the preset eye movement.

[0093] Optionally, the head motion recognition unit 3023 is specifically used for:

[0094] Map the head pose angle values ​​of the target person in each frame of the face video to a Cartesian coordinate system, and connect the head pose coordinate points corresponding to the head pose angle values ​​in each frame of the video to generate head action line segments of the target person. The horizontal coordinate of the head pose coordinate points is used to indicate the video frame, and the vertical coordinate is used to indicate the head pose angle value.

[0095] Template matching of head movement segments is performed using a pre-defined head pose detection model;

[0096] If the matching distance between any line segment in the head movement line segment and the preset head movement curve template is greater than or equal to the preset head movement matching distance, then it is determined that the head movement of the target person conforms to the preset head movement. The preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right.

[0097] If the matching distance between each line segment in the head movement line segment and the preset head movement curve template is less than the preset head movement matching distance, then it is determined that the target person's head movement does not conform to the preset head movement.

[0098] Optionally, the hand motion recognition unit 3024 is specifically used for:

[0099] Generate a bounding box containing the facial region of the target person based on the facial video;

[0100] Hand detection in facial videos;

[0101] If a hand is present in the face video, a bounding box corresponding to the hand's location is generated.

[0102] Calculate the intersection value between the face region bounding box and the hand bounding box. The intersection value is used to indicate the ratio of the area of ​​the overlapping area between the face region bounding box and the hand bounding box to the total area of ​​the face region bounding box and the hand bounding box.

[0103] If the intersection value is greater than or equal to the preset value, it is determined that the target person's hand movements are obscuring the target person's face.

[0104] If the intersection value is less than the preset value, it is determined that the target person's hand movements do not obscure the target person's face.

[0105] Optionally, the audio detection module 301 is specifically used for:

[0106] When the approver and the target personnel are in the preset audio and video detection area, the corresponding audio and video data is acquired;

[0107] The audio data is extracted from the audio and video data to obtain the audio data.

[0108] Voiceprint features are extracted from audio data to obtain a voiceprint feature sequence;

[0109] If the voiceprint feature sequence matches the preset voiceprint feature sequence of the approver, then it is determined that the audio of the approver exists in the audio and video data;

[0110] If the voiceprint feature sequence does not match the preset voiceprint feature sequence of the approver, it is determined that there is no audio of the approver in the audio and video data.

[0111] Optionally, video interview assistance devices may also include:

[0112] The acquisition module 305 is used to acquire the face video of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data.

[0113] Color detection module 306 is used to perform color detection on the ears of a target person based on a facial video.

[0114] The second determination module 307 is used to determine that the target person has engaged in fraudulent face-to-face interview behavior if the color of the ear matches the preset color.

[0115] In this embodiment of the invention, when the approver and the target are in a preset audio and video detection area, corresponding audio and video data are acquired, and audio detection is performed based on the audio and video data. If the audio and video data contains the approver's audio, then when the approver's audio ends, multiple actions of the target are identified, including eye movements, head movements, and hand movements. If the eye movements match preset eye movements, the head movements match preset head movements, or the hand movements obscure the target's face, then it is determined that the target has engaged in fraudulent behavior during the face-to-face interview. The preset eye movements include slow glances, fast glances, and trembling eyes, and the preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. A reminder message corresponding to the fraudulent behavior during the face-to-face interview is generated and sent to the face-to-face interview reminder terminal, thereby improving the accuracy of fraud identification in video interviews.

[0116] above Figure 3 and Figure 4 The video interview assistance device in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The video interview assistance device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0117] Figure 5This is a schematic diagram of a video interview assistance device 500 provided in an embodiment of the present invention. The video interview assistance device 500 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 510 (e.g., one or more processors) and a memory 520, and one or more storage media 530 (e.g., one or more mass storage devices) for storing application programs 533 or data 532. The memory 520 and storage media 530 can be temporary or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the video interview assistance device 500. Furthermore, the processor 510 may be configured to communicate with the storage media 530 and execute the series of instruction operations in the storage media 530 on the video interview assistance device 500.

[0118] The video interview support device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 5 The structure of the video interview assistive device shown does not constitute a limitation on the video interview assistive device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0119] The present invention also provides a video interview assistance device, wherein the computer device includes a memory and a processor, the memory storing computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the video interview assistance method in the above embodiments.

[0120] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the video interview assistance method.

[0121] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0122] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0123] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video interview assistance method, characterized in that, The video interview assistance method includes: When the approver and the target personnel are in the preset audio and video detection area, the corresponding audio and video data is acquired, and audio detection is performed based on the audio and video data; If the audio of the approver is present in the audio and video data, then when the audio of the approver ends, multiple actions of the target person are identified, including eye movements, head movements, and hand movements. If the eye movement matches a preset eye movement, the head movement matches a preset head movement, or the hand movement obstructs the face of the target person, then it is determined that the target person has engaged in face-to-face fraud. The preset eye movements include slow glances, fast glances, and trembling eyes. The preset head movements include rapid head rotation, head rotation to the left, and head rotation to the right. Generate a reminder message corresponding to the face-to-face interview fraud, and send the reminder message to the face-to-face interview reminder terminal; If the audio of the approver is present in the audio / video data, then when the audio of the approver ends, multiple actions of the target person are identified. These actions include eye movements, head movements, and hand movements. Specifically: if the audio of the approver is present in the audio / video data, then when the audio of the approver ends, a facial video of the target person is acquired; eye movements are identified based on the facial video to obtain an eye movement recognition result; the angle value of the target person's gaze in each frame of the facial video is mapped to a Cartesian coordinate system, and the eye coordinate points corresponding to the angle values ​​in each frame are connected to generate a line segment representing the target person's eye movements. The horizontal coordinate of the eye coordinate point is used to indicate the video frame. The vertical axis indicates the angle value of the gaze point. A preset gaze point detection model is invoked to perform template matching on the gaze action line segments. If the matching distance between any line segment of the gaze action line segment and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, it is determined that the gaze action of the target person conforms to the preset gaze action. The preset gaze action includes slow gaze, fast gaze, and gaze jitter. If the matching distance between each line segment of the gaze action line segment and the preset gaze action curve template is less than the preset gaze action matching distance, it is determined that the gaze action of the target person does not conform to the preset gaze action. Head action recognition is performed based on the face video to obtain head action recognition results. Hand action recognition is performed based on the face video to obtain hand action recognition results.

2. The video interview assistance method according to claim 1, characterized in that, The step of performing head motion recognition based on the facial video to obtain head motion recognition results includes: The head posture angle value of the target person in each frame of the face video is mapped to a Cartesian coordinate system, and the head posture coordinate points corresponding to the head posture angle value in each frame of the video are connected to generate the head action line segment of the target person. The horizontal coordinate of the head posture coordinate point is used to indicate the video frame, and the vertical coordinate is used to indicate the head posture angle value. Template matching of the head movement segments is performed using a pre-set head pose detection model; If the matching distance between any line segment in the head movement line segment and the preset head movement curve template is greater than or equal to the preset head movement matching distance, then it is determined that the head movement of the target person conforms to the preset head movement, which includes rapid head rotation, head rotation to the left, and head rotation to the right. If the matching distance between each line segment in the head movement line segment and the preset head movement curve template is less than the preset head movement matching distance, then it is determined that the head movement of the target person does not conform to the preset head movement.

3. The video interview assistance method according to claim 1, characterized in that, The step of performing hand motion recognition based on the facial video to obtain hand motion recognition results includes: Generate a bounding box of the target person's face region based on the facial video; Hand detection is performed on the face video; If a hand is present in the face video, a bounding box corresponding to the hand is generated; Calculate the intersection value between the face region location box and the hand location box. The intersection value is used to indicate the ratio of the area of ​​the overlapping area between the face region location box and the hand location box to the total area of ​​the face region location box and the hand location box. If the intersection value is greater than or equal to a preset value, it is determined that the target person's hand movements are obscuring the target person's face. If the intersection value is less than a preset value, it is determined that the target person's hand movements do not obstruct the target person's face.

4. The video interview assistance method according to claim 1, characterized in that, When the approver and the target personnel are in the preset audio and video detection area, the corresponding audio and video data is acquired, and audio detection is performed based on the audio and video data, including: When the approver and the target personnel are in the preset audio and video detection area, the corresponding audio and video data is acquired; The audio data is extracted from the audio and video data to obtain the audio data. Voiceprint features are extracted from the audio data to obtain a voiceprint feature sequence; If the voiceprint feature sequence matches the preset voiceprint feature sequence of the approver, then it is determined that the audio of the approver exists in the audio and video data; If the voiceprint feature sequence does not match the preset voiceprint feature sequence of the approver, it is determined that the audio of the approver does not exist in the audio and video data.

5. The video interview assistance method according to any one of claims 1-4, characterized in that, Before generating a reminder message corresponding to the face-to-face audit fraud behavior and sending the reminder message to the face-to-face audit reminder terminal, the method further includes: When the approver and the target person are in the preset audio and video detection area, acquiring the corresponding audio and video data, and performing audio detection based on the audio and video data. If the audio of the approver is present in the audio and video data, then the facial video of the target person is obtained when the audio of the approver ends. Color detection of the target person's ears is performed based on the facial video; If the color of the ear matches a preset color, it is determined that the target person has engaged in fraudulent face-to-face interview behavior.

6. A video interview assistance device, characterized in that, The video interview assistance device includes: The audio detection module is used to acquire corresponding audio and video data when the approver and the target person are in a preset audio and video detection area, and to perform audio detection based on the audio and video data; The action recognition module is used to identify multiple actions of the target person when the audio of the approver ends if the audio of the approver is present in the audio and video data. The multiple actions include eye movements, head movements and hand movements. The first determining module is used to determine that the target person has engaged in face-to-face fraud if the eye movement matches a preset eye movement, the head movement matches a preset head movement, or the hand movement obscures the face of the target person. The preset eye movement includes slow eye glance, fast eye glance, and eye trembling, and the preset head movement includes rapid head rotation, head rotation to the left, and head rotation to the right. The information sending module is used to generate a reminder message corresponding to the face-to-face fraud behavior and send the reminder message to the face-to-face reminder terminal; The action recognition module is specifically used for: if the audio of the approver is present in the audio-visual data, then acquiring the face video of the target person when the audio of the approver ends; performing gaze action recognition based on the face video to obtain gaze action recognition results; mapping the gaze angle value of the target person in each frame of the face video to a Cartesian coordinate system, and connecting the gaze coordinate points corresponding to the gaze angle values ​​in each frame of the video to generate the gaze action line segment of the target person, where the horizontal coordinate of the gaze coordinate point is used to indicate the video frame, the vertical coordinate is used to indicate the gaze angle value, and calling a preset gaze point detection model to perform gaze action recognition. Template matching is performed on line segments. If the matching distance between any line segment of the gaze action line segment and the preset gaze action curve template is greater than or equal to the preset gaze action matching distance, then the gaze action of the target person is determined to conform to the preset gaze action. The preset gaze action includes slow glance, fast glance, and gaze tremor. If the matching distance between each line segment of the gaze action line segment and the preset gaze action curve template is less than the preset gaze action matching distance, then the gaze action of the target person is determined to not conform to the preset gaze action. Head action recognition is performed based on the face video to obtain head action recognition results. Hand action recognition is performed based on the face video to obtain hand action recognition results.

7. A video interview auxiliary device, characterized in that, The video interview auxiliary device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the video interview assistance device to perform the video interview assistance method as described in any one of claims 1-5.

8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the video interview assistance method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, computer equipment and storage medium

    CN112820297A

  • Abnormal behavior detection method and device, computer equipment and storage medium

    CN113435362A