Methods, systems, devices, and storage media for detecting flick action

By detecting facial key points and gaze angles in video streams, and combining sliding time windows and gaze tracking models, the problem of real-time identification of cheating actions in video review was solved, achieving accurate video review and elimination of misjudgments.

CN116189287BActive Publication Date: 2026-05-05CREDIT CARD CENT OF GUANGFA BANK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CREDIT CARD CENT OF GUANGFA BANK CO LTD
Filing Date
2022-12-27
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing video review processes struggle to accurately identify instances of cheating in real time, increasing the difficulty and cost of the review process.

Method used

By detecting facial key points and gaze angles in the video stream, combined with a sliding time window and gaze tracking model, it can determine whether the user is speaking and whether there is any cheating gaze. By integrating facial key points and gaze tracking, false judgments can be eliminated.

Benefits of technology

It enables real-time and accurate video review, reduces misjudgments, improves recognition accuracy, and provides important reference for subsequent approvals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189287B_ABST
    Figure CN116189287B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, device, and storage medium for detecting whether a person in a video is making cheating gestures. The method includes: detecting whether a user's face exists in the current video stream; when a face exists in the current video stream, acquiring facial key point information; determining whether the user is speaking based on the acquired facial key point information; detecting the user's gaze angle in the current video stream; determining whether the user is making cheating gestures based on the detected user's gaze angle; and confirming that the user is making cheating gestures when it is determined that the user is speaking and making cheating gestures. This method integrates facial key point and user gaze tracking to identify whether a user is making cheating gestures in a video in real time, providing important reference for subsequent video review and solving the problem of misjudgment caused by user gaze or head posture drift.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video detection, and more specifically, to methods, systems, devices, and storage media for detecting cheating. Background Technology

[0002] Credit card activation verification is typically conducted via video. This requires the bank's verification system to process a large volume of videos. During video verification, individuals may be seen looking at cheat sheets. If such behavior is confirmed, the activation application should be carefully reviewed, and relevant records should be kept, as this can provide valuable information for the overall verification process. Currently, video verification relies heavily on human agents using their experience and visual inspection. However, the skill levels of human agents vary, leading to potential misjudgments or omissions in practice.

[0003] Secondly, the aforementioned video review scenarios require real-time review, as it is impossible to identify cheating actions in the video afterward. Real-time review increases the difficulty of the review process. Furthermore, based on the requirement for model accuracy, directly using footage containing cheating actions for modeling also places high demands on the data samples. Summary of the Invention

[0004] The present invention aims to overcome at least one of the defects of the prior art and provides a method, system, device and storage medium for detecting cheating actions, in order to solve the problem that the current real-time video review cannot automatically identify whether people in the video are cheating, which increases the difficulty and cost of review.

[0005] The technical solution adopted in this invention includes:

[0006] In a first aspect, the present invention provides a method for detecting whether a person in a video is making cheating gestures, comprising: detecting whether a user's face exists in the current video stream; when a face exists in the current video stream, acquiring facial key point information; determining whether the user is speaking based on the acquired facial key point information; detecting the user's gaze angle in the current video stream; determining whether the user is making cheating gestures based on the detected user's gaze angle; and determining that the user is making cheating gestures when it is determined that the user is speaking and making cheating gestures.

[0007] The method provided by this invention detects faces and acquires facial key point information in real time within a video stream. Based on this information, it analyzes whether the user is speaking; if so, it indicates the user is answering questions for video review. Simultaneously, it detects the user's gaze angle within the video stream to determine the area of ​​focus. This allows it to determine if the user is not looking at the camera area but rather at the area where cheat sheets are located. If so, it confirms the user is using cheat sheets. If the user is already determined to be answering questions for video review and is using cheat sheets, then the presence of cheat sheets can be definitively confirmed. This result provides crucial information for real-time video review, allowing for real-time approval of user applications based on the detection of cheat sheet activity, achieving real-time and accurate video review. This method integrates facial key point detection and user gaze tracking, resolving misjudgments caused by user gaze or head posture drift.

[0008] Furthermore, the facial key point information includes at least facial mouth key point information; determining whether the user is speaking based on the acquired facial key point information specifically includes: determining the user's mouth width and mouth height based on the acquired facial mouth key point information; determining the user's mouth opening ratio based on the user's mouth width and mouth height; and determining whether the user is speaking based on the determined user's mouth opening ratio.

[0009] Based on the mouth opening ratio determined by the user's mouth width and mouth height, and unaffected by head posture drift, it can accurately determine whether the user is speaking, thereby determining whether the user is answering the video review questions.

[0010] Furthermore, the mouth opening ratio of the user is determined based on the user's mouth width and mouth height. The determination of whether the user is speaking is then based on the determined mouth opening ratio. Specifically, this includes: determining a preset sliding time window covering several frames of video stream, where each frame consists of the current video stream frame and several frames preceding and following it; determining the user's mouth opening ratio in each frame based on the user's mouth width and mouth height within the sliding time window; determining the standard deviation of the user's mouth opening ratio across the frames covered by the sliding time window based on the user's mouth opening ratio in each frame; and determining whether the standard deviation of the user's mouth opening ratio is greater than a preset threshold. If so, it is determined that the user is speaking.

[0011] To reduce the possibility of false positives, such as avoiding situations where a user opens their mouth but does not speak in a certain frame of the video stream, a preset sliding time window is used to detect speech in the current video stream and several frames before and after it. Based on the standard deviation of the user's mouth opening ratio in all frames covered by the window, it is determined whether the user is speaking.

[0012] Furthermore, the mouth opening ratio of the user in each frame is determined based on the width and height of the user's mouth in each frame covered by the sliding time window. Specifically, this includes: according to the formula... Determine the user's mouth width and height in each frame to determine the user's mouth opening ratio in each frame; where i is the user's mouth opening ratio, w is the user's mouth width, and h is the user's mouth height; based on the user's mouth opening ratio in each frame, determine the standard deviation of the user's mouth opening ratio across several frames covered by the sliding time window, specifically including: according to the formula... Determine the standard deviation of the user's mouth opening ratio in several frames covered by the sliding time window tw; where s is the standard deviation of the mouth opening ratio, w is the user's mouth width, h is the user's mouth height, and n is... t represents the average mouth opening ratio of the user across the N frames covered by the sliding time window tw; t is the first frame covered by the sliding time window tw; t+N is the last frame covered by the sliding time window tw; i tw The ratio of the user's mouth opening to the number of frames covered by the sliding time window tw.

[0013] Furthermore, the system detects the user's gaze angle in the video stream and determines whether the user is engaging in cheating based on the detected gaze angle. Specifically, this includes: determining the user's horizontal and vertical gaze angles in the current video stream; a first criterion: determining whether the user's horizontal gaze angle is less than a preset minimum horizontal gaze angle or greater than a preset maximum horizontal gaze angle; a second criterion: determining whether the user's vertical gaze angle is less than a preset minimum vertical gaze angle or greater than a preset maximum vertical gaze angle; if the current video stream meets either the first or second criterion, then it determines whether the subsequent frames all meet either the first or second criterion. If so, it is determined that the user is engaging in cheating; otherwise, it is determined that the user is not engaging in cheating.

[0014] The first and second judgment conditions are used to determine whether the user's horizontal or vertical gaze deviation angle exceeds the preset range. If so, it can be preliminarily considered that the user is cheating. In order to avoid misjudgment, it is further determined whether the following frames of the current video stream also meet the condition of exceeding the preset range, that is, whether the user's gaze deviates from the normal range for a period of time. If so, it can be finally determined that the user is cheating.

[0015] Furthermore, determining the horizontal and vertical offset angles of the user's gaze in the current video stream involves: extracting the face and eye images from the current video stream and inputting them into the gaze tracking model, so that the gaze tracking model can extract feature vectors from the face and eye images respectively, encode the extracted feature vectors, compress the encoded feature vectors, and obtain the horizontal and vertical offset angles of the user's gaze in the current video stream.

[0016] Furthermore, the gaze tracking model includes a sequentially connected backbone feature extraction network, a secondary feature extraction network, a feature encoding module, and several fully connected layers. The backbone feature extraction network consists of a focus module and a CSP network, used to extract feature vectors from the image. The secondary feature extraction network consists of an FPN network, a PAN network, and fully connected layers, used to further process the feature vectors extracted by the backbone feature extraction network. The feature encoding module is a Transformer encoder module, used to encode the feature vectors processed by the secondary feature extraction network. Several fully connected layers are used to compress the output of the feature encoding module, ultimately outputting two-dimensional features to represent the horizontal and vertical gaze offset angles.

[0017] The eye-tracking model provided by this invention can obtain the angle at which the user's gaze falls on the surface of interest, determine the horizontal and vertical deviation angles of the user's gaze in real time, and subsequently determine whether the user's gaze point falls outside the normal area based on the gaze deviation angle, as well as whether the gaze duration exceeds a preset value, and obtain the result of whether the user has made any cheating actions.

[0018] Secondly, the present invention provides a system for detecting whether a person in a video is making cheating gestures, comprising: a face detection module for detecting whether a user's face exists in the current video stream; when a face exists in the current video stream, acquiring facial key point information; a user speaking detection module for determining whether the user is speaking based on the acquired facial key point information; a gaze angle detection module for detecting the user's gaze angle in the current video stream; determining whether the user is making cheating gestures based on the detected user's gaze angle; and a cheating gesture detection module for determining that the user is making cheating gestures when it is determined that the user is speaking and making cheating gestures.

[0019] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-mentioned method for detecting whether a person in a video is making cheating gestures.

[0020] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for detecting whether a person in a video is making cheating gestures.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] The method provided by this invention detects faces and acquires facial key point information in real time within a video stream. Based on the facial mouth key point information and using a sliding time window, it accurately determines the user's mouth opening ratio and standard deviation in each frame. The mouth opening ratio and standard deviation are then used to analyze whether the user is speaking. Secondly, an eye-tracking model is used to detect the user's gaze angle in the video stream in real time. This gaze angle determines the user's gaze area, thus judging whether the user's gaze is within the normal area. If not, it can be determined that the user is engaging in cheating. If it has been determined that the user is speaking and answering questions for video review, and cheating gaze is observed, then cheating is definitively confirmed. This result provides an important reference for real-time video review, allowing for real-time approval of user applications based on the detection results of cheating, achieving real-time and accurate video review. This method integrates facial key points and user eye-tracking, eliminating misjudgments caused by user gaze or head posture drift and improving recognition accuracy. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating steps S110 to S160 of the method provided in Embodiment 1 of the present invention.

[0024] Figure 2 This is a schematic diagram of the composition of facial key points in Embodiment 1 of the present invention.

[0025] Figure 3 This is a flowchart illustrating the method steps S131 to S132 provided in Embodiment 1 of the present invention.

[0026] Figure 4 This is a flowchart illustrating steps S1311 to S1323 of the method provided in Embodiment 1 of the present invention.

[0027] Figure 5 This is a schematic diagram of the eye-tracking model in Embodiment 1 of the present invention.

[0028] Figure 6 This is a flowchart illustrating steps S151 to S160 of the method provided in Embodiment 1 of the present invention.

[0029] Figure 7 This is a schematic diagram of the eye-tracking safety area in Embodiment 1 of the present invention.

[0030] Figure 8 This is a schematic diagram of the system module composition provided in Embodiment 2 of the present invention. Detailed Implementation

[0031] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention. To better illustrate the following embodiments, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions; it is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0032] Example 1

[0033] This embodiment provides a method for detecting whether a person in a video is making cheating gestures, applicable to real-time video review, especially for videos requiring users to answer questions in real time. This method combines eye tracking and facial landmark recognition to detect cheating gestures in the video stream in real time, thus providing important reference for the overall video review process.

[0034] like Figure 1 As shown, the method includes the following steps:

[0035] S110. Detect whether there is a user's face in the current video stream. If there is a face in the current video stream, execute step S120.

[0036] In a specific implementation, this step utilizes the get frontal facedetector algorithm from the existing Dlib library to detect face bounding boxes in the current video stream. When a face bounding box is detected, it indicates that the user's face exists in the video stream.

[0037] To detect cheating in video tasks in real time, this step and subsequent steps will be repeated at certain time intervals.

[0038] S120. Obtain facial key point information;

[0039] In a specific implementation, this step utilizes the shape predictor algorithm from the existing Dlib library to detect facial landmark information within the current video stream. This facial landmark information includes at least the key points of the mouth, and may also include key points of other facial features, such as eyebrows, eyes, nose, and facial contours, depending on actual needs. Facial landmark information generally refers to the coordinate information of these key points. For example,... Figure 2 The image shows commonly used facial landmarks. Facial landmarks, particularly those related to the mouth, generally refer to... Figure 2 Key points 51-68.

[0040] S130. Determine whether the user is speaking based on the obtained facial key point information;

[0041] In this step, the specific determination of whether the user is speaking is based on the acquired facial mouth key point information. Based on this, such as... Figure 3 As shown, this step specifically includes the following steps:

[0042] S131. Determine the user's mouth width w and mouth height h based on the obtained facial mouth key point information;

[0043] like Figure 2 As shown, the user's mouth height h is the distance between key points 52 and 58, and the user's mouth width w is the distance between key points 49 and 55.

[0044] S132. Determine the user's mouth opening ratio i based on the user's mouth width w and mouth height h, and determine whether the user is speaking based on the determined user's mouth opening ratio i.

[0045] The formula for calculating the mouth opening ratio i is: The larger the value of i, the closer the user is to a closed mouth state, meaning the user is not speaking; the smaller the value of i, the closer the user is to an open mouth state, meaning the user is speaking. The mouth opening ratio i changes continuously with the frequency of the user's speech.

[0046] To improve the accuracy of identification and judgment and avoid misjudgments, such as Figure 4 As shown, steps S131 to S132 specifically include the following steps:

[0047] S1311. Determine the number of video stream frames covered by the preset sliding time window tw;

[0048] The sliding time window tw has a size of N, which means it can cover N frames of video stream. Generally speaking, N frames of video stream consist of the current video stream and several frames before and after it. The first frame covered by the sliding time window tw can be denoted as frame t, and the last frame can be denoted as frame t+N.

[0049] S1312. Obtain the user's mouth width w and mouth height h in each frame of the video stream determined in step S1311;

[0050] S1321. Determine the user's mouth opening ratio i in each frame based on the user's mouth width w and mouth height h in each frame covered by the sliding time window tw. tw ;

[0051] In this step, the user's mouth opening ratio i in each frame tw All based on the formula Calculated.

[0052] S1322, Based on the user's mouth opening ratio i in each frame tw Determine the standard deviation s of the user's mouth opening ratio in the N frames covered by the sliding time window tw;

[0053] In this step, the formula for calculating the standard deviation s of the mouth opening ratio is: in, t is the average mouth opening ratio of the user in the N frames covered by the sliding time window tw; t is the first frame covered by the sliding time window tw; t+N is the last frame covered by the sliding time window tw. The calculation formula is:

[0054] Using the standard deviation of the mouth opening ratio (s) as the data basis for subsequent judgments on whether a user is speaking can more evenly reflect changes in the user's mouth, thus avoiding misjudgments caused by special circumstances.

[0055] S1323. Determine whether the standard deviation s of the user's mouth opening ratio is greater than the preset speech detection threshold k. If yes, determine that the user is speaking; otherwise, determine that the user is not speaking.

[0056] S140. Detect the user's viewing angle in the current video stream;

[0057] In this step, the user's line of sight angle is determined by the horizontal offset angle x. ° The angle y between the line of sight and the vertical offset ° Composition, horizontal offset angle x of the line of sight ° The angle y between the line of sight and the vertical offset ° The combination can reflect the area that the user is focusing on.

[0058] Based on this, the specific execution process of step S140 is as follows: by inputting the current video stream image into the eye-tracking model, the eye-tracking model outputs the user's horizontal eye-tracking offset angle x. ° And its vertical offset angle y ° ;

[0059] In this step, before inputting the gaze tracking model, the video stream needs to be preprocessed. Then, the face, left eye, and right eye images are extracted from the preprocessed video stream. Finally, the three extracted images are input into the gaze tracking model.

[0060] The eye-tracking model, upon receiving three input frames, extracts feature vectors from each frame, encodes these feature vectors, compresses them, and finally obtains the user's horizontal eye-tracking offset angle x. °And its vertical offset angle y ° .

[0061] Specifically, such as Figure 5 As shown, the gaze tracking model provided in this embodiment is a lightweight gaze tracking model based on the iTracker model (tiny-iTracker). This model includes a sequentially connected backbone feature extraction network, a secondary feature extraction network, a feature encoding module, and several fully connected layers.

[0062] The backbone feature extraction network consists of a focus module and a CSP (Cross Partial Network) network, which is used to extract feature vectors from images.

[0063] like Figure 5 As shown, after the above three images are input into the model, the focus module of the backbone feature extraction network performs a focus operation on the input images. The purpose is to reduce the amount of computation and improve the processing speed of the model. The images after focus processing will be input into a 5-layer CSP network. The 5-layer CSP network will extract the feature vector of each input image.

[0064] The secondary feature extraction network consists of FPN (Feature Pyramid Networks), PAN (Path Aggregation Network), and fully connected layers, and is used to further process the feature vectors extracted by the backbone feature extraction network.

[0065] The feature vectors of each input image, processed by the 5-layer CSP network, are sequentially input into the FPN and PAN networks for further processing. After passing through fully connected layers, they yield 2048-dimensional facial gaze feature vectors, 2048-dimensional gaze feature vectors for the left eye, and 2048-dimensional gaze feature vectors for the right eye.

[0066] The feature encoding module is a Transformer encoder module, which is used to encode the feature vectors processed by the secondary feature extraction network.

[0067] The three 2048-dimensional feature vectors output by the feature extraction network are input into the feature encoding module for encoding. In the feature encoding module, a multi-head attention mechanism is used to encode the features of the face, left eye, and right eye.

[0068] Several fully connected layers are used to compress the output of the feature encoding module, ultimately outputting two-dimensional features to represent the horizontal offset angle x of the gaze. ° The angle y between the line of sight and the vertical offset ° .

[0069] The three 2048-dimensional feature vectors encoded and output by the feature encoding module are input into a fully connected layer, compressing the feature dimension from 2048 to 256. These vectors are then further input into a second fully connected layer, ultimately compressing the feature vectors to 2 dimensions to obtain the final horizontal gaze offset angle x. ° (pitch) and vertical offset angle of the line of sight y ° (yaw). Specifically, there are 3 fully connected layers.

[0070] In a specific implementation, the training method for the gaze tracking model includes the following steps:

[0071] Create an eye-tracking dataset that meets the requirements of production scenarios.

[0072] Using 20,000 labeled gaze tracking images, and dividing them into training, validation, and test data in a 7:2:1 ratio, a gaze tracking model was trained using the training, validation, and test data.

[0073] S150. Determine whether the user is looking at a cheat sheet based on the detected user's line of sight.

[0074] like Figure 6 As shown, the horizontal offset angle x of the gaze is output by the gaze tracking model. ° (pitch) and vertical offset angle of the line of sight y ° The specific process by which (yaw) determines whether a user is using cheating gaze includes the following steps:

[0075] S151. Determine the horizontal deviation angle x of the line of sight obtained in step S140. ° The angle y between the line of sight and the vertical offset ° Does the first or second judgment condition meet? If yes, proceed to step S152; if no, proceed to step S154.

[0076] First judgment condition: Determine the user's horizontal deviation angle x ° Is it less than the preset minimum horizontal line of sight? Or greater than the preset maximum horizontal line of sight

[0077] Second judgment condition: Determine the user's vertical deviation angle y. ° Is it less than the preset minimum vertical line of sight? Or greater than the preset maximum vertical line of sight.

[0078] In this step, such as Figure 7 As shown, if either the first or second judgment condition is met, it indicates that the user's visual focus has deviated from the normal area. There is a possibility that the user is observing the cheat sheet, so it can be preliminarily determined that the user is observing the cheat sheet. If neither of these two conditions is met, it can be determined that the user is not observing the cheat sheet, i.e., there is no cheat sheet activity. In a specific implementation, the normal area... It can be adjusted according to business strategy.

[0079] S152. Determine whether the following frames of the current video stream all meet the first or second judgment condition. If yes, proceed to step S153; otherwise, proceed to step S154.

[0080] This step is to determine whether the user's cheating gaze action has lasted for a certain period of time T. That is, to continuously detect several frames after the video stream. If the subsequent frames all meet the first judgment condition or the second judgment condition, it means that the user has been cheating gaze action for a certain period of time T, and it can be finally determined that the user has cheating gaze action.

[0081] S153. Determine that the user is observing the cheat sheet, and proceed to step S160.

[0082] S154. Determine that the user is not observing any cheat sheet behavior;

[0083] After performing this step, step S110 and subsequent steps can be re-executed after a certain time interval to detect subsequent video stream images.

[0084] S160. Determine whether the user is speaking and whether there is a cheating gaze. If yes, then determine that the user is cheating; if no, then record the user's cheating gaze behavior.

[0085] In this step, if a user is found to be using cheat sheets, an alert will be sent to the human customer service representative who will then conduct a more rigorous review of the user's application. If the user makes a glancing motion while using cheat sheets but does not speak, this behavior will be recorded for future review strategies.

[0086] The method provided in this embodiment utilizes Dlib's get frontal face detector algorithm and shapepredictor algorithm to detect face bounding boxes and acquire facial key point information in real time within the video stream. The user's mouth width and height are obtained from the facial key point information, and a sliding time window (tw) is used to accurately determine the user's mouth opening ratio and standard deviation in each frame. The mouth opening ratio and standard deviation are then used to analyze whether the user is speaking. Secondly, a self-developed gaze tracking model is used to detect the user's horizontal and vertical gaze offset angles in the video stream in real time. The user's gaze area is determined by the gaze offset angle, thereby judging whether the user's gaze focus is within the normal area. This helps determine if the user is engaging in cheating. If it is determined that the user is speaking while answering questions for video review and exhibits cheating gaze behavior, then cheating behavior is definitively confirmed, and an alert is issued to remind human customer service to focus on the user's application review. The method provided in this embodiment integrates facial key point and user gaze tracking to achieve real-time and accurate video review, eliminating misjudgments caused by user gaze or head posture drift, and improving recognition accuracy.

[0087] Example 2

[0088] Based on the same concept as Embodiment 1, this embodiment provides a system for detecting whether a person in a video is making cheating gestures, such as... Figure 8 As shown, the system includes:

[0089] The face detection module 210 is used to detect whether a user's face exists in the current video stream; when a face exists in the current video stream, it acquires facial key point information.

[0090] Facial key point information includes at least the key point information of the mouth area.

[0091] The user speaking detection module 220 is used to determine whether the user is speaking based on the acquired facial key point information.

[0092] The user speech detection module 220 specifically includes:

[0093] The mouth width and height calculation unit 221 determines several frames of video stream images covered by a preset sliding time window, and determines the user's mouth width and mouth height based on the key point information of the human face mouth in each frame.

[0094] A video stream consists of the current video stream frame and several video stream frames before and after it.

[0095] Mouth opening ratio determination unit 222 is used to determine the mouth opening ratio of the user in each frame based on the mouth width and mouth height of the user in each frame covered by the sliding time window; and to determine the standard deviation of the mouth opening ratio of the user in several frames covered by the sliding time window based on the mouth opening ratio of the user in each frame.

[0096] According to the formula Determine the user's mouth width and mouth height in each frame to determine the user's mouth opening ratio in each frame; where i is the user's mouth opening ratio, w is the user's mouth width, and h is the user's mouth height.

[0097] According to the formula Determine the standard deviation of the user's mouth opening ratio across several frames covered by the sliding time window tw; where s is the standard deviation of the mouth opening ratio, w is the user's mouth width, and h is the user's mouth height. The average mouth opening ratio of the user across N frames covered by the sliding time window tw; t is the first frame covered by the sliding time window tw; t+N is the last frame covered by the sliding time window tw; i tw The ratio of the user's mouth opening to the number of frames covered by the sliding time window tw.

[0098] The speaking detection unit 223 is used to determine whether the standard deviation of the user's mouth opening ratio is greater than a preset threshold. If so, it is determined that the user is speaking; otherwise, it is determined that the user is not speaking.

[0099] The gaze angle detection module 230 is used to detect the user's gaze angle in the current video stream; and to determine whether the user is making a cheating gaze based on the detected user's gaze angle.

[0100] The line-of-sight angle detection module 230 specifically includes:

[0101] The image preprocessing unit 231 is used to preprocess the current video stream image and extract the face image, the left eye image, and the right eye image from the current video stream image, and input the three extracted images into the eye tracking model.

[0102] The gaze tracking model 232 includes a sequentially connected backbone feature extraction network, a secondary feature extraction network, a feature encoding module, and several fully connected layers.

[0103] The backbone feature extraction network consists of a focus module and a CSP network, used to extract feature vectors from images;

[0104] The secondary feature extraction network consists of an FPN network, a PAN network, and fully connected layers, and is used to further process the feature vectors extracted by the backbone feature extraction network.

[0105] The feature encoding module is a Transformer encoder module, which is used to encode the feature vectors processed by the secondary feature extraction network;

[0106] Several fully connected layers are used to compress the output of the feature encoding module, and finally output two-dimensional features to represent the horizontal and vertical offset angles of the line of sight.

[0107] The eye tracking module 232 outputs the final horizontal and vertical eye movement offset angles to the cheat sheet gaze detection unit 233.

[0108] The cheat sheet gazing action detection unit 233 is used to determine whether the current video stream image meets the first judgment condition or the second judgment condition. If it does, it determines whether the next few frames of the current video stream image all meet the first judgment condition or the second judgment condition. If so, it determines that the user has a cheat sheet gazing action; otherwise, it determines that the user does not have a cheat sheet gazing action.

[0109] First judgment condition: Determine whether the user's horizontal line of sight offset angle is less than the preset minimum horizontal line of sight, or greater than the preset maximum horizontal line of sight.

[0110] The second judgment condition is to determine whether the user's vertical line of sight offset angle is less than the preset minimum vertical line of sight or greater than the preset maximum vertical line of sight.

[0111] The cheating action detection module 240 is used to determine that the user is cheating when the user is speaking and there is a cheating gaze action, and to issue an alarm to human customer service; when the user is not speaking but there is a cheating gaze action, the user's cheating gaze action behavior is recorded.

[0112] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A method for detecting whether a person in a video is making cheating gestures, characterized in that, include: Detect whether a user's face is present in the current video stream; When a face is present in the current video stream, obtain facial landmark information; Determine whether the user is speaking based on the obtained facial landmark information; Detect the user's viewing angle in the current video stream; Determine whether the user is looking at a cheat sheet based on the angle of their gaze. Once it is determined that the user is speaking and there is a glancing action indicating that they are using a cheat sheet, it is confirmed that the user is using a cheat sheet. The facial key point information includes at least the key point information of the mouth area; Determining whether a user is speaking based on the acquired facial landmark information includes: Based on the obtained facial mouth key point information, determine the user's mouth width and mouth height; The user's mouth opening ratio is determined based on the user's mouth width and mouth height, and whether the user is speaking is determined based on the determined mouth opening ratio. The user's mouth opening ratio is determined based on the width and height of their mouth. Based on this ratio, it is then used to determine whether the user is speaking. This process includes: Determine the number of video stream frames covered by the preset sliding time window. The number of video stream frames consists of the current video stream frame and the video stream frames before and after it. The mouth opening ratio of the user in each frame is determined based on the width and height of the user's mouth in each frame covered by the sliding time window; Based on the user's mouth opening ratio in each frame, determine the standard deviation of the user's mouth opening ratio across several frames covered by the sliding time window. Determine if the standard deviation of the user's mouth opening ratio is greater than a preset threshold; if so, determine that the user is speaking.

2. The method for detecting whether a person in a video is cheating, as described in claim 1, is characterized in that... The mouth opening ratio of the user in each frame is determined based on the width and height of the user's mouth in each frame covered by the sliding time window, specifically including: According to the formula Determine the user's mouth width and mouth height in each frame; determine the user's mouth opening ratio in each frame; where i is the user's mouth opening ratio, w is the user's mouth width, and h is the user's mouth height; Based on the user's mouth opening ratio in each frame, determine the standard deviation of the user's mouth opening ratio across several frames covered by the sliding time window, specifically including: According to the formula Determine the standard deviation of the user's mouth opening ratio across several frames covered by the sliding time window tw; where s is the standard deviation of the mouth opening ratio, w is the user's mouth width, and h is the user's mouth height. t is the average mouth opening ratio of the user in the N frames covered by the sliding time window tw; t is the first frame covered by the sliding time window tw; t+N is the last frame covered by the sliding time window tw. The ratio of the user's mouth opening to the number of frames covered by the sliding time window tw.

3. The method for detecting whether a person in a video is cheating, as described in claim 1 or 2, is characterized in that... Detect the user's gaze angle in the video stream, and determine whether the user is engaging in cheating based on the detected gaze angle. Specifically, this includes: Determine the horizontal and vertical offset angles of the user's gaze in the current video stream. First judgment condition: Determine whether the user's horizontal line of sight offset angle is less than the preset minimum horizontal line of sight, or greater than the preset maximum horizontal line of sight; The second judgment condition is to determine whether the user's vertical line of sight offset angle is less than the preset minimum vertical line of sight or greater than the preset maximum vertical line of sight. If the current video stream frame meets the first or second judgment condition, then it is determined whether the next few frames of the video stream frame all meet the first or second judgment condition. If so, it is determined that the user is performing a cheating look; otherwise, it is determined that the user is not performing a cheating look.

4. The method for detecting whether a person in a video is cheating, as described in claim 3, is characterized in that... Determine the horizontal and vertical offset angles of the user's gaze in the current video stream, specifically including: The face and eyes frames from the current video stream are captured and input into the gaze tracking model. The gaze tracking model then extracts feature vectors from the face and eyes frames, encodes the extracted feature vectors, compresses the encoded feature vectors, and obtains the user's horizontal and vertical gaze offset angles in the current video stream.

5. The method for detecting whether a person in a video is cheating, as described in claim 4, is characterized in that... The gaze tracking model consists of a sequentially connected backbone feature extraction network, a secondary feature extraction network, a feature encoding module, and several fully connected layers. The backbone feature extraction network consists of a focus module and a CSP network, used to extract feature vectors from images; The sub-feature extraction network consists of an FPN network, a PAN network, and a fully connected layer, and is used to further process the feature vectors extracted by the backbone feature extraction network. The feature encoding module is a Transformer encoder module, which is used to encode the feature vector processed by the secondary feature extraction network; Several fully connected layers are used to compress the output of the feature encoding module, and finally output two-dimensional features to represent the horizontal and vertical offset angles of the line of sight.

6. A system for detecting whether a person in a video is making cheating gestures, used to implement the method for detecting whether a person in a video is making cheating gestures as described in claim 1, characterized in that, include: The face detection module is used to detect whether a user's face exists in the current video stream; when a face exists in the current video stream, it acquires facial key point information. The user speech detection module is used to determine whether a user is speaking based on the acquired facial key point information. The gaze angle detection module is used to detect the user's gaze angle in the current video stream; based on the detected user's gaze angle, it determines whether the user is making a cheating gaze action; The cheat sheet detection module is used to determine if a user is cheating when they are speaking and there is a gaze action indicating that they are looking at a cheat sheet.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method for detecting whether a person in a video is cheating, as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for detecting whether a person in a video is cheating, as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Examination cheating behavior identification method, electronic equipment and storage medium

    CN111611865A

  • Sight tracking method and device, computer equipment and storage medium

    CN112749655A

  • Speaking recognition method based on video analysis, system and equipment and medium

    CN113177531A