Anti-cheating method, anti-cheating system and anti-cheating equipment for online examination system, and medium

By collecting and deeply analyzing multimodal data and combining environmental anomaly confidence levels, the cheating risk level of the online examination system is generated, which solves the problems of low efficiency and high false alarm rate in existing technologies and achieves a high accuracy and low false alarm rate in anti-cheating.

CN121456782APending Publication Date: 2026-02-03广州三七极耀网络科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511404268.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing anti-cheating solutions for online examination systems are inefficient, subjective, and prone to misjudgment. They also fail to effectively perceive and interpret the context of the examination environment, making it difficult to distinguish between genuine cheating and reasonable interference in a home environment.

Method used

By collecting real-time video streams of candidates' behavior, audio streams of the environment, screen content streams, and human-computer interaction event streams, multimodal data is generated and in-depth analysis is performed to generate anomaly confidence levels and cheating risk levels. The data is then dynamically adjusted based on the environmental anomaly confidence levels to trigger corresponding levels of intervention mechanisms.

Benefits of technology

It improves the accuracy of cheating detection, reduces the false alarm rate, can capture complex cheating methods that disguise or circumvent traditional detection, effectively avoids misjudgment from a single data source, and reduces false alarms caused by environmental factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456782A_ABST
    Figure CN121456782A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computers, and provides an anti-cheating method of an online examination system, which comprises the following steps: in an examination process, collecting a behavior video stream, an environment audio stream, a screen content stream and a man-machine interaction event stream of an examinee in real time, and generating multi-modal data; performing deep analysis on the multi-modal data to generate an abnormal confidence coefficient, and generating a cheating risk level in combination with the abnormal confidence coefficient; according to the cheating risk level, an intervention mechanism of the corresponding level is triggered, false alarms caused by misjudgment of a single data source are fundamentally avoided, and meanwhile the false alarms caused by environmental factors are avoided by dynamically adjusting the decision threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and in particular relates to a method, system, device and medium for preventing cheating in an online examination system. Background Technology

[0002] Online examination systems are unproctored systems where students take exams online, and preventing cheating is key to ensuring the fairness of the exam.

[0003] Most existing online examination systems employ anti-cheating measures that rely on the client's webcam to monitor candidate behavior, such as periodically taking and transmitting images of the candidate's exam status to the server. The examiners then subjectively review these images to determine if the candidate is cheating. This approach is inefficient, subjective, and difficult to scale. Furthermore, with the advancement of artificial intelligence, AI technologies are often introduced to monitor candidate behavior in real time through image recognition, behavioral analysis, and natural language processing. However, existing AI-based anti-cheating solutions typically suffer from the following drawbacks: 1) They rely on a single analytical dimension, analyzing only the candidate's facial expressions or mouse movements, lacking cross-validation of multi-dimensional evidence, making them prone to misjudgments; 2) They cannot effectively perceive and interpret the exam environment context, failing to distinguish between genuine cheating and legitimate interference in a home environment. Summary of the Invention

[0004] This application provides an anti-cheating method, system, device, and medium for an online examination system, which can solve one of the above-mentioned problems in the prior art.

[0005] In a first aspect, embodiments of this application provide a method for preventing cheating in an online examination system, including: During the examination, the system collects real-time video streams of examinees' behavior, ambient audio streams, screen content streams, and human-computer interaction event streams to generate multimodal data. The multimodal data is subjected to in-depth analysis to generate anomaly confidence levels, and the cheating risk level is generated by combining the anomaly confidence levels. Based on the level of cheating risk, the corresponding level of intervention mechanism will be triggered.

[0006] Furthermore, the real-time acquisition of the examinee's behavioral video stream, environmental audio stream, screen content stream, and human-computer interaction event stream to generate multimodal data includes: During the examination, multiple parallel data acquisition threads are initiated to generate data units of multiple data streams. For each data unit, the time of the master clock is used as the acquisition timestamp.

[0007] Furthermore, the step of performing in-depth analysis on the multimodal data to generate anomaly confidence levels, and combining these anomaly confidence levels to generate a cheating risk level, includes: The behavioral video stream is processed in real time to calculate the behavioral anomaly confidence score, which includes facial expression anomaly score, gaze deviation score, and body movement anomaly score. The screen content stream is identified and window activity is analyzed. Logical correlation analysis is performed by combining the human-computer interaction event stream and the behavior video stream to calculate the behavior-content correlation anomaly confidence. An initial environment paradigm is constructed, and the environmental audio stream is compared with the initial environment paradigm to calculate the environmental anomaly confidence level. By combining the confidence levels of the behavioral anomalies and the confidence levels of the behavioral-content association anomalies, and by combining the confidence levels of the environmental anomalies, a cheating risk level is generated.

[0008] Furthermore, the behavior video stream is processed in real time to calculate the behavior anomaly confidence score, which includes facial expression anomaly score, gaze deviation score, and body movement anomaly score, including: Facial region images are obtained from the behavioral video stream, and the facial region images are standardized. The standardized facial region image is input into the core classification model, which outputs the candidate's basic confidence score for different emotions. The facial expression abnormality score is calculated by using a preset facial abnormality calculation function and combining the basic confidence score. The eye region image is extracted from the standardized facial region image. The eye region image is input into the gaze recognition model, and the gaze direction vector of the examinee is output. The gaze direction vector is mapped to the screen coordinate system to obtain the gaze point coordinates. The degree of deviation of the gaze point coordinates from the preset legal area is calculated to obtain the gaze deviation degree. From the behavioral video stream, based on continuous time frames, the coordinates of the main joints of the human body are extracted, the coordinates of the temporal key points are generated, the sub-feature indicators of abnormal actions are calculated, and the weighted calculation of multiple sub-feature indicators is performed to obtain the abnormality degree of body action. The abnormality of facial expression, the deviation of gaze, and the abnormality of body movement are weighted and fused to obtain the confidence level of behavioral abnormality.

[0009] Furthermore, the process of identifying and analyzing the screen content stream and window activity, and combining the human-computer interaction event stream and the behavioral video stream to perform logical correlation analysis and calculate the behavior-content correlation anomaly confidence level, includes: Receive the gaze point coordinates and the corresponding gaze timestamp, extract the screenshot of the gaze timestamp from the screen content stream, and crop an image of a preset size with the gaze point coordinates as the center; Analyze the content of the image in the region to identify gaze-to-screen anomaly events; Real-time analysis of human-computer interaction event streams to identify high-risk events; analysis of screen content streams before and after the timestamp of each high-risk event; and determination of interaction-screen abnormal events by combining with a preset verification rule base. Define base scores and decay functions for the gaze-screen anomaly events and the interaction-screen anomaly events respectively, and generate behavior-content association anomaly confidence scores.

[0010] Furthermore, the construction of the initial environment paradigm, which involves comparing the environmental audio stream with the initial environment paradigm and calculating the environmental anomaly confidence level, includes: During the exam preparation phase, video images and ambient audio are captured using the camera and microphone on the candidate's device, generating a static background model and an audio baseline model. The static background model and the audio baseline model are then stored as the initial environment paradigm. The image frames of the behavioral video stream are compared with the static background model to generate a foreground mask. The foreground mask is then preprocessed to obtain an effective foreground region. The effective foreground region is input into the target detection network to identify risk targets related to cheating, and a visual anomaly score is calculated based on the risk targets. Feature extraction is performed on the audio frames of the environmental audio stream to obtain audio feature vectors. The similarity between the audio feature vectors and the audio baseline model is calculated to obtain a spectral anomaly score. The audio feature vector is input into a keyword recognition model to identify keywords related to cheating and obtain a speech anomaly score. The environmental anomaly confidence level is calculated based on the visual anomaly score, the spectral anomaly score, and the speech anomaly score.

[0011] Furthermore, the process of fusing the behavioral anomaly confidence level and the behavioral-content association anomaly confidence level, combined with the environmental anomaly confidence level, to generate a cheating risk level includes: The environmental anomaly confidence level is used as a dynamic adjustment factor to generate a decision threshold for judging abnormal behavior. The confidence scores for behavioral anomalies and behavioral-content association anomalies are fused together to obtain a cheating risk index. The cheating risk index is compared with the decision threshold to determine the cheating risk level of the examinee.

[0012] Secondly, embodiments of this application provide an anti-cheating system for an online examination system, including: The first processing module is used to collect real-time video streams of examinees' behavior, ambient audio streams, screen content streams, and human-computer interaction event streams during the examination, and generate multimodal data. The second processing module is used to perform in-depth analysis on the multimodal data, generate anomaly confidence scores, and combine the anomaly confidence scores to generate a cheating risk level. The third processing module is used to trigger the corresponding level of intervention mechanism based on the cheating risk level.

[0013] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the anti-cheating method of the online examination system described above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium, including a computer program stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the anti-cheating method of the online examination system described above.

[0015] The beneficial effects of the embodiments in this application compared with the prior art are: This application discloses an anti-cheating method for an online examination system. By collaboratively verifying behavioral video streams, screen content streams, and human-computer interaction event streams, the accuracy of cheating identification is improved. Specifically, through logical correlation analysis of multi-source information fusion, false alarms caused by misjudgment from a single data source are fundamentally avoided. It can also capture complex cheating methods that attempt to disguise or circumvent traditional detection, effectively reducing the false alarm rate. In addition, environmental anomaly confidence is introduced as a regulating factor to dynamically adjust the decision threshold, avoiding false alarms caused by environmental factors. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating an anti-cheating method for an online examination system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an anti-cheating system for an online examination system provided in an embodiment of the present invention; Figure 3This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0019] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0020] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0021] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0022] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0023] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0024] Please see Figure 1 As shown, this invention provides a method for preventing cheating in an online examination system, comprising the following steps: S100. During the examination, the system collects real-time video streams of the examinee's behavior, audio streams of the environment, screen content streams, and human-computer interaction event streams to generate multimodal data. In some embodiments, step S100 above includes: During the examination, multiple parallel data acquisition threads are initiated to generate data units of multiple data streams. For each data unit, the time of the master clock is used as the acquisition timestamp.

[0025] In this application, the collected multimodal data needs to be aligned on the timeline in order to make accurate association judgments in the future. For example, a keyboard event, screen content and eye movement at the same time point need to be fully aligned on the timeline. Specifically, a high-precision master clock is used as a unified collection timestamp for all collected data units. The accuracy of the collection timestamp is at least at the millisecond level to ensure strict alignment of different data streams on the timeline.

[0026] Furthermore, during the examination, multiple parallel data acquisition threads are initiated to synchronously acquire multiple independent data streams from the examination process and generate data units. These independent data streams include behavioral video streams, environmental audio streams, screen content streams, and human-computer interaction event streams. Specifically, the behavioral video stream is a sequence of continuous image frames acquired through the candidate's device camera, with each frame timestamped for the acquisition time. This behavioral video stream is used for subsequent analysis of the candidate's facial expressions, gaze direction, and body movements. The environmental audio stream is a continuous audio signal acquired through the candidate's device microphone, used for analyzing environmental anomalies. The screen content stream is acquired through a screen capture interface, generating continuous screen image frames, specifically recording or... Screenshots are taken periodically and recorded at the same or higher frame rate as the behavioral video stream. The capture frequency is increased when screen events such as window switching are detected to improve the accuracy of screen content extraction. For the human-computer interaction event stream, specifically the candidate's keyboard keystrokes, mouse movements, clicks, and scroll wheel events, global hooks or monitoring device files are installed on the candidate's computer device to capture these input events. The type, value, and timestamp of each event are recorded. Specifically, when the candidate is typing, the current time is recorded as a keyboard event, along with the corresponding key code, mouse coordinates, and timestamp. This human-computer interaction event stream provides the most direct and accurate evidence of the candidate's interaction with the examination system, facilitating subsequent analysis of the candidate's behavior.

[0027] In this embodiment, each data unit includes a behavioral video stream, an ambient audio stream, a screen content stream, and a human-computer interaction event stream, as well as the current time value of the master clock when each of the above independent data streams is collected. It can be understood that the collection timestamps of each independent data stream in a data unit are the same. For the collection delays that exist between different devices, such as camera exposure delays and audio buffer delays, benchmark tests are conducted during the exam preparation stage to measure the inherent delays of each data source, and corresponding compensation calibrations are performed when determining the collection timestamps. Furthermore, data units with unified timestamps are sorted and encapsulated according to the time frame order to generate multimodal data, which serves as the data foundation for subsequent data analysis.

[0028] S200. Perform in-depth analysis on the multimodal data to generate anomaly confidence levels, and combine the anomaly confidence levels to generate cheating risk levels; This application improves the accuracy of cheating detection by collaboratively verifying behavioral video streams, screen content streams, and human-computer interaction event streams. Specifically, through logical correlation analysis of multi-source information fusion, it fundamentally avoids false alarms caused by misjudgment from a single data source and can capture complex cheating methods that attempt to disguise or circumvent traditional detection, effectively reducing the false negative rate. In addition, it introduces environmental anomaly confidence as a moderating factor to dynamically adjust the decision threshold, avoiding false alarms caused by environmental factors.

[0029] In some embodiments, step S200 above includes: The behavioral video stream is processed in real time to calculate the behavioral anomaly confidence score, which includes facial expression anomaly score, gaze deviation score, and body movement anomaly score. The screen content stream is identified and window activity is analyzed. Logical correlation analysis is performed by combining the human-computer interaction event stream and the behavior video stream to calculate the behavior-content correlation anomaly confidence. An initial environment paradigm is constructed, and the environmental audio stream is compared with the initial environment paradigm to calculate the environmental anomaly confidence level. By combining the confidence levels of the behavioral anomalies and the confidence levels of the behavioral-content association anomalies, and by combining the confidence levels of the environmental anomalies, a cheating risk level is generated.

[0030] In this embodiment, the environment is introduced to combine and analyze multimodal data. In this process, the cheating risk level of candidates is determined by co-verifying behavioral video streams, screen content streams, and human-computer interaction event streams. The decision threshold for candidate behavior is adjusted by using the environmental anomaly confidence level obtained from the environmental audio stream, thereby achieving intelligent proctoring with high accuracy and low false alarm rate.

[0031] In some embodiments, the real-time processing of the behavioral video stream to calculate behavioral anomaly confidence levels includes facial expression anomaly levels, gaze deviation levels, and body movement anomaly levels, including: Facial region images are obtained from the behavioral video stream, and the facial region images are standardized. The standardized facial region image is input into the core classification model, which outputs the candidate's basic confidence score for different emotions. The facial expression abnormality score is calculated by using a preset facial abnormality calculation function and combining the basic confidence score. The eye region image is extracted from the standardized facial region image. The eye region image is input into the gaze recognition model, and the gaze direction vector of the examinee is output. The gaze direction vector is mapped to the screen coordinate system to obtain the gaze point coordinates. The degree of deviation of the gaze point coordinates from the preset legal area is calculated to obtain the gaze deviation degree. From the behavioral video stream, based on continuous time frames, the coordinates of the main joints of the human body are extracted, the coordinates of the temporal key points are generated, the sub-feature indicators of abnormal actions are calculated, and the weighted calculation of multiple sub-feature indicators is performed to obtain the abnormality degree of body action. The abnormality of facial expression, the deviation of gaze, and the abnormality of body movement are weighted and fused to obtain the confidence level of behavioral abnormality.

[0032] In this embodiment, after receiving the synchronized behavioral video stream, each frame of the image is copied and subjected to multi-threaded analysis of facial expressions, gaze deviation, and body movements. Each analysis process is processed in parallel to improve the efficiency of data processing.

[0033] Furthermore, for facial expression analysis, facial region images are obtained from each image frame and standardized. Specifically, the computationally inexpensive BlazeFace model is used to quickly locate the face bounding boxes in each image frame. Then, the landmark detection model is used to locate key points such as eyes, nose, and mouth within the face bounding boxes. Based on these key points, face alignment and cropping are performed to obtain standardized facial region images. It is worth noting that the BlazeFace model and the Landmark detection model are two models used for different tasks in the field of computer vision. Specifically, the BlazeFace model is a lightweight face detection model used for real-time face localization, while the Landmark detection model is used to identify the key point locations of target objects in images or videos, such as the node coordinates of faces, bodies, and hands.

[0034] Furthermore, the standardized facial region images are analyzed to output a facial expression anomaly score representing the deviation of the expression from a neutral state. Specifically, the standardized facial images are input into the core classification model. It's worth noting that this core classification model is a MobileNetV3-Small network that has undergone knowledge distillation. Specifically, the knowledge from the large teacher model is transferred to the lightweight student model to improve its performance. For the teacher model, a ResNet-50 is pre-trained on a large dataset to output the probability distribution of various emotions. The student model learns the probability distribution output by the teacher model and ultimately outputs a probability vector containing each emotion. Each element represents the probability of the corresponding emotion, with a probability range of [0,1]. The sum of the probabilities of all emotions is 1. In one embodiment, the emotion types include anger, disgust, fear, happiness, sadness, surprise, and neutrality. An exemplary embodiment of the corresponding probability vector is: [0.02, 0.01, 0.03, 0.85, 0.05, 0.03, ... [0.01], where the probability of "happy" is 85%, which corresponds to the basic confidence level of "happy".

[0035] For the facial abnormality calculation function, "neutral" emotion is used as the benchmark. It is understandable that neutral expression is usually a normal state, while some positive or negative emotions may be temporarily reasonable in the exam, such as sadness when thinking or joy when solving a problem. Furthermore, the higher the abnormality, the more the expression deviates from the common reasonable state. Therefore, the facial abnormality calculation function is specifically: Facial expression abnormality = 1 - P(neutral) - max(P(happy), P(sad), P(surprised)).

[0036] Furthermore, regarding gaze deviation, the coordinates of key eye points obtained through the Landmark detection model are used to determine the corresponding eye region image in the standardized facial region image. Combined with the gaze recognition model, the candidate's gaze direction is determined, thereby assessing the degree of gaze deviation during the examination. Specifically, based on the eye key point coordinates extracted from the Landmark detection model, its bounding rectangle is calculated and appropriately expanded to obtain a complete eye region image. The eye region image is then grayscaled and its size normalized to reduce computational complexity and minimize the impact of illumination changes. The eye region image includes both left and right eye images. Further, the processed left and right eye images are concatenated along the channel dimension to form a multi-channel input tensor, providing binoculars... The collaborative information helps improve estimation accuracy. Then, the input tensor is fed into the gaze recognition model, which outputs a two-dimensional gaze direction vector g=(g_x,g_y), where g_x and g_y represent the sine values ​​of the gaze direction's deflection angles in the horizontal and vertical planes, respectively. This two-dimensional vector is not an absolute screen coordinate, but rather a geometric vector representing the gaze direction. Specifically, g_x and g_y represent the projection components of the gaze vector onto the plane of the computer screen. Further, the two-dimensional gaze direction vector g is converted into two-dimensional gaze point coordinates p=(p_x,p_y) in the candidate's computer screen coordinate system. Specifically, p_x = screen_width * (0.5 + k_x * (g_x - g_center_x)), p_y... = screen_height*(0.5+k_y*(g_center_y-g_y)), where k_x and k_y represent scaling factors, g_center_x and g_center_y are the center values ​​of the viewing direction vector g=(g_x, g_y), that is, the component values ​​when the examinee looks directly at the center of the screen, and screen_height and screen_width represent the physical size of the screen.

[0037] More specifically, the gaze recognition model includes an input layer, a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a Dropout layer, a flattening layer, a fully connected layer, and an output layer. The input layer receives the stitched binocular grayscale images. The first convolutional layer extracts low-level features, such as features related to edges and corners, which are then reduced in dimensionality by the first pooling layer to increase translation invariance. The second convolutional layer extracts mid-level features, which are further reduced in dimensionality by the second pooling layer. The Dropout layer prevents overfitting during training. The flattening layer flattens the feature map into vectors. Further, a fully connected layer performs high-level feature fusion, and the output layer outputs the final gaze direction vector g=(g_x,g_y). In this embodiment, a legal region is predefined. This region is a rectangular area in the screen coordinate system corresponding to the examination application window or computer examination device. The spatial relationship between the gaze point coordinates and the aforementioned legal region is calculated in real time. Specifically, the spatial relationship is the shortest distance d between the gaze point and the boundary of the rectangular region, and the deviation angle θ between the gaze point and the center of the region. The spatial deviation S_spatial is calculated from this. Simultaneously, within a sliding time window T_window, the instantaneous value of spatial deviation S_spatial is integrated, and a time decay factor is introduced to calculate the cumulative time deviation S_temporal. The instantaneous value of spatial deviation S_spatial and the cumulative time deviation S_temporal are weighted and fused to generate a gaze deviation score, which is used to measure the gaze deviation. Specifically... ,in, The distance attenuation term is defined as follows: when the fixation point is within the legal area, d=0, and the distance attenuation term is 1, having no effect on spatial deviation. However, when the fixation point is outside the legal area and d gradually increases, this term attenuates exponentially, reflecting the non-linear relationship of "the greater the deviation, the more points are deducted." Furthermore, a risk direction term is introduced. ,in, It is a preset high-risk direction. When the deviation direction is consistent with the high-risk direction, that is... , Then the directional risk item is ,like If the value is close to 1, then the risk factor in that direction is close to 0. The degree of abnormality increases, and when the deviation direction is perpendicular to the high-risk direction, The directional risk item is 0, and the directional risk item is 1. Deviation from the direction has no impact on the spatial deviation, which is understandable. It is the direction weight, 0 <= <1 is used to control the proportion of directional factors in spatial deviation. The above formula ensures that even if the shortest distance d is the same, candidates will get a higher spatial deviation when looking in a high-risk direction.

[0038] Furthermore, since cheating is often continuous, it is necessary to integrate the spatial deviation S_spatial generated by the instantaneous deviation over time to obtain the cumulative temporal deviation S_temporal. Specifically, ,in, The length of the integration time window, This represents the time decay factor, specifically... ,in, It is the attenuation coefficient.

[0039] Furthermore, the gaze deviation score S_gaze is obtained by fusing spatial deviation and temporal cumulative deviation. Specifically, S_gaze = [S_spatial]^α * [S_temporal]^β, where α + β = 1. In other words, if a candidate suddenly glances away, S_spatial is high, but if they immediately look away, S_temporal is low, and S_gaze will not be very high. This may be an unintentional behavior of the candidate. If a candidate continues to look away slightly, both S_spatial and S_temporal will slowly increase, and S_gaze will increase. If a candidate stares at something outside the screen for a long time at a large angle, both S_spatial and S_temporal will increase, and to some extent, S_gaze will rise to close to 1, which can be used to determine abnormal gaze.

[0040] Furthermore, the stability of body posture is analyzed. Specifically, the landmark detection model detects key human joints such as nose, shoulder, elbow, and wrist in real time on each frame of the behavioral video stream. The coordinates of each key human joint are stored in continuous time frames. Then, features that can characterize suspicious movements are calculated, including head orientation, hand visibility, and trunk stability. For head orientation, the head deflection angle relative to the screen is calculated based on the key points of the nose, left ear, and right ear. For hand visibility, it is calculated whether the wrist key point is continuously outside the image boundary or is invisible for a long time. For trunk stability, the variance of the movement speed of the shoulder key points is calculated. If the variance is too large, it indicates frequent body shaking or turning. This generates sub-feature indicators for each abnormal movement. Furthermore, each sub-feature is mapped to a score of 0-1 according to its degree of abnormality. Finally, S_pose= The abnormality of body movements is calculated by adding w1*S_angle, w2*S_hand, and w3*S_shoulder. Here, w1, w2, and w3 are weight coefficients, and w1+w2+w3=1. The values ​​of each weight coefficient can be adjusted according to the actual scenario. S_angle represents the head orientation feature, S_hand represents the hand visibility feature, and S_shoulder represents the trunk stability.

[0041] Furthermore, the specific formula for calculating the confidence level of behavioral abnormality is C1=α*S_face+β*S_gaze+γ*S_pose, where α, β, and γ are weight coefficients adjusted based on experience, S_face represents the degree of facial expression abnormality, S_gaze represents the degree of gaze deviation, and S_pose represents the degree of body movement abnormality.

[0042] In some embodiments, the step of identifying and analyzing the screen content stream and window activity, and combining the human-computer interaction event stream and the behavioral video stream to perform logical correlation analysis and calculate the behavior-content correlation anomaly confidence level, includes: Receive the gaze point coordinates and the corresponding gaze timestamp, extract the screenshot of the gaze timestamp from the screen content stream, and crop an image of a preset size with the gaze point coordinates as the center; Analyze the content of the image in the region to identify gaze-to-screen anomaly events; Real-time analysis of human-computer interaction event streams to identify high-risk events; analysis of screen content streams before and after the timestamp of each high-risk event; and determination of interaction-screen abnormal events by combining with a preset verification rule base. Define base scores and decay functions for the gaze-screen anomaly events and the interaction-screen anomaly events respectively, and generate behavior-content association anomaly confidence scores.

[0043] In this embodiment, when the examinee's gaze coordinates fall within the computer screen, the gaze coordinates and the corresponding gaze timestamp are received. Based on the gaze timestamp, a corresponding screenshot is extracted from the screen content stream. Furthermore, a region image of a preset size is cropped with the gaze coordinates as the center. In one embodiment, the size of the region image is 400x300 pixels, which can at least cover a complete word or application icon. Further, based on the region image, optical character recognition and application window icon recognition are performed. If the recognition result contains predefined illegal keywords or interface elements of unauthorized applications, a gaze-screen abnormality event is triggered.

[0044] Furthermore, real-time analysis of the human-computer interaction event flow during the examination process is required to filter out high-risk events. Understandably, this necessitates classifying events within the human-computer interaction event flow as high-risk events. For example, triggering system shortcut keys such as Ctrl+C, Ctrl+V, Alt+Tab, and Win+D, or window switching events, can be set as high-risk events. Further analysis is needed of the screen content before and after each high-risk event, such as using OCR technology to recognize the title bar text of the active window and using computer vision technology to detect instantaneous changes in screen content. Combined with the triggered events, a pre-defined verification rule base is used to determine whether a corresponding interaction-screen anomaly event is triggered. That is, if there is a logical contradiction between the intent of the interaction event recorded in the pre-defined rules and the actual state reflected in the screen content flow, an interaction-screen anomaly event is triggered.

[0045] In some possible implementations, there is rule 1: if a Ctrl+C event is detected, a Ctrl+V event should immediately follow, and the screen content should change accordingly during the two events, with text area highlighting and content addition. If there is only a copy event without subsequent pasting, or if the content does not change after pasting, an interaction-screen anomaly event is triggered. Rule 2: if an Alt+Tab event is detected, the screen content stream must capture a clear change in the foreground window, such as a change in the window title or interface layout. If a switching event is recorded but the screen content shows the foreground window is always the examination system, it is determined to be a "virtual switch" that may be simulated by cheating software, triggering an interaction-screen anomaly event. Understandably, each rule can be set according to the actual needs in the examination scenario, ultimately generating a verification rule base for the determination of interaction anomalies.

[0046] Furthermore, the frequency, duration, and severity of the aforementioned triggering events within a unit of time are statistically analyzed, and a fusion calculation is performed to obtain the behavior-content association anomaly confidence score. Specifically, a base score and a decay function are defined for each triggered anomaly event. Within the current time window, for each type of triggering event, an event score is calculated based on its base score and decay function, and then weighted and summed according to weights to obtain the behavior-content association anomaly confidence score. Where n is the number of events triggered. Let i be the weight of the i-th type of event. This is the base score for the i-th type of event. Let be the decay function of the i-th type of event, and t be the time interval after the event occurs.

[0047] In some embodiments, the construction of an initial environment paradigm, comparing the environmental audio stream with the initial environment paradigm, and calculating the environmental anomaly confidence level includes: During the exam preparation phase, video images and ambient audio are captured using the camera and microphone on the candidate's device, generating a static background model and an audio baseline model. The static background model and the audio baseline model are then stored as the initial environment paradigm. The image frames of the behavioral video stream are compared with the static background model to generate a foreground mask. The foreground mask is then preprocessed to obtain an effective foreground region. The effective foreground region is input into the target detection network to identify risk targets related to cheating, and a visual anomaly score is calculated based on the risk targets. Feature extraction is performed on the audio frames of the environmental audio stream to obtain audio feature vectors. The similarity between the audio feature vectors and the audio baseline model is calculated to obtain a spectral anomaly score. The audio feature vector is input into a keyword recognition model to identify keywords related to cheating and obtain a speech anomaly score. The environmental anomaly confidence level is calculated based on the visual anomaly score, the spectral anomaly score, and the speech anomaly score.

[0048] In this embodiment, during the exam preparation phase, candidates are guided to place their faces in a designated area to ensure they are part of the foreground rather than the background. An environmental modeling paradigm acquisition process lasting T seconds is then performed. This includes acquiring F consecutive video images using the candidate's device's camera, performing lens distortion correction and color space conversion on each frame to reduce the impact of lighting changes, and constructing a static background model representing the scene's static elements using a background modeling algorithm. Simultaneously, a continuous environmental audio segment is acquired via a microphone, and after frame segmentation, windowing, and short-time Fourier transform, the MFCC coefficients for each audio frame are calculated. The MFCC coefficients are then analyzed using cepstral analysis to separate the audio signal's spectrum into spectral envelope and spectral details. The audio signal consists of two parts: the spectral envelope, which represents the main frequency components and their energy distribution, and the spectral details, which contain more refined information about frequency changes. Thus, the MFCC coefficients can effectively extract the spectral envelope information to characterize the spectral features of the audio signal. Based on this, the statistical characteristics of the MFCC feature vectors of all frames within the T-second audio are calculated to obtain an audio baseline model that can characterize the background noise of the environment. Specifically, the audio baseline model includes the mean vector, variance vector, and covariance matrix of the MFCC features in the environment. Then, the static background model and the audio baseline model are stored as the candidate's initial environmental paradigm so that abnormal changes in the candidate's surrounding environment can be keenly perceived during the examination.

[0049] In this embodiment, during the examination, image frames from the behavioral video stream and audio frames from the environmental audio stream are received in real time and compared with the static background model and audio baseline model in the initial environment paradigm, respectively, to measure the degree of environmental anomaly during the examination.

[0050] Specifically, for the received real-time image frames, the same background modeling algorithm as in the exam preparation stage is used. The background model of the current frame is compared with the stored static background model to generate a binary foreground mask, where white pixels represent moving or changing foreground regions. Morphological operations, such as opening or closing operations, are applied to the foreground mask to remove small noise points and connect broken foreground regions. Then, connected regions are calculated, and regions with too small an area are filtered out to remove irrelevant interference. For the remaining effective foreground regions, the proportion of their total area to the entire image area is calculated, and the effective foreground regions are cropped out and input into the object detection network. This object detection network is trained to identify risky targets related to cheating, such as faces, bodies, mobile phones, and books. Furthermore, based on the risky targets, a visual anomaly score S_visual is calculated. ,in, Indicates the proportion of the foreground area. The threshold representing the proportion of the foreground region. This refers to anomaly scores where the object detection network identifies a "face" or "human body," but the target is not the test taker. These scores are understandable. This is a fixed value, representing an abnormal score recorded when the person in question is not the actual test taker. This represents the anomaly score when identifying risky items such as "cell phones" or "books," where c represents the confidence level for different risky items. For a time-dependent function, that is, to introduce the duration of the occurrence of the risk target. ,in, This represents the time contribution coefficient.

[0051] More specifically, for the received audio frames, MFCC features are extracted using the same method as in the exam preparation stage. Then, the similarity between the feature vector and the stored audio baseline model is calculated. In some embodiments, Euclidean distance is used for calculation to obtain a spectral anomaly score. Understandably, the larger the distance, the greater the difference between the current audio and the initial ambient noise. Further, the ambient audio stream is segmented into human voice segments using VAD technology. Speaker separation is performed on the human voice segments to distinguish the candidate's voice from the voices of potential others. For each separated human voice track, cheating keyword identification is performed. Specifically, the human voice track is input into a preset keyword recognition model. Understandably, this model focuses on detecting cheating-related keywords, such as "choose A" or "what is the answer," and presets a risk value for each keyword. The risk values ​​of multiple keywords are weighted and calculated to obtain a speech anomaly score.

[0052] Furthermore, when calculating the confidence level of environmental anomalies, persistent anomaly scores are given higher weight than transient anomaly scores. ,in, Indicates the confidence level of environmental anomalies. Indicates the visual abnormality score. Indicates the spectral anomaly score. Indicates the speech abnormality score. a, b, and c represent the weights of each outlier score, a + b + c = 1, and in the process of calculating the weights, the following is introduced: This is used to assign higher weight to anomaly scores based on the duration of an event, such as... = _0 ,in, _0 represents the initial weight. Indicates the duration.

[0053] In some embodiments, the step of fusing the behavioral anomaly confidence level and the behavior-content association anomaly confidence level, combined with the environmental anomaly confidence level, to generate a cheating risk level includes: The environmental anomaly confidence level is used as a dynamic adjustment factor to generate a decision threshold for judging abnormal behavior. The confidence scores for behavioral anomalies and behavioral-content association anomalies are fused together to obtain a cheating risk index. The cheating risk index is compared with the decision threshold to determine the cheating risk level of the examinee.

[0054] In this embodiment, the environmental anomaly confidence level is not directly involved in the fusion calculation of the cheating risk index, but rather serves as a control signal for dynamically adjusting the decision threshold. This is used to determine the degree of abnormality in the candidate's behavior. Specifically, the decision threshold = basic threshold × (1 + η × environmental anomaly confidence level), where η represents the adjustment coefficient. In particular, when the examination environment is stable, the decision threshold is lowered, making the online examination system more sensitive to abnormal candidate behavior and thus reducing the risk of missed reports. When the examination environment is highly disruptive, the decision threshold is raised, relaxing the criteria for judging abnormal behavior in the online examination system and thus reducing the risk of false reports.

[0055] Furthermore, by using a weighted calculation method, the confidence level of behavioral anomalies and the confidence level of behavioral-content association anomalies are fused together to obtain a cheating risk index. If the cheating risk index is greater than the decision threshold, it indicates that the candidate has a risk of cheating, triggering the corresponding intervention mechanism. If the cheating risk index is less than the decision threshold, it indicates that the candidate's behavior during the examination is normal.

[0056] Furthermore, when a candidate is determined to have a risk of cheating, the difference between the cheating risk index and the decision threshold is calculated to determine the candidate's cheating risk level, which includes low risk, medium risk, and high risk levels. Different risk thresholds are set for different cheating risk levels.

[0057] S300. Based on the cheating risk level, trigger the corresponding level of intervention mechanism.

[0058] In this embodiment, different levels of intervention mechanisms are triggered based on the level of cheating risk. These intervention mechanisms include front-end local alerts, asynchronous back-end flagging, and real-time manual intervention. This tiered intervention mechanism avoids excessive interference with the examination process.

[0059] Specifically, when a candidate is at a low risk level, a local front-end alert is triggered. This alert displays a gentle message on the exam interface reminding the candidate to follow exam rules, such as "Please maintain the correct exam posture and avoid unnecessary fidgeting." This alert method does not significantly disrupt the exam process while making the candidate aware that their behavior may be under scrutiny, thus encouraging them to behave appropriately. For candidates at a medium risk level, in addition to possible front-end alerts, the online exam system will flag the candidate in the background. This flag is generated by combining the video stream of their behavior and the ambient audio stream during the exam and recorded in the exam system's database. After the exam, invigilators can review the flag to further analyze whether the candidate's behavior constitutes cheating. Furthermore, the online exam system can be used in subsequent exam sessions... The system monitors the candidate more closely but does not interrupt the exam in real time. When a candidate is at a high risk level, a real-time manual intervention mechanism is immediately triggered. The online examination system automatically notifies the online proctors, who can communicate directly with the candidate through real-time video monitoring and voice dialogue to understand the situation and stop potential cheating. For example, the proctor can remind the candidate via voice: "Attention, your behavior has attracted attention. Please immediately stop any possible inappropriate behavior and abide by the examination rules." If cheating is confirmed, the proctor can take appropriate action according to the examination regulations. Through this tiered intervention mechanism, corresponding measures can be taken based on the different levels of cheating risk, effectively preventing cheating while avoiding excessive interference with the examination process and ensuring the smooth conduct of the examination.

[0060] Please see Figure 2 As shown, the present invention also provides an anti-cheating system for an online examination system, the system comprising: First processing module 201: used to collect the candidate's behavior video stream, environmental audio stream, screen content stream and human-computer interaction event stream in real time during the examination, and generate multimodal data; The second processing module 202 is used to perform in-depth analysis on the multimodal data, generate anomaly confidence scores, and generate a cheating risk level based on the anomaly confidence scores. The third processing module 203 is used to trigger the corresponding level of intervention mechanism based on the cheating risk level.

[0061] It is understandable that, such as Figure 1 The content of the anti-cheating method embodiments of the online examination system shown is applicable to the anti-cheating system embodiments of this online examination system. The specific functions implemented by the anti-cheating system embodiments of this online examination system are the same as those shown below. Figure 1 The anti-cheating method of the online examination system shown is the same as that implemented in this example, and the beneficial effects achieved are the same as those described above. Figure 1 The anti-cheating method of the online examination system shown in the embodiment achieves the same beneficial effect.

[0062] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0063] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0064] Please see Figure 3 As shown, this embodiment of the invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the anti-cheating method of the online examination system as described in any of the above methods.

[0065] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0066] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0067] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0068] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the anti-cheating method for an online examination system as described in any of the above methods.

[0069] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / computer device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0070] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for preventing cheating in an online examination system, characterized in that, include: During the examination, the system collects real-time video streams of examinees' behavior, ambient audio streams, screen content streams, and human-computer interaction event streams to generate multimodal data. The multimodal data is subjected to in-depth analysis to generate anomaly confidence levels, and the cheating risk level is generated by combining the anomaly confidence levels. Based on the level of cheating risk, the corresponding level of intervention mechanism will be triggered.

2. The method as described in claim 1, characterized in that, The real-time acquisition of examinee behavior video streams, environmental audio streams, screen content streams, and human-computer interaction event streams generates multimodal data, including: During the examination, multiple parallel data acquisition threads are initiated to generate data units of multiple data streams. For each data unit, the time of the master clock is used as the acquisition timestamp.

3. The method as described in claim 1, characterized in that, The process of performing deep analysis on the multimodal data to generate anomaly confidence levels, and combining these anomaly confidence levels to generate a cheating risk level, includes: The behavioral video stream is processed in real time to calculate the behavioral anomaly confidence score, which includes facial expression anomaly score, gaze deviation score, and body movement anomaly score. The screen content stream is identified and window activity is analyzed. Logical correlation analysis is performed by combining the human-computer interaction event stream and the behavior video stream to calculate the behavior-content correlation anomaly confidence. An initial environment paradigm is constructed, and the environmental audio stream is compared with the initial environment paradigm to calculate the environmental anomaly confidence level. By combining the confidence levels of the behavioral anomalies and the confidence levels of the behavioral-content association anomalies, and by combining the confidence levels of the environmental anomalies, a cheating risk level is generated.

4. The method as described in claim 3, characterized in that, The behavior video stream is processed in real time to calculate the behavior anomaly confidence score, which includes facial expression anomaly score, gaze deviation score, and body movement anomaly score, including: Facial region images are obtained from the behavioral video stream, and the facial region images are standardized. The standardized facial region image is input into the core classification model, which outputs the candidate's basic confidence score for different emotions. The facial expression abnormality score is calculated by using a preset facial abnormality calculation function and combining the basic confidence score. The eye region image is extracted from the standardized facial region image. The eye region image is input into the gaze recognition model, and the gaze direction vector of the examinee is output. The gaze direction vector is mapped to the screen coordinate system to obtain the gaze point coordinates. The degree of deviation of the gaze point coordinates from the preset legal area is calculated to obtain the gaze deviation degree. From the behavioral video stream, based on continuous time frames, the coordinates of the main joints of the human body are extracted, the coordinates of the temporal key points are generated, the sub-feature indicators of abnormal actions are calculated, and the weighted calculation of multiple sub-feature indicators is performed to obtain the abnormality degree of body action. The abnormality of facial expression, the deviation of gaze, and the abnormality of body movement are weighted and fused to obtain the confidence level of behavioral abnormality.

5. The method as described in claim 4, characterized in that, The process of identifying and analyzing the screen content stream and window activity, and combining the human-computer interaction event stream and the behavioral video stream to perform logical correlation analysis and calculate the behavior-content correlation anomaly confidence level, includes: Receive the gaze point coordinates and the corresponding gaze timestamp, extract the screenshot of the gaze timestamp from the screen content stream, and crop an image of a preset size with the gaze point coordinates as the center; Analyze the content of the image in the region to identify gaze-to-screen anomaly events; Real-time analysis of human-computer interaction event streams to identify high-risk events; analysis of screen content streams before and after the timestamp of each high-risk event; and determination of interaction-screen abnormal events by combining with a preset verification rule base. Define base scores and decay functions for the gaze-screen anomaly events and the interaction-screen anomaly events respectively, and generate behavior-content association anomaly confidence scores.

6. The method as described in claim 3, characterized in that, The construction of the initial environment paradigm involves comparing the environmental audio stream with the initial environment paradigm and calculating the environmental anomaly confidence level, including: During the exam preparation phase, video images and ambient audio are captured using the camera and microphone on the candidate's device, generating a static background model and an audio baseline model. The static background model and the audio baseline model are then stored as the initial environment paradigm. The image frames of the behavioral video stream are compared with the static background model to generate a foreground mask. The foreground mask is then preprocessed to obtain an effective foreground region. The effective foreground region is input into the target detection network to identify risk targets related to cheating, and a visual anomaly score is calculated based on the risk targets. Feature extraction is performed on the audio frames of the environmental audio stream to obtain audio feature vectors. The similarity between the audio feature vectors and the audio baseline model is calculated to obtain a spectral anomaly score. The audio feature vector is input into a keyword recognition model to identify keywords related to cheating and obtain a speech anomaly score. The environmental anomaly confidence level is calculated based on the visual anomaly score, the spectral anomaly score, and the speech anomaly score.

7. The method as described in claim 3, characterized in that, The method of fusing the behavioral anomaly confidence level and the behavior-content association anomaly confidence level, combined with the environmental anomaly confidence level, generates a cheating risk level, including: The environmental anomaly confidence level is used as a dynamic adjustment factor to generate a decision threshold for judging abnormal behavior. The confidence scores for behavioral anomalies and behavioral-content association anomalies are fused together to obtain a cheating risk index. The cheating risk index is compared with the decision threshold to determine the cheating risk level of the examinee.

8. An anti-cheating system for an online examination system, characterized in that, include: The first processing module is used to collect real-time video streams of examinees' behavior, ambient audio streams, screen content streams, and human-computer interaction event streams during the examination, and generate multimodal data. The second processing module is used to perform in-depth analysis on the multimodal data, generate anomaly confidence scores, and combine the anomaly confidence scores to generate a cheating risk level. The third processing module is used to trigger the corresponding level of intervention mechanism based on the cheating risk level.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.