Examination cheating detection method and system based on fusion of local detection features and global semantic features
By improving the illegal product detection model and combining global semantic features with local detection features for spatial and channel alignment, the problem of low detection accuracy of the traditional yolo-world model is solved, and higher-precision cheating detection in exams is achieved.
Patent Information
- Application Number
- CN202510705250.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing cheating detection methods rely on manual video monitoring, which has problems such as blind spots in monitoring, high labor costs, and high misjudgment rates. In addition, the traditional yolo-world model only extracts local detection features, resulting in low detection accuracy.
An improved illegal product detection model is adopted. Global semantic features of three different scales are extracted from the test area image, and then spatial and channel dual alignment is performed with the local detection features of the traditional yolo-world model. Feature fusion is then performed to combine the global semantic features for target recognition.
It improves the accuracy of illegal product detection, effectively identifies illegal products used in exams, reduces false detections and missed detections, and improves the accuracy of cheating detection.
Smart Images

Figure CN120635407A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting cheating in an examination based on the fusion of local detection features and global semantic features, belonging to the technical field of cheating detection. Background Art
[0002] Existing cheating detection in online exams primarily relies on manual video monitoring, which presents challenges such as blind spots, high labor costs, and a high rate of false positives. Traditional anti-cheating devices, however, often utilize single-view cameras, making them ineffective at identifying covert cheating attempts (such as checking a cheat sheet on the leg or using a miniature headset) or the use of illegal items (such as cheat sheets and miniature headsets) used in these attempts. Detecting cheating in exams encompasses both detecting illegal items used for cheating and detecting cheating attempts.
[0003] To detect illicit items used for cheating in exams, a Chinese invention patent application with application publication number CN116537601A and publication date August 4, 2023, discloses a single-person exam cabin. The cabin comprises a square housing that isolates the cabin from the outside world and contains exam tables and chairs. The housing also incorporates an identity recognition module for determining the examinee's identity, a wireless signal shielding module for shielding wireless signals, a file storage module for storing confidential exam papers and answer sheets, a video surveillance module for real-time monitoring of the examinee's status and recording the answering process, and an anti-cheating module for identifying examinee behavior and detecting suspicious behavior. The video surveillance module includes several cameras, both inside and outside the housing, that monitor the exam situation. Each camera transmits real-time footage to a cloud monitor via a wireless network card for viewing by the invigilator. This solution, when applied to cheating detection in exams, still primarily relies on manual video monitoring, which still presents problems such as blind spots, high labor costs, and a high rate of false positives.
[0004] To this end, artificial intelligence learning is currently being used to detect illegal substances used for cheating in exams.
[0005] For example, a Chinese invention patent application with application publication number CN118470645A and publication date August 9, 2024, discloses a system and method for intelligent written exam monitoring based on visual detection. This solution utilizes a visual anti-cheating module to detect cheating behaviors such as illegal items, illegal actions, and out-of-bounds behavior. The visual anti-cheating module uses the MaskRCNN instance segmentation algorithm and the Transformer-based video understanding model ViViT to control a panoramic camera to scan the examination room. The module receives a sequence of video frames from the panoramic camera and first detects illegal items such as mobile phones and documents in each frame using the MaskRCNN instance segmentation algorithm. False detections are filtered out using prior knowledge such as area and aspect ratio. The high-precision ViViT video classification model then performs semantic understanding of multiple consecutive frames of examinee behavior, identifying typical cheating behaviors such as passing notes, whispering, and looking around. Simultaneously, the visual anti-cheating module periodically calls the examinee location recording module to obtain the latest examinee spatial coordinates. By comparing these coordinates with the seat reference coordinates, the module determines whether the examinee has engaged in suspicious out-of-bounds behavior. To further improve cheating detection accuracy, this module also utilizes algorithms such as FaceBoxes face detection, EfficientPose body pose estimation, and optical flow motion estimation to analyze abnormal student behavior through multimodal fusion. Upon detecting suspected cheating, the module immediately records the relevant video clips, extracts keyframe screenshots, labels the cheating type, and sends real-time alerts to exam invigilators.
[0006] Unlike the previous solution, the Chinese invention patent application with application publication number CN117640896A and application publication date 2024.03.01 discloses an unmanned proctoring system based on visual anti-cheating technology. This solution uses the YOLOv8 detection algorithm to detect whether there are abnormal targets. If an abnormal target is detected, it is determined to be cheating behavior.
[0007] In addition, there is also the possibility of using reference Figure 5 The traditional yolo-world target detection model is used to detect and identify illegal items used for cheating in the examination area. In this solution, the traditional yolo-world inputs the local detection features of three different scales extracted by its yolo backbone network from the input examination area image into the visual language path aggregation network in the traditional yolo-world. Because the yolo backbone network of the traditional yolo-world relies on convolution operations, it tends to capture detailed features such as local textures and edges. Although high-level features will gradually integrate semantic information, the ability to understand the global scene is limited. In complex scenes (such as occlusion, small targets, and dense targets), relying solely on the extracted local detection features may lead to false detections or missed detections. In this way, there may be situations where illegal items as targets may not be fully identified or identified incorrectly, and the accuracy of detecting illegal items is low. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for detecting cheating in examinations based on the fusion of local detection features and global semantic features, so as to solve the problem that the existing detection of illegal products for cheating in examinations adopts traditional yolo-world, and the detection accuracy is low due to the extraction of only local detection features; and also to provide a system for detecting cheating in examinations based on the fusion of local detection features and global semantic features, so as to solve the problem that the existing detection of illegal products for cheating in examinations adopts traditional yolo-world, and the detection accuracy is low due to the extraction of only local detection features.
[0009] To achieve the above object, the solution of the present invention includes:
[0010] The present invention provides a method for detecting cheating in an exam based on the fusion of local detection features and global semantic features, comprising the following steps:
[0011] Obtain an image of the examination area and input it into a trained illegal product detection model to obtain a detection result of whether there are illegal products in the examination area image;
[0012] The illegal product detection model is an improvement on the traditional yolo-world model, and its improvements include:
[0013] Three global semantic features of different scales are extracted from the input of the detection model, and the global semantic features are aligned so that the three aligned global semantic features of different scales can be dually aligned in space and channel with the three local detection features of different scales extracted by the yolo backbone network in the traditional yolo-world. The three local detection features of different scales are then fused with the aligned global semantic features of three different scales, and the result of the feature fusion is input into the visual language path aggregation network in the traditional yolo-world.
[0014] Furthermore, the global semantic features are extracted from the input of the detection model by using the image segmentation and position encoding module of dinov2 to perform image segmentation and position encoding on the input of the detection model, and then using the Transformer encoder module of dinov2 to process the output of the image segmentation and position encoding module to extract global semantic features of three different scales.
[0015] Furthermore, the spatial alignment in the dual alignment includes dimensional alignment and size alignment. Accordingly, the alignment processing method includes the following steps: first, using spatial structure reorganization to adjust the dimensions of the global semantic features of three different scales to the same dimension as the local detection features, thereby achieving dimensional alignment; then, through 1×1 convolution, the channels of the dimensionally aligned global semantic features are adjusted to the same channels as the local detection features, thereby achieving channel alignment; finally, through upsampling, the size of the dimension-channel aligned global semantic features is adjusted to the same size as the local detection features, thereby achieving size alignment.
[0016] Furthermore, the outputs of the 3rd, 9th, and 12th transformer encoding blocks in the Transformer encoder module are global semantic features at three different scales.
[0017] Furthermore, feature fusion adopts the deformable convolution feature fusion method. The processing of deformable convolution feature fusion includes:
[0018] The extracted local detection features of three different scales are used as input for deformable convolution to calculate the offset, the aligned global semantic features of three different scales are used as input for bilinear difference processing in deformable convolution, and the output of deformable convolution is concatenated with the extracted local detection features of three different scales as the output of deformable convolution feature fusion.
[0019] Furthermore, feature fusion adopts a splicing approach.
[0020] Furthermore, the examination area image adopts images of the examination area taken at at least two different shooting angles.
[0021] The present invention provides an examination cheating detection system based on the fusion of local detection features and global semantic features, including a processing terminal, a cabin for providing an examination area, and a camera installed in the cabin for acquiring an image of the examination area. The processing terminal includes a processor, and the processor is used to execute a computer program to implement the steps of the above-mentioned examination cheating detection method based on the fusion of local detection features and global semantic features.
[0022] Furthermore, the camera is installed at at least two installation positions in the cabin to obtain images of the examination area at at least two different shooting angles.
[0023] Furthermore, the installation positions of the camera include directly above, above the left rear, and above the right rear of the monitored person in the examination area.
[0024] Beneficial effects of the present invention:
[0025] The present invention is an improved invention, which provides a method for detecting cheating in exams based on the fusion of local detection features and global semantic features. Taking into account that the local detection features extracted by traditional yolo-world are easily affected by surrounding environmental factors, such as occlusion and lighting; and the global semantic features can provide prior knowledge of scene semantics and targets, such as desktop texture and lighting changes in the exam scene, which can make up for the shortcomings of local detection features, so that local detection features can be combined with the surrounding environment to perform target detection. Therefore, the present invention improves the traditional yolo-world and adopts a trained illegal product detection model obtained by improving the traditional yolo-world to detect whether there are illegal products in the image of the exam area, thereby realizing accurate detection of illegal products used for cheating in exams. The improved illegal product detection model is different from the existing local detection features input into the visual language path aggregation network in the traditional yolo-world. Specifically, the space and channel of the global semantic features of three different scales extracted from the test area image are doubly aligned with the local detection features of three different scales extracted from the same test area image in space and channel respectively. The local detection features after doubly alignment are fused with the global semantic features and then input into the visual language path aggregation network in the traditional yolo-world. This makes the target recognition no longer limited to the local detection features, but combines the global semantic features for target recognition, so that the detected features can take the surrounding environment (i.e., the global semantic features) into account, effectively improving the detection accuracy of illegal products. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is the network architecture diagram of yolo-world integrated with dinov2;
[0027] Figure 2 This is the network architecture diagram for extraction and alignment of global semantic features;
[0028] Figure 3 This is the architecture diagram of the Encoder Block;
[0029] Figure 4 This is the architecture diagram of MLP Block;
[0030] Figure 5 This is the traditional yolo-world network architecture diagram. DETAILED DESCRIPTION
[0031] In order to solve the problems in the background technology, the present invention, when detecting illegal items used for cheating in exams, not only extracts local detection features related to the illegal items as the detection target, but also extracts global semantic features related to the environment in which the detection target is located. The detection of illegal items is achieved based on local detection features and global semantic features, which can effectively improve the detection accuracy of illegal items.
[0032] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and implementation methods.
[0033] An implementation method of an exam cheating detection system based on the fusion of local detection features and global semantic features:
[0034] A cheating detection system for exams based on the fusion of local detection features and global semantic features is used to detect whether there are any illegal items used for cheating in a cabin. The detection system includes basic hardware facilities related to the exam, such as a cabin providing an exam area, an all-in-one exam machine for examinees to perform exam operations, special exam seats for examinees to sit in front of the all-in-one exam machine, a camera for acquiring images of the exam area for illegal item detection, and a processing terminal for detecting whether there are any illegal items used for cheating in the cabin.
[0035] Among them, the processing terminal mainly detects whether the examinees in the cabin use any illegal items for cheating during the examination stage. The processing terminal includes a processor, which is used to execute a computer program to implement the steps of an examination cheating detection method based on the fusion of local detection features and global semantic features.
[0036] The basic hardware facilities also include a broadband electromagnetic shielding device, which is activated during the examination phase to shield wireless signals in the examination area and automatically cut off the Bluetooth / WiFi signals inside the cabin from communication outside the cabin.
[0037] In order to conduct all-round and multi-angle cross-monitoring of the examination area where the candidates are located, no less than two cameras can be installed in the cabin. The installation position can be selected according to actual needs to obtain images of the examination area from multiple different shooting angles.
[0038] Specifically, the installation position of the camera includes but is not limited to directly above, above the left rear, and above the right rear of the monitored person in the examination area.
[0039] Taking four cameras as an example, they should form a three-dimensional monitoring network covering front, top, left rear, and right rear. This multi-angle monitoring improves the accuracy of identifying illegal items used for cheating. The specific installation locations and functions are as follows:
[0040] ① The camera is installed in the middle position above the display of the all-in-one examination machine. The camera is used for face recognition and face tracking by facing the examinee's face, and captures the direction of sight and head posture in real time. It is used to analyze cheating, replacement of candidates in the middle of the exam, the concentration level of the examinee, and identify abnormal behaviors such as lowering the head for a long time or turning the head frequently.
[0041] ②The camera is installed directly above the cabin where the all-in-one examination machine is located. The camera covers the desktop area vertically, focusing on monitoring whether the examinee's hands are placed in the designated position (such as the keyboard operation area), and at the same time capturing whether there are any cheat sheets, books, electronic devices and other illegal items on and around the desktop.
[0042] ③ The camera is installed at the upper left rear position of the cabin. This camera forms a cross-monitoring with the upper right rear camera, covering the entire cabin space through the side and rear perspective, focusing on identifying changes in the number of people (such as substitute test takers entering), abnormal limb movements of candidates (such as rummaging for items sideways), and illegal items used by candidates.
[0043] ④ The camera is installed at the upper right rear position of the cabin. This camera works together with the upper left camera to build a monitoring network with no blind spots. Through dual-view comparison, it eliminates monitoring blind spots and focuses on checking the placement of electronic equipment (such as the brightness of mobile phone screens) and paper items in the cabin.
[0044] Among them, a method for detecting cheating in an exam based on the fusion of local detection features and global semantic features includes the following steps: obtaining an image of the exam area, and inputting the obtained image of the exam area into a trained illegal product detection model to obtain a detection result of whether there are illegal products in the exam area image.
[0045] The construction process of the illegal product detection model is as follows:
[0046] 1) A large number of images of the test area are obtained through a camera installed in the cabin, and the illegal items in these images of the test area are marked as an image set. Most of the images in the image set are used as the training set, and the remaining images are used as the test set.
[0047] 2) Build a machine learning model and train and test it using the training and test sets to obtain a model for detecting illegal products. The machine learning model uses an improved yolo-world network model, and the model for detecting illegal products is constructed using the improved yolo-world network model.
[0048] In this embodiment, the data is randomly divided into a training set and a test set in a ratio of 9:1. The specific division ratio can also be adjusted, such as 8:2. Generally, the amount of data in the training set is greater than the amount of data in the test set.
[0049] Among them, the examination area image uses images of the examination area taken from at least two different shooting angles, which can detect illegal items in the examination area from multiple angles to avoid missed inspections.
[0050] The improved yolo-world network model is obtained by improving the traditional YOLO-World network model.
[0051] Among them, the traditional YOLO-World network model architecture, reference Figure 5 , consists of a YOLO detector, a text encoder, and a reparameterizable visual-linguistic path aggregation network (RepVL-PAN). Given an input text, the text is encoded into a text embedding through the text encoder in the traditional YOLO-World. The image encoder in the YOLO detector extracts multi-scale features from the input image. Finally, RepVL-PAN is used to perform cross-modal fusion between image features and text embeddings to enhance the representation of text and images.
[0052] like Figure 5As shown in the figure, the traditional YOLO-World network model includes Training (Online Vocabulary), Deployment (Offline Vocabulary), Vocabulary Embeddings, Image-aware Embeddings, Region-Text Matching, Object Embeddings, Input Image, Yolo Backbone, Vision-Language PAN (Visual Language Path Aggregation Network), Text Contrastive Head, and Box Head. The Extract Nouns function in the Training part extracts nouns from "A man and a woman are skiing with a dog" and inputs them into the Text Encoder in the Training part. The User function in the Deployment part inputs the User's Vocabulary into the Text Encoder. Vocabulary Embeddings: man, woman, dog; Image-aware Embeddings: man, woman, dog. The TextContrastive Head is a key module for open-word detection. It aligns text descriptions with image features through cross-modal contrastive learning, enabling the model to understand arbitrary text instructions and detect corresponding objects. The Box Head is the core component responsible for predicting object bounding boxes. Traditional YOLO-World network models are existing technologies and will not be discussed in detail here.
[0053] Improvements to the traditional YOLO-World network model include: extracting global semantic features of three different scales from the input of the illegal product detection model, and aligning the three global semantic features of different scales so that the three aligned global semantic features of different scales are respectively aligned with the three local detection features of different scales extracted by the YOLO backbone network in the traditional YOLO-World to achieve dual spatial and channel alignment, and then fusing the three local detection features of different scales with the aligned global semantic features of three different scales, and then inputting the feature fusion results into the visual language path aggregation network in the traditional YOLO-World.
[0054] Improved yolo-world network model, such as Figure 1 As shown, Figure 1 Training in YOLO-World uses dynamic input of arbitrary text, while deployment / inference uses a predefined vocabulary, text encoding, word embeddings for person, cellphone, and cheat note, and image visualization embeddings for person, cellphone, and cheat note. The target embedding, region-text similarity calculation, bounding box prediction, reparameterized visual-linguistic path aggregation network, and YOLO backbone are all the same as those in the traditional YOLO-World model. The improved YOLO-World network model adapts its input image to the test area, which can be captured by cameras mounted directly above, to the upper left of the rear, or to the upper right of the rear. Images of the test area from these three angles are simultaneously fed into the YOLO backbone for parallel processing. Figure 1 The feature layers and their sizes output by the yolo backbone are C3 (height H × width W × channel C: 80 × 80 × 256), C4 ((H × W × C: 40 × 40 × 512)), and C5 (H × W × C: 20 × 20 × 512). These are the local detection features of three different scales extracted from the input image by the yolo backbone in the traditional yolo-world. Specifically, Figure 5 Multi-scale Image Features in
[15] . Figure 1 EB3, EB9, and EB12 are global semantic features of three different scales after spatial-channel dual alignment with local detection features of three different scales, which are output by the dinov2vit-b / 14distilled distillation model as the dinov2vit-b / 14backbone. Figure 1 EB3, EB9 and EB12 are Figure 2 EB3: 80×80×256, EB9: 40×40×512 and EB12: 20×20×512.
[0055] Among them, the global semantic features are extracted from the input of the detection model as follows:
[0056] like Figure 2 As shown, the image block and position encoding module of dinov2 is used to perform image block and position encoding on the input of the detection model. For example, it can be divided into 16×16 image blocks, and then the Transformer encoder module of dinov2 is used ( Figure 2The Transformer encoder × 12 in the Transformer encoder module processes the outputs of the image block and position encoding module to extract global semantic features at three different scales. The transformer encoder block in the Transformer encoder module is also called the Encoder Block. Encoder Block layers 1-3 are used to extract shallow local features such as edges and textures; Encoder Block layers 4-9 are used to extract mid-level global correlation features such as object regions and contours; and Encoder Block layers 10-12 are used to extract deep global semantic features such as object categories and scene structures. The Transformer encoder × 12 outputs the encoded features ((position) 1 + 16 × 16) × 768.
[0057] Specifically, the global semantic features of three different scales can be extracted and selected as the output of the 3rd, 9th, and 12th layer transformer encoding blocks in the Transformer encoder module, that is, Figure 2 The outputs of Encoder Block 3 (EB3: 257×768), Encoder Block 9 (EB9: 257×768), and Encoder Block 12 (EB12: 257×768) in .
[0058] In order to perform feature fusion, dual consistency of space and channel is required. Therefore, the global semantic features of three different scales output by Encoder Block 3, EncoderBlock9 and Encoder Block 12 are aligned so that the three global semantic features of different scales after alignment are respectively aligned with the local detection features of three different scales extracted by the Yolo backbone network in the traditional Yolo-World to achieve dual alignment of space and channel. Among them, the spatial alignment in the dual alignment includes dimension alignment and size alignment. Accordingly, the alignment processing method includes the following steps:
[0059] like Figure 2As shown in the figure, the vit-b / 14 features ((1+256)×768), i.e., the global semantic features of three different scales output by Encoder Block 3, Encoder Block 9, and Encoder Block 12, are first adjusted to the same dimension (16×16×768) as the local detection features by spatial structure reorganization, thus achieving dimension alignment. The channels of the dimensionally aligned global semantic features are then adjusted to the same channel (16×16×C) as the local detection features by 1×1 convolution, thus achieving channel alignment. Finally, the size of the dimension-channel aligned global semantic features is adjusted to the same size as the local detection features (resolution matching H×W×C) by upsampling, thus achieving size alignment.
[0060] Figure 2 The architecture of the Encoder Block is as follows Figure 3 As shown, the architecture of the MLP Block in the Encoder Block is as follows Figure 4 As shown, Figure 3 Layer Norm, Multi-Head Attentation, Dropout / Dropath, MLP Block and Figure 4 Linear, GELU, Droupout, and Dropout are all existing technologies and will not be introduced in detail here.
[0061] Specifically, feature fusion can be achieved by splicing. Figure 1 C3 and EB3 ( Figure 2 EB3 in: 80×80×256) for splicing, Figure 1 C4 and EB9 ( Figure 2 EB9: 40×40×512) for splicing, Figure 1 C5 and EB12 ( Figure 2 The EB12 in
[15] is spliced together, and the spliced features are input into the reparameterizable visual language path aggregation network.
[0062] As other embodiments, Figure 1 As shown in the figure, feature fusion adopts the deformable convolution feature fusion method. The processing of deformable convolution feature fusion includes:
[0063] The three local detection features of different scales extracted by the yolo backbone network in the traditional yolo-world ( Figure 1 The yolo feature is used as the input for calculating the offset in the deformable convolution, and the global semantic features of three different scales after alignment ( Figure 1 The dinov2 feature is used as the input for bilinear difference processing in the deformable convolution, and the output of the deformable convolution is spliced with the extracted local detection features of three different scales as the output of the deformable convolution feature fusion (i.e., the result of feature fusion).
[0064] The deformable convolution includes conv2d for calculating the offset. The output of conv2d for calculating the offset is the offset (x, y). Conv2d is a normal convolution layer. Figure 1 The offset (x, y) and the result of the bilinear difference processing are input together to conv2d, and the output obtained is the output result of the deformable convolution. It should be noted that, in general, the input for calculating the offset in the deformable convolution and the input for performing the bilinear difference processing are the same, but in the present invention, they are no longer the same. Instead, the local detection features extracted by the yolo backbone network in the traditional yolo-world are used as the input for calculating the offset in the deformable convolution, and the global semantic features after the alignment processing are used as the input for performing the bilinear difference processing in the deformable convolution. It is intended that while performing feature fusion, the deformable convolution can be used to enhance the edge information processing and improve the final target detection accuracy.
[0065] Specifically, illegal items include any one or any combination of cheat sheets, books, and electronic devices.
[0066] In summary, this solution adds a learnable cross-modal adapter to YOLO-World's RepVL-PAN network, uses DINOv2's vit-b / 14distilled distillation model as the backbone, extracts global semantic features (such as desktop texture and lighting changes in exam scenes), and aligns them with YOLO-World's local detection features (such as cheat sheet edges and mobile phone reflective points) in a spatial-channel dual alignment to enhance the distinction between small targets and backgrounds. DINOv2's global attention mechanism is combined with YOLO-World's real-time detection capabilities to detect illegal items such as cheat sheets, books, and electronic devices in real time. This approach is no longer limited to local detection features, but takes the surrounding environment into account, combining local detection features with global semantic features for target recognition and detection, effectively improving the detection accuracy of illegal items.
[0067] An implementation method for detecting cheating in an exam based on the fusion of local detection features and global semantic features:
[0068] A method for detecting cheating in an exam based on the fusion of local detection features and global semantic features. The specific implementation steps and effects of this method have been described in detail in an implementation of an exam cheating detection system based on the fusion of local detection features and global semantic features, and will not be repeated here.
Claims
1. A cheating detection method for exams based on the fusion of local detection features and global semantic features, characterized in that: The steps include: Obtain an image of the examination area and input it into a trained illegal product detection model to obtain a detection result of whether there are illegal products in the examination area image; The illegal product detection model is an improvement on the traditional yolo-world model, and its improvements include: Three global semantic features of different scales are extracted from the input of the detection model, and the global semantic features are aligned so that the three aligned global semantic features of different scales are respectively aligned with the three local detection features of different scales extracted by the yolo backbone network in the traditional yolo-world in terms of space and channels; then the three local detection features of different scales are respectively fused with the three aligned global semantic features of different scales, and the result of the feature fusion is input into the visual language path aggregation network in the traditional yolo-world.
2. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to claim 1, characterized in that: The global semantic features are extracted from the input of the detection model by using the image segmentation and position encoding module of dinov2 to perform image segmentation and position encoding on the input of the detection model, and then using the Transformer encoder module of dinov2 to process the output of the image segmentation and position encoding module to extract global semantic features of three different scales.
3. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to claim 2, characterized in that: The spatial alignment in the dual alignment includes dimensional alignment and size alignment. Accordingly, the alignment processing method includes the following steps: first, using spatial structure reorganization to adjust the dimensions of the global semantic features of three different scales to the same dimension as the local detection features, thereby achieving dimensional alignment; then, through 1×1 convolution, the channels of the dimensionally aligned global semantic features are adjusted to the same channel as the local detection features, thereby achieving channel alignment; finally, through upsampling, the size of the dimension-channel aligned global semantic features is adjusted to the same size as the local detection features, thereby achieving size alignment.
4. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to claim 2, characterized in that: The outputs of the 3rd, 9th, and 12th transformer encoding blocks in the Transformer encoder module are the global semantic features of the three different scales, respectively.
5. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to any one of claims 1 to 4, characterized in that: The feature fusion adopts the deformable convolution feature fusion method, and the processing process of the deformable convolution feature fusion includes: The extracted local detection features of three different scales are used as input for deformable convolution to calculate the offset, the aligned global semantic features of three different scales are used as input for bilinear difference processing in deformable convolution, and the output of deformable convolution is concatenated with the extracted local detection features of three different scales as the output of deformable convolution feature fusion.
6. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to any one of claims 1 to 4, characterized in that: Feature fusion adopts splicing method.
7. The method for detecting cheating in an exam based on the fusion of local detection features and global semantic features according to any one of claims 1 to 4, characterized in that: The test area image uses images of the test area taken at at least two different angles.
8. A system for detecting cheating in an exam based on the fusion of local detection features and global semantic features, comprising a processing terminal, a cabin for providing an exam area, and a camera installed in the cabin for acquiring an image of the exam area, wherein the processing terminal includes a processor and is characterized in that: The processor is used to execute a computer program to implement the steps of the method for detecting cheating in an examination based on the fusion of local detection features and global semantic features as described in any one of claims 1 to 7.
9. The examination cheating detection system based on the fusion of local detection features and global semantic features according to claim 8 is characterized in that: The camera is installed at at least two installation positions in the cabin to obtain images of the examination area at at least two different shooting angles.
10. The examination cheating detection system based on the fusion of local detection features and global semantic features according to claim 8 or 9, characterized in that: The installation positions of the camera include directly above, upper left rear and upper right rear of the monitored person in the examination area.
Citation Information
Patent Citations
Examination cabin for single person
CN116537601A
Unmanned invigilation system based on visual anti-cheating technology
CN117640896A
Intelligent pen test monitoring system and method based on visual inspection
CN118470645A
Cited By
Multi-mode micro-cheating behavior real-time identification system based on edge calculation
CN121545120A
Real-time identification system for multimodal micro-cheating behaviors based on edge computing
CN121545120B