A visual-aid proctoring system for offline multi-person examinations
Through the convolutional neural network assisted in the invigilator system, automated portrait positioning and cheating behavior recognition in offline multi-person exams are realized, solving the problem of labor-consuming and low accuracy in manual judgments, and improving the efficiency and accuracy of invigilator.
Patent Information
- Application Number
- CN202111607021.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-20
AI Technical Summary
In offline multi-person exams, the existing technology relies on manual judgment of cheating, which consumes a lot of human resources and is limited in accuracy. Especially in large examination environments, cameras cannot effectively identify cheating.
A visual assisted proctoring system based on convolutional neural network is adopted to automatically analyze through portrait positioning and cheating behavior recognition, using existing monitoring equipment to reduce the amount of calculation and improve accuracy.
It reduces the burden of manual invigilance, improves the accuracy and efficiency of cheating, adapts to multi-person examination environments and is not affected by light and camera angles.
Smart Images

Figure CN114267016B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of machine learning and computer vision, and specifically to a visual-aided invigilation system for offline multi-person examinations. Background Art
[0002] Currently, monitoring devices are generally installed in primary and secondary schools, universities, etc., and basically cover each classroom. These classrooms are the examination rooms for many examinations, and generally, video data will be transmitted to the examination organizers during large-scale examinations for secondary determination of cheating behavior. In actual use, the determination of cheating behavior from these video data relies entirely on manual labor, requiring a huge human resource cost. Moreover, it is greatly affected by the pixel and size of the monitor, and there may be cases of missed judgment.
[0003] In recent years, in online examinations and competitions, facial recognition and comparison of single individuals are mostly carried out through the cameras on the devices to confirm the identities of candidates and prevent cheating behaviors such as proxy examinations and multi-person examinations. Offline examinations need to be held in examination rooms. There are dozens of candidates in one examination room, and there is usually only one or two wide-angle cameras in each examination room. The optimal facial recognition range of these cameras is generally within 3 meters, and the distance from the monitoring in most examination rooms to the last candidate has exceeded 3 meters. Therefore, it is impossible to detect cheating behaviors in offline examinations through facial recognition like in online examinations. Summary of the Invention
[0004] The purpose of the present invention is to provide a technical solution to utilize the middle part of the video recorded and displayed on the monitor. Through computer-aided data analysis and accurate search for cheating behaviors, the human resource cost can be reduced and the accuracy of finding cheating behaviors can be improved.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A visual-aided invigilation system for offline multi-person examinations provided by an embodiment of the present invention is divided into two parts: portrait positioning and cheating behavior recognition. The portrait positioning part includes the following steps:
[0007] Obtain the monitoring frame images of this examination room, and cyclically obtain the monitoring frame images and execute this part at fixed time intervals;
[0008] Crop the edge pixels of the obtained frame images according to a preset size, traverse and extract multiple small images from the cropped frame images according to a fixed pixel size, and record the coordinates of each small image;
[0009] Input all the extracted small images into the pre-trained convolutional neural network 1 to determine whether the content of this small image is a portrait. If the content of this small image is determined to be a portrait by the pre-trained convolutional neural network 1, record the coordinates of this portrait image;
[0010] Concatenate the coordinates of all adjacent portrait images to synthesize and record large coordinates, and delete the coordinates before concatenation. If there are no adjacent portrait images, use their original coordinates. Return all the processed coordinates (hereinafter referred to as coordinate data).
[0011] As a possible implementation of this embodiment, at least one monitoring device should be installed in the examination room, and the monitoring range should cover every student in the examination room.
[0012] As a possible implementation of this embodiment, the loop to obtain monitoring frame images and execute this part at fixed time intervals means that within the examination time, this part is run at regular intervals, such as every 1 minute, every 5 minutes, etc. Each time this part is run, the coordinate data is updated. When this part is not run, the coordinate data obtained during the last run is used.
[0013] As a possible implementation of this embodiment, the calculation method of the edge pixels is as follows:
[0014]
[0015] Among them, x represents the length or width of the frame image; y represents the side length of the fixed pixels used to extract small images; z is the pixel size that needs to be additionally cropped. The output of this formula is the distance from the upper left starting point of the cropped image to the upper side and the left side of the original frame image. The units are all pixels.
[0016] As a possible implementation of this embodiment, the fixed pixel size refers to a square pixel block with a smaller size, such as pixel blocks with sizes of 15*15 pixels, 30*30 pixels, and 50*50 pixels. In a normal examination room environment, in a 1920*1080 pixel image, the size of this pixel block should not be greater than 50*50 pixels.
[0017] As a possible implementation of this embodiment, the coordinates refer to a tensor with the shape of (X, Y, W, H) or a tensor with the shape of (X, Y, L). Among them, X and Y represent the upper left pixel coordinates of this pixel block, W is the pixel width, H is the pixel height, and L is the pixel side length. The units are all pixels.
[0018] As a possible implementation of this embodiment, the pre-trained convolutional neural network 1 internally has three convolutional layers, a max pooling layer, a flattening layer, a regularization layer, and two fully connected layers. The specific network parameters are determined by the training and debugging results of the corresponding training samples.
[0019] As a possible implementation of this embodiment, the adjacent portrait images refer to that if the ordinate or abscissa of the starting point in the upper left corner of a portrait image, plus or minus a fixed pixel side length, is equal to the ordinate or abscissa of the starting point in the upper left corner of another portrait image, then these two portrait images are adjacent portrait images.
[0020] The cheating behavior recognition part of a visual auxiliary invigilation system for offline multi-person examinations provided by an embodiment of the present invention includes the following steps:
[0021] Obtain the monitoring frame images in real time at the video frame rate or interval and execute this part;
[0022] Obtain the coordinate data of the portrait positioning part, and extract the images from the frame images according to this coordinate data;
[0023] Input all the extracted images into the pre-trained convolutional neural network 2 to determine whether the content of this image is a cheating behavior. If the content of this image is determined by the pre-trained convolutional neural network 2 to be a cheating behavior, record the time at this moment and the coordinates of this image, and display a red rectangular prompt box on the monitor according to the coordinates of this image, for prompting the invigilator to check this place.
[0024] As a possible implementation of this embodiment, the frame extraction at intervals means to execute once every certain number of frames. For example, if the existing video frame rate is 30 frames per second and it is executed once every 10 frames, then it is only executed 3 times in 1 second.
[0025] As a possible implementation of this embodiment, the pre-trained convolutional neural network 2 internally has three convolutional layers and max pooling layers, a flattening layer, a regularization layer, and two fully connected layers. The specific network parameters are determined by the training and debugging results of the corresponding training samples.
[0026] The technical solution of the embodiment of the present invention can have the following beneficial effects:
[0027] When the present invention has enough sample data trained for this examination room or uses the sample data of a similar examination room, it can adapt to the environments of most examination rooms, can use the original monitoring cameras, and is almost not affected by light and the angle of the monitoring cameras; on this basis, a visual auxiliary invigilation system for offline multi-person examinations is proposed, which involves two parts: candidate positioning and cheating behavior judgment. This system converts the complex video data recognition problem into two binary classification problems, greatly reducing the calculation amount, making better use of the existing equipment, improving the invigilation efficiency, and reducing the burden of manual invigilation. Description of the Drawings
[0028] Figure 1It is a flowchart of a visual-aid proctoring system for offline multi-person examinations shown according to an exemplary embodiment;
[0029] Figure 2 It is a schematic diagram (top view) of monitoring at 6 different angles in the examination room;
[0030] Figure 3 It is a schematic diagram of video frame processing;
[0031] Figure 4 It is a schematic diagram for determining adjacent portrait images (in coordinate form). Detailed implementation
[0032] The present invention will be further described below in conjunction with the drawings and embodiments:
[0033] To clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific implementation manners and in conjunction with its drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. In addition, the present invention may repeatedly refer to numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present invention omits the description of well-known components and processing technologies and processes to avoid unnecessarily limiting the present invention.
[0034] Figure 1 It is a flowchart of a visual-aid proctoring system for offline multi-person examinations shown according to an exemplary embodiment. As Figure 1 shown, a visual-aid proctoring system for offline multi-person examinations provided by an embodiment of the present invention includes the following steps:
[0035] The first part is the portrait positioning part, including the following steps:
[0036] 1. Obtain the monitoring frame of this examination room, and cyclically obtain the monitoring frame and execute this part at fixed time intervals.
[0037] As a possible implementation manner of this embodiment, at least one monitoring device should be installed in the examination room, and its position should be the same as that of any one of the 6 monitors shown in Figure 2 . The monitoring range should cover every student in the examination room.
[0038] Cyclically obtaining the monitoring frame and executing this part at fixed time intervals means that within the examination time, this part is run once every certain time interval, such as an interval of 1 minute, an interval of 5 minutes, etc.
[0039] 2. As shown in Figure 3 , crop the edge pixels of the obtained frame image according to a preset size, traverse and extract multiple small images from the cropped frame image according to a fixed pixel size, and record the coordinates of each small image.
[0040] In the figure, 1 is the cropped edge pixel area, and 2 is the extracted small image.
[0041] The calculation method of the edge pixels is as follows:
[0042]
[0043] Among them, x represents the length or width of the frame image; y represents the side length of the fixed pixel used to extract the small image; z is the pixel size that needs to be additionally cropped. The output of this formula is the distance between the upper left starting point of the cropped image and the upper side and left side of the original frame image. The units are all pixels.
[0044] Cropping is not a real image cropping, but according to the result of the above formula, the next operation is carried out at a certain distance from the left side and the upper side of the original frame image.
[0045] The sum of the areas of all small images should be equal to the area of the cropped frame image. The fixed pixel size refers to a small square pixel block, such as a pixel block with a size of 15*15 pixels, 30*30 pixels, or 50*50 pixels. In a normal examination room environment, in a standard 1920*1080 pixel image, the size of this pixel block should not be greater than 50*50 pixels.
[0046] The coordinates refer to a tensor with the shape of (X, Y, W, H) or a tensor with the shape of (X, Y, L). Among them, X and Y represent the upper left pixel coordinates of this pixel block, W is the pixel width, H is the pixel height, and L is the pixel side length. The units are all pixels.
[0047] 3. Input all the extracted small images into the pre-trained convolutional neural network 1 to determine whether the content of this small image is a portrait. If the content of this small image is determined to be a portrait by the pre-trained convolutional neural network 1, record the coordinates of this portrait image.
[0048] As a possible implementation manner of this embodiment, there should be enough sample data of this examination room or use the sample data of a similar examination room to train the convolutional neural network 1. The pre-trained convolutional neural network 1 has three convolutional layers and max-pooling layers, a flattening layer, a regularization layer, and two fully connected layers inside. The specific network parameters are determined by the training and debugging results of the corresponding training samples.
[0049] 4. Concatenate the coordinates of all adjacent portrait images to synthesize and record large coordinates, and delete the coordinates before concatenation. If there are no adjacent portrait images, use their original coordinates. Return all the processed coordinates (hereinafter referred to as coordinate data).
[0050] The determination method of adjacent portrait images is as Figure 4 shown. Each square in the figure represents the coordinates of the upper left vertex of this square in the portrait coordinates, and the format is (X, Y); the side length of all squares is 15. Among them, 3, 4, and 5 are portrait images. Since the difference in the X coordinates between 3 and 4 is 15, and the difference in the Y coordinates between 3 and 5 is 15, it is determined that 3 and 4, and 3 and 5 are adjacent portrait coordinates.
[0051] The second part is the cheating behavior recognition part, including the following steps:
[0052] 1. Obtain the monitoring frame images in real time by taking frames at the video frame rate or interval and execute this part.
[0053] Taking frames at intervals means executing once every certain number of frames. For example, if the existing video frame rate is 30 frames per second and it is executed once every 10 frames, it will only be executed 3 times in 1 second.
[0054] 2. Obtain the coordinate data of the portrait positioning part, and extract images from the frame images according to this coordinate data.
[0055] 3. Input all the extracted images into the pre-trained convolutional neural network 2 to determine whether the content of this image is a cheating behavior. If the content of this image is determined to be a cheating behavior by the pre-trained convolutional neural network 2, record the time at this time and the coordinates of this image, and display a red rectangular prompt box on the monitor according to the coordinates of this image to prompt the invigilator to check this place.
[0056] As a possible implementation of this embodiment, there should be enough sample data of this examination room or use the sample data of similar examination rooms to train the convolutional neural network 2. The pre-trained convolutional neural network 2 has three convolutional layers and max pooling layers, a flattening layer, a regularization layer, and two fully connected layers inside. The specific network parameters are determined by the training and debugging results of the corresponding training samples.
[0057] The present invention proposes a visual assistant invigilation system for offline multi-person examinations based on a convolutional neural network for offline examinations. Among them, 2 convolutional neural networks are applied in the system, and their network architectures are the same, and the specific parameters are determined by the training and debugging results of different training samples. The convolutional neural network has excellent performance in processing computer vision data. This system converts the complex video data recognition problem into two binary classification problems, greatly reducing the computational complexity. In the case of having enough sample data trained for this examination room or using sample data from similar examination rooms, it can adapt to the environments of most examination rooms, can use the original monitoring cameras, and is hardly affected by light and the angles of the monitoring cameras, and is an effective visual assistant invigilation system.
[0058] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a machine for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or blocks.
[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the processes Figure 1 or blocks.
[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or blocks.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A visual-aided invigilation method for offline multi-person examinations, characterized in that, It includes the following steps: Obtain the monitoring frame images of this examination room, and cyclically obtain the monitoring frame images at fixed time intervals; Crop the edge pixels of the obtained frame images according to a preset size, traverse and extract multiple small images from the cropped frame images according to a fixed pixel size, and record the coordinates of each small image; Input all the extracted small images into the pre-trained convolutional neural network 1 to determine whether the content of this small image is a portrait; if the content of this small image is determined to be a portrait by the pre-trained convolutional neural network 1, record the coordinates of this portrait image; Stitch the coordinates of all adjacent portrait images, synthesize and record the large coordinates, and delete the coordinates before stitching; if there are no adjacent portrait images, use its original coordinates; return all the processed coordinate data; Obtain the monitoring frame images in real time by taking frames at the video frame rate or interval; Obtain the coordinate data of the portrait images, and extract the images from the frame images according to this coordinate data; Input all the extracted images into the pre-trained convolutional neural network 2 to determine whether the content of this image is a cheating behavior; If the content of this image is determined to be a cheating behavior by the pre-trained convolutional neural network 2, record the time at this time and the coordinates of this image, and display a red rectangular prompt box on the monitor according to the coordinates of this image to prompt the invigilator to check this place.
2. The method according to claim 1, characterized in that, During the examination time, this part is run once every certain time interval; each time this part is run, the coordinate data is updated; when this part is not run, the coordinate data obtained during the previous run is used.
3. The method according to claim 1, wherein The calculation method of the edge pixels is: Where, x represents the length or width of the frame image; y represents the side length of the fixed pixel used to extract the small image; z is the pixel size that needs to be additionally cropped; the output of this formula is the distance between the upper left starting point of the cropped image and the upper side and left side of the original frame image; the unit is all pixels.
4. The method according to claim 1, wherein [[ID= 5. The method according to claim 1, characterized in that, 6. The method according to claim 1, characterized in that, 7. The method according to claim 1, wherein 8. The method according to claim 1, characterized in that,
Citation Information
Patent Citations
Examination cheating behavior detection method and system in standard examination room environment
CN112036299A
Online examination anti-cheating implementation method based on AI face recognition technology
CN113657300A