Video operation and maintenance analysis system based on artificial intelligence face capture
The video monitoring and analysis system, which uses artificial intelligence to capture facial images, collects and analyzes candidates' facial images in real time. This solves the problem that existing invigilation systems cannot identify abnormal behavior, enabling timely identification of candidates and early warning of abnormal behavior, thereby improving invigilation efficiency and the fairness of the examination.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU LUOXIANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-06-10
- Publication Date
- 2026-07-21
AI Technical Summary
The existing proctoring system cannot identify and alert test takers with abnormal behavior through video analysis, resulting in a high degree of dependence on proctors and an inability to effectively prevent cheating.
A video operation and maintenance analysis system based on artificial intelligence facial capture is adopted, including a verification layer, a monitoring layer and an alert layer. By collecting candidates' facial images in real time, the system performs identity verification and facial spatial position change morphology analysis, and issues early warning prompts.
It enables accurate identification of candidates and timely detection of abnormal behavior, reducing the workload of invigilators and ensuring the fairness, impartiality and seriousness of the examination.
Smart Images

Figure CN120673337B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a video operation and maintenance analysis system based on artificial intelligence facial capture. Background Technology
[0002] Surveillance technology is widely used in examination invigilation. By installing cameras and other equipment in the examination room, it can monitor the examinees' examination status in real time, effectively preventing cheating and ensuring the fairness and impartiality of the examination. At the same time, its recording function can also provide evidence for subsequent review.
[0003] Patent application No. 202011558036.3 discloses a verification method for a remote proctoring system. The remote proctoring system includes a first smart terminal, a second smart terminal, and an examination terminal. The method includes: acquiring a verification command input by a candidate; responding to the verification command, controlling the first smart terminal, the second smart terminal, and the examination terminal to display verification information according to preset rules, and controlling the first smart terminal and the second smart terminal to acquire verification images; determining whether the first smart terminal, the second smart terminal, and the examination terminal meet preset position rules based on the verification images and the verification information; if the preset position rules are met, confirming that the remote proctoring system has passed verification. This application aims to solve the problem that "existing remote proctoring systems mainly collect examination videos from the examination room through cameras and monitor whether the candidates' examination behavior is compliant based on the video content. However, in practical applications, candidates often deliberately damage or affect the remote proctoring system, such as deliberately moving or damaging the camera, causing the remote proctoring system to malfunction and affecting the smooth conduct of the examination."
[0004] However, in examination invigilation scenarios, monitoring technology is currently still used as an auxiliary device to assist invigilators in carrying out their work. It cannot identify and warn examinees with abnormal behavior through image analysis, resulting in a high degree of dependence on invigilators in the examination room.
[0005] To address this, we propose a video operation and maintenance analysis system based on artificial intelligence facial capture. Summary of the Invention
[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a video operation and maintenance analysis system based on artificial intelligence facial capture, which can effectively solve the problems of the existing technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions;
[0008] This invention discloses a video operation and maintenance analysis system based on artificial intelligence facial capture, comprising: a verification layer, a monitoring layer, and an alert layer;
[0009] The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer simultaneously compares the candidates' faces with the collected face images in the acquired candidate identity information to complete the candidate face verification. The monitoring layer collects examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the changes in the spatial position of the candidates' faces based on the captured face images. The warning layer receives the changes in the spatial position of the candidates' faces from the monitoring layer, confirms the warning targets based on the changes in the spatial position of the candidates' faces, and issues warning prompts based on the warning targets.
[0010] The monitoring layer includes a data acquisition module, a configuration module, and an analysis module. The data acquisition module is used to acquire examination room images in real time. The configuration module is used to configure the candidate's facial image capture frame in the examination room images. The analysis module is used to segment and obtain the examination room images corresponding to each candidate based on the candidate's facial image capture frame in the examination room images, pick up the screen frames in the examination room images corresponding to each candidate, and analyze the spatial position change shape of each candidate's face based on the picked screen frames.
[0011] Using the frame of the same candidate in the examination room as the target for analyzing the spatial position change of the candidate's face, the candidate's outline image in each frame is identified. Each frame is segmented based on the vertical and horizontal midlines, so that each frame is divided into four sub-frames. Each sub-frame is marked with the upper left, upper right, lower left and lower right labels. The sub-frames are distinguished based on the labels to obtain four sets of sub-frames.
[0012] ;
[0013] In the formula: Parameters representing the morphological changes in the spatial position of the examinee's face; This indicates the set of sub-frames marked as top left and bottom left; , Let be the area of the candidate's face image in the s-th sub-frame, and the area of the area outside the candidate's face image in the s-th sub-frame.
[0014] Among them, the facial spatial position transformation morphological representation parameters of the examinee are calculated based on the above formula, respectively, using the upper left and lower left; upper left and upper right; upper right and lower right; and lower right and lower left as denoted as Based on the continuous acquisition of examination room images to complete frame picking and updating, we have:
[0015] This is recorded as the result of the morphological analysis of the spatial position changes of the faces of historical test takers.
[0016] Furthermore, the verification layer includes a camera module, an upload module, and a verification module. The camera module is used to capture a global image of the examinee's face in the examination room, the upload module is used to upload the examinee's identity information, and the verification module is used to receive the global image of the examinee's face in the examination room captured by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison between the examinee's identity information and the global image of the examinee's face in the examination room, the verification module verifies whether the examinee in the examination room is correct.
[0017] The upload module uploads candidate identity information including: admission ticket number, seat number, candidate ID number, and candidate facial image. After candidates take their seats in the examination room according to their seat numbers, the camera module captures a global image of the candidates' faces in the examination room, displaying all candidates' facial images in the same image. After the global image of the candidates' faces in the examination room is captured, an image segmentation operation is performed simultaneously, so that each segmented sub-image contains only one candidate's facial image. Based on the sub-images, a set is formed. The candidate's facial image is obtained from the candidate identity information uploaded in the upload module, and the obtained candidate facial images are used as a set. The two sets are sent to the verification module.
[0018] Furthermore, the verification logic for whether the examinee in the examination room is correct in the verification module is represented as follows:
[0019] ;
[0020] In the formula: This is the determination value; For decision functions; , The total number of candidate face images obtained based on global image segmentation of candidate faces, and the total number of candidate face images in candidate identity information; Let be the similarity between the i-th candidate's facial image and the j-th candidate's facial image; The number of candidates is determined based on the number of seats in the examination room or the number of candidates' identity information. This refers to the number of students who were absent from the exam.
[0021] In cases where there are absentees in the examination room, the assembly will take place. , The corresponding number of candidates is reduced synchronously. When formula (2) is true, the verification module verifies that the candidates in the examination room are correct; otherwise, it is incorrect, and the judgment function is used. If the condition within the parentheses is true, then =1, otherwise take zero.
[0022] Furthermore, when the verification module performs a verification that results in a negative result, it obtains the result in the determination function. The candidate facial images in the set that do not make the determination function true are selected as the candidates whose facial images are sourced from. The candidates whose facial images are sourced from these candidates are then used as the offline processing targets for the invigilators. The invigilators verify the candidates' identity information offline again. If the verification is successful, the process jumps to the monitoring layer. If the verification fails, the candidates whose facial images are sourced from these candidates are excluded from the monitoring layer's monitoring scope.
[0023] Furthermore, the logic for calculating the similarity of the candidate's facial images is as follows: For example:
[0024] ;
[0025] In the formula: The dynamic weights of facial geometric features, texture features, and deep learning features are calculated based on the attention mechanism; The total number of facial feature key points; Let be the Euclidean distance between the v-th corresponding key point in candidate face image i and candidate face image j; These are normalization parameters; The texture feature vectors extracted from candidate face image i and candidate face image j using local binary mode; Let be the magnitude of the vector; The high-dimensional deep learning feature vectors extracted from candidate face image i and candidate face image j using a pre-trained face recognition model; Let be the magnitude of the vector;
[0026] in, ∈[0,1], The sum is 1, and all values are positive, and are user-defined on the system side.
[0027] Furthermore, the aforementioned The value of follows:
[0028] An attention network is constructed, taking facial geometric feature differences, texture feature distribution entropy, and intra-class scatter of deep learning features as inputs, and the weights are calculated by a multilayer perceptron.
[0029] ;
[0030] In the formula: This is used to measure the distance difference of facial key points, the difference in entropy of texture feature distribution, and the difference in intra-class discreteness of deep learning features. These are the multilayer perceptron mapping functions for the corresponding features; This is an index used to iterate through different feature types.
[0031] Furthermore, the acquisition module is integrated into the camera module in the verification layer. The examination room images are acquired based on the camera module. When the configuration module configures the candidate's face image capture frame in the examination room image, it divides the area contained in the examination room image equally based on each capture frame, and the area corresponding to each capture frame is the same size. In addition, the examination table is centered in each capture frame. When the analysis module picks up the screen frame in the examination room image, the time interval between each adjacent screen frame is equal.
[0032] Furthermore, the warning layer includes a judgment module, a warning module, and a visualization module. The judgment module is used to receive the analysis results of the changes in the spatial position of the candidates' faces in the historical monitoring layer, and to determine whether the candidates have abnormal behavior based on the analysis results. The warning module is used to receive the judgment result of whether the candidates have abnormal behavior in the judgment module, and to issue a warning voice when the judgment result is yes. The visualization module is used to create a layer on the examination room image, place the candidate's face image capture frame on the surface of the created layer, obtain the judgment result of whether the candidates have abnormal behavior in the judgment module, and to render the candidate's face image capture frame with a specified color transparently for candidates whose judgment result is yes.
[0033] The logic for determining whether a candidate exhibits abnormal behavior within the judgment module is as follows:
[0034] Obtain the morphological representation parameters of the candidate's facial spatial position transformation from the latest three calculations in the combined results of the labeled directions for each sub-image, namely:
[0035] ;
[0036] In the formula, x represents the calculation sequence number of the facial spatial position change morphological representation parameter of the examinee. If any set of the facial spatial position change morphological representation parameters of the examinee obtained from the three most recent calculations above shows a continuous upward trend, it is determined that the examinee has abnormal behavior.
[0037] Furthermore, the warning module is integrated with a speaker, which has a preset warning voice message. The format of the warning voice message is: "Please abide by the examination rules for candidate number XX". In the message "Please abide by the examination rules for candidate number XX", "XX" is the seat number in the candidate's identity information.
[0038] The visualization module is integrated with computer equipment that has display capabilities. The computer equipment displays real-time rendered images of the examination room for invigilators to read visually.
[0039] Furthermore, the acquisition module is interconnected with a configuration module and an analysis module via a wireless network. The acquisition module is interconnected with a camera module via a wireless network. The camera module is interconnected with an upload module and a verification module via a wireless network. The analysis module is interconnected with a judgment module via a wireless network. The judgment module is interconnected with an alert module and a visualization module via a wireless network.
[0040] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects:
[0041] 1. The system in this invention can globally collect the facial images of candidates in the same examination room and compare and verify them with the facial images in the candidates' identity information. After the candidates are seated according to their seat numbers, the collected facial images are segmented and matched and verified with the identity information. This can accurately identify the candidates' identities, effectively prevent cheating behaviors such as impersonation, maintain the order of the examination room, ensure the fairness and impartiality of the examination, and provide a solid guarantee for the smooth conduct of the examination.
[0042] 2. In this invention, the system collects examination room images in real time and monitors behavior by analyzing the changes in the spatial position of the examinee's face. It picks up and analyzes image frames at fixed intervals. Once it detects abnormalities in the parameters of the examinee's facial position change, such as a continuous increase, it can promptly detect abnormal behavior. This helps invigilators to detect examinee's violations as soon as possible, stop them in time, and ensure the seriousness of the examination environment.
[0043] 3. In this invention, when the system determines that a candidate has abnormal behavior, it will issue a targeted warning with a preset voice prompt. At the same time, the capture frame of the abnormal candidate will be rendered transparently in a specified color on the image. This allows invigilators to quickly locate the abnormal candidate and identify the violator, which greatly improves invigilation efficiency and reduces the workload of invigilators. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0045] Figure 1 This is a schematic diagram of a video operation and maintenance analysis system based on artificial intelligence facial capture. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0047] The present invention will be further described below with reference to embodiments.
[0048] Example:
[0049] This embodiment presents a video operation and maintenance analysis system based on artificial intelligence facial capture, such as... Figure 1 As shown, it includes: a verification layer, a monitoring layer, and an alert layer;
[0050] The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer simultaneously compares the candidates' faces with the collected face images in the acquired candidate identity information to complete the candidate face verification. The monitoring layer collects examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the changes in the spatial position of the candidates' faces based on the captured face images. The warning layer receives the changes in the spatial position of the candidates' faces from the monitoring layer, confirms the warning targets based on the changes in the spatial position of the candidates' faces, and issues warning prompts based on the warning targets.
[0051] The verification layer includes a camera module, an upload module, and a verification module. The camera module is used to capture a global image of the examinee's face in the examination room. The upload module is used to upload the examinee's identity information. The verification module is used to receive the global image of the examinee's face in the examination room captured by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison between the examinee's identity information and the global image of the examinee's face in the examination room, the verification module verifies whether the examinee in the examination room is correct.
[0052] The candidate identity information uploaded in the upload module includes: admission ticket number, seat number, candidate ID number, and candidate facial image. After candidates take their seats in the examination room according to their seat numbers, the camera module captures a global image of the candidates' faces in the examination room, so that all candidates' facial images are displayed in the same image. After the global image of the candidates' faces in the examination room is captured, an image segmentation operation is performed simultaneously so that each segmented sub-image contains only one candidate's facial image. Based on the sub-images, a set is formed. The candidate's facial image is obtained from the candidate identity information uploaded in the upload module, and the obtained candidate facial images are used as a set. The two sets are sent to the verification module.
[0053] The verification logic for checking whether the examinee in the examination room is correct in the verification module is represented as follows:
[0054] ;
[0055] In the formula: This is the determination value; For decision functions; , The total number of candidate face images obtained based on global image segmentation of candidate faces, and the total number of candidate face images in candidate identity information; Let be the similarity between the i-th candidate's facial image and the j-th candidate's facial image; The number of candidates is determined based on the number of seats in the examination room or the number of candidates' identity information. This refers to the number of students who were absent from the exam.
[0056] In cases where there are absentees in the examination room, the assembly will take place. , The corresponding number of candidates is reduced synchronously. When formula (2) is true, the verification module verifies that the candidates in the examination room are correct; otherwise, it is incorrect, and the judgment function is used. If the condition within the parentheses is true, then =1, otherwise take zero;
[0057] The above logical formula is used to perform a one-time accurate verification of the candidates' identities in the examination room, which largely replaces the invigilators' work of verifying candidates' identities and is more accurate than manual verification.
[0058] When the verification module returns a negative result, the result is obtained in the decision function. The candidate facial images in the set that do not make the judgment function true are selected as the candidates whose facial images are sourced by the candidates. The candidates whose facial images are sourced are used as the offline processing targets of the invigilators. The invigilators verify the candidates' identity information offline again. If the verification is successful, the process jumps to the monitoring layer. If the verification fails, the candidates whose facial images are sourced are excluded from the monitoring scope of the monitoring layer.
[0059] The logic for calculating the similarity of candidates' facial images is as follows: For example:
[0060] ;
[0061] In the formula: The dynamic weights of facial geometric features, texture features, and deep learning features are calculated based on the attention mechanism; The total number of facial feature key points; Let be the Euclidean distance between the v-th corresponding key point in candidate face image i and candidate face image j; These are normalization parameters; The texture feature vectors extracted from candidate face image i and candidate face image j using local binary mode; Let be the magnitude of the vector; The high-dimensional deep learning feature vectors extracted from candidate face image i and candidate face image j using a pre-trained face recognition model; Let be the magnitude of the vector;
[0062] in, ∈[0,1], The sum is 1, and all values are positive, and are user-defined on the system side;
[0063] The value of follows:
[0064] An attention network is constructed, taking facial geometric feature differences, texture feature distribution entropy, and intra-class scatter of deep learning features as inputs, and the weights are calculated by a multilayer perceptron.
[0065] ;
[0066] In the formula: This is used to measure the distance difference of facial key points, the difference in entropy of texture feature distribution, and the difference in intra-class discreteness of deep learning features. These are the multilayer perceptron mapping functions for the corresponding features; This serves as an index for iterating through different feature types;
[0067] in express The range of index values is a set; when calculating the weight of each feature, the denominator is... By analyzing j in The values are iterated through to calculate the sum of the exponential function results of the output values of the multilayer perceptron (MLP) corresponding to different features. This sum is then combined with the numerator and the principle of the Softmax function is used to convert the MLP output of each feature into a probabilistic weight. This allows for the dynamic allocation of the importance of each feature in the overall similarity calculation through an attention mechanism.
[0068] The above logical formula provides support for verifying whether the candidates in the examination room are correct.
[0069] The monitoring layer includes an acquisition module, a configuration module, and an analysis module. The acquisition module is used to acquire examination room images in real time. The configuration module is used to configure the candidate's face image capture frame in the examination room image. The analysis module is used to segment and obtain the examination room image corresponding to each candidate based on the candidate's face image capture frame in the examination room image, pick up the screen frame in the examination room image corresponding to the candidate, and analyze the change shape of the facial spatial position of each candidate based on the picked screen frame.
[0070] Using the frame of the same candidate in the examination room as the target for analyzing the spatial position change of the candidate's face, the candidate's outline image in each frame is identified. Each frame is segmented based on the vertical and horizontal midlines, so that each frame is divided into four sub-frames. Each sub-frame is marked with the upper left, upper right, lower left and lower right labels. The sub-frames are distinguished based on the labels to obtain four sets of sub-frames.
[0071] ;
[0072] In the formula: Parameters representing the morphological changes in the spatial position of the examinee's face; This indicates the set of sub-frames marked as top left and bottom left; , Let be the area of the candidate's face image in the s-th sub-frame, and the area of the area outside the candidate's face image in the s-th sub-frame.
[0073] Among them, the facial spatial position transformation morphological representation parameters of the examinee are calculated based on the above formula, respectively, using the upper left and lower left; upper left and upper right; upper right and lower right; and lower right and lower left as denoted as Based on the continuous acquisition of examination room images to complete frame picking and updating, we have:
[0074] Record the results of the facial spatial position transformation morphology analysis of the historical examinee;
[0075] The above formula is used to calculate and analyze the changes in the spatial position of each candidate's face, providing support for the further operation of the warning layer of the system in this embodiment.
[0076] Within the camera module of the acquisition module integration and verification layer, the examination room images are acquired based on the camera module. When the configuration module configures the candidate's face image capture frame in the examination room image, it divides the area contained in the examination room image equally based on each capture frame, and the area corresponding to each capture frame is the same size. In addition, the examination table is centered in each capture frame. When the analysis module picks up the frame in the examination room image, the time interval between each adjacent frame is equal.
[0077] The warning layer includes a judgment module, a warning module, and a visualization module. The judgment module receives the analysis results of the changes in the spatial position of the candidates' faces in the historical monitoring layer and determines whether the candidates have abnormal behavior based on the analysis results. The warning module receives the judgment result of whether the candidates have abnormal behavior in the judgment module and issues a warning voice when the judgment result is yes. The visualization module creates a layer on the examination room image, places the candidate's face image capture frame on the created layer surface, obtains the judgment result of whether the candidates have abnormal behavior in the judgment module, and renders the candidate's face image capture frame with a specified color for candidates whose judgment result is yes.
[0078] The logic for determining whether a candidate exhibits abnormal behavior within the judgment module is as follows:
[0079] Obtain the morphological representation parameters of the candidate's facial spatial position transformation from the latest three calculations in the combined results of the labeled directions for each sub-image, namely:
[0080] ;
[0081] In the formula, x represents the calculation sequence number of the facial spatial position change morphological representation parameter of the examinee. If any set of the facial spatial position change morphological representation parameters of the examinee obtained from the three most recent calculations above shows a continuous upward trend, it is determined that the examinee has abnormal behavior.
[0082] The warning module is integrated into the speaker, which has a preset warning voice. The format of the warning voice is: "Please abide by the examination rules for candidate number XX". In "Please abide by the examination rules for candidate number XX", "XX" is the seat number in the candidate's identity information.
[0083] The visualization module is integrated with computer equipment that has display capabilities. The computer equipment displays real-time rendered images of the examination room for invigilators to read visually.
[0084] The acquisition module interacts with the configuration module and analysis module via a wireless network. The acquisition module interacts with the camera module via a wireless network. The camera module interacts with the camera module via a wireless network. The camera module interacts with the camera module via a wireless network. The analysis module interacts with the judgment module via a wireless network. The judgment module interacts with the camera module via a wireless network. The camera ...
[0085] In this embodiment, the camera module collects a global image of the examinee's face within the examination room, while the upload module simultaneously uploads the examinee's identity information. The verification module further receives the global image of the examinee's face within the examination room collected by the camera module, as well as the examinee's identity information uploaded by the upload module. Based on the comparison between the examinee's identity information and the global image of the examinee's face within the examination room, the module verifies whether the examinee is correct. The acquisition module collects examination room images in real time, and the configuration module further configures the examinee's face image capture frame in the examination room images. The analysis module then segments the examination room images based on the examinee's face image capture frame to obtain the examination room image corresponding to each examinee. Frames are picked from the corresponding examination room images of each examinee, and based on the picked frames... The system analyzes the spatial changes in the facial positions of each examinee. The judgment module then receives the analysis results of historical examinee facial spatial changes from the monitoring layer. Based on these results, it determines whether an examinee exhibits any abnormal behavior. Simultaneously, the warning module receives the judgment result from the judgment module. If the judgment result is yes, it issues a warning voice. Finally, the visualization module creates a layer on the examination room image and places the examinee's facial image capture frame on the created layer surface. It then obtains the judgment result from the judgment module regarding whether an examinee exhibits any abnormal behavior. For examinees with a yes judgment result, the corresponding facial image capture frame is rendered transparently with a specified color.
[0086] Through the system operation in the above embodiments, visual detection technology brings a brand-new, more intelligent monitoring system to the invigilation scene, which can largely replace the invigilation work of invigilators and monitor examinees with abnormal behavior in the examination room.
[0087] In summary, the system described above can globally capture the facial images of candidates within the same examination room and compare them with the facial images in the candidates' identity information. After candidates are seated according to their seat numbers, the captured facial images are segmented and matched with their identity information for verification. This accurately identifies candidates, effectively prevents cheating behaviors such as impersonation, maintains examination room order, ensures fairness and impartiality in the examination, and provides a solid guarantee for the smooth conduct of the examination. Simultaneously, the system can capture examination room images in real time and monitor behavior by analyzing the changes in the spatial position of candidates' faces. It picks up and analyzes image frames at fixed intervals. Once abnormal facial position change parameters are detected, such as a continuous increase, abnormal behavior can be detected in a timely manner. This helps invigilators to discover candidates' violations immediately, stop them promptly, and ensure the seriousness of the examination environment. Furthermore, when the system determines that a candidate has exhibited abnormal behavior, it will issue a targeted warning with a preset voice prompt and render the capture frame of the abnormal candidate transparently in a specified color on the image. This allows invigilators to quickly locate the abnormal candidate and identify the violator, greatly improving invigilation efficiency and reducing the workload of invigilators.
[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video operation and maintenance analysis system based on artificial intelligence facial capture, characterized in that, include: Verification layer, monitoring layer, and alert layer; The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer simultaneously compares the candidates' faces with the collected face images in the acquired candidate identity information to complete the candidate face verification. The monitoring layer collects examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the changes in the spatial position of the candidates' faces based on the captured face images. The warning layer receives the changes in the spatial position of the candidates' faces from the monitoring layer, confirms the warning targets based on the changes in the spatial position of the candidates' faces, and issues warning prompts based on the warning targets. The monitoring layer includes a data acquisition module, a configuration module, and an analysis module. The data acquisition module is used to acquire examination room images in real time. The configuration module is used to configure the candidate's facial image capture frame in the examination room images. The analysis module is used to segment and obtain the examination room images corresponding to each candidate based on the candidate's facial image capture frame in the examination room images, pick up the screen frames in the examination room images corresponding to each candidate, and analyze the spatial position change shape of each candidate's face based on the picked screen frames. Using the frame of the same candidate in the examination room as the target for analyzing the spatial position change of the candidate's face, the candidate's outline image in each frame is identified. Each frame is segmented based on the vertical and horizontal midlines, so that each frame is divided into four sub-frames. Each sub-frame is marked with the upper left, upper right, lower left and lower right labels. The sub-frames are distinguished based on the labels to obtain four sets of sub-frames. ; In the formula: Parameters representing the spatial positional changes of the examinee's face; This indicates the set of sub-frames marked as top left and bottom left; , Let be the area of the candidate's face image in the s-th sub-frame, and the area of the area outside the candidate's face image in the s-th sub-frame. Among them, the facial spatial position transformation morphological representation parameters of the examinee are calculated based on the above formula, respectively, using the upper left and lower left; upper left and upper right; upper right and lower right; and lower right and lower left as denoted as Based on the continuous acquisition of examination room images to complete frame picking and updating, we have: This is recorded as the result of the morphological analysis of the spatial position changes of the faces of historical test takers.
2. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1, characterized in that, The verification layer includes a camera module, an upload module, and a verification module. The camera module is used to collect a global image of the examinee's face in the examination room. The upload module is used to upload the examinee's identity information. The verification module is used to receive the global image of the examinee's face in the examination room collected by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison between the examinee's identity information and the global image of the examinee's face in the examination room, the verification module verifies whether the examinee in the examination room is correct. The upload module uploads candidate identity information including: admission ticket number, seat number, candidate ID number, and candidate facial image. After candidates take their seats in the examination room according to their seat numbers, the camera module captures a global image of the candidates' faces in the examination room, displaying all candidates' facial images in the same image. After the global image of the candidates' faces in the examination room is captured, an image segmentation operation is performed simultaneously, so that each segmented sub-image contains only one candidate's facial image. Based on the sub-images, a set is formed. The candidate's facial image is obtained from the candidate identity information uploaded in the upload module, and the obtained candidate facial images are used as a set. The two sets are sent to the verification module.
3. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 2, characterized in that, The verification logic for checking whether the examinee in the examination room is correct in the verification module is represented as follows: ; In the formula: This is the determination value; For decision functions; , The total number of candidate face images obtained based on global image segmentation of candidate faces, and the total number of candidate face images in candidate identity information; Let be the similarity between the i-th candidate's facial image and the j-th candidate's facial image; The number of candidates is determined based on the number of seats in the examination room or the number of candidates' identity information. This refers to the number of students who were absent from the exam. In cases where there are absentees in the examination room, the assembly will take place. , The corresponding number of candidates is reduced synchronously. When formula (2) is true, the verification module verifies that the candidates in the examination room are correct; otherwise, it is incorrect, and the judgment function is used. If the condition within the parentheses is true, then =1, otherwise take zero.
4. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 3, characterized in that, When the verification module performs a verification and the result is negative, it obtains the result in the decision function. The candidate facial images in the set that do not make the determination function true are selected as the candidates whose facial images are sourced from. The candidates whose facial images are sourced from these candidates are then used as the offline processing targets for the invigilators. The invigilators verify the candidates' identity information offline again. If the verification is successful, the process jumps to the monitoring layer. If the verification fails, the candidates whose facial images are sourced from these candidates are excluded from the monitoring layer's monitoring scope.
5. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 3, characterized in that, The logic for calculating the similarity of the candidates' facial images is as follows: ; In the formula: The dynamic weights of facial geometric features, texture features, and deep learning features are calculated based on the attention mechanism; The total number of facial feature key points; Let V be the Euclidean distance between the v-th corresponding key point in candidate face image i and candidate face image j; These are normalization parameters; The texture feature vectors extracted from candidate face image i and candidate face image j using local binary mode; Let be the magnitude of the vector; The high-dimensional deep learning feature vectors extracted from candidate face image i and candidate face image j using a pre-trained face recognition model; Let be the magnitude of the vector; in, ∈[0,1], The sum is 1, and all values are positive, and are user-defined on the system side.
6. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 5, characterized in that, The The value of follows: An attention network is constructed, taking facial geometric feature differences, texture feature distribution entropy, and intra-class scatter of deep learning features as inputs, and the weights are calculated by a multilayer perceptron. ; In the formula: This is used to measure the difference in distance between facial key points, the difference in entropy of texture feature distribution, and the difference in intra-class discreteness of deep learning features. These are the multilayer perceptron mapping functions for the corresponding features; This is an index used to iterate through different feature types.
7. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1, characterized in that, The acquisition module is integrated into the camera module in the verification layer. The examination room images are acquired based on the camera module. When the configuration module configures the candidate's face image capture frame in the examination room image, it divides the area contained in the examination room image equally based on each capture frame, and the area corresponding to each capture frame is the same size. In addition, the examination table in each capture frame is centered in the capture frame. When the analysis module picks up the screen frame in the examination room image, the time interval between each adjacent screen frame is equal.
8. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1, characterized in that, The warning layer includes a judgment module, a warning module, and a visualization module. The judgment module is used to receive the analysis results of the changes in the spatial position of the candidates' faces in the historical monitoring layer, and to determine whether the candidates have abnormal behavior based on the analysis results. The warning module is used to receive the judgment result of whether the candidates have abnormal behavior in the judgment module, and to issue a warning voice when the judgment result is yes. The visualization module is used to create a layer on the examination room image, place the candidate's face image capture frame on the surface of the created layer, obtain the judgment result of whether the candidates have abnormal behavior in the judgment module, and render the candidate's face image capture frame with a specified color transparently for candidates whose judgment result is yes. The logic for determining whether a candidate exhibits abnormal behavior within the judgment module is as follows: Obtain the morphological representation parameters of the candidate's facial spatial position transformation from the latest three calculations in the combined results of the labeled directions for each sub-image, namely: ; In the formula, x represents the calculation sequence number of the facial spatial position change morphological representation parameter of the examinee. If any set of the facial spatial position change morphological representation parameters of the examinee obtained from the three most recent calculations above shows a continuous upward trend, it is determined that the examinee has abnormal behavior.
9. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 8, characterized in that, The warning module is integrated with a speaker, which has a preset warning voice. The format of the warning voice is: "Please abide by the examination rules for candidate number XX". In the "Please abide by the examination rules for candidate number XX", "XX" is the seat number in the candidate's identity information. The visualization module is integrated with computer equipment that has display capabilities. The computer equipment displays real-time rendered images of the examination room for invigilators to read visually.
10. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1, characterized in that, The acquisition module is interconnected with a configuration module and an analysis module via a wireless network. The acquisition module is interconnected with a camera module via a wireless network. The camera module is interconnected with an upload module and a verification module via a wireless network. The analysis module is interconnected with a judgment module via a wireless network. The judgment module is interconnected with an alert module and a visualization module via a wireless network.