Video operation and maintenance analysis system based on artificial intelligence face capture

The video operation and maintenance analysis system using artificial intelligence facial capture solves the problem that the existing invigilation system is unable to determine abnormal behavior, realizes candidate identity authentication and abnormal behavior monitoring, and ensures exam fairness and invigilation efficiency.

CN120673337AActive Publication Date: 2025-09-19JIANGSU LUOXIANG INFORMATION TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510766379.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing invigilation system is unable to identify and warn abnormal candidates through image analysis, resulting in a high degree of dependence on invigilation work and an inability to effectively prevent candidates from damaging cameras and other behaviors, affecting the smooth progress of the exam.

Method used

A video operation and maintenance analysis system based on artificial intelligence facial capture is adopted, including a verification layer, a monitoring layer and a warning layer. By globally collecting facial images of candidates, the facial spatial position transformation is analyzed in real time, and early warning prompts are issued to achieve candidate identity verification and abnormal behavior monitoring.

Benefits of technology

It achieves accurate identification of examinees’ identities, prevents cheating, ensures fairness and justice in examinations, promptly detects and stops abnormal behavior, reduces pressure on invigilators, and improves invigilation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673337A_ABST
    Figure CN120673337A_ABST
Patent Text Reader

Abstract

The invention discloses a video operation and maintenance analysis system based on artificial intelligence face capture, which relates to the field of data processing and comprises a verification layer, a monitoring layer and a warning layer. The examinee faces in the same examination room are globally collected through the verification layer, the verification layer synchronizes the examinee faces in the obtained examinee identity information and compares the examinee faces with the collected examinee face images so as to complete examinee face verification, the monitoring layer collects examination room images in real time, the faces of all examinees in the examination room images are captured, and the examinee face verification is completed. According to the method, the face of the examinee in the same examination room can be globally collected and compared with the face image in the identity information of the examinee for verification, after the examinee sits according to the seat number, the collected face image is segmented and matched with the identity information for verification, and the face space position transformation form of the examinee is analyzed. The identity of the examinee can be accurately identified, cheating behaviors such as an alternative examination are effectively prevented, the examination room order is maintained, fairness and justice of the examination are ensured, and a solid guarantee is provided for smooth proceeding of the examination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a video operation and maintenance analysis system based on artificial intelligence facial capture. Background Art

[0002] Surveillance technology is widely used in exam invigilation scenarios. By installing cameras and other equipment in the examination room, it can monitor candidates' status in real time, effectively preventing cheating and ensuring exam fairness. The recording function also provides a basis for subsequent review.

[0003] The invention patent application with application number 202011558036.3 discloses a remote proctoring system verification method, wherein the remote proctoring system includes: a first smart terminal, a second smart terminal and an examination terminal, and the method includes: obtaining a verification instruction input by the examinee; responding to the verification instruction, controlling the first smart terminal, the second smart terminal and the examination terminal to display verification information according to preset rules, and controlling the first smart terminal and the second smart terminal to collect verification images; judging whether the first smart terminal, the second smart terminal and the examination terminal meet the preset position rules based on the verification image and the verification information; if the preset position rules are met, confirming that the remote proctoring system has passed the verification, the application aims to solve the problem that "the existing remote proctoring system mainly collects the examination video of the examination room through the camera, and monitors the examinee's examination behavior based on the video content to see whether it is compliant. However, in actual applications, examinees often deliberately destroy or affect the remote proctoring system, such as deliberately moving the camera, damaging the camera, etc., which causes the remote proctoring system to fail to work normally and affects the smooth progress of the examination."

[0004] However, in the invigilation scenario, monitoring technology is currently still used as auxiliary equipment to assist invigilators in carrying out their invigilation work. It is unable to judge and warn abnormal candidates through image analysis, resulting in a high degree of reliance on invigilators in the implementation of examination invigilation work.

[0005] To this end, we propose a video operation and analysis system based on artificial intelligence facial capture. Summary of the Invention

[0006] In response to the above-mentioned shortcomings of the prior art, the present invention provides a video operation and maintenance analysis system based on artificial intelligence facial capture, which can effectively solve the problems of the prior art.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0008] The present invention discloses a video operation and maintenance analysis system based on artificial intelligence facial capture, comprising: a verification layer, a monitoring layer and a warning layer;

[0009] The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer synchronously obtains the faces of candidates in the candidate identity information and compares them with the collected candidate facial images to complete the candidate facial verification. The monitoring layer collects the examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the spatial position transformation of the candidate's face based on the captured face results. The warning layer receives the spatial position transformation of the candidate's face from the monitoring layer, confirms the warning target based on the spatial position transformation of the candidate's face, and issues an early warning prompt based on the warning target;

[0010] The monitoring layer includes an acquisition module, a configuration module, and an analysis module. The acquisition module is used to acquire examination room images in real time. The configuration module is used to configure a facial image capture frame of the examinee in the examination room image. The analysis module is used to segment the examination room image based on the examinee's facial image capture frame to obtain the examination room image corresponding to each examinee, pick up picture frames from the examinee's corresponding examination room image, and analyze the spatial position transformation of each examinee's face based on the picked picture frames.

[0011] The frames in the examination room image corresponding to the capture frames of the same examinee are used as the target for the spatial position transformation morphological analysis of the examinee's face. The examinee's contour image in each frame is identified, and each frame is segmented based on the vertical and horizontal midlines. Each frame is segmented into four sub-frames, and each sub-frame is marked with upper left, upper right, lower left, and lower right markers. The sub-frames are distinguished based on the markers to obtain four sub-frame sets.

[0012]

[0013] Where: ADFV is the parameter representing the spatial position transformation of the candidate's face; Q 左上 , Q 左下 Indicates the sub-picture frame set marked as the upper left and lower left; S(por) s 、S(unpor) s is the area of ​​the examinee's facial image in the s-th sub-picture frame and the area outside the examinee's facial image in the s-th sub-picture frame;

[0014] Among them, based on the above formula, the face position transformation morphological representation parameters of the examinee's face are calculated by upper left, lower left; upper left, upper right; upper right, lower right; lower right, lower left respectively. Based on the continuous acquisition of examination room images, to complete the picture frame picking and updating, we have:

[0015] Recorded as the result of the morphological analysis of the spatial position transformation of the face of the history candidate.

[0016] Furthermore, the verification layer includes a camera module, an upload module, and a verification module. The camera module is used to collect a global image of the examinee's face in the examination room. The upload module is used to upload the examinee's identity information. The verification module is used to receive the global image of the examinee's face in the examination room collected by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison of the examinee's identity information with the global image of the examinee's face in the examination room, the verification module verifies whether the examinee in the examination room is correct.

[0017] Among them, the candidate identity information uploaded in the upload module includes: admission ticket number, seat number, candidate ID number, and candidate facial image. After the candidates take their seats in the examination room in order according to their seat numbers, the camera module collects the global facial image of the candidates in the examination room, so that all candidate facial images are displayed in the same image. After the global facial image of the candidates in the examination room is collected, the image segmentation operation is performed synchronously so that each sub-image obtained by segmentation contains only the facial image of one candidate. A set is formed based on the sub-images, and the candidate facial image is obtained from the identity information of each candidate uploaded in the upload module. The obtained candidate facial images are taken as a set, and the two sets are sent to the verification module.

[0018] Furthermore, the verification logic of whether the examinee in the examination room is correct in the verification module is expressed as follows:

[0019]

[0020] Where: P is the judgment value; f(SIMM(i,j)≥98%) is the judgment function; n and m are two sets determined based on the global image segmentation of the examinee's face and the examinee's face image obtained from the examinee's identity information; SIMM(i,j) is the similarity between the i-th examinee's face image and the j-th examinee's face image; P0 is the number of examinees determined based on the number of seats in the examination room or the number of examinee's identity information; p miss The number of candidates who failed to take the exam;

[0021] Among them, when there is an absence in the examination room, the corresponding candidate objects in the sets n and m are synchronously eliminated. When formula (2) is established, the verification module verifies that the candidates in the examination room are correct, otherwise, it is incorrect. In the judgment function f(SIMM(i,j)≥98%), if the conditions in the brackets are established, then f(SIMM(i,j)≥98%)=1, otherwise it is zero.

[0022] Furthermore, when the verification result of the verification module is negative, the facial image of the examinee that does not make the judgment function valid is obtained in the judgment function, and the examinee from whom the facial image of the examinee is obtained is used as the offline processing target of the examination room invigilator. The examination room invigilator verifies the examinee's identity information again offline. After the verification is passed, the monitoring layer is jumped to run. If the verification fails, the examinee from whom the facial image of the examinee is obtained is excluded from the monitoring range of the monitoring layer.

[0023] Furthermore, the logic for obtaining the similarity of the examinee's facial images is as follows, taking SIMM(i,j) as an example:

[0024]

[0025] Where: α geo , α tex , α deep is the dynamic weight of facial geometric features, texture features, and deep learning features calculated based on the attention mechanism; 68 represents the set of key points lost in the candidate's facial image; d v is the Euclidean distance between the vth corresponding key point in the candidate's face image i and the candidate's face image j; max_distance is the normalization parameter; T i 、T j is the texture feature vector extracted from the local binary pattern of the candidate's face image i and the candidate's face image j; || T i ||、||T j || is the modulus of the vector; F i 、F j is the high-dimensional deep learning feature vector extracted from the candidate's facial image i and the candidate's facial image j through the pre-trained face recognition model; ||F i ||、||F j || is the modulus of the vector;

[0026] Among them, SIMM(i,j)∈[0,1],α geo , α tex , α deep The sum is 1, and they are all positive numbers, and are customized by the system user.

[0027] Furthermore, the α geo , α tex , α deep The value or obedience of:

[0028] Construct an attention network, taking the difference of facial geometric features, the distribution entropy of texture features, and the intra-class dispersion of deep learning features as input, and calculate the weights through a multi-layer perceptron;

[0029]

[0030] Where: Δgeo, Δtex, Δdeep are the distance difference measures of facial key points, the difference in texture feature distribution entropy, and the difference in intra-class discreteness of deep learning features; MLP geo 、MLP tex 、MLP deep are the multi-layer perceptron mapping functions of the corresponding features; q is the index used to traverse different feature types.

[0031] Furthermore, the acquisition module is integrated into the camera module in the verification layer, and the examination room image is acquired based on the camera module. When the configuration module configures the examinee's facial image capture frame in the examination room image, the area contained in the examination room image is equally divided based on each capture frame, and the corresponding area sizes of each capture frame are the same, and the examination table in each capture frame is located in the center of the capture frame. When the analysis module picks up the picture frames in the examination room image, the corresponding interval time of each adjacent picture frame picked up is equal.

[0032] Furthermore, the warning layer includes a determination module, a warning module, and a visualization module. The determination module is used to receive the results of the spatial position transformation morphology analysis of the historical examinees' faces in the monitoring layer, and determine whether the examinees have abnormal behavior based on the results of the spatial position transformation morphology analysis of the historical examinees' faces. The warning module is used to receive the determination result of whether the examinees have abnormal behavior in the determination module, and when the determination result is yes, issue a warning voice. The visualization module is used to create a layer on the examination room image, place the examinee's facial image capture frame on the created layer surface, obtain the determination result of whether the examinee has abnormal behavior in the determination module, and transparently render the examinee's facial image capture frame corresponding to the examinee with a yes determination result in a specified color.

[0033] The logic for determining whether a candidate has abnormal behavior in the judgment module is as follows:

[0034] Obtain the facial spatial position transformation morphological representation parameters of the examinee's face obtained by the three latest calculations in the combined results of the corresponding annotation directions of each sub-image, namely:

[0035] N01:

[0036] N02:

[0037] N03:

[0038] N04:

[0039] In the formula, x represents the calculation time sequence number of the candidate's facial spatial position transformation morphological representation parameters. When any group of the candidate's facial spatial position transformation morphological representation parameters obtained from the three most recent calculations shows a continuous upward trend, it is determined that the candidate has abnormal behavior.

[0040] Furthermore, the warning module is integrated with a speaker, and a warning voice is preset in the speaker. The format of the warning voice is: "Candidate No. XX, please comply with the examination room rules", where "XX" in "Candidate No. XX, please comply with the examination room rules" is the seat number in the candidate's identity information;

[0041] Among them, the visualization module is integrated by a computer device with display function, and the computer device displays the real-time rendered examination room image for visual reading by the invigilator.

[0042] Furthermore, the acquisition module is interactively connected to the configuration module and the analysis module through a wireless network, the acquisition module is interactively connected to the camera module through a wireless network, the camera module is interactively connected to the upload module and the verification module through a wireless network, the analysis module is interactively connected to the determination module through a wireless network, and the determination module is interactively connected to the warning module and the visualization module through a wireless network.

[0043] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:

[0044] 1. The system of the present invention can globally collect the faces of candidates in the same examination room and compare and verify them with the facial images in the candidate's identity information. After the candidate takes a seat according to the seat number, the collected facial image is segmented and matched with the identity information for verification. This can accurately identify the candidate's identity, effectively prevent cheating such as taking the exam on behalf of others, maintain order in the examination room, ensure the fairness of the examination, and provide a solid guarantee for the smooth progress of the examination.

[0045] 2. The system of the present invention collects examination room images in real time and monitors behavior by analyzing the spatial position transformation of the examinee's face. It picks up image frames at fixed intervals and analyzes them. Once the parameters of the examinee's facial position transformation are found to be abnormal, such as a continuous increase, the abnormal behavior can be detected in time. This helps the invigilator to discover the examinee's illegal behavior at the first time and stop it in time to ensure the seriousness of the examination environment.

[0046] 3. When the system of the present invention determines that a candidate has abnormal behavior, it will issue a targeted warning with a preset voice prompt, and at the same time render the capture frame of the abnormal candidate in a specified color transparently on the image. This allows the invigilator to quickly locate the abnormal candidate and identify the violator, greatly improving the invigilation efficiency and reducing the work pressure of the invigilator. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0048] Figure 1 This is a structural diagram of the video operation and maintenance analysis system based on artificial intelligence facial capture. DETAILED DESCRIPTION

[0049] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0050] The present invention will be further described below with reference to the embodiments.

[0051] Example:

[0052] The video operation and maintenance analysis system based on artificial intelligence facial capture in this embodiment is as follows: Figure 1 As shown, it includes: verification layer, monitoring layer and warning layer;

[0053] The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer synchronously obtains the faces of candidates in the candidate identity information and compares them with the collected candidate facial images to complete the candidate facial verification. The monitoring layer collects the examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the spatial position transformation of the candidate's face based on the captured face results. The warning layer receives the spatial position transformation of the candidate's face from the monitoring layer, confirms the warning target based on the spatial position transformation of the candidate's face, and issues an early warning prompt based on the warning target;

[0054] The verification layer includes a camera module, an upload module, and a verification module. The camera module is used to capture a global image of the examinee's face in the examination room. The upload module is used to upload the examinee's identity information. The verification module is used to receive the global image of the examinee's face captured by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison of the examinee's identity information with the global image of the examinee's face in the examination room, it verifies whether the examinee in the examination room is correct.

[0055] Among them, the candidate identity information uploaded in the upload module includes: admission ticket number, seat number, candidate ID number, and candidate facial image. After the candidates take their seats in the examination room in order according to their seat numbers, the camera module collects a global facial image of the candidates in the examination room, so that the facial images of all candidates are displayed in the same image. After the global facial image of the candidates in the examination room is collected, the image segmentation operation is performed synchronously so that each sub-image obtained by segmentation contains only the facial image of one candidate. Based on the sub-images, a set is formed, and the candidate facial image is obtained from the identity information of each candidate uploaded in the upload module. The obtained candidate facial images are taken as a set, and the two sets are sent to the verification module;

[0056] The verification logic of whether the candidates in the examination room are correct in the verification module is expressed as follows:

[0057]

[0058] Where: P is the judgment value; f(SIMM(i,j)≥98%) is the judgment function; n and m are two sets determined based on the global image segmentation of the examinee's face and the examinee's face image obtained from the examinee's identity information; SIMM(i,j) is the similarity between the i-th examinee's face image and the j-th examinee's face image; P0 is the number of examinees determined based on the number of seats in the examination room or the number of examinee's identity information; p miss The number of candidates who failed to take the exam;

[0059] Among them, when there is an absence in the examination room, the corresponding candidate objects in the sets n and m are synchronously eliminated. When formula (2) is established, the verification module verifies that the candidates in the examination room are correct, otherwise, it is incorrect. In the judgment function f(SIMM(i,j)≥98%), if the condition in the brackets is established, then f(SIMM(i,j)≥98%)=1, otherwise it is zero;

[0060] Through the calculation of the above logical formula, the identity of the candidates in the examination room can be accurately proofread once, which largely replaces the candidate identity verification work of the invigilator and has higher accuracy than manual verification.

[0061] When the verification module runs the verification result and the result is negative, the determination function obtains the candidate's facial image within n that does not make the determination function true, and the candidate who is the source of the candidate's facial image is used as the offline processing target of the examination room invigilator. The examination room invigilator verifies the candidate's identity information again offline. If the verification is passed, the monitoring layer is jumped to run. If the verification fails, the candidate who is the source of the candidate's facial image is excluded from the monitoring range of the monitoring layer;

[0062] The logic for calculating the similarity of the candidate's facial images is as follows, taking SIMM(i,j) as an example:

[0063]

[0064] Where: α geo , α tex , α deep is the dynamic weight of facial geometric features, texture features, and deep learning features calculated based on the attention mechanism; 68 represents the set of key points lost in the candidate's facial image; d v is the Euclidean distance between the vth corresponding key point in the candidate's face image i and the candidate's face image j; max_distance is the normalization parameter; T i 、T j is the texture feature vector extracted from the local binary pattern of the candidate's face image i and the candidate's face image j; || T i ||、||T j || is the modulus of the vector; F i 、Fj is the high-dimensional deep learning feature vector extracted from the candidate's facial image i and the candidate's facial image j through the pre-trained face recognition model; ||F i ||、||F j || is the modulus of the vector;

[0065] Among them, SIMM(i,j)∈[0,1],α geo , α tex , α deep The sum is 1, and they are all positive numbers, and are customized by the system user;

[0066] α geo , α tex , α deep The value or obedience of:

[0067] Construct an attention network, taking the difference of facial geometric features, the distribution entropy of texture features, and the intra-class dispersion of deep learning features as input, and calculate the weights through a multi-layer perceptron;

[0068]

[0069] Where: Δgeo, Δtex, Δdeep are the distance difference measures of facial key points, the difference in texture feature distribution entropy, and the difference in intra-class discreteness of deep learning features; MLP geo 、MLP tex 、MLP deep are the multi-layer perceptron mapping functions of the corresponding features; q is the index used to traverse different feature types;

[0070] Among them, (geo,tex,deep) indicates that the value range of q index is a set. When calculating the weight of each feature, the denominator part By looping through the values ​​of j in (geo, tex, deep), the sum of the exponential function results of the multi-layer perceptron (MLP) output values ​​corresponding to different features is calculated, and then combined with the numerator part. The principle of the Softmax function is used to convert the MLP output of each feature into a probabilistic weight, thereby dynamically allocating the importance of each feature in the overall similarity calculation through the attention mechanism.

[0071] The above logic formula is used to calculate and support the verification logic of whether the candidates in the examination room are correct.

[0072] The monitoring layer includes an acquisition module, a configuration module, and an analysis module. The acquisition module is used to collect examination room images in real time. The configuration module is used to configure the candidate's facial image capture frame in the examination room image. The analysis module is used to segment the examination room image based on the candidate's facial image capture frame to obtain the examination room image corresponding to each candidate, pick up the picture frame in the candidate's corresponding examination room image, and analyze the spatial position transformation of each candidate's face based on the picked picture frame.

[0073] The frames in the examination room image corresponding to the capture frames of the same examinee are used as the target for the spatial position transformation morphological analysis of the examinee's face. The examinee's contour image in each frame is identified, and each frame is segmented based on the vertical and horizontal midlines. Each frame is segmented into four sub-frames, and each sub-frame is marked with upper left, upper right, lower left, and lower right markers. The sub-frames are distinguished based on the markers to obtain four sub-frame sets.

[0074]

[0075] Where: ADFV is the parameter representing the spatial position transformation of the candidate's face; Q 左上 , Q 左下 Indicates the sub-picture frame set marked as the upper left and lower left; S(por) s 、S(unpor) s is the area of ​​the examinee's facial image in the s-th sub-picture frame and the area outside the examinee's facial image in the s-th sub-picture frame;

[0076] Among them, based on the above formula, the face position transformation morphological representation parameters of the examinee's face are calculated by upper left, lower left; upper left, upper right; upper right, lower right; lower right, lower left respectively. Based on the continuous acquisition of examination room images, to complete the picture frame picking and updating, we have:

[0077] Record the results of the face spatial position transformation morphological analysis of the history examinee;

[0078] By calculating with the above formula, the spatial position transformation of each examinee's face is analyzed to provide support for the further operation of the warning layer of the system in this embodiment.

[0079] The acquisition module is integrated with the camera module in the verification layer. The examination room image is collected based on the camera module. When the configuration module configures the examinee's facial image capture frame in the examination room image, the area contained in the examination room image is evenly divided based on each capture frame. The corresponding area of ​​each capture frame is the same size, and the examination table is centered in each capture frame. When the analysis module picks up the picture frames in the examination room image, the corresponding interval time of each picked up adjacent picture frames is equal.

[0080] The warning layer includes a judgment module, a warning module and a visualization module. The judgment module is used to receive the facial spatial position transformation morphological analysis results of historical examinees in the monitoring layer, and judge whether the examinee has abnormal behavior based on the facial spatial position transformation morphological analysis results of historical examinees. The warning module is used to receive the judgment result of whether the examinee has abnormal behavior in the judgment module, and when the judgment result is yes, issue a warning voice. The visualization module is used to create a layer on the examination room image, place the examinee's facial image capture frame on the created layer surface, obtain the judgment result of whether the examinee has abnormal behavior in the judgment module, and render the examinee's facial image capture frame corresponding to the examinee with a yes judgment result transparently in a specified color;

[0081] The logic for determining whether a candidate has abnormal behavior in the judgment module is as follows:

[0082] Obtain the facial spatial position transformation morphological representation parameters of the examinee's face obtained by the three latest calculations in the combined results of the corresponding annotation directions of each sub-image, namely:

[0083] N01:

[0084] N02:

[0085] N03:

[0086] N04:

[0087] Where x represents the calculation sequence number of the candidate's facial spatial position transformation morphological representation parameters. If any group of the candidate's facial spatial position transformation morphological representation parameters obtained from the three most recent calculations shows a continuous upward trend, the candidate is judged to have abnormal behavior;

[0088] The warning module is integrated with a loudspeaker, which has a preset warning voice. The format of the warning voice is: "Candidate No. XX, please comply with the examination room rules." "XX" in "Candidate No. XX, please comply with the examination room rules" is the seat number in the candidate's identity information;

[0089] The visualization module is integrated with a computer device with display function, which displays the real-time rendered examination room image for visual reading by the invigilator;

[0090] The acquisition module is interactively connected to the configuration module and the analysis module through a wireless network. The acquisition module is interactively connected to the camera module through a wireless network. The camera module is interactively connected to the upload module and the verification module through a wireless network. The analysis module is interactively connected to the judgment module through a wireless network. The judgment module is interactively connected to the warning module and the visualization module through a wireless network.

[0091] In this embodiment, the camera module is operated to collect the global image of the examinee's face in the examination room, the upload module synchronously uploads the examinee's identity information, the verification module further receives the global image of the examinee's face in the examination room collected by the camera module and the examinee's identity information uploaded by the upload module, and verifies whether the examinee in the examination room is correct based on the comparison between the examinee's identity information and the global image of the examinee's face in the examination room, the acquisition module collects the examination room image in real time, the configuration module further configures the examinee's face image capture frame in the examination room image, and obtains the examination room image corresponding to each examinee based on the examinee's face image capture frame segmentation in the examination room image through the analysis module, picks up the picture frame in the examinee's corresponding examination room image, and based on the picked up picture frame The facial frame analyzes the spatial position transformation of each candidate's face, and then the judgment module receives the historical analysis results of the facial spatial position transformation of candidates in the monitoring layer. Based on the historical analysis results of the facial spatial position transformation of candidates, it is determined whether the candidate has abnormal behavior. The warning module synchronously receives the judgment result of whether the candidate has abnormal behavior in the judgment module. When the judgment result is yes, a warning voice is issued. Finally, a layer is created on the examination room image through the visualization module, and the candidate's facial image capture frame is placed on the surface of the created layer. The judgment result of whether the candidate has abnormal behavior in the judgment module is obtained, and the corresponding candidate's facial image capture frame for the candidate with a judgment result of yes is transparently rendered with a specified color.

[0092] Through the operation of the system in the above embodiment, visual detection technology brings a new monitoring system to the invigilation scene that is more intelligent and can largely take over the invigilation work of invigilators, and monitor candidates with abnormal behavior in the examination room.

[0093] In summary, the system in the above embodiment can globally capture the faces of all examinees in the same examination room and compare and verify them with the facial images in the examinee's identity information. After the examinee takes his or her seat according to his or her seat number, the captured facial image is segmented and matched with the identity information for verification. This can accurately identify the examinee's identity, effectively prevent cheating such as impersonation, maintain order in the examination room, ensure the fairness and impartiality of the examination, and provide a solid guarantee for the smooth conduct of the examination. At the same time, the system can capture examination room images in real time and monitor behavior by analyzing the spatial position transformation morphology of the examinee's face. It will pick up image frames at fixed intervals and analyze them. Once an abnormality in the examinee's facial position transformation morphology parameters is detected, such as a continuous increase, the abnormal behavior can be detected in a timely manner. This helps invigilators immediately detect examinee violations and stop them in a timely manner, ensuring the seriousness of the examination environment. If the system determines that a examinee has engaged in abnormal behavior, it will issue a targeted warning with a preset voice prompt and simultaneously render a capture frame of the abnormal examinee in a specified color transparently on the image. This allows invigilators to quickly locate the abnormal examinee and clearly identify the offending examinee, greatly improving invigilation efficiency and reducing their workload.

[0094] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. Video operation and maintenance analysis system based on artificial intelligence facial capture, characterized by: include: Verification layer, monitoring layer and warning layer; The faces of candidates in the same examination room are globally collected through the verification layer. The verification layer synchronously obtains the faces of candidates in the candidate identity information and compares them with the collected candidate facial images to complete the candidate facial verification. The monitoring layer collects the examination room images in real time, captures the faces of each candidate in the examination room images, and analyzes the spatial position transformation of the candidate's face based on the captured face results. The warning layer receives the spatial position transformation of the candidate's face from the monitoring layer, confirms the warning target based on the spatial position transformation of the candidate's face, and issues an early warning prompt based on the warning target; The monitoring layer includes an acquisition module, a configuration module, and an analysis module. The acquisition module is used to acquire examination room images in real time. The configuration module is used to configure a facial image capture frame of the examinee in the examination room image. The analysis module is used to segment the examination room image based on the examinee's facial image capture frame to obtain the examination room image corresponding to each examinee, pick up picture frames from the examinee's corresponding examination room image, and analyze the spatial position transformation of each examinee's face based on the picked picture frames. The frames in the examination room image corresponding to the capture frames of the same examinee are used as the target for the spatial position transformation morphological analysis of the examinee's face. The examinee's contour image in each frame is identified, and each frame is segmented based on the vertical and horizontal midlines. Each frame is segmented into four sub-frames, and each sub-frame is marked with upper left, upper right, lower left, and lower right markers. The sub-frames are distinguished based on the markers to obtain four sub-frame sets. Where: ADFV is the parameter representing the spatial position transformation of the candidate's face; Q 左上 , Q 左下 Indicates the sub-picture frame set marked as the upper left and lower left; S(por) s 、S(unpor) s is the area of ​​the examinee's facial image in the s-th sub-picture frame and the area outside the examinee's facial image in the s-th sub-picture frame; Among them, based on the above formula, the face position transformation morphological representation parameters of the examinee are calculated by upper left, lower left; upper left, upper right; upper right, lower right; lower right, lower left respectively, and recorded as ADFV ↖↙ 、ADFV ↖↗ 、ADFV ↗↘ 、ADFV ↘↙ , based on the continuous acquisition of examination room images to complete the picture frame pickup and update, we have: Recorded as the result of the morphological analysis of the spatial position transformation of the face of the history candidate.

2. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1 is characterized in that The verification layer includes a camera module, an upload module, and a verification module. The camera module is used to collect a global image of the examinee's face in the examination room. The upload module is used to upload the examinee's identity information. The verification module is used to receive the global image of the examinee's face in the examination room collected by the camera module and the examinee's identity information uploaded by the upload module. Based on the comparison of the examinee's identity information with the global image of the examinee's face in the examination room, the verification module verifies whether the examinee in the examination room is correct. Among them, the candidate identity information uploaded in the upload module includes: admission ticket number, seat number, candidate ID number, and candidate facial image. After the candidates take their seats in the examination room in order according to their seat numbers, the camera module collects the global facial image of the candidates in the examination room, so that all candidate facial images are displayed in the same image. After the global facial image of the candidates in the examination room is collected, the image segmentation operation is performed synchronously so that each sub-image obtained by segmentation contains only the facial image of one candidate. A set is formed based on the sub-images, and the candidate facial image is obtained from the identity information of each candidate uploaded in the upload module. The obtained candidate facial images are taken as a set, and the two sets are sent to the verification module.

3. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 2 is characterized in that: The verification logic of whether the examinee in the examination room is correct in the verification module is expressed as follows: Where: P is the judgment value; f(SIMM(i,j)≥98%) is the judgment function; n and m are two sets determined based on the global image segmentation of the examinee's face and the examinee's face image obtained from the examinee's identity information; SIMM(i,j) is the similarity between the i-th examinee's face image and the j-th examinee's face image; P0 is the number of examinees determined based on the number of seats in the examination room or the number of examinee's identity information; p miss The number of candidates who failed to take the exam; Among them, when there is an absence in the examination room, the corresponding candidate objects in the sets n and m are synchronously eliminated. When formula (2) is established, the verification module verifies that the candidates in the examination room are correct, otherwise, it is incorrect. In the judgment function f(SIMM(i,j)≥98%), if the conditions in the brackets are established, then f(SIMM(i,j)≥98%)=1, otherwise it is zero.

4. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 3 is characterized in that When the verification result of the verification module is negative, the facial image of the examinee that does not make the judgment function valid is obtained in the judgment function, and the examinee from whom the facial image of the examinee is obtained is used as the offline processing target of the examination room invigilator. The examination room invigilator verifies the examinee's identity information again offline. After the verification is passed, the monitoring layer is jumped to run. If the verification fails, the examinee from whom the facial image of the examinee is obtained is excluded from the monitoring range of the monitoring layer.

5. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 3 is characterized in that: The logic for obtaining the similarity of the candidate's facial images is as follows, taking SIMM(i,j) as an example: Where: α geo , α tex , α deep Dynamic weights of facial geometric features, texture features, and deep learning features calculated based on the attention mechanism; 68 represents the set of key points lost in the candidate's face image; d v is the Euclidean distance between the vth corresponding key point in the candidate's face image i and the candidate's face image j; max_distance is the normalization parameter; T i 、T j is the texture feature vector extracted from the local binary pattern of the candidate's face image i and the candidate's face image j; || T i ||、||T j || is the modulus of the vector; F i 、F j is the high-dimensional deep learning feature vector extracted from the candidate's facial image i and the candidate's facial image j through the pre-trained face recognition model; ||F i ||、||F j || is the modulus of the vector; Among them, SIMM(i,j)∈[0,1],α geo , α tex , α deep The sum is 1, and they are all positive numbers, and are customized by the system user.

6. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 5 is characterized in that: The α geo , α tex , α deep The value or compliance of: Construct an attention network, taking the difference of facial geometric features, the distribution entropy of texture features, and the intra-class dispersion of deep learning features as input, and calculate the weights through a multi-layer perceptron; Where: Δgeo, Δtex, Δdeep are the distance difference measures of facial key points, the difference in texture feature distribution entropy, and the difference in intra-class discreteness of deep learning features; MLP geo 、MLP tex 、MLP deep are the multi-layer perceptron mapping functions of the corresponding features; q is the index used to traverse different feature types.

7. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1 is characterized in that: The acquisition module is integrated into the camera module in the verification layer, and the examination room image is acquired based on the camera module. When the configuration module configures the examinee's facial image capture frame in the examination room image, the area contained in the examination room image is evenly divided based on each capture frame, and the corresponding area size of each capture frame is the same, and the examination table in each capture frame is located in the center of the capture frame. When the analysis module picks up the picture frame in the examination room image, the corresponding interval time of each adjacent picture frame picked up is equal.

8. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1 is characterized in that: The warning layer includes a determination module, a warning module and a visualization module. The determination module is used to receive the results of the spatial position transformation morphology analysis of the historical examinees' faces in the monitoring layer, and determine whether the examinees have abnormal behavior based on the results of the spatial position transformation morphology analysis of the historical examinees' faces. The warning module is used to receive the determination result of whether the examinees have abnormal behavior in the determination module, and when the determination result is yes, issue a warning voice. The visualization module is used to create a layer on the examination room image, place the examinee's facial image capture frame on the created layer surface, obtain the determination result of whether the examinee has abnormal behavior in the determination module, and transparently render the examinee's facial image capture frame corresponding to the examinee with a yes determination result in a specified color. The logic for determining whether a candidate has abnormal behavior in the judgment module is as follows: Obtain the facial spatial position transformation morphological representation parameters of the examinee's face obtained by the three latest calculations in the combined results of the corresponding annotation directions of each sub-image, namely: N01:ADFV ↖↙ x-2、ADFV ↖↙ x-1、ADFV ↖↙ x N02:ADFV ↖↗ x-2、ADFV ↖↗ x-1、ADFV ↖↗ x N03:ADFV ↗↘ x-2、ADFV ↗↘ x-1、ADFV ↗↘ x; N04:ADFV ↘↙ x-2、ADFV ↘↙ x-1、ADFV ↘↙ x In the formula, x represents the calculation time sequence number of the candidate's facial spatial position transformation morphological representation parameters. When any group of the candidate's facial spatial position transformation morphological representation parameters obtained from the three most recent calculations shows a continuous upward trend, it is determined that the candidate has abnormal behavior.

9. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 8 is characterized in that: The warning module is integrated with a speaker, and a warning voice is preset in the speaker. The format of the warning voice is: "Candidate No. XX, please comply with the examination room rules." "XX" in "Candidate No. XX, please comply with the examination room rules" is the seat number in the candidate's identity information; Among them, the visualization module is integrated by a computer device with display function, and the computer device displays the real-time rendered examination room image for visual reading by the invigilator.

10. The video operation and maintenance analysis system based on artificial intelligence facial capture according to claim 1 is characterized in that: The acquisition module is interactively connected to the configuration module and the analysis module through a wireless network, the acquisition module is interactively connected to the camera module through a wireless network, the camera module is interactively connected to the upload module and the verification module through a wireless network, the analysis module is interactively connected to the determination module through a wireless network, and the determination module is interactively connected to the warning module and the visualization module through a wireless network.

Citation Information

Patent Citations

  • Whole-journey identity verification method and system for driving examinee based on face identification

    CN102800024A

  • Examination room monitoring data processing method and automatic monitoring system thereof

    CN105959624A

  • Method and device for face recognition of candidate substitutes by using real-time monitoring

    CN107330419A

  • Intelligent monitoring system suitable for time series data operation and maintenance

    CN117292330A

  • Examination cheating behavior detection method based on deep learning algorithm

    CN118397536A