Method for detecting illegal behaviors after examination based on cross-mirror tracking and identity authentication
By employing cross-camera tracking and identity authentication methods, the problem of cross-camera tracking and identity verification in post-exam violation detection was solved, achieving high-precision violation identification and real-time monitoring, and improving the automated detection capability of post-exam violations.
Patent Information
- Application Number
- CN202511002704.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-31
AI Technical Summary
Post-exam violation detection suffers from poor continuity in cross-camera tracking, low accuracy in identity verification, and insufficient recognition of specific violation patterns, resulting in inadequate robustness, accuracy, and real-time performance in violation identification.
A method based on cross-camera tracking and identity authentication is adopted. Target detection and Reid identity matching are performed through camera monitoring video to establish personnel identity mapping relationship. The quality of target bounding boxes is dynamically evaluated and trajectory features are constructed. Combined with GAN generator and Bi-LSTM discriminator, violations are identified, realizing real-time binding of identity ID and behavior and violation alarm.
It achieves high-precision identity matching and violation identification in complex scenarios, improves the robustness, accuracy and real-time performance of violation identification, and builds a full-process real-name intelligent supervision system.
Smart Images

Figure CN120877378A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video behavior analysis technology, specifically to a method for detecting post-exam violations based on cross-camera tracking and identity authentication. Background Technology
[0002] With the increasing standardization of large-scale examination management and the widespread adoption of video surveillance systems, the demand for automated monitoring of post-examination processes using intelligent video analytics technology has significantly increased. In the core area of the examination center, staff may engage in unauthorized activities when handling sensitive materials such as exam papers and answer sheets, such as making phone calls, taking photos, or smoking. Traditional manual monitoring methods struggle to detect these hidden violations in a timely manner during the post-examination process.
[0003] To address these shortcomings, deep learning-driven computer vision technologies, particularly in object detection, multi-object tracking, and behavior recognition, have made significant progress in recent years, providing a technological foundation for the automated detection of such violations. However, the examination center setting has its unique characteristics: the personnel density is relatively low, but the behavior is more purposeful; the environment is relatively controllable, but violations are often more covert and attempt to evade monitoring; and the cameras are located in different functional areas within the examination center. This results in shortcomings in existing solutions regarding the continuity of cross-camera tracking, the accuracy of identity authentication in complex scenarios, and the specificity of identifying examination-specific violation patterns, affecting the robustness, accuracy, and real-time performance of violation identification. Summary of the Invention
[0004] To address the shortcomings of existing technologies, such as poor continuity in cross-camera tracking, low accuracy in identity verification in complex scenes, and inadequacy in identifying specific violation patterns, this invention aims to provide a post-exam violation detection method based on cross-camera tracking and identity authentication. The specific technical solution adopted is as follows:
[0005] Several cameras were deployed based on the examination scenario, and monitoring videos from all cameras were collected;
[0006] The system identifies personnel entering the examination area as targets, captures image sequences based on monitoring videos from cameras at the entrance, uses Reid identity matching for authentication, and establishes and maintains personnel identity mapping relationships.
[0007] Based on the monitoring video, all personnel entering the area are located, target bounding boxes are determined, the quality scores of the target bounding boxes are dynamically evaluated and enhanced, optimized trajectory features are constructed, the trajectories are tracked through hierarchical clustering, and corresponding identity IDs are obtained.
[0008] A historical feature database is established by combining the identity ID and the corresponding trajectory. The historical feature database is dynamically updated according to the actions of the people entering the area. Identity conflict detection is performed, and identity IDs are bound or reset.
[0009] Based on the identity ID and the corresponding target bounding box, a continuous image patch sequence corresponding to the person entering is cropped, and the sequence is input into the GAN generator to synthesize an enhanced spatiotemporal feature sequence. The image patch sequence is then combined with the image patch sequence and input into the Bi-LSTM discriminator to perform adversarial discrimination and action classification to identify violations.
[0010] When a violation is detected, a real-name alarm message containing the identity ID and the type of violation is generated by combining the bound identity ID.
[0011] Preferably, target detection is performed to capture image sequences based on monitoring video from cameras at the entrance of the examination venue, and Reid identity matching is used for identity authentication to form and maintain personnel identity mapping relationships, including:
[0012] Analyze the image sequence, use a key point face detector to locate the face region of the person entering, extract the feature vector and perform normalization processing, compare it with the pre-stored face database to generate a unique identity ID;
[0013] An identity feature vector is established by fusing feature vectors from multiple frames, denoted as . Establish and maintain personnel identity mapping relationships, denoted as Among them, ID k This represents the ID of the kth person entering the room. This represents the identity feature vector of the kth person entering the room.
[0014] Preferably, all personnel entering the area are located based on monitoring video, target bounding boxes are determined, the quality scores of the target bounding boxes are dynamically evaluated and enhanced, optimized trajectory features are constructed, trajectories are tracked through hierarchical clustering, and corresponding identity IDs are obtained, including:
[0015] Preprocess the monitoring video, calculate and enhance the quality score of the target bounding box, and obtain the fusion feature of the trajectory based on the enhanced quality score and the feature vector of each frame of the trajectory of the person entering the vehicle.
[0016] The fusion features of the normalized trajectory are processed, the cosine similarity between each pair of trajectories is calculated to establish a similarity matrix, the cosine distance is obtained, hierarchical clustering is performed based on the cosine distance to determine the inter-class distance, the activity trajectory of the personnel entering the examination scene is generated, and the corresponding identity ID is obtained.
[0017] Preferably, the monitoring video is preprocessed, the quality score of the target bounding box is calculated and enhanced, and the fusion features of the trajectory are obtained by combining the enhanced quality score with the feature vector of each frame of the trajectory of the person entering the vehicle, including:
[0018] The examination administration scenario involves multiple rooms. A set of room types is established based on the monitoring videos of each room, denoted as . The set of target bounding boxes in any frame of the monitoring video of any room is obtained, denoted as B = {b1, b2, ..., b}. N}, where any target bounding box is denoted as b. i =(x min ,y min ,x max ,y max ), calculate the area, aspect ratio, and height of the target bounding box respectively;
[0019] By determining the constraints, and combining the area, aspect ratio, and height of the target bounding box, we can obtain the rate of change of the target bounding box area, the distance the center point moves, and the rate of change of the aspect ratio between consecutive frames.
[0020] The corresponding frame image is cropped based on the target bounding box to obtain the cropped image. The cropped image is then converted to grayscale, and the quality score of the target bounding box is calculated and enhanced. The corresponding calculation formula is as follows:
[0021]
[0022] q′ i =q i ·exp(-λ1R area (i)-λ2D center (i)-λ3R ratio (i))
[0023] Where, q i q′ represents the quality score of the i-th target bounding box in the current frame. i Indicates the corresponding q i Enhanced quality score; η represents the quality score threshold; W and H represent the width and height of the cropped image, respectively; I gray Represents the cropped image converted to grayscale; (k,j) represents the coordinates of a pixel in the grayscale image; R area (i) represents the rate of change of the area of the i-th target bounding box between consecutive frames; D center (i) represents the distance the center point of the i-th target bounding box moves between consecutive frames; R ratio (i) represents the aspect ratio change rate of the i-th target bounding box between consecutive frames; λ1, λ2 and λ3 are all examination scenario scaling coefficients;
[0024] The trajectory fusion features are obtained by combining the enhanced quality score with the feature vector of each frame of the incoming person's trajectory. The corresponding calculation formula is as follows:
[0025]
[0026] Among them, f track Indicates the fusion features of the current trajectory; I T={t1,t2,…,t k} represents the set of frames contained in the current trajectory T, for each t∈I T The current trajectory T corresponds to the i-th target bounding box in the t-th frame; f t Let represent the feature vector of frame t.
[0027] Preferably, the fusion features of the trajectories are normalized, the cosine similarity between pairwise trajectories is calculated to establish a similarity matrix, the cosine distance is obtained, hierarchical clustering is performed based on the cosine distance, the inter-class distance is determined, the activity trajectories of personnel entering the examination scenario are generated, and the corresponding identity IDs are obtained, including:
[0028] The set of trajectories of the entering personnel is determined, denoted as Ω = {T1, T2, T3, ..., T...} N}, and establish a feature set for all trajectories based on the fusion characteristics of the trajectories, denoted as {f i |i=1,…,T N The normalization process for the feature set is calculated using the following formula:
[0029]
[0030] in, f represents the fusion feature after normalization of the i-th trajectory; i This represents the fusion feature of the i-th trajectory;
[0031] To construct a similarity matrix, calculate the cosine similarity between each pair of trajectories. The corresponding formula is:
[0032]
[0033] in, Let represent the cosine similarity between the i-th trajectory and the j-th trajectory; S represents the fusion feature after normalization of the j-th trajectory; i,j Let represent the similarity matrix established between the i-th trajectory and the j-th trajectory;
[0034] The cosine distance between trajectories is obtained based on cosine similarity, and the corresponding calculation formula is:
[0035]
[0036] Where, d i,j Let represent the cosine distance between the i-th trajectory and the j-th trajectory;
[0037] First-stage hierarchical clustering and second-stage hierarchical clustering are performed separately. Based on the first-stage hierarchical clustering and cosine distance, the inter-cluster distance is calculated to model the spatiotemporal relationship of trajectories and enhance the inter-cluster distance. The corresponding calculation formula is as follows:
[0038]
[0039] d enhanced (A,B)=α·d(A,B)+(1-α)·Φ temporal
[0040] Where d(A,B) represents the inter-class distance between cluster a and cluster B; Φ temporal It represents the spatiotemporal relationship of the trajectory; Δt represents the time interval between two trajectory segments; β represents the sensitivity to control the time interval; This indicates the time corresponding to the last frame of trajectory in cluster A; d represents the time corresponding to the first frame trajectory in cluster B; enhanced (A,B) represents the enhanced inter-class distance d(A,B); α represents the weight coefficient;
[0041] Set a cosine similarity threshold, and iteratively merge cross-camera clusters with cosine similarity exceeding the threshold and no overlap in the monitored videos according to the second-stage hierarchical clustering. Generate the activity trajectory of personnel entering the examination scene and obtain the corresponding identity ID.
[0042] Preferably, a historical feature database is established by combining the identity ID and the corresponding trajectory. The historical feature database is dynamically updated based on the actions of the people entering the premises. Identity conflict detection is performed, and the identity ID is bound or reset, including:
[0043] For any identity ID's trajectory, a historical feature database is established. A sliding window update strategy is used to generate a new trajectory feature vector. This new feature vector is then compared with all trajectory features in the historical feature database to calculate the maximum similarity. The corresponding calculation formula is as follows:
[0044]
[0045] Among them, s k Indicates personnel identity ID k Maximum similarity; ID k f represents the identity ID of the k-th person; query Represents the feature vector of the new trajectory; This represents the personnel identity feature vector of the k-th person;
[0046] Personnel identity matching is performed based on the maximum similarity score, and the corresponding calculation formula is:
[0047]
[0048] Among them, ID * Represents the maximum similarity s k Matched personnel ID;
[0049] Set a similarity threshold; when the maximum similarity is greater than the similarity threshold, the identity ID is successfully matched.
[0050] A conflict detection mechanism is set up so that when a conflict is detected between the identity ID of an entrant and the trajectory, the similarity between the trajectory features of the conflict and the corresponding trajectory features in the historical feature database is calculated and sorted in descending order. The identity ID with the highest similarity is retained, and the identity IDs corresponding to the remaining conflicting trajectories are canceled or reset.
[0051] Preferably, based on the identity ID and the corresponding target bounding box, a continuous sequence of image patches corresponding to the entrant is cropped, input into a GAN generator to synthesize an enhanced spatiotemporal feature sequence, and combined with the image patch sequence, input into a Bi-LSTM discriminator to perform adversarial discrimination and action classification to identify violations, including:
[0052] Based on the identity ID and the corresponding target bounding box, a target detection architecture is established, including adaptive geometric-aware convolution and multi-scale feature fusion, which enhances the geometric modeling and multi-scale adaptation of the target bounding box and crops the continuous image block sequence corresponding to the person entering.
[0053] The image patch sequence is input into the GAN generator, and a 2D convolution-1D temporal convolution cascade structure is used to construct a spatial-temporal dual-path structure to synthesize an enhanced spatiotemporal feature sequence.
[0054] An intermediate representation sequence is generated based on the image patch sequence. The enhanced spatiotemporal feature sequence is then input into a Bi-LSTM to capture the motion evolution of the person entering the scene. Adversity discrimination and action classification are performed to identify violations.
[0055] The system integrates adversarial discrimination, behavior classification supervision, instantaneous action focusing, and occlusion-free distortion loss to enhance and optimize the identification of erroneous violations, thereby overcoming the interference of identifying erroneous violations.
[0056] Preferably, based on the identity ID and the corresponding target bounding box, a target detection architecture is established, including adaptive geometrically aware convolution and multi-scale feature fusion, to enhance the geometric modeling and multi-scale adaptation of the target bounding box, and to crop the corresponding continuous image patch sequence of the entering person, including:
[0057] Standard convolution kernel It is decomposed into a global principal component B and a local geometric component V, and an adaptive convolutional kernel is generated by dynamic fusion. The corresponding calculation formula is:
[0058]
[0059] Among them, W * β represents the adaptive convolution kernel; β represents the basic scaling factor; reshape represents the matrix dimension reshaping operation; D represents the total number of feature transformation dimensions; S represents the geometric transformation matrix; d represents the feature transformation dimension.
[0060] Multi-level feature maps are generated based on the target bounding box, denoted as... L represents the total number of levels. Learnable weights are introduced to weightedly fuse features from different levels. The corresponding calculation formula is:
[0061]
[0062] in, This indicates the fusion of features from different levels; Conv * Represents an adaptive geometry-aware convolution operation; I l Represents the feature set participating in the fusion of the l-th layer; Resize represents the upsampling / downsampling strategy; w i ∈ represents the learnable weight; ∈ represents the adjustment coefficient, used to prevent the denominator from being 0;
[0063] After enhancement processing, a sequence of consecutive image blocks corresponding to the entering personnel is cropped, denoted as... in, Let H represent the target bounding box after enhancement and cropping in frame t, where T represents the number of frames in the monitored video; H×W represents the spatial size; and C represents the number of channels.
[0064] Preferably, the image patch sequence is input into the GAN generator, and a 2D convolution-1D temporal convolution cascade structure is used to construct a spatial-temporal dual-path structure to synthesize an enhanced spatiotemporal feature sequence, including:
[0065] The calculation formula for processing image patch sequences frame by frame based on spatial pathways is as follows:
[0066]
[0067] in, W represents the spatial features obtained in the l-th layer for the t-th frame; init Represents the convolution kernel; Represents a 2D convolution kernel; b represents the target bounding box after enhancement and cropping in the t-th frame of the image patch sequence; init b (l) Indicates the bias term; Represents a 2D convolution operation; BN represents a normalization operation; ReLU represents an activation function.
[0068] Based on the temporal pathway, spatial features are reshaped and input into a temporal convolutional network to synthesize an enhanced spatiotemporal feature sequence. The corresponding calculation formula is as follows:
[0069]
[0070] in, Indicates enhanced spatiotemporal features, This represents the final residual connection feature of the k-th layer; Represents the temporal characteristics of the k-th layer; Represents a one-dimensional dilated convolution with a dilation rate of d; This represents the spatial characteristics after the (k-1)th layer has been reshaped; Both represent 1D convolution kernels; W down This represents the downsampling convolution kernel.
[0071] Preferably, an intermediate representation sequence is generated based on the image patch sequence, and combined with an enhanced spatiotemporal feature sequence, it is input into a Bi-LSTM to capture the motion evolution of the person entering the scene. Adversarial discrimination and action classification are then performed to identify violations, including:
[0072] Based on image patch sequences, features are extracted using the MobileNet network to generate intermediate representation sequences, denoted as...
[0073] By combining the enhanced spatiotemporal feature sequence with Bi-LSTM, the evolutionary patterns of the movement of people entering the area are captured through bidirectional temporal modeling. The corresponding calculation formula is as follows:
[0074]
[0075] Among them, h t This represents the timing state of frame t. These represent the hidden states of the forward and backward LSTMs, respectively; LSTM represents the Bi-LSTM operation. represents the hidden states of the forward and backward LSTMs of the adjacent previous and adjacent next frames of frame t, respectively; LayerNorm represents normalization processing; || represents the concatenation operation;
[0076] The process involves adversarial detection and action classification to identify violations; the corresponding calculation formula is as follows:
[0077] P cls =Softmax(W cls ·h T +b cls )
[0078] P d =σ(W d ·h T +b d )
[0079] Among them, P cls P represents the probability distribution of action categories. d W represents the probability of sequence authenticity in adversarial discrimination; Softmax represents the classification function; cls W represents the classification weight matrix; dh represents the adversarial discrimination weight matrix; T Indicates the timing status of the last frame of the monitored video; b cls b d Both represent bias terms; σ represents the activation function.
[0080] The present invention has the following beneficial effects:
[0081] Identity authentication is achieved through identity matching, combined with personnel identity mapping relationships and conflict detection mechanisms. This maintains millisecond-level trajectory-identity association under occlusion and lighting changes, clearly defining the responsible party to ensure real-time and accurate identity binding, achieving high-precision identity matching in complex scenarios. The system dynamically evaluates and enhances the quality score of target bounding boxes, constructing optimized trajectory features. Through a hierarchical clustering algorithm incorporating spatiotemporal and geometric constraints, it generates complete personnel trajectories across functional areas such as the test paper library, card library, and scanning room, eliminating trajectory interruptions caused by viewpoint switching and enabling continuous tracking across cameras. By combining a GAN generator and a Bi-LSTM discriminator to analyze the spatiotemporal context of violation action sequences, it accurately detects violations, specifically addressing the challenges of the instantaneousness, concealment, and occlusion inherent in violations during examinations, thus improving the robustness, accuracy, and real-time performance of violation identification. In short, by using identity IDs throughout the entire tracking and behavior detection process, a closed-loop intelligent supervision system of "identity registration - trajectory binding - behavior association" is constructed, significantly enhancing the automated discovery and real-name tracing capabilities of post-examination violations. Attached Figure Description
[0082] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0083] Figure 1 The flowchart illustrates a method for detecting post-exam violations based on cross-camera tracking and identity authentication, as provided in one embodiment of the present invention.
[0084] Figure 2 This is a framework diagram of a post-exam violation detection method based on cross-camera tracking and identity authentication, provided as an embodiment of the present invention. Detailed Implementation
[0085] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a post-exam violation detection method based on cross-camera tracking and identity authentication proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0086] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0087] The following description, in conjunction with the accompanying drawings, details a specific scheme for a post-exam violation detection method based on cross-camera tracking and identity authentication provided by this invention.
[0088] Please combine Figure 1 and Figure 2 The first embodiment of the present invention provides a flowchart of the steps of a post-exam violation detection method based on cross-camera tracking and identity authentication, the method comprising:
[0089] Step S1: Deploy several cameras based on the examination scenario and collect monitoring videos from all cameras;
[0090] Step S2: Define the personnel entering as the target. Based on the monitoring video from the camera at the entrance of the examination scene, target detection is performed to capture image sequences. Reid identity matching is used for identity authentication to form and maintain the personnel identity mapping relationship.
[0091] Step S3: Locate all personnel entering the area based on the monitoring video, determine the target bounding box, dynamically evaluate and enhance the quality score of the target bounding box, construct optimized trajectory features, track the trajectory through hierarchical clustering, and obtain the corresponding identity ID;
[0092] Step S4: Establish a historical feature database by combining the identity ID and the corresponding trajectory, dynamically update the historical feature database based on the actions of the people entering, perform identity conflict detection, and bind or reset the identity ID;
[0093] Step S5: Based on the identity ID and the corresponding target bounding box, crop the continuous image patch sequence corresponding to the person entering, input it into the GAN generator to synthesize the enhanced spatiotemporal feature sequence, and combine it with the image patch sequence to input into the Bi-LSTM discriminator to perform adversarial discrimination and action classification to identify violations;
[0094] Step S6: When a violation is detected, generate a real-name alarm message containing the identity ID and the type of violation by combining the bound identity ID.
[0095] To better illustrate, the core challenges of existing methods for detecting violations by examination center staff are: First, continuous tracking across multiple cameras is crucial. Staff move between different functional areas within the examination setting, such as the exam paper library, card library, or scanning room, requiring seamless integration of different monitoring perspectives to accurately reconstruct the complete activity trajectory of each individual. Second, precise staff identification is essential. Detected behaviors in the monitoring footage must be linked to the specific staff member's identity in real time to clarify the responsible party, heavily reliant on facial recognition technology that can function reliably under complex lighting, changing angles, and partial obstruction. Third, refined violation modeling is necessary. Violations such as taking photos, making phone calls, and smoking often possess specific semantic meanings, requiring high-level understanding by considering the target location, actions, interaction objects, and temporal context. Existing technologies have limitations in addressing these core challenges. Therefore, this application proposes a post-examination violation detection method based on cross-camera tracking and identity authentication to solve the problem of post-examination supervision of examination center staff, achieving full-process real-name intelligent monitoring.
[0096] It can be explained that in step S1, several cameras are deployed based on the examination scenario, and monitoring videos from all cameras are collected. That is, cameras are arranged at the entrance and in different functional areas such as examination rooms, test paper storage, card storage or scanning room according to the overall examination scenario, and then monitoring videos from the cameras are collected to comprehensively cover and monitor the entire examination behavior.
[0097] Understandably, by monitoring the video from the camera at the entrance of the examination scene, facial key point detection is used to compare with the pre-stored database to complete identity registration, construct a stable identity feature vector, and assign a unique identity identifier to each individual to ensure the reliability of the identity recognition baseline. That is, by combining the multi-view Reid identity matching method, the consistency binding of identity and trajectory is dynamically maintained, laying a solid initial identity reference foundation for the entire cross-camera tracking and identity binding process.
[0098] Further, step S2 includes:
[0099] Step S21: Analyze the image sequence, use a key point face detector to locate the face region of the person entering, extract the feature vector and normalize it, and compare it with the pre-stored face database to generate a unique identity ID;
[0100] Step S22: Establish an identity feature vector by fusing feature vectors from multiple frames, denoted as... Establish and maintain personnel identity mapping relationships, denoted as Among them, ID k This represents the ID of the kth person entering the room. This represents the identity feature vector of the kth person entering the room.
[0101] Specifically, the sequence of image frames from the surveillance video captured by the camera at the entrance is defined as I = {I1, I2, I3, ..., I...} N}, where N represents the total number of frames in the surveillance video; for each frame image I i The face region R is obtained using a key-point face detector. i =D(I i In this context, D represents a key-point face detector; that is, using advanced algorithms and deep learning technology, it accurately locates various key feature points of a face, such as the corners of the eyes, the corners of the mouth, and the nose, to accurately define the boundaries of the face and outline its contours; for each detected face region R... i Extracting feature vectors d represents the dimension of the face region; then L2 normalization is performed, and the feature vector of the normalized face region is compared with a pre-established face information database, that is, compared with the pre-stored face database to generate a unique identity ID, thus completing the identity ID recognition, denoted as ID. k Among them, the pre-stored face database refers to a set of face image data that is stored in advance for identity verification or recognition.
[0102] To obtain more stable identity features, a feature fusion strategy is adopted, resulting in the final identity feature vector denoted as:
[0103]
[0104] in, f represents the identity feature vector of the kth person entering; i This represents the feature vector of the i-th frame.
[0105] This leads to the formation and maintenance of identity mapping relationships. This serves as a benchmark, providing a reference for subsequent cross-border tracking and identity consistency comparison.
[0106] Furthermore, step S3 includes:
[0107] Step S31: Preprocess the monitoring video, calculate and enhance the quality score of the target bounding box, and obtain the fusion feature of the trajectory based on the enhanced quality score and the feature vector of each frame of the trajectory of the person entering the vehicle.
[0108] Understandably, to address the issues of bounding box noise and deformation caused by the differences in multiple perspectives in examination scenarios and the interaction between people and objects, the monitoring video is preprocessed, and the target is optimized through adaptive bounding box filtering and adaptive quality enhancement, providing high-quality and robust input data for subsequent cross-camera trajectory association.
[0109] Further, step S31 includes:
[0110] Step S311: The examination scenario involves multiple rooms. Based on the monitoring video of each room, a set of room types is established, denoted as... The set of target bounding boxes in any frame of the monitoring video of any room is obtained, denoted as B = {b1, b2, ..., b}. N}, where any target bounding box is denoted as b. i =(x min ,y min ,x max ,y max ), calculate the area, aspect ratio and height of the target bounding box respectively.
[0111] Specifically, geometric feature analysis is performed on the target bounding box, i.e., the detected human bounding box, to calculate the area, aspect ratio, and height of the target bounding box. The corresponding calculation formulas are as follows:
[0112] Area(b i )=(x max -x min )×(y max -y min )
[0113]
[0114] Height(b i )=y max -y min
[0115] Step S312: Determine the constraints, and combine the area, aspect ratio and height of the target bounding box to obtain the rate of change of the area, the distance the center point moves and the rate of change of the aspect ratio of the target bounding box between consecutive frames.
[0116] Specifically, an independent threshold is set for each room, denoted as α(room). k ), γ (room k ), β low (room k ), β high (room k ), where α(room k ) represents the area threshold of the target bounding box of the k-th room; γ(room k ) represents the threshold for the height of the target bounding box of the k-th room; β low (room k ), β high (room k ) represent the minimum and maximum aspect ratio thresholds of the target bounding box of the k-th room, respectively; and then determine the constraint conditions, denoted as Area(b i)>α(room k ), β low (room k ) <AspectRatio(b i )<β high (room k ), Height(b i )>γ(room k This effectively filters out noisy detection boxes and incomplete human body areas, ensuring that the target bounding boxes in subsequent processing have sufficient visual information.
[0117] Next, based on the target bounding box b i Combine the target bounding box b from the adjacent previous frame i-1 To conduct the analysis, the target bounding box b is first determined. i The formula for calculating the center point is:
[0118] c i =((x) min,i +x max,i ) / 2,(y min,i +y max,i ) / 2)
[0119] Then, the rate of change of the target bounding box area, the center point movement distance, and the rate of change of aspect ratio are calculated between consecutive frames. The corresponding calculation formulas are as follows:
[0120]
[0121] D center (i)=||c i -c i-1 ||2
[0122]
[0123] Among them, R area (i) represents the rate of change of the target bounding box area; Area(b i Area(b) i-1 ) represent the target bounding box b respectively i and b i-1 The area of D; center (i) represents the distance the center point moved; c i c i-1 These represent the target bounding box b. i and b i-1 The center point; R ratio (i) represents the aspect ratio change rate; AspectRatio(b i ), AspectRatio(b i-1 ) represent the target bounding box b respectivelyi and b i-1 The aspect ratio.
[0124] Step S313: Crop the corresponding frame image based on the target bounding box to obtain the cropped image. Convert the cropped image to grayscale, calculate the quality score of the target bounding box, and enhance it. The corresponding calculation formula is as follows:
[0125]
[0126] q′ i =q i ·exp(-λ1R area (i)-λ2D center (i)-λ3R ratio (i))
[0127] Where, q i q′ represents the quality score of the i-th target bounding box in the current frame. i Indicates the corresponding q i Enhanced quality score; η represents the quality score threshold; W and H represent the width and height of the cropped image, respectively; I gray Represents the cropped image converted to grayscale; (k,j) represents the coordinates of a pixel in the grayscale image; R area (i) represents the rate of change of the area of the i-th target bounding box between consecutive frames; D center (i) represents the distance the center point of the i-th target bounding box moves between consecutive frames; R ratio (i) represents the aspect ratio change rate of the i-th target bounding box between consecutive frames; λ1, λ2 and λ3 are all examination scenario scale coefficients.
[0128] It can be explained that interactions between personnel and exam papers or equipment in examination scenarios can easily lead to abrupt changes in the geometric features of the target bounding box. For example, a sudden drop in height when bending over to organize an exam paper or a dramatic change in aspect ratio when turning around can negatively impact the trajectory features of the target. To improve the robustness of trajectory features, a dynamic quality scoring function is used; that is, the corresponding frame image is cropped based on the target bounding box to obtain the cropped image, denoted as I. crop , will I crop Convert to grayscale image I gray It then calculates and enhances the quality score of the target bounding box to cope with drastic changes in the target's geometric features, suppresses the degradation of trajectory features caused by interactions such as bending over and turning around between people and objects, and ensures stable and accurate tracking and identification of targets in complex examination scenarios.
[0129] Step S314: Based on the enhanced quality score and the feature vector of each frame of the incoming personnel trajectory, the fusion feature of the trajectory is obtained. The corresponding calculation formula is as follows:
[0130]
[0131] Among them, f track Indicates the fusion features of the current trajectory; I T ={t1,t2,…,t k} represents the set of frames contained in the current trajectory T, for each t∈I T The current trajectory T corresponds to the i-th target bounding box in the t-th frame; f t Let represent the feature vector of frame t.
[0132] This is explained by calculating representative feature vectors of the trajectory to enhance the feature vector representation of the trajectory, where, d represents the dimension of the feature vector; that is, the weighted fusion strategy retains the important information of high-quality frames while suppressing the negative impact of low-quality frames.
[0133] Step S32: Normalize the fusion features of the trajectories, calculate the cosine similarity between each pair of trajectories to establish a similarity matrix, obtain the cosine distance, perform hierarchical clustering based on the cosine distance, determine the inter-class distance, generate the activity trajectory of the personnel entering the examination scene, and obtain the corresponding identity ID.
[0134] Understandably, based on the fusion features of the trajectory, in order to eliminate feature scale differences and suppress the influence of low-quality frames, while blocking erroneous associations with the same camera trajectory, a strategy of feature normalization and cosine distance calculation is adopted to achieve accurate similarity measurement between cross-camera trajectories. That is, by using cross-border trajectory hierarchical clustering association, accurate association of the same target under different cameras can be achieved.
[0135] Further, step S32 includes:
[0136] Step S321: Determine the set of trajectories of the entering personnel, denoted as Ω = {T1, T2, T3, ..., T...} N}, and establish a feature set for all trajectories based on the fusion characteristics of the trajectories, denoted as {f i |i=1,…,T N The normalization process for the feature set is calculated using the following formula:
[0137]
[0138] in, f represents the fusion feature after normalization of the i-th trajectory; i This represents the fusion feature of the i-th trajectory.
[0139] The explanation is that the fusion features of the trajectory are subjected to L2 normalization, that is, the fusion features are scaled so that the L2 norm of each fusion feature, that is, the Euclidean length of the fusion feature, is equal to 1. The fusion features of the trajectory are mapped onto the unit hypersphere, eliminating the influence of feature amplitude differences on similarity calculation.
[0140] Step S322: Calculate the cosine similarity between each pair of trajectories to establish a similarity matrix. The corresponding calculation formula is:
[0141]
[0142] in, Let represent the cosine similarity between the i-th trajectory and the j-th trajectory; S represents the fusion feature after normalization of the j-th trajectory; i,j Let represent the similarity matrix established between the i-th trajectory and the j-th trajectory.
[0143] It is explained that quality-weighted cosine similarity is used as the metric for fusion feature similarity, and a similarity matrix is constructed, denoted as . N×N represents the size of the similarity matrix.
[0144] Step S323: Obtain the cosine distance between trajectories based on cosine similarity. The corresponding calculation formula is:
[0145]
[0146] Where, d i,j This represents the cosine distance between the i-th and j-th trajectories; that is, converting cosine similarity into a distance metric to more intuitively measure the difference between two fused features.
[0147] Preferably, in this embodiment, to avoid erroneous association of different trajectories within the same camera, a camera constraint mask mechanism is established. Specifically, for trajectory pairs from the same camera, the mask value is set to 0, indicating that association is prohibited; for trajectory pairs from different cameras, the mask value is set to 1, allowing association calculation. Therefore, the final similarity matrix is equal to the element-wise product of the cosine distance matrix and the camera constraint mask value, with diagonal elements set to 0 to eliminate autocorrelation between trajectories and themselves, ensuring that the final similarity matrix considers both the similarity between trajectories and the camera's observation constraints.
[0148] Understandably, in order to overcome the fragmentation of trajectories caused by high-frequency and brief off-camera behavior in examination scenarios and to achieve seamless connection across functional areas, a two-stage hierarchical clustering algorithm is adopted to achieve accurate mapping of identity IDs from local trajectories to global trajectories, and finally generate complete and continuous personnel activity trajectories covering the examination scenarios.
[0149] Step S324: Perform first-stage hierarchical clustering and second-stage hierarchical clustering respectively. Based on the first-stage hierarchical clustering and cosine distance, calculate the inter-class distance, model the spatiotemporal relationship of the trajectory, and enhance the inter-class distance. The corresponding calculation formula is:
[0150]
[0151] d enhanced (A,B)=α·d(A,B)+(1-α)·Φ temporal
[0152] Where d(A,B) represents the inter-class distance between cluster A and cluster B; Φ temporal It represents the spatiotemporal relationship of the trajectory; Δt represents the time interval between two trajectory segments; β represents the sensitivity to control the time interval; This indicates the time corresponding to the last frame of trajectory in cluster A; d represents the time corresponding to the first frame trajectory in cluster B; enhanced (A,B) represents the enhanced inter-class distance d(A,B); α represents the weight coefficient.
[0153] Specifically, based on the first-stage hierarchical clustering, the trajectory within each camera is pre-clustered using cosine distance, i.e., using the agglomerative hierarchical clustering algorithm with average linking. In examination scenarios, there are frequent, brief off-camera behaviors, such as personnel entering blind spots to organize exam papers or moving items behind shelves, causing the trajectory of the same person within the same camera to be segmented. Therefore, a time interval is introduced to address the time blind spot problem and trajectory feature degradation problem, thereby explicitly modeling the spatiotemporal relationship of the trajectory, i.e., establishing a time penalty term. Here, β represents the sensitivity of controlling the time interval to accurately perceive and adjust the time interval, ensuring efficient task execution and smooth process flow; α represents the weight coefficient, which is used to balance the importance of inter-class distance and the time penalty term.
[0154] Step S325: Set a cosine similarity threshold, and merge cross-camera clusters with cosine similarity exceeding the threshold and no overlap in the monitored videos according to the second-stage hierarchical clustering iteration, to generate the activity trajectory of the personnel entering the examination scene and obtain the corresponding identity ID.
[0155] Specifically, the second-stage hierarchical clustering performs cross-camera cluster merging operations based on cluster center features. A screening threshold is set, and when the cosine similarity of two cluster centers from different cameras exceeds the screening threshold, the merging operation is performed. The merging condition also includes the non-intersection constraint of the camera sets to ensure that the merged clusters come from completely different cameras, so as to avoid erroneously merging similar segments from the same camera. Then, an iterative merging strategy is adopted to optimize the clustering results. Each time, the cluster pair with the highest cosine similarity among all the cluster pairs formed in the first-stage hierarchical clustering is found and merged. The merged cluster center is recalculated, and the next round of merging continues until there are no cluster pairs that meet the conditions.
[0156] Understandably, in examination scenarios, the appearance of personnel changes over time due to changes in posture, lighting, and occlusion. Therefore, a historical feature enhancement matching module is proposed. This module maintains a dynamically updated historical feature library and uses it for cross-batch matching, which significantly improves the accuracy and robustness of continuous identity tracking under complex conditions.
[0157] Further, step S4 includes:
[0158] Step S41: For the trajectory of any identity ID, establish a historical feature database, and use a sliding window update strategy to generate a new trajectory feature vector. Compare the new trajectory feature vector with all trajectory features in the historical feature database, and calculate the maximum similarity. The corresponding calculation formula is as follows:
[0159]
[0160] Among them, s k Indicates personnel identity ID k Maximum similarity; ID k f represents the identity ID of the k-th person; query Represents the feature vector of the new trajectory; This represents the personnel identity feature vector of the kth person.
[0161] The explanation is as follows: First, for each trajectory with a confirmed identity ID, a historical feature library is established, which involves segmenting and maintaining the features of the trajectory. Specifically, the trajectory feature sequence corresponding to the identity ID is evenly divided into four segments according to time, and representative features of each segment are extracted. This segmentation process captures changes in the appearance of a person at different time periods, such as posture adjustment, occlusion recovery, and lighting changes. The features of each segment are obtained by weighted averaging of the features corresponding to the target bounding boxes of all frames within that segment, and the weights are determined according to the enhanced quality score in step S313 above. This establishes and maintains a dynamically updated historical feature library, storing the corresponding feature vectors according to the identity ID combination. That is, each identity ID corresponds to a feature list containing multiple feature representations of that identity ID at different time periods. Next, the historical feature library adopts a sliding window update strategy. Preferably, in this embodiment, each identity ID stores a maximum of 2000 features. When the storage limit is exceeded as new features are added to the historical feature library, the oldest feature is deleted to maintain the timeliness of the historical feature library. At the same time, the historical feature library supports deduplication based on similarity to avoid storing overly similar redundant features.
[0162] Step S42: Match the identities of individuals based on the maximum similarity score. The corresponding calculation formula is as follows:
[0163]
[0164] Among them, ID * Represents the maximum similarity s k Matched personnel ID.
[0165] To explain, when processing new features, the maximum similarity is first calculated and matched with the identity ID in the historical feature database to find the best match between the trajectory and the identity ID. Preferably, the matching process adopts the nearest neighbor search strategy, that is, by calculating the distance between the feature corresponding to the target bounding box and all reference features, the reference feature with the smallest distance is selected as the matching result.
[0166] Step S43: Set a similarity threshold. When the maximum similarity is greater than the similarity threshold, the identity ID is successfully matched.
[0167] As an optional implementation, in this embodiment, the similarity threshold is set to 0.6 to determine whether the new feature matches the trajectory features stored in the historical feature database. When the maximum similarity is greater than 0.6, the identity ID is successfully matched; otherwise, when the match fails, a loop processing process is entered, that is, waiting for the next 2-3 frames, recording them as new features, accumulating and saving them in the historical feature database to enrich the feature information, and realizing the update of the historical feature database; then, matching is performed again with the new features determined in the previous steps according to the updated historical feature database. If the match still fails, the matching evaluation is performed in a loop until the maximum similarity of the new feature is greater than 0.6. After a successful match, the loop is exited and the trajectory status of the currently analyzed person is updated.
[0168] Understandably, although the matching of identity ID and updated trajectory status is completed based on the aforementioned steps, it is still necessary to solve problems such as multiple trajectories competing for the same identity ID or spatiotemporal contradictions in allocation. Therefore, an identity conflict detection mechanism is proposed, which effectively ensures the uniqueness, accuracy and spatiotemporal rationality of identity ID and trajectory binding through multi-level conflict identification and similarity-based conflict resolution strategies, so as to provide stable and reliable identity recognition and trajectory tracking.
[0169] Step S44: Set up a conflict detection mechanism. When a conflict is detected between the identity ID of an entrant and the trajectory, calculate the similarity between the conflicting trajectory features and the corresponding trajectory features in the historical feature database, sort them in descending order, retain the identity ID with the highest similarity, and cancel or reset the identity IDs corresponding to the remaining conflicting trajectories.
[0170] As an optional implementation, this embodiment designs a three-level conflict identification mechanism. Specifically, the first level is camera-level conflict detection, which identifies situations where multiple trajectories within the same camera are assigned the same identity ID; the second level is video-level conflict detection, which identifies duplicate identity IDs within the same monitored video segment; and the third level is spatiotemporal conflict detection, which identifies unreasonable identity ID allocations in spatiotemporal terms, such as the same identity ID appearing at two locations far apart at the same time. Based on each different level of conflict identification mechanism, different constraints and detection strategies are adopted to ensure comprehensive monitoring and accurate identification from local to global perspectives, forming a progressive conflict identification system to reduce the possibility of erroneous identification between identity IDs and trajectories, thereby clarifying the responsible party.
[0171] Next, when an identity conflict is detected, for the set of conflicting trajectory features, the similarity between the conflicting trajectory features and the corresponding trajectory features in the historical feature database is calculated. The similarity is sorted from high to low, and the trajectory with the highest similarity and the corresponding identity ID are retained. The matching identity IDs of the remaining trajectories are canceled or reassigned to other identity IDs. That is, based on the principle of "best match first", the most likely correct identity ID and trajectory match are retained.
[0172] Understandably, in response to the challenges of the instantaneous and concealed nature of violations in examination scenarios, as well as the occlusion present in these scenarios, a temporal action recognition method based on GAN (Generative Adversarial Networks) adversarial enhancement Bi-LSTM (Bi-directional Long Short-Term Memory) is proposed. This method integrates a target detection architecture, a GAN generator, and a Bi-LSTM discriminator, and performs adversarial enhancement optimization to achieve high-precision action recognition in complex scenarios such as occlusion, changes in lighting, and rapid movement.
[0173] Furthermore, step S5 includes:
[0174] Step S51: Based on the identity ID and the corresponding target bounding box, establish the target detection architecture, including adaptive geometric-aware convolution and multi-scale feature fusion, to enhance the geometric modeling and multi-scale adaptation of the target bounding box, and crop the continuous image block sequence corresponding to the person entering.
[0175] It can be explained that, based on the identity ID and the corresponding target bounding box, that is, the target bounding box that completes the identity ID and trajectory matching without conflict, a target detection architecture is established to accurately locate human targets in the examination scenario that are easily affected by occlusion and deformation. In this embodiment, an improved YOLOv5 target detection architecture is adopted, including adaptive geometric perception convolution and multi-scale feature fusion, which enhances geometric modeling capabilities and multi-scale adaptability, providing reliable and high-quality input for subsequent refined behavior analysis and ensuring the stability of data analysis.
[0176] Further, step S51 includes:
[0177] Step S511: Apply standard convolution kernels It is decomposed into a global principal component B and a local geometric component V, and an adaptive convolutional kernel is generated by dynamic fusion. The corresponding calculation formula is:
[0178]
[0179] Among them, W *β represents the adaptive convolution kernel; β represents the basic scaling factor; reshape represents the matrix dimension reshaping operation; D represents the total number of feature change dimensions; S represents the geometric transformation matrix; d represents the feature change dimension.
[0180] The explanation states that, in response to the problem of human targets being easily affected by occlusion and deformation in complex scenes, an improved YOLOv5 architecture is used for accurate positioning.
[0181] Specifically, this method introduces adaptive geometry-aware convolution to enhance geometric deformation modeling capabilities. Firstly, the core of this method lies in introducing adaptive geometry-aware convolution to enhance geometric deformation modeling capabilities: firstly, the standard convolution kernel... It is decomposed into a global principal component B and a local geometric component V, and the corresponding calculation formula is:
[0182]
[0183] V = WB
[0184] Among them, C out C in These represent the number of output channels and input channels of the standard convolutional kernel, respectively; K = M × L represents the size of the standard convolutional kernel; :,:,k indicates that the number of output channels and input channels of the standard convolutional kernel remain constant.
[0185] Then, an adaptive convolution kernel is generated through dynamic fusion, where the geometric transformation matrix... Used for learning spatial location correlation; basic scaling factor Used to regulate the intensity of the global principal component B; it can be explained that, based on the improved YOLOv5, the receptive field distribution is automatically adjusted according to changes in human posture, that is, the geometric transformation matrix S realizes spatial position correlation transformation through einsum to enhance the modeling ability of human limb deformation; that is, tensor operations are implemented through the einsum function to more accurately capture the relative positional relationship of various human limb parts in space, as well as the deformation during movement; β·B maintains the overall semantic information and alleviates local occlusion interference. Among them, local occlusion may lead to information loss or misunderstanding, so the influence of such interference is effectively reduced by combining the basic scaling factor with the global principal component.
[0186] Step S512: Generate multi-level feature maps based on the target bounding box, denoted as... L represents the total number of levels. Learnable weights are introduced to weightedly fuse features from different levels. The corresponding calculation formula is:
[0187]
[0188] in, This indicates the fusion of features from different levels; Conv *Represents an adaptive geometry-aware convolution operation; I l Represents the feature set participating in the fusion of the l-th layer; Resize represents the upsampling / downsampling strategy; w i ∈ represents the learnable weight; ∈ represents the adjustment coefficient, used to prevent the denominator from being 0.
[0189] As explained, step S512 achieves multi-scale feature fusion to address the problem of multi-scale changes in human targets in complex scenes, significantly improving the accuracy of human detection and tracking, especially when facing complex backgrounds and occlusions, exhibiting stronger robustness and adaptability.
[0190] Step S513: After enhancement processing, crop the continuous image block sequence corresponding to the entering personnel, denoted as... in, Let H represent the target bounding box after enhancement and cropping in frame t, where T represents the number of frames in the monitored video; H×W represents the spatial size; and C represents the number of channels.
[0191] The explanation is as follows: after enhancement processing, namely after geometric modeling and multi-scale adaptation of the enhanced target bounding box, the image patch sequence with the matched identity ID is cropped, and the target bounding box of each frame is cropped to construct the image patch sequence for subsequent operations, ensuring efficient data processing.
[0192] Understandably, adversarial-enhanced spatiotemporal feature learning is introduced, which enhances the bidirectional long short-term memory network (Bi-LSTM) through the generative adversarial network (GAN) architecture to overcome the feature ambiguity caused by rapid movement in complex scenes and strengthen the discriminative representation ability of action sequences. That is, the adversarial training mechanism is used to drive the GAN generator to synthesize high-fidelity temporal features, i.e., enhance the spatiotemporal feature sequence. This forces the Bi-LSTM discriminator, i.e., the subject that identifies the occurrence of the violation, to learn a more robust spatiotemporal feature representation that represents the core pattern of examination violation, so as to obtain the motion evolution law of personnel in the examination scene.
[0193] Step S52: Input the image patch sequence into the GAN generator, and use a 2D convolution-1D temporal convolution cascade structure to construct a spatial-temporal dual-path structure to synthesize an enhanced spatiotemporal feature sequence.
[0194] It can be explained that the GAN generator is used for adversarial training to generate data samples that are closer to the real situation; the 2D convolution-1D temporal convolution cascade structure combines the characteristics of two-dimensional convolution and one-dimensional temporal convolution to help construct a spatial-temporal dual-path structure to process image patch sequences with spatial and temporal features. That is, it uses two-dimensional convolution to extract the spatial features of the image patch sequence, and then uses one-dimensional temporal convolution to capture the temporal features of the image patch sequence, effectively utilizing the spatial and temporal information of the image patch sequence to obtain an enhanced spatiotemporal feature sequence.
[0195] Further, step S52 includes:
[0196] Step S521: Process the image patch sequence frame by frame based on the spatial path. The corresponding calculation formula is:
[0197]
[0198] in, W represents the spatial features obtained in the l-th layer for the t-th frame; init Represents the convolution kernel; Represents a 2D convolution kernel; b represents the target bounding box after enhancement and cropping in the t-th frame of the image patch sequence; init b (l) Indicates the bias term; Represents a 2D convolution operation; BN represents a normalization operation; ReLU represents an activation function.
[0199] Step S522: Reshape the spatial features based on the temporal pathway, input them into a temporal convolutional network, and synthesize an enhanced spatiotemporal feature sequence. The corresponding calculation formula is as follows:
[0200]
[0201] in, Indicates enhanced spatiotemporal features, This represents the final residual connection feature of the k-th layer; Represents the temporal characteristics of the k-th layer; Represents a one-dimensional dilated convolution with a dilation rate of d; This represents the spatial characteristics after the (k-1)th layer has been reshaped; Both represent 1D convolution kernels; W down This represents the downsampling convolution kernel.
[0202] Specifically, for spatial pathways, image patch sequences are processed frame by frame; for temporal pathways, the obtained spatial features are... Remodeling That is, T×H′×W′×C′ represents the spatial features. Because the features of different layers are fused, and the size of each layer is different, the size of the resulting feature map is different. After processing, the enhanced spatiotemporal feature sequence is synthesized, that is, high-fidelity temporal features, to ensure that the information in the image block sequence is not lost or distorted during the processing.
[0203] Step S53: Generate an intermediate representation sequence based on the image patch sequence, combine it with the enhanced spatiotemporal feature sequence and input it into Bi-LSTM to capture the motion evolution law of the entering personnel, perform adversarial discrimination and action classification, and identify violations.
[0204] It can be explained that the Bi-LSTM discriminator, through its bidirectional long short-term memory network structure, effectively captures the temporal dependencies in image patch sequences, while simultaneously undertaking the dual tasks of adversarial discrimination and action recognition.
[0205] Furthermore, step S53 includes:
[0206] Step S531: Based on the image patch sequence, extract features using the MobileNet network to generate an intermediate representation sequence, denoted as...
[0207] Specifically, the corresponding implementation method is as follows:
[0208]
[0209] Where T represents the temporal length of the image patch sequence, i.e. the number of frames corresponding to the monitoring video; D represents the dimension of the enhanced spatiotemporal feature sequence.
[0210] Step S532: Combine the enhanced spatiotemporal feature sequence with the Bi-LSTM, and capture the evolutionary pattern of the movement of the people entering through bidirectional temporal modeling. The corresponding calculation formula is as follows:
[0211]
[0212] Among them, h t This represents the timing state of frame t. These represent the hidden states of the forward and backward LSTMs, respectively; LSTM represents the Bi-LSTM operation. represents the hidden states of the forward and backward LSTMs of the adjacent previous and adjacent next frames of frame t, respectively; LayerNorm represents normalization processing; || represents the splicing operation.
[0213] It is explained that LayerNorm (Layer Normalization) is used to normalize the spliced features. To better control the distribution of hidden states in forward and backward LSTMs and reduce the problems of vanishing or exploding gradients.
[0214] Step S533: Perform adversarial discrimination and action classification to identify violations. The corresponding calculation formula is as follows:
[0215] P cls =Softmax(W cls ·h T +b cls )
[0216] P d =σ(W d ·h T +b d )
[0217] Among them, P cls P represents the probability distribution of action categories. d W represents the probability of sequence authenticity in adversarial discrimination; Softmax represents the classification function; cls W represents the classification weight matrix; d h represents the adversarial discrimination weight matrix; T Indicates the timing status of the last frame of the monitored video; b cls b d Both represent bias terms; σ represents the activation function.
[0218] Provide an explanation, h T This indicates the timing status of the last frame of the monitored video, i.e. Where H represents the height of the temporal state; the classification weight matrix K = M × N; to determine the size of the weight matrix.
[0219] Step S54: Perform adversarial enhancement optimization on the violation behavior fusion adversarial discrimination, behavior classification supervision, instantaneous action focusing and occlusion non-deformation loss to overcome the interference of erroneous violation behavior identification.
[0220] Understandably, adversarial enhancement optimization based on violation behavior aims to accurately address the identification of transient and concealed violations in examination scenarios, while overcoming interference from occlusion and deformation. This involves fusing multiple supervisory signals: adversarial discrimination, behavior classification supervision, instantaneous action focusing, and an adversarial optimization loss function that avoids occlusion and deformation, and includes the discriminator loss L. D And generator loss L G By jointly optimizing the method, the accuracy and stability of the overall behavior recognition method for identifying violations in complex scenarios are significantly improved.
[0221] Specifically, adversarial enhancement optimization based on the discriminator involves, firstly, introducing an adversarial discriminant loss in the adversarial discrimination process. Used to distinguish image patch sequences, i.e., enhanced spatiotemporal feature sequences synthesized from real data and GAN generators, this forces the discriminator to accurately capture subtle spatiotemporal dynamic patterns generated during interactions between examination staff and exam papers or devices. The corresponding calculation formula is:
[0222]
[0223] in, This represents the score by which the discriminator identifies the intermediate representation sequence generated by the image patch sequence as true; This represents the score by which the discriminator identifies the enhanced spatiotemporal feature sequence as false.
[0224] Next, regarding behavioral classification and supervision, since various violations occur in examination scenarios, such as making phone calls, taking photos, and smoking, a multi-classification loss system is introduced. Using the multi-class cross-entropy function, the behavior category is correctly classified to guide the accurate differentiation of violations. The corresponding calculation formula is as follows:
[0225]
[0226] in, L represents the probability that the discriminator correctly predicts the category of the violation; CE λ represents the cross-entropy loss for multi-class classification; cls y represents the classification balance weight; y represents the true label distribution of the violation behavior action classification.
[0227] Furthermore, regarding instantaneous action focusing, due to the problem of identifying transient and covert violations in examination scenarios—that is, the difficulty in identifying brief and difficult-to-detect violations in examination scenarios—instantaneous action focusing loss is introduced. To ensure rapid identification of diverse data and address the issue of instantaneous violation signals being diluted over long time series, the corresponding calculation formula is as follows:
[0228]
[0229] Where, λ temp Indicates the weight of instantaneous behavior; The gradient magnitude of the feature vector in frame t is represented by τ; τ represents the temperature coefficient used to control the focus intensity; y t The category label representing the violation in frame t; This represents the probability distribution of the violation action categories in frame t.
[0230] Then, regarding the issue of distortion prevention under occlusion, the partial occlusion caused by examination staff bending over to arrange the exam papers may lead to the loss or distortion of image information in some areas, thus introducing a loss of distortion prevention under occlusion. To ensure the overall stability of violation behavior recognition is maintained based on the occlusion invariance constraint of adversarial training, thus preventing any impact on recognition accuracy, the corresponding calculation formula is as follows:
[0231]
[0232] Where, λ occ Z represents the weight of the occlusion scene. occ Z represents the image sequence after random spatial occlusion; Z represents the enhanced spatiotemporal feature sequence.
[0233] Finally, by integrating four mechanisms—adversarial discrimination, behavior classification supervision, instantaneous action focusing, and occlusion invariance loss—the discriminator architecture is jointly optimized to obtain the total discriminator loss L. D The corresponding calculation formula is:
[0234]
[0235] Preferably, the GAN generator uses adversarial training to drive the synthesized enhanced spatiotemporal feature sequence to approximate the real distribution. Simultaneously, it introduces feature consistency constraints to weaken the constraints of non-keyframes. That is, the generator continuously adjusts parameters to reduce the difference between the synthesized enhanced spatiotemporal feature sequence and the real image patch sequence, making the synthesized enhanced spatiotemporal feature sequence more realistic and natural. To further improve the ability to identify violations, behavior classification is integrated, allowing direct feedback from the classification loss when optimizing trajectory features. This not only improves the understanding of complex scenes but also enhances adaptability to different examination scenarios. Furthermore, a generator classification loss is established to effectively identify and classify violations. The corresponding calculation formula is as follows:
[0236]
[0237] Among them, L G This represents the generator classification loss; This represents the adversarial loss of the generator; λ represents the feature consistency loss; feat λ represents the feature balance weights; cls L CE (D cls (Z),y) represents the classification loss; λ cls y represents the category balance weights; y represents the category label of the violation; D, D cls Both represent discriminator operations; Z represents the enhanced spatiotemporal feature sequence; This represents the intermediate representation sequence obtained from the image patch sequence.
[0238] It is explained that the GAN generator and the Bi-LSTM discriminator are executed synchronously and trained through a multi-task joint adversarial optimization strategy to accurately identify momentary concealed violations in complex examination scenarios.
[0239] It can be explained that in step S6, when a violation is identified, a real-name alarm message containing the identity ID and the type of violation is generated by combining the bound identity ID; that is, the identity ID bound to the location at the time of the violation and the corresponding category of the violation are determined, and a detailed real-name alarm message is generated to ensure that the information is highly targeted and traceable, effectively improving the efficiency and accuracy of reliable monitoring.
[0240] Understandably, identity authentication is achieved through identity matching, combined with personnel identity mapping relationships and conflict detection mechanisms. This maintains millisecond-level trajectory-identity association under occlusion and lighting changes, clearly defining the responsible party to ensure real-time and accurate identity binding, achieving high-precision identity matching in complex scenarios. The quality score of the target bounding box is dynamically evaluated and enhanced to construct optimized trajectory features. A hierarchical clustering algorithm incorporating spatiotemporal and geometric constraints generates complete personnel trajectories across functional areas such as the test paper library, card library, and scanning room, eliminating trajectory interruptions caused by perspective switching and enabling continuous tracking across cameras. A GAN generator and Bi-LSTM discriminator are combined to analyze the spatiotemporal context of the action sequence of violations, accurately detecting violations. This specifically addresses the challenges of the instantaneousness, concealment, and occlusion inherent in violations during examinations, improving the robustness, accuracy, and real-time performance of violation identification. In short, by using identity ID throughout the entire tracking and behavior detection process, a closed-loop intelligent supervision system of "identity registration - trajectory binding - behavior association" is constructed, significantly enhancing the automated discovery and real-name tracing capabilities of post-examination violations.
[0241] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0242] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for detecting post-exam violations based on cross-camera tracking and identity authentication, characterized in that, The method includes: Several cameras were deployed based on the examination scenario, and monitoring videos from all cameras were collected; The system identifies personnel entering the examination area as targets, captures image sequences based on monitoring videos from cameras at the entrance, uses Reid identity matching for authentication, and establishes and maintains personnel identity mapping relationships. Based on the monitoring video, all personnel entering the area are located, target bounding boxes are determined, the quality scores of the target bounding boxes are dynamically evaluated and enhanced, optimized trajectory features are constructed, the trajectories are tracked through hierarchical clustering, and corresponding identity IDs are obtained. A historical feature database is established by combining the identity ID and the corresponding trajectory. The historical feature database is dynamically updated according to the actions of the people entering the area. Identity conflict detection is performed, and identity IDs are bound or reset. Based on the identity ID and the corresponding target bounding box, a continuous image patch sequence corresponding to the person entering is cropped, and the sequence is input into the GAN generator to synthesize an enhanced spatiotemporal feature sequence. The image patch sequence is then combined with the image patch sequence and input into the Bi-LSTM discriminator to perform adversarial discrimination and action classification to identify violations. When a violation is detected, a real-name alarm message containing the identity ID and the type of violation is generated by combining the bound identity ID.
2. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 1, characterized in that, Based on the monitoring video from the cameras at the entrance of the examination venue, target detection is performed to capture image sequences. Reid identity matching is used for identity authentication to form and maintain personnel identity mapping relationships, including: Analyze the image sequence, use a key point face detector to locate the face region of the person entering, extract the feature vector and perform normalization processing, compare it with the pre-stored face database to generate a unique identity ID; An identity feature vector is established by fusing feature vectors from multiple frames, denoted as . Establish and maintain personnel identity mapping relationships, denoted as Among them, ID k This represents the ID of the kth person entering the room. This represents the identity feature vector of the kth person entering the room.
3. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 1, characterized in that, Based on the monitoring video, all personnel entering the area are located, target bounding boxes are determined, the quality scores of the target bounding boxes are dynamically evaluated and enhanced, optimized trajectory features are constructed, trajectories are tracked through hierarchical clustering, and corresponding identity IDs are obtained, including: Preprocess the monitoring video, calculate and enhance the quality score of the target bounding box, and obtain the fusion feature of the trajectory based on the enhanced quality score and the feature vector of each frame of the trajectory of the person entering the vehicle. The fusion features of the normalized trajectory are processed, the cosine similarity between each pair of trajectories is calculated to establish a similarity matrix, the cosine distance is obtained, hierarchical clustering is performed based on the cosine distance to determine the inter-class distance, the activity trajectory of the personnel entering the examination scene is generated, and the corresponding identity ID is obtained.
4. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 3, characterized in that, Preprocess the monitoring video, calculate and enhance the quality score of the target bounding box, and obtain the trajectory fusion features by combining the enhanced quality score with the feature vector of each frame of the incoming person's trajectory, including: The examination administration scenario involves multiple rooms. A set of room types is established based on the monitoring videos of each room, denoted as . The set of target bounding boxes in any frame of the monitoring video of any room is obtained, denoted as B = {b1, b2, ..., b}. N }, where any target bounding box is denoted as b. i =(x min ,y min ,x max ,y max ), calculate the area, aspect ratio, and height of the target bounding box respectively; By determining the constraints, and combining the area, aspect ratio, and height of the target bounding box, we can obtain the rate of change of the target bounding box area, the distance the center point moves, and the rate of change of the aspect ratio between consecutive frames. The corresponding frame image is cropped based on the target bounding box to obtain the cropped image. The cropped image is then converted to grayscale, and the quality score of the target bounding box is calculated and enhanced. The corresponding calculation formula is as follows: q′ i =q i ·exp(-λ1R area (i)-λ2D center (i)-λ3R ratio (i)) Where, q i q′ represents the quality score of the i-th target bounding box in the current frame. i Indicates the corresponding q i Enhanced quality score; η represents the quality score threshold; W and H represent the width and height of the cropped image, respectively; I gray Represents the cropped image converted to grayscale; (k,j) represents the coordinates of a pixel in the grayscale image; R area (i) represents the rate of change of the area of the i-th target bounding box between consecutive frames; D center (i) represents the distance the center point of the i-th target bounding box moves between consecutive frames; R ratio (i) represents the aspect ratio change rate of the i-th target bounding box between consecutive frames; λ1, λ2 and λ3 are all examination scenario scaling coefficients; The trajectory fusion features are obtained by combining the enhanced quality score with the feature vector of each frame of the incoming person's trajectory. The corresponding calculation formula is as follows: Among them, f track Indicates the fusion features of the current trajectory; I T ={t1,t2,…,t k } represents the set of frames contained in the current trajectory T, for each t∈I T The current trajectory T corresponds to the i-th target bounding box in the t-th frame; f t Let represent the feature vector of frame t.
5. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 3, characterized in that, The fusion features of the normalized trajectories are processed, and the cosine similarity between pairwise trajectories is calculated to establish a similarity matrix, obtaining the cosine distance. Hierarchical clustering is then performed based on the cosine distance to determine the inter-class distance, generating the activity trajectories of personnel entering the examination scenario and obtaining their corresponding identity IDs, including: The set of trajectories of the entering personnel is determined, denoted as Ω = {T1, T2, T3, ..., T...} N }, and establish a feature set for all trajectories based on the fusion characteristics of the trajectories, denoted as {f i |i=1,…,T N The normalization process for the feature set is calculated using the following formula: in, f represents the fusion feature after normalization of the i-th trajectory; i This represents the fusion feature of the i-th trajectory; To construct a similarity matrix, calculate the cosine similarity between pairwise trajectories. The corresponding calculation formula is: in, Let represent the cosine similarity between the i-th trajectory and the j-th trajectory; S represents the fusion feature after normalization of the j-th trajectory; i,j Let represent the similarity matrix established between the i-th trajectory and the j-th trajectory; The cosine distance between trajectories is obtained based on cosine similarity, and the corresponding calculation formula is: Where, d i,j Let represent the cosine distance between the i-th trajectory and the j-th trajectory; First-stage hierarchical clustering and second-stage hierarchical clustering are performed separately. Based on the first-stage hierarchical clustering and cosine distance, the inter-cluster distance is calculated to model the spatiotemporal relationship of trajectories and enhance the inter-cluster distance. The corresponding calculation formula is as follows: d enhanced (A,B)=α·d(A,B)+(1-α)·Φ temporal Where d(A,B) represents the inter-class distance between cluster A and cluster B; Φ temporal It represents the spatiotemporal relationship of the trajectory; Δt represents the time interval between two trajectory segments; β represents the sensitivity to control the time interval; This indicates the time corresponding to the last frame of trajectory in cluster A; d represents the time corresponding to the first frame trajectory in cluster B; enhanced (A,B) represents the enhanced inter-class distance d(A,B); α represents the weight coefficient; Set a cosine similarity threshold, and iteratively merge cross-camera clusters with cosine similarity exceeding the threshold and no overlap in the monitored videos according to the second-stage hierarchical clustering. Generate the activity trajectory of personnel entering the examination scene and obtain the corresponding identity ID.
6. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 1, characterized in that, A historical feature database is established by combining the individual's identity ID and corresponding trajectory. This database is dynamically updated based on the actions of individuals entering the premises. Identity conflict detection is performed, and identity IDs are bound or reset, including: For any identity ID's trajectory, a historical feature database is established. A sliding window update strategy is used to generate a new trajectory feature vector. This new feature vector is then compared with all trajectory features in the historical feature database to calculate the maximum similarity. The corresponding calculation formula is as follows: Among them, s k Indicates personnel identity ID k Maximum similarity; ID k f represents the identity ID of the k-th person; query Represents the feature vector of the new trajectory; This represents the personnel identity feature vector of the k-th person; Personnel identity matching is performed based on the maximum similarity score, and the corresponding calculation formula is: Among them, ID * Represents the maximum similarity s k Matched personnel ID; Set a similarity threshold; when the maximum similarity is greater than the similarity threshold, the identity ID is successfully matched. A conflict detection mechanism is set up so that when a conflict is detected between the identity ID of an entrant and the trajectory, the similarity between the trajectory features of the conflict and the corresponding trajectory features in the historical feature database is calculated and sorted in descending order. The identity ID with the highest similarity is retained, and the identity IDs corresponding to the remaining conflicting trajectories are canceled or reset.
7. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 1, characterized in that, Based on the identity ID and the corresponding target bounding box, a continuous sequence of image patches corresponding to the entrants is cropped and input into a GAN generator to synthesize an enhanced spatiotemporal feature sequence. This sequence is then combined with the image patch sequence and input into a Bi-LSTM discriminator to perform adversarial discrimination and action classification, identifying violations, including: Based on the identity ID and the corresponding target bounding box, a target detection architecture is established, including adaptive geometric-aware convolution and multi-scale feature fusion, which enhances the geometric modeling and multi-scale adaptation of the target bounding box and crops the continuous image block sequence corresponding to the person entering. The image patch sequence is input into the GAN generator, and a 2D convolution-1D temporal convolution cascade structure is used to construct a spatial-temporal dual-path structure to synthesize an enhanced spatiotemporal feature sequence. An intermediate representation sequence is generated based on the image patch sequence. The enhanced spatiotemporal feature sequence is then input into a Bi-LSTM to capture the motion evolution of the person entering the scene. Adversity discrimination and action classification are performed to identify violations. The system integrates adversarial discrimination, behavior classification supervision, instantaneous action focusing, and occlusion-free distortion loss to enhance and optimize the identification of erroneous violations, thereby overcoming the interference of identifying erroneous violations.
8. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 7, characterized in that, Based on the identity ID and the corresponding target bounding box, an object detection architecture is established, including adaptive geometry-aware convolution and multi-scale feature fusion, to enhance the geometric modeling and multi-scale adaptation of the target bounding box. This involves cropping consecutive image patch sequences corresponding to the entering personnel, including: Standard convolution kernel It is decomposed into a global principal component B and a local geometric component V, and an adaptive convolutional kernel is generated by dynamic fusion. The corresponding calculation formula is: Among them, W * β represents the adaptive convolution kernel; β represents the basic scaling factor; reshape represents the matrix dimension reshaping operation; D represents the total number of feature transformation dimensions; S represents the geometric transformation matrix; d represents the feature transformation dimension. Multi-level feature maps are generated based on the target bounding box, denoted as... L represents the total number of levels. Learnable weights are introduced to weightedly fuse features from different levels. The corresponding calculation formula is: in, This indicates the fusion of features from different levels; Conv * Represents an adaptive geometry-aware convolution operation; I l Represents the feature set participating in the fusion of the l-th layer; Resize represents the upsampling / downsampling strategy; w i ∈ represents the learnable weight; ∈ represents the adjustment coefficient, used to prevent the denominator from being 0; After enhancement processing, a sequence of consecutive image blocks corresponding to the entering personnel is cropped, denoted as... in, Let H represent the target bounding box after enhancement and cropping in frame t, where T represents the number of frames in the monitored video; H×W represents the spatial size; and C represents the number of channels.
9. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 7, characterized in that, The image patch sequence is input into the GAN generator, and a 2D convolutional-1D temporal convolutional cascade structure is used to construct a spatial-temporal dual-path structure to synthesize enhanced spatiotemporal feature sequences, including: The calculation formula for processing image patch sequences frame by frame based on spatial pathways is as follows: in, W represents the spatial features obtained in the l-th layer for the t-th frame; init Represents the convolution kernel; Represents a 2D convolution kernel; b represents the target bounding box after enhancement and cropping in the t-th frame of the image patch sequence; init b (l) Indicates the bias term; Represents a 2D convolution operation; BN represents a normalization operation; ReLU represents an activation function. Based on the temporal pathway, spatial features are reshaped and input into a temporal convolutional network to synthesize an enhanced spatiotemporal feature sequence. The corresponding calculation formula is as follows: in, Indicates enhanced spatiotemporal features, This represents the final residual connection feature of the k-th layer; Represents the temporal characteristics of the k-th layer; Represents a one-dimensional dilated convolution with a dilation rate of d; This represents the spatial characteristics after the (k-1)th layer has been reshaped; Both represent 1D convolution kernels; W down This represents the downsampling convolution kernel.
10. The method for detecting post-exam violations based on cross-camera tracking and identity authentication according to claim 7, characterized in that, Intermediate representation sequences are generated based on image patch sequences. These sequences, combined with enhanced spatiotemporal feature sequences, are input into a Bi-LSTM to capture the motion evolution patterns of incoming personnel. Adversarial discrimination and action classification are then performed to identify violations, including: Based on image patch sequences, features are extracted using the MobileNet network to generate intermediate representation sequences, denoted as... By combining the enhanced spatiotemporal feature sequence with Bi-LSTM, the evolutionary patterns of the movement of people entering the area are captured through bidirectional temporal modeling. The corresponding calculation formula is as follows: Among them, h t This represents the timing state of frame t. These represent the hidden states of the forward and backward LSTMs, respectively; LSTM represents the Bi-LSTM operation. represents the hidden states of the forward and backward LSTMs of the adjacent previous and adjacent next frames of frame t, respectively; LayerNorm represents normalization processing; || represents the concatenation operation; The process involves adversarial detection and action classification to identify violations; the corresponding calculation formula is as follows: P cls =Softmax(W cls ·h T +b cls ) P d =σ(W d ·h T +b d ) Among them, P cls P represents the probability distribution of action categories. d W represents the probability of sequence authenticity in adversarial discrimination; Softmax represents the classification function; cls W represents the classification weight matrix; d Represents the adversarial discrimination weight matrix; h T Indicates the timing status of the last frame of the monitored video; b cls b d Both represent bias terms; σ represents the activation function.
Citation Information
Cited By
Same-track gang identification method and device for antagonistic behaviors
CN121692068A
A method and apparatus for identifying gangs with the same trajectory in response to adversarial behavior.
CN121692068B