Campus intelligent alarm linkage method and system based on face recognition
By deploying cameras in key areas of the campus to collect video streams in real time and perform facial recognition and behavior analysis, threat levels are generated and device linkage is triggered, which solves the problem of insufficient proactive early warning and emergency response in the existing campus security system and realizes intelligent security linkage and rapid response.
Patent Information
- Application Number
- CN202511443684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing campus security systems lack proactive early warning capabilities and emergency response efficiency, and cannot achieve device linkage, resulting in an inability to effectively cope with complex security threats.
By deploying cameras in key areas of the campus to collect video streams in real time, using face detection algorithms to extract facial images, and comparing them in real time with a whitelist database, threat levels are generated by combining behavioral analysis, automatically triggering device linkage commands, and generating security incident handling reports.
It has improved the proactive early warning capabilities and emergency response efficiency of the campus security system, and enabled intelligent linkage and rapid handling of equipment.
Smart Images

Figure CN121214616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of campus security, and particularly relates to a campus intelligent alarm linkage method and system based on face recognition. Background Art
[0002] Campus security is an important part of social public security. Traditional security systems mainly rely on manual video monitoring and patrols, which have problems such as lagging response and insufficient early warning capabilities. With the popularization of face recognition technology, existing solutions have tried to apply it to scenarios such as campus access control, but the functions are relatively single. These solutions usually can only complete simple identity recognition and comparison, lack in-depth analysis of personnel behavior intentions and scenario risks, and cannot achieve early warning. At the same time, each security device in the system (such as cameras, access control, alarms) operates independently, making it difficult to form an effective linkage, resulting in the breakage of the closed-loop chain from detecting abnormalities to taking disposal measures, and the overall security efficiency is low. In the face of complex security threats that may occur on campus, existing technologies are difficult to achieve active early warning, accurate judgment, and rapid collaborative disposal, and cannot meet the urgent needs of modern intelligent campuses for active and intelligent security protection. Summary of the Invention
[0003] The purpose of the present invention is to provide a campus intelligent alarm linkage method and system based on face recognition to solve the deficiencies in the prior art and improve the active early warning ability and emergency response efficiency of the campus security system.
[0004] An embodiment of the present application provides a campus intelligent alarm linkage method based on face recognition, and the method includes: Real-time collect video streams through cameras deployed in key areas of the campus, and extract face images in video frames using a face detection algorithm; Compare the face images with the pre-stored white list database in the campus in real time to generate an identity recognition result and a confidence score of the face; Based on the identity recognition result and the confidence score, combined with real-time analysis of the personnel behavior scenario, calculate the threat level of the current behavior scenario through a dynamic early warning determination model and generate a classified early warning signal; Automatically trigger corresponding device linkage instructions according to the classified early warning signal, where the device linkage instructions include pushing warning information to the security terminal, starting an audible and visual alarm device, or locking the access control of the relevant area; Generate a complete security event disposal report including the location of the incident, the images of the involved personnel, and disposal suggestions based on the execution status of the device linkage instructions and subsequent video analysis results.
[0005] Another embodiment of the present application provides a campus intelligent alarm linkage system based on face recognition, and the system includes: The extraction module is used to collect video streams in real time through cameras deployed in key areas of the campus and to extract face images from video frames using a face detection algorithm; The comparison module is used to compare the face image with the whitelist database pre-stored on campus in real time to generate the face identification result and confidence score. The analysis module is used to calculate the threat level of the current behavior scenario and generate a graded early warning signal based on the identity recognition results and confidence scores, combined with real-time analysis of personnel behavior scenarios, through a dynamic early warning judgment model. The early warning module is used to automatically trigger corresponding device linkage instructions based on the graded early warning signals. The device linkage instructions include pushing alarm information to the security terminal, activating the audible and visual alarm device, or locking the access control of the relevant area. The generation module is used to generate a complete security incident handling report, including the location of the incident, images of the people involved, and handling suggestions, based on the execution status of the device linkage command and the subsequent video analysis results.
[0006] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.
[0007] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.
[0008] Compared with existing technologies, the present invention provides a campus intelligent alarm linkage method based on facial recognition, which can improve the proactive early warning capability and emergency response efficiency of campus security systems. Attached Figure Description
[0009] Figure 1 A hardware structure block diagram of a computer terminal for a campus intelligent alarm linkage method based on face recognition provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a campus intelligent alarm linkage method based on face recognition, provided as an embodiment of the present invention; Figure 3 This is a schematic diagram of a campus intelligent alarm linkage system based on face recognition, provided as an embodiment of the present invention. Detailed Implementation
[0010] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0011] Figure 1This is a hardware structure block diagram of a computer terminal for a campus intelligent alarm linkage method based on face recognition, provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0012] See Figure 2 The present invention provides a campus intelligent alarm linkage method based on face recognition, which may include the following steps: S201 uses cameras deployed in key areas of the campus to collect video streams in real time and uses face detection algorithms to extract face images from video frames; Specifically, network cameras can be deployed in key areas of the campus, including at least entrances, corridors, and playgrounds, to generate a camera deployment plan; The deployment of cameras in key areas of the campus needs to balance "complete coverage" and "rational resource allocation." First, the different security needs of each area must be clearly defined: entrance areas need to clearly capture facial features (for identity verification), corridors need to cover people's movement trajectories (to prevent congestion or unusual lingering), and the playground needs to monitor large-scale activities (to avoid gatherings and conflicts). Based on this, the technical parameters and deployment density of the cameras should be determined. Entrance area: Use 4-megapixel network cameras (2560×1440 resolution, supports H.265 encoding, high compression ratio and clear image quality), 25fps frame rate (to ensure blur-free capture of dynamic faces), 90° field of view (one camera covers the width of a single school gate entrance, approximately 3-5 meters), with a deployment density of 1 unit per school gate entrance, installed at a height of 2.5 meters (to avoid obstruction and ensure the face is captured from close-up), for example, 1 unit at the main entrance of the campus and 1 unit at each of the side gates, for a total of 3 units; Corridor area: Use 2-megapixel cameras (1920×1080 resolution, meeting the requirements for trajectory tracking), 20fps frame rate, 120° field of view (wide-angle coverage of the corridor width, about 2-3 meters), with a deployment density of 1 unit every 50 meters and an installation height of 2.2 meters (installed along the center of the corridor top to avoid obstruction by pillars). For example, in a 100-meter-long main corridor of the teaching building, deploy 1 unit at 25 meters and 75 meters, for a total of 2 units. Playground area: Use 8-megapixel panoramic cameras (resolution 3840×2160, supports 360° panoramic stitching), frame rate 15fps (high frame rate is not required for large-area monitoring), field of view 360°, deployment density is 1 unit per 2000 square meters, installation height 8 meters (top of the playground stands, covering the entire playground area). For example, for a playground with an area of 4000 square meters, deploy 1 unit on each of the east and west stands, for a total of 2 units.
[0013] The final camera deployment plan must include "area name, camera model, quantity, installation location, parameter configuration, and coverage area," for example, "Main entrance: Model DS-2CD3T46WD-I5, quantity 1, installation location 2.5 meters from the right pillar of the main gate, parameters 4 megapixels / 25fps / 90° field of view, coverage area 3 meters wide at the main entrance; Teaching building main corridor: Model DS-2CD3T26WD-I3, quantity 2, installation locations at 25 meters and 75 meters at the top center, parameters 2 megapixels / 20fps / 120° field of view, coverage area 0-50 meters and 50-100 meters of the corridor section," ensuring that the plan can directly guide the construction and deployment.
[0014] Based on the camera deployment scheme, video stream data is collected in real time via the RTSP protocol, and the video stream is preprocessed to generate standardized video stream data; RTSP (Real-Time Streaming Protocol) is the standard video transmission protocol for network cameras, enabling low-latency (≤300ms) real-time data acquisition, meeting the real-time requirements of campus security. First, configure the RTSP service for the cameras: enable RTSP port 554 (default port) for each camera, set the transmission method to TCP (ensuring no data loss and avoiding packet loss issues associated with UDP transmission), use username + password authentication (to prevent unauthorized access, e.g., username admin, password 123456), and use H.265 video stream encoding format (50% higher compression efficiency than H.264, reducing bandwidth usage; campus LAN bandwidth needs to be ≥100Mbps, supporting simultaneous transmission from all cameras).
[0015] Video stream acquisition is achieved through the security platform's streaming media server (using an NVR, a network video recorder, supporting 32 simultaneous access channels). The server periodically sends RTSP request commands to each camera (e.g., DESCRIBErtsp: / / admin:123456@192.168.1.101:554 / Streaming / Channels / ). 101, where 192.168.1.101 is the camera IP and 101 is the channel number), after the camera responds, it starts transmitting video stream data. The server receives the data and stores it on the local hard drive (capacity 4TB, supports 7-day cycle overwrite), and forwards it to the subsequent processing module.
[0016] Preprocessing aims to eliminate noise and format differences in the video stream, generating a video stream with a unified standard. The specific steps are as follows: Noise Removal: A Gaussian filtering algorithm is used with a 3×3 filter kernel size and a standard deviation σ=1.0 (the smaller the σ value, the smoother the filtering, which can effectively eliminate snow noise in the video, such as intra-frame noise caused by flickering corridor lights. After processing, the proportion of noise pixels is reduced from 5% to 0.5%). Frame rate unification: The video stream frame rate of different cameras is uniformly adjusted to 25fps (the original 25fps of the entrance camera does not need to be adjusted, the 20fps of the corridor camera is interpolated to 25fps, and the 15fps of the playground camera is interpolated to 25fps through repeated frames to ensure that the input frame rate of the subsequent detection algorithm is consistent). Resolution standardization: The resolution of all video streams was adjusted to 1920×1080 (the original 2560×1440 at the entrance was reduced by bilinear interpolation, and the original 3840×2160 at the playground was reduced to 1920×1080 by segmenting and splicing, while retaining key monitoring areas). Color space conversion: Convert RGB format to YUV420P format (the standard input format for subsequent face detection algorithms; the Y channel stores luminance information, and the UV channel stores chrominance information, reducing the amount of data while preserving facial features).
[0017] The final generated standardized video stream data has a frame rate of 25fps, a resolution of 1920×1080, a YUV420P format, and a latency of ≤300ms. For example, the standardized video stream of the main entrance at a certain moment contains one frontal face per frame, with uniform brightness and no obvious noise, and can be directly input into the face detection algorithm.
[0018] Using standardized video stream data, a YOLOv5-based face detection algorithm is used to analyze video frames in real time, locate and extract face regions, and generate an initial set of face images. YOLOv5 is a real-time object detection algorithm. The lightweight YOLOv5s model (approximately 7.5M parameters, approximately 14.1 GFLOPs computation) was selected to adapt to the computing resources of the campus security server (CPU is Intel Xeon E3-1230v5, GPU is NVIDIA GTX1660, single frame detection time ≤20ms, meeting the 25fps real-time requirement). The model was fine-tuned using the WIDERFace dataset (containing 32203 face images, covering different poses and lighting scenes), achieving a face detection accuracy of ≥98% and a false detection rate of ≤1%.
[0019] The face detection process consists of two steps: Model Inference: Each frame of the normalized video stream (1920×1080×3) is scaled to the YOLOv5s input size of 640×640 (using letterbox padding to maintain aspect ratio and avoid face stretching and distortion) and input into the model for inference. The model outputs "category (face), confidence score, and bounding box coordinates (x1, y1, x2, y2)" for each detected target. The confidence score threshold is set to 0.5 (to filter low-confidence false detections, such as removing "suspected face" regions with a confidence score of 0.4), and the IOU (Intersection over Union) threshold is set to 0.45 (to handle overlapping faces, such as the overlapping area when two people walk side by side; bounding boxes with high confidence are retained when IOU>0.45). For example, in a certain frame of an image, the model detects two valid faces. The confidence score of the first face is 0.92, and the bounding box coordinates are (300, 200, 500, 400) (x1=300, y1=200 are the coordinates of the top left corner, x2=500, y2=400 are the coordinates of the bottom right corner); the confidence score of the second face is 0.88, and the bounding box coordinates are (600, 220, 800, 420). Face region cropping: Based on the bounding box coordinates, the corresponding face region is cropped from the original 1920×1080 frame image. To avoid truncating the face edge features, the bounding box is extended outward by 5 pixels (for example, the extended coordinates of the first face are (295,195,505,405)). After cropping, two initial face images are obtained, with sizes of 210×210 pixels (505-295=210, 405-195=210) and 210×210 pixels (805-595=210, 425-215=210), respectively.
[0020] The detected and captured face images from all video frames are aggregated to generate an initial face image set. Each image in the set is labeled with "source camera IP, frame timestamp, confidence score, and bounding box coordinates", for example, "IP: 192.168.1.101, timestamp: 20251001080000000, confidence score: 0.92, coordinates: (295, 195, 505, 405), image size: 210×210", providing relevant information for subsequent quality assessment.
[0021] The initial set of face images is quality-assessed, and face images with a sharpness higher than a preset sharpness threshold are selected. Size normalization and illumination compensation are then performed to output standardized face images.
[0022] Initial facial images may have problems such as blurriness (e.g., people moving quickly) or uneven lighting (e.g., backlighting causing dark areas on the face). Valid images need to be screened through quality assessment and then preprocessed to unify the format to provide high-quality input for subsequent identity comparison.
[0023] Quality Assessment: The BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator) algorithm is used for quality assessment. This algorithm calculates a quality score by analyzing the natural scene statistical features (such as contrast and texture) of the image. The score ranges from 0 to 100, with lower scores indicating clearer images. The preset clarity threshold is 30 (based on campus scene testing, images with a score ≤30 can meet the requirements for facial feature extraction). A BRISQUE score is calculated for each image in the initial set: for example, the first facial image has a score of 25 (clear, retained), while the second facial image is blurred due to rapid movement of people, with a score of 45 (above the threshold, discarded). Finally, one valid facial image is selected. Size normalization: The filtered face images are uniformly scaled to 112×112 pixels (the standard input size for subsequent deep convolutional neural networks, balancing feature preservation and computational efficiency). The scaling uses a bilinear interpolation algorithm to ensure that the relative positions of key facial features (such as eyes, nose, and mouth) remain unchanged. For example, after scaling a 210×210 image, the position of the left eye changes from (50,80) to (23,40), and the position of the right eye changes from (160,80) to (89,40), keeping the eyes horizontally aligned. Illumination compensation: The CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm is used to divide the image into 8×8 sub-blocks. Histogram equalization is performed on each sub-block, and the contrast limit is set to 40 (to avoid excessive enhancement that would amplify noise). The illumination of backlit or overexposed areas is compensated: For example, in a face image, the left side is backlit (grayscale value 50-100), and the right side is overexposed (grayscale value 200-255). After CLAHE processing, the grayscale value on the left side is increased to 80-150, and the right side is reduced to 150-200. The overall brightness of the face is uniform, and details in dark areas (such as eyebrows and chin contours) are clearly visible.
[0024] The final output standardized face image is 112×112 pixels in size, with a BRISQUE score ≤30, uniform illumination, and labeled with "source information, quality score, and processing algorithm parameters", such as "IP: 192.168.1.101, timestamp: 20251001080000000, BRISQUE score: 25, size: 112×112, illumination compensation parameters: CLAHE8×8 sub-block / contrast limit 40", which can be directly input into the subsequent face feature extraction module.
[0025] S202, compare the face image with the whitelist database pre-stored on campus in real time to generate the face identification result and confidence score; Specifically, a deep convolutional neural network can be used to extract facial feature vectors from standardized facial images to generate multidimensional feature representations for each facial image. The standardized face images have been normalized to size (112×112 pixels) and have undergone illumination compensation. High-discrimination feature vectors need to be extracted using a lightweight deep convolutional neural network (CNN). The MobileFaceNet model was chosen because it is designed for mobile / embedded scenarios. It uses depthwise separable convolutions instead of standard convolutions, with a total of only 4.8M parameters (about 1 / 20 of VGG16). The single-frame inference time is ≤10ms (suitable for the real-time requirements of campus security and supports 25fps video stream processing). Moreover, the face recognition accuracy on the LFW (Labeled Faces in the Wild) dataset reaches 99.86%, which meets the accuracy requirements for campus whitelist comparison.
[0026] The MobileFaceNet network structure needs to be adapted to the requirements of facial feature extraction. The specific layers and functions are as follows: Input layer: Receives a 112×112×3 RGB normalized face image, with pixel values normalized to the range [-1,1] (formula: x_norm=(x / 255-0.5)×2, for example, if the original value of a pixel is 200, after normalization it is (200 / 255-0.5)×2≈0.57). Depthwise separable convolutional blocks (6 blocks): Each block contains "depthmial convolution (3×3 kernel, extracting spatial features per channel) + pointwise convolution (1×1 kernel, fusing channel information) + BatchNorm + ReLU6". For example, the first convolutional block outputs 64 channels with stride=2, compressing the input size from 112×112 to 56×56 and extracting low-level features (such as edges and textures). Subsequent blocks gradually increase the number of channels to 512 and compress the size to 7×7, extracting high-level semantic features (such as combined features of the eyes, nose, and mouth). Global average pooling layer: compresses the 7×7×512 feature map into a 1×1×512 vector, eliminating spatial dimension and retaining channel dimension feature information; Fully connected layer (no activation): outputs a 512-dimensional feature vector, with vector elements ranging from [-1, 1]. Each dimension corresponds to a key facial feature (e.g., the 10th dimension corresponds to the "horizontal distance between the outer corner of the left eye and the outer corner of the right eye", and the 200th dimension corresponds to the "curvature of the bridge of the nose").
[0027] Taking a standardized face image (a frontal photo of student Zhang San, with a BRISQUE quality score of 25) as an example, after extraction by MobileFaceNet, a 512-dimensional feature vector fragment is obtained as: [0.23, -0.15, 0.42, -0.08, 0.31, -0.12, 0.28, -0.05, ..., 0.19, -0.07]. This vector needs to be stored as a floating-point array, labeled with "source camera IP and frame timestamp", for example, "IP: 192.168.1.101, timestamp: 20251001080000000, feature vector: [0.23, ..., -0.07]", to provide input for subsequent similarity calculations.
[0028] Load the facial feature vectors of registered individuals from the pre-stored whitelist database on campus to construct a feature retrieval database; The campus whitelist database stores the facial features and identity information of registered personnel (students, faculty, staff, and visitors). Data security and efficient retrieval must be ensured. The database adopts a "local storage + cloud backup" architecture. Local storage is performed on a MySQL database (version 8.0, supporting transactions and encryption) on the campus security server, and cloud backup is performed on the campus private cloud (using AES-256 encryption to prevent data leakage).
[0029] The core fields of the whitelist database include: Unique personnel ID (e.g., "STU-2025001" represents student number 001 of the 2025 cohort, "TCH-001" represents faculty / staff member number 001); personnel name (e.g., "Zhang San"); 512-dimensional face feature vector (with the same dimensions as the MobileFaceNet output, stored as a string, such as "0.23,-0.15,0.42,...,-0.07"); Registration time (e.g., "20250901"); Identity type (e.g., "student", "faculty / staff", "temporary visitor"); Feature vector update time (e.g., "20250910", supports periodic updates to adapt to facial changes, such as hairstyle changes).
[0030] When building the feature retrieval database, the feature vectors in MySQL need to be loaded into memory, and an efficient retrieval index needs to be built using the FAISS (FacebookAISimilaritySearch) library. FAISS is specifically designed for high-dimensional vector retrieval, supporting millisecond-level retrieval of millions of vectors, and is suitable for campus whitelists (assuming 10,000 registered users, corresponding to 10,000 512-dimensional feature vectors). The index type should be selected as "IVF_FLAT" (inverted index + flat search). Specific parameter configurations are as follows: The number of cluster centers, nlist, is 100 (10,000 vectors are clustered into 100 centers; during retrieval, the center is matched first, and then the search is performed precisely within the cluster to which the center belongs, balancing speed and accuracy); the distance metric is cosine similarity (consistent with the similarity algorithm used for subsequent real-time comparison); and the vector dimension, d, is 512 (matching the output dimension of MobileFaceNet).
[0031] Example of index construction process: After loading 10,000 feature vectors, FAISS uses the K-means algorithm to cluster the vectors into 100 centers, each center corresponding to a cluster (with an average of 100 vectors per cluster), generating the index file "face_index.faiss" and storing it in memory. The average response time during retrieval is ≤5ms, meeting the requirements for real-time comparison. For example, when loading the feature vector of faculty member Li Si (ID: TCH-001, vector: [0.21,-0.17,0.40,...,-0.09]), FAISS assigns it to cluster number 15. During subsequent real-time vector retrieval, the similarity with the 100 centers is calculated first, the most similar center (such as cluster number 15) is found, and then 100 vectors are compared within that cluster, significantly reducing the amount of computation.
[0032] The facial feature vectors extracted in real time are compared with the feature retrieval database to calculate similarity and generate a list of similarity scores. The core of similarity calculation is to measure the similarity between the "real-time feature vector" and the "whitelist feature vector". The cosine similarity algorithm is used. This algorithm calculates the cosine value of the angle between the two vectors. The value ranges from -1 to 1. The closer the value is to 1, the more consistent the vector direction is (the more similar the faces are). It is not affected by the vector length (avoiding feature vector scaling due to differences in image brightness, which affects the comparison results).
[0033] The cosine similarity calculation formula is: sim(V_real,V_white)=(V_real·V_white) / (||V_real||×||V_white||), where V_real is the 512-dimensional feature vector extracted in real time, V_white is a 512-dimensional feature vector in the whitelist; V_real・V_white is the dot product of the two vectors (i.e. Σ(V_real[i]×V_white[i]), i from 1 to 512); ||V_real|| is the L2 norm of V_real (i.e. √Σ(V_real[i]²)), and ||V_white|| is the L2 norm of V_white.
[0034] The calculation process needs to be accelerated by the FAISS index. Specific steps are as follows: Input the real-time feature vector V_real (e.g., [0.23, -0.15, 0.42, ..., -0.07]) into the FAISS index. FAISS first calculates the cosine similarity between V_real and 100 cluster centers, and finds the most similar center (e.g., cluster number 15, similarity 0.85). Within cluster 15, calculate the cosine similarity between V_real and 100 whitelist vectors within that cluster to obtain 100 similarity scores; Sort the similarity scores of all whitelisted vectors (a total of 10,000, with vectors not in cluster 15 having a similarity score of 0) in descending order, and take the top 10 highest scores (to avoid outputting too much redundant data) to generate a similarity score list.
[0035] Example similarity score list: Person ID: STU-2025001 (Zhang San), similarity: 0.92; Person ID: STU-2025002 (Li Si), similarity: 0.65; Person ID: TCH-001 (Wang Wu), similarity: 0.43; Person ID: STU-2025003 (Zhao Liu), similarity: 0.38; Person ID: STU-2025004 (Sun Qi), similarity: 0.35; ... (the following 5 scores are all <0.3).
[0036] This list needs to be labeled with "Real-time Vector ID, Calculation Time, and Score Sorting Criteria", for example, "Real-time Vector ID: REAL-20251001080000, Calculation Time: 20251001080000005, Sorting Criteria: Cosine Similarity Descending Order", to provide data support for subsequent dynamic threshold settings.
[0037] Based on the similarity score list, a dynamic similarity threshold is set to determine whether a face matches a whitelisted identity and generate a preliminary identity recognition result. A fixed similarity threshold (e.g., 0.8) is prone to misjudgment (e.g., high similarity in low-quality face images may be due to noise) or missed judgment (e.g., high-quality faces may have a similarity slightly lower than the threshold due to angle differences). Therefore, a dynamic threshold needs to be set based on the distribution characteristics of the similarity score list. The core logic is to "adjust the threshold according to the difference between the highest and second-highest scores and the overall distribution density of scores" to ensure that the threshold can both filter noise and retain effective matches.
[0038] The dynamic threshold setting uses the "Elbow Method" combined with "score difference judgment". Specific steps are as follows: Extract the top 20 non-zero scores from the similarity score list (covering all possible valid matches) and plot the curve of score change with ranking; Find the "elbow" of the curve (i.e. the inflection point where the rate of decrease in the score slows down), and use the score corresponding to this inflection point as a candidate for the initial threshold. Calculate the difference ΔS = S1 - S2 between the highest score (S1) and the second highest score (S2). If ΔS ≥ 0.2 (indicating that the highest score is significantly different from other scores and the match is unique), then lower the initial threshold candidate by 5% (e.g., if the inflection point score is 0.85, lower it to 0.8075) to improve the matching success rate. If ΔS < 0.2 (indicating that there are multiple similar candidate faces and strict screening is required), then raise the initial threshold candidate by 5% (e.g., if the inflection point score is 0.85, raise it to 0.8925) to reduce the false positive rate.
[0039] Taking the example score list from the aforementioned steps as an example: Of the first 20 scores, the first 5 scores were 0.92, 0.65, 0.43, 0.38, and 0.35, and the subsequent scores were all <0.3. The rate of decline of the curve increased sharply between the first and second places (from 0.92 to 0.65), and the score corresponding to the elbow was 0.85. The highest score S1 = 0.92, the second highest score S2 = 0.65, ΔS = 0.27 ≥ 0.2, therefore the dynamic threshold = 0.85 × (1 - 5%) = 0.8075, which is 0.81 after rounding to two decimal places.
[0040] The logic for judging the preliminary identity recognition results: If the highest score S1 in the similarity score list is greater than or equal to the dynamic threshold, then it is determined that "the identity of the whitelist is matched", and the preliminary result is "Personnel ID: [ID corresponding to the highest score], Matching status: Success"; If S1 < dynamic threshold, then it is determined that "the whitelist identity is not matched", and the preliminary result is "matching status: failed, highest similarity: [S1]".
[0041] In the example, S1=0.92≥0.81, so the preliminary identity recognition result is "Personnel ID: STU-2025001 (Zhang San), Matching status: Successful, Similarity score: 0.92, Dynamic threshold: 0.81"; If the highest score of a real-time vector S1=0.78<0.81, the preliminary result is: "Matching status: Failed, Highest similarity: 0.78, Dynamic threshold: 0.81".
[0042] Based on the preliminary identity recognition results, combined with the face image quality score and similarity score, the final identity recognition result and confidence score are output through the confidence calculation model.
[0043] Specifically, the step of outputting the final identity recognition result and confidence score based on the preliminary identity recognition result, combined with the face image quality score and similarity score, through a confidence calculation model, may include: Based on the preliminary identity recognition results, the highest similarity score in the similarity score list and the corresponding whitelist identity are extracted to generate a candidate identity set and associated similarity scores; The initial identity recognition result is a preliminary judgment of the face matching status based on a dynamic similarity threshold, including core information such as "match successful / failed / pending verification," "corresponding whitelist identity (if matched)," and "highest similarity score." To ensure the comprehensiveness of subsequent confidence calculations, key information needs to be extracted from the similarity score list to generate a candidate identity set. This set focuses on the identity corresponding to the "highest similarity score," while also including the second-highest similarity identity as an alternative, avoiding misjudgments due to the randomness of a single matching result.
[0044] First, clarify the typical form of the preliminary identity recognition result: If the preliminary result is "Match successful, Person ID: STU-2025001 (Zhang San), Highest similarity score: 0.92, Dynamic threshold: 0.81", then the first two in the similarity score list are "STU-2025001 (0.92), STU-2025002 (Li Si, 0.65)"; if the preliminary result is "Match failed, Highest similarity score: 0.72, Dynamic threshold: 0.75", then the first in the list is "No corresponding identity, 0.72", and the second is "STU-2025003 (Wang Wu, 0.68)".
[0045] The rules for generating the candidate identity set are as follows: Regardless of whether the initial results match, identities with a similarity score ≥ 0.5 are extracted from the list (identities with a similarity score below 0.5 are too different and have no reference value), and at most the top 3 are retained (to balance computational efficiency and accuracy). Each candidate identity needs to be associated with a corresponding similarity score and marked with "whether it is a preliminary match identity" (e.g., STU-2025001 is marked "yes", STU-2025002 is marked "no"); If the initial result is a match failure, the candidate identity set must include "no matching identity" as the first one, with the association score being the highest similarity (e.g., 0.72), followed by the identity corresponding to the second highest score.
[0046] The following example of "Successful Match" generates a candidate identity set: "Candidate Identity Set: 1. Identity ID: STU-2025001 (Zhang San), Association Similarity Score: 0.92, Initial Match: Yes; 2. Identity ID: STU-2025002 (Li Si), Association Similarity Score: 0.65, Initial Match: No; 3. Identity ID: TCH-001 (Wang Wu), Association Similarity Score: 0.52, Initial Match: No." This set needs to be labeled with "Source Similarity List ID, Extraction Time," for example, "Source List ID: SIM-20251001080000, Extraction Time: 20251001080001," to ensure traceability to the original comparison data.
[0047] The quality score of a face image is calculated using an image quality assessment algorithm, and a standardized quality score is generated. The quality of facial images directly affects the reliability of feature extraction and comparison. Blurry, unevenly lit, or excessively posed images may lead to "high similarity but incorrect matching" (e.g., noise in a low-resolution image is misjudged as a feature). Therefore, image quality assessment algorithms are needed to quantify quality and provide a basis for correction in subsequent confidence calculations. The BRISQUE (Blind / Referenceless Image Spatial Quality Evaluator) algorithm was selected. This algorithm does not require an original clear image as a reference; it directly assesses quality by analyzing the natural scene statistical features of the image (such as local mean, variance, and gradient). It is suitable for scenarios in campus security where there are "no preset clear templates," and its fast computation speed (single frame processing time ≤ 5ms) meets real-time requirements.
[0048] The BRISQUE algorithm scores images from 0 to 100, with lower scores indicating better image quality (e.g., a clear frontal photo scores 20-30, while a blurry profile photo scores 60-70). The calculation process is as follows: The standardized face image (112×112 pixels, with illumination compensation) is preprocessed to remove black borders (if any) from the image edges and retain the effective face area of 100×100 pixels in the center. Extracting statistical features of natural scenes: Calculating the local normalized brightness of the image (the difference between each pixel and the mean of the 3×3 neighborhood divided by the neighborhood standard deviation) and gradient magnitude distribution (Sobel gradient in the horizontal and vertical directions), generating a total of 36-dimensional feature vectors. The feature vectors are regressed using a support vector regression (SVR) model to output a BRISQUE quality score (the model has been pre-trained on the TID2013 image quality dataset with an error ≤ 3 points).
[0049] For example, a standardized face image (a frontal photo of Zhang San, without blur and with uniform lighting) scored 25 after being calculated by BRISQUE; another image (a side profile photo of Li Si, slightly blurry) scored 58.
[0050] To facilitate inputting the similarity score (0-1) into the confidence model, the BRISQUE score needs to be standardized—converting "low score, high quality" into a positive indicator of "high score, high quality." The formula is: Standardized Quality Score = (100 - BRISQUE score) / 100. Here, 100 is the highest BRISQUE score; subtracting the original score yields the positive score, which is then normalized to the 0-1 range by dividing by 100. In the example above, Zhang San's image standardized quality score is (100-25) / 100 = 0.75, and Li Si's is (100-58) / 100 = 0.42. The standardized score should be labeled with "Original BRISQUE score, normalization formula," for example, "Standardized Quality Score: 0.75 (Original BRISQUE: 25, Formula: (100-25) / 100)."
[0051] The association similarity scores and standardized quality scores of the candidate identity sets are input into a pre-trained confidence calculation model, which uses a gradient boosting decision tree algorithm to output an initial confidence score. The core of the confidence calculation model is to integrate "similarity score (reflecting the degree of feature matching)" and "standardized quality score (reflecting image reliability)" to output a quantified confidence score (0-100 points), measuring the reliability of the identity recognition result. The Gradient Boosting Decision Tree (GBDT) algorithm is selected. This algorithm, through ensemble learning of multiple decision trees, can capture the non-linear relationship between two input features (e.g., "images with high similarity but low quality should have their confidence score reduced"), and has higher prediction accuracy compared to linear models (such as logistic regression).
[0052] 1. Model Training Preparation Training dataset construction: 100,000 samples were collected. Each sample contained "candidate identity association similarity score (0-1), standardized quality score (0-1), and true label (1 = correct match, 0 = incorrect match)". For example, "Similarity 0.92, quality 0.75, label 1 (correct match)" and "Similarity 0.85, quality 0.30, label 0 (incorrect match, due to image blur)". Data partitioning: The data was divided into a training set (80,000 sets) and a validation set (20,000 sets) in an 8:2 ratio to ensure a consistent data distribution; Model parameter configuration: Number of decision trees n_estimators=150 (determined through validation set tuning; too many trees can lead to overfitting, while too few can lead to underfitting), depth of a single tree max_depth=5 (to limit complexity and avoid learning noise), learning rate learning_rate=0.05 (to control the contribution weight of each tree and optimize it step by step), and the loss function is log loss (adapted to binary classification confidence prediction).
[0053] 2. Model Training and Inference The training process minimizes the log loss through gradient descent: initially, a decision tree is built to predict the confidence of the sample. After calculating the prediction error, the next decision tree is trained with the goal of "reducing the error". This process is repeated 150 times. Finally, the model achieves an accuracy of 97% on the validation set, with a confidence prediction error of ≤2 points.
[0054] During inference, the association similarity score and standardized quality score of the "preliminary matching identity" in the candidate identity set are used as input (if the match fails, the highest similarity score and the corresponding image quality score are input). Taking the example of the previous steps: if the input is "similarity 0.92, quality 0.75", the model will output a preliminary confidence score of 95 (indicating that the matching result is reliable) through ensemble calculation of 150 decision trees; if the input is "similarity 0.85, quality 0.30", the model will output a preliminary confidence score of 62 (indicating that the matching result should be approached with caution).
[0055] The initial confidence score should be labeled with "input feature value and model version", for example, "Initial confidence score: 95 points (input: similarity 0.92, quality 0.75; model version: GBDT_v2.1)".
[0056] The initial identity recognition results are adjusted based on the preliminary confidence score and the preset confidence threshold, and the confidence is smoothed by combining time series analysis to generate the final identity recognition results and the corresponding confidence score.
[0057] The initial confidence score may fluctuate due to occasional noise in a single frame image (such as momentary blur), and needs to be processed in two steps: "threshold adjustment + time smoothing" to ensure the stability and accuracy of the final result.
[0058] 1. Adjust the preliminary identity recognition results based on a preset confidence threshold. The pre-set reliability thresholds are set at three levels based on campus security needs, balancing the "false positive rate" and the "false negative rate": High threshold (80 points): If the score is ≥80 points, it is judged as "reliable result" and the preliminary identity recognition result is directly retained; Medium threshold (60-79 points): If the score is in this range, it is judged as "result to be verified", the preliminary result is retained but marked "requires manual review"; Low threshold (<60 points): If the score is <60 points, it is judged as "unreliable result" and the preliminary result is corrected to "unmatched whitelist identity". Even if the preliminary result is a successful match, it must be overturned.
[0059] For example, if the initial confidence score is 95 or higher than 80, the adjusted result is "Match successful, identity: STU-2025001"; if the initial confidence score is 62, which falls between 60 and 79, the adjusted result is "Match successful, identity: STU-2025001, tag: pending verification"; if the initial confidence score is 55 or lower than 60, the adjusted result is "No whitelist identity matched, original matched identity: STU-2025001 (unreliable)".
[0060] 2. Smooth the confidence level using time series analysis. The identification results of the same person in consecutive video frames (e.g., 50 frames within 10 seconds) should be continuous. Abrupt changes in confidence scores in a single frame (e.g., a sudden drop from 95 to 60 points) may be caused by noise and need to be smoothed through time series analysis. A moving average method is used: the average of the initial confidence scores for N consecutive frames (N=5, covering 0.2 seconds, balancing real-time performance and stability) is calculated as the smoothed confidence score.
[0061] The calculation formula is: Smoothed confidence score = (score of frame t + score of frame t-1 + ... + score of frame t-N+1) / N. For example, if the initial confidence scores of 5 consecutive frames are 95, 94, 96, 95, and 94, the smoothed score is (95 + 94 + 96 + 95 + 94) / 5 = 95. If a frame suddenly drops to 60 due to blurring, and the scores of 5 consecutive frames are 95, 94, 60, 95, and 94, the smoothed score is (95 + 94 + 60 + 95 + 94) / 5 = 87.6 (rounded to 88), thus mitigating the impact of the sudden drop.
[0062] 3. Generate the final identity verification result The system integrates the adjusted identity results with the smoothed confidence scores, outputting a final result that includes "Identity ID, Name, Identity Type, Final Confidence Score, Result Label, and Smoothing Processing Explanation". For example: "Final Identity Recognition Result: Identity ID: STU-2025001, Name: Zhang San, Identity Type: Student, Final Confidence Score: 95 (smoothed from 5 consecutive frames, original scores 95 / 94 / 96 / 95 / 94), Result Label: Reliable".
[0063] If the result is to be verified: "Final identity recognition result: Identity ID: STU-2025001, Name: Zhang San, Identity type: Student, Final confidence score: 63 points (smoothed from 5 consecutive frames, original score 65 / 62 / 60 / 64 / 64), Result label: to be verified, suggestion: manually review image quality (BRISQUE score 58)".
[0064] The final results need to be stored in the campus security database and synchronized to the real-time monitoring interface to provide a reliable identity basis for subsequent early warning judgments.
[0065] S203, Based on the identity recognition results and confidence scores, and combined with real-time analysis of personnel behavior scenarios, the threat level of the current behavior scenario is calculated through a dynamic early warning judgment model, and a graded early warning signal is generated. Specifically, based on the identity recognition results and confidence scores, an identity credibility index can be extracted to generate an identity credibility vector. The identity recognition results contain three core information categories: "matching status, corresponding identity information, and confidence score." It is necessary to extract quantitative indicators that reflect "identity reliability" from these. Different matching statuses and identity types have significantly different threat risks in campus scenarios (e.g., the credibility of faculty and staff appearing in the classroom during class time is much higher than that of strangers). Therefore, the indicator design needs to take into account both "matching quality" and "identity suitability," and finally generate a fixed-dimensional vector to ensure that the input format of subsequent models is consistent.
[0066] First, we need to clarify the three typical scenarios for identity recognition results and the rules for extracting indicators: Successful matching scenario: For example, if the result is "ID: STU-2025001 (Zhang San), Identity type: Student, Confidence score: 95 points, Result label: Reliable", extract three indicators: Identity matching status value: 1 is recorded for a successful match (0 for no match, 0.5 for pending verification), which quantifies the validity of the match; Normalized confidence score: The original confidence score (0-100) is normalized to the interval [0,1]. The formula is norm_conf = confidence score / 100. In this example, norm_conf = 95 / 100 = 0.95, which reflects the reliability of the matching result. Identity type weight: The weight is set according to the suitability of the identity and the scenario (faculty and staff 1.0, students 0.8, temporary visitors 0.5, no match 0). In this example, the identity type is student, and the weight is 0.8, which quantifies the reasonableness of the identity in the current scenario.
[0067] Matching failure scenario: If the result is "Matching status: Failed, highest similarity: 0.72, no corresponding identity", the extracted metrics are "Identity matching status value = 0, confidence normalization value = 0.72 (take the highest similarity), identity type weight = 0". Scenario to be verified: If the result is "ID: TCH-001 (Li Si), identity type: faculty and staff, confidence score: 65 points, result label: to be verified", the extracted indicators are "identity matching status value = 0.5, confidence normalization value = 0.65, identity type weight = 1.0".
[0068] The identity credibility vector is a 3-dimensional vector, concatenated in the order of "matching status value - confidence normalized value - identity type weight". For example, in a successful match scenario, the vector is [1, 0.95, 0.8]; in a failed match scenario, the vector is [0, 0.72, 0]. This vector needs to be labeled with "generation timestamp and corresponding personnel ID (if any)," for example, "timestamp: 20251001143005000, personnel ID: STU-2025001, identity credibility vector: [1, 0.95, 0.8]", to ensure temporal consistency with subsequent behavioral and scenario features.
[0069] The YOLOv4 behavior detection algorithm is used to extract human behavior features from the video stream, including at least motion trajectory, speed and abnormal posture, and generate behavior feature vectors. YOLOv4 is a mainstream algorithm for real-time object detection and behavior analysis. Its "backbone network (CSPDarknet53) + neck network (PANet) + head network (YOLOHead)" architecture can ensure high detection accuracy (43.5% mAP on the COCO dataset) and meet the real-time requirements of campus video streams (frame rate ≥20fps at 1920×1080 resolution). It is suitable for extracting dynamic behavioral features such as motion trajectory, speed, and abnormal posture.
[0070] 1. Behavioral Feature Extraction Process Human Target Detection: A standardized video stream (1920×1080, 25fps) is input into the YOLOv4 model. The model outputs the bounding box coordinates (x1, y1, x2, y2), confidence score (≥0.5 is considered a valid target), and category (human) for each human target. For example, the bounding box of a detected person is (300, 200, 500, 450) (x1=300, y1=200 is the top left corner, x2=500, y2=450 is the bottom right corner), with a confidence score of 0.98.
[0071] Motion trajectory extraction: Every 5 frames (time interval Δt=5 / 25=0.2 seconds), the center coordinates of the bounding box of the human target (x_center=(x1+x2) / 2, y_center=(y1+y2) / 2) are recorded. Ten sets of coordinates are recorded continuously (covering a 2-second duration) to form the motion trajectory. For example, the 10 sets of coordinates are (300,225),(320,230),(350,235),(380,240),(420,245),(450,250),(480,255),(520,260),(550,265),(580,270), and the trajectory shows a rapid movement along the positive x-axis.
[0072] Speed calculation: Speed is divided into "pixel speed" and "actual speed" - pixel speed is the Euclidean distance between the end point and the start point of the trajectory divided by the total time. The formula is pixel_speed=√[(x_end-x_start)²+(y_end-y_start)²] / (total time). In this example, x_end=580, x_start=300, y_end=270, y_start=225, and the total time is 2 seconds. Therefore, pixel_speed=√[(280)²+(45)²] / 2≈283.7 / 2≈141.8 pixels / second. Through camera calibration (pre-calibrated using a checkerboard pattern to determine that 1 pixel corresponds to an actual distance of 0.02 meters), the actual speed is converted to real_speed=141.8×0.02≈2.84 meters / second (the normal walking speed in the campus corridor is about 1.2-1.5 meters / second, and this speed is abnormal).
[0073] Abnormal posture recognition: Utilizing the YOLOv4 keypoint detection branch (outputting 17 human keypoints, such as head, torso, and limbs), "abnormal posture" is defined as: ① Running (leg keypoint angle < 30°, and speed > 2 m / s); ② Falling (head keypoint y-coordinate > torso keypoint y-coordinate, and speed drops sharply by > 50%); ③ Climbing (hand keypoint y-coordinate < shoulder keypoint y-coordinate, and close to a wall boundary). In this example, the leg keypoint angle is 25°, and the speed is 2.84 m / s, which is judged as "abnormal running," with an abnormal posture score of 0.9 (1 for completely abnormal, 0 for normal).
[0074] 2. Generation of behavioral feature vectors The behavioral feature vector is a 4-dimensional vector, concatenated according to the following parameters: "Actual speed (m / s) - Trajectory variance - Abnormal posture score - Trajectory duration (seconds)". Variance of the motion trajectory: Calculate the variance of 10 sets of x-coordinates (reflecting the stability of the trajectory; the larger the variance, the more chaotic the trajectory). In this example, the variance of the x-coordinate = √[(300-440)²+(320-440)²+...+(580-440)²] / 10≈126.5, which becomes 0.85 after normalization to [0,1] (the maximum variance is set to 150). Trajectory duration: 2 seconds (fixed to 2 seconds to ensure consistent vector dimensions); The final behavioral feature vector is [2.84, 0.85, 0.9, 2], labeled with "target ID and extraction time", for example, "target ID: PER-001, extraction time: 20251001143007000, behavioral feature vector: [2.84, 0.85, 0.9, 2]".
[0075] By combining the camera's geographic location information, the current scene type, such as a classroom or corridor, is identified, and the scene context, including time point and crowd density, is analyzed to generate scene context features; The context of a scene directly affects the "threat attribute" of a behavior. For example, "running" is normal on the playground (during breaks) but abnormal in the corridor (during class time). Therefore, features need to be extracted from three dimensions: "space (scene type), time (time period), and environment (crowd density)" to quantify the impact of the scene on threat determination.
[0076] 1. Scene type recognition Camera location information is pre-stored in the campus security database, and each camera IP corresponds to a unique scene type and code: Code 1: Classroom (e.g., IP: 192.168.1.103, corresponding to classroom 302 in the teaching building); Code 2: Corridor (e.g., IP: 192.168.1.102, corresponding to the main corridor of the teaching building); Code 3: Playground (e.g., IP: 192.168.1.104, corresponding to the east playground of the campus); Code 4: Entry point (e.g., IP: 192.168.1.101, corresponding to the main campus entrance).
[0077] In this example, the person involved is from IP: 192.168.1.102, and the scene type code is 2 (corridor).
[0078] 2. Time point analysis Divide the day into 6 time periods and code them according to "scenario adaptability": Code 1: During class hours (08:00-12:00, 14:00-17:30), the corridor should be kept quiet and running is prohibited. Code 2: During breaks between classes (12:00-14:00, 17:30-18:30), there is high pedestrian traffic in the corridors; slow walking is permitted. Code 3: Evening self-study time (19:00-21:30), similar to class time; Code 4: It is abnormal for non-residents to be on campus during the nighttime period (21:30-06:00).
[0079] In this example, the time is 14:30, which is class time, so the time slot code is 1.
[0080] 3. Crowd density calculation The number of human targets in the current video frame is detected using YOLOv4, and the crowd density is calculated by combining this with the scene area. Scene area: Pre-store the actual area of each scene, such as the area of the main corridor of the teaching building = 50 meters × 3 meters = 150 square meters; Target count: YOLOv4 detected 8 people in the current frame; Crowd density = number of targets / scene area = 8 / 150 ≈ 0.053 people / square meter, coded by level: ① low density (<0.03 people / ㎡, code 1), ② medium density (0.03-0.08 people / ㎡, code 2), ③ high density (>0.08 people / ㎡, code 3), in this example code = 2.
[0081] 4. Scene Context Feature Generation The scene context features are 5-dimensional vectors, concatenated according to the following sequence: "Scene type encoding - Time period encoding - Crowd density encoding - Scene risk weight - Time period risk weight". Scene risk weights: These are set according to the likelihood of accidents occurring in each scene (classroom 0.6, corridor 0.9, playground 0.3, entrance 0.8). In this example, the corridor weight is 0.9. Risk weighting by time period: set according to the safety of the time period (class time 0.8, break time 0.5, evening self-study 0.7, nighttime 1.0), in this example, the weight of class time = 0.8; The final scene context feature vector is [2,1,2,0.9,0.8], labeled with "camera IP and analysis time", for example, "camera IP: 192.168.1.102, analysis time: 20251001143008000, scene context feature vector: [2,1,2,0.9,0.8]".
[0082] The identity credibility vector, behavioral feature vector, and scene context features are input into the dynamic early warning and judgment model based on XGBoost to calculate the threat probability score. XGBoost (ExtremeGradientBoosting) is an optimized version of gradient boosting trees. It has the advantages of handling high-dimensional non-linear features, resisting overfitting, and fast training speed. It is suitable for fusing three types of heterogeneous features: identity, behavior, and scene, and outputs a threat probability score in the range of 0-1 (the higher the score, the higher the threat level).
[0083] 1. Model Training Preparation Training dataset construction: 100,000 sets of security scene samples were collected on campus. Each set of samples contains a 12-dimensional concatenated feature consisting of an "identity credibility vector (3-dimensional) + behavioral feature vector (4-dimensional) + scene context feature vector (5-dimensional)" and a corresponding "threat label" (0=no threat, 1=low threat, 2=medium threat, 3=high threat). For example, "a stranger (identity vector [0,0.6,0]) running in a nighttime corridor (scene vector [2,4,1,0.9,1.0]) (behavior vector [3.5,0.9,1.0,2])" is labeled as high threat (label 3). Data preprocessing: Normalize all features (e.g., speed normalized to [0,5] m / s, weight normalized to [0,1]), and divide them into training set (70,000 sets) and validation set (30,000 sets) in a 7:3 ratio; Model parameter configuration: n_estimators=100: Build 100 decision trees to balance accuracy and computational cost; max_depth=6: Limits the depth of each tree to 6 to avoid overfitting (12-dimensional features do not require overly deep tree structures); learning_rate=0.1: The contribution weight of each tree is 0.1, and the loss is optimized step by step; objective='multi:softprob': Multi-class probability output mode, outputs the probability of each threat tag.
[0084] 2. Model Training and Inference Training process: The cross-entropy loss function is used, and the loss is minimized through gradient descent. After 100 iterations, the multi-class classification accuracy on the validation set reaches 92%, and the threat probability prediction error is ≤5%. Reasoning example: Concatenate the three types of vectors in this example into a 12-dimensional input feature [1,0.95,0.8,2.84,0.85,0.9,2,2,1,2,0.9,0.8], input it into the trained model, and output the threat label probability distribution as "no threat 0.02, low threat 0.08, medium threat 0.25, high threat 0.65". Take the label probability corresponding to the highest probability as the threat probability score = 0.65.
[0085] Threat probability scores must be labeled with "inference time and feature source", for example, "inference time: 20251001143009000, threat probability score: 0.65, feature source: PER-001+192.168.1.102".
[0086] Threat levels are classified based on threat probability scores, and corresponding graded early warning signals are generated and updated in real time.
[0087] Threat level classification should be combined with the actual needs of campus security, taking into account both "timeliness of early warning" and "avoidance of excessive early warning". The probability score of 0-1 is mapped to four levels, with each level corresponding to specific early warning measures and signal formats. At the same time, it should support real-time updates of the level based on the feature changes of subsequent video frames to ensure the dynamic nature of the early warning.
[0088] 1. Threat Level Classification Rules Level 1 (No Threat): Probability score 0-0.2, corresponding to normal behavior (such as students walking slowly in the corridor during class), no warning required; Level 2 (Low Threat): Probability score 0.2-0.4, corresponding to minor anomalies (such as students jogging in the corridor during breaks), the warning measure is "only recorded on the security platform, not pushed to the terminal"; Level 3 (Medium Threat): Probability score 0.4-0.7, corresponding to moderate abnormality (such as running in the corridor during class time in this example). The warning measure is to "push alarm information to the nearest security terminal and activate the corridor voice reminder ('Please keep quiet and do not run')". Level 4 (High Threat): Probability score 0.7-1.0, corresponding to serious anomalies (such as strangers climbing the corridor at night). The warning measures are to "push the alarm to all security terminals and the campus management platform, activate the sound and light alarm device (red light flashing + 80dB buzzer), and lock the access control at both ends of the corridor".
[0089] In this example, the threat probability score is 0.65, which falls under level three (medium threat).
[0090] 2. Generation and updating of graded early warning signals The warning signal uses JSON format and includes "warning level, incident information, action instructions, and generation time". In this example, the signal is: {"Warning Level":"Level 3 (Medium Threat)","Incident Information":{"Location":"Main corridor of the teaching building (camera IP: 192.168.1.102)","Time":"20251001143010000","Person Involved":"STU-2025001 (Zhang San, student, 95% confidence level)","Behavior Description":"Running at 2.84 m / s during class time, abnormal posture score 0.9"},"Action Instructions":{"Push Terminal":"Security Terminal (ID: SEC-003, located on the 1st floor of the teaching building)","Voice Reminder":"Activate the corridor voice module (ID: VOICE-002)","Record Storage":"Save the last 10 seconds of video to the security database"},"Generation Time":"20251001143010000"}.
[0091] Real-time update mechanism: Features are re-extracted and threat probability is calculated every 1 second (corresponding to 25 frames of video). If the speed of the person involved in the incident drops to 1.3 m / s in subsequent frames, the abnormal posture score is 0.1, the threat probability drops to 0.25, the warning level is updated to level 2 (low threat), and the warning signal is updated to "record only, stop voice reminder" simultaneously. If the speed rises to 3.5 m / s, the probability rises to 0.75, the level is updated to level 4 (high threat), and the command "start sound and light alarm + lock access control" is added.
[0092] S204, based on the graded early warning signal, the corresponding device linkage command is automatically triggered, wherein the device linkage command includes pushing alarm information to the security terminal, activating the sound and light alarm device, or locking the access control of the relevant area; Specifically, it can analyze graded early warning signals to generate a mapping table of device linkage instructions corresponding to different threat levels; The tiered early warning signal contains core fields such as "early warning level, incident location, information of personnel involved, and behavior description". Parsing requires first extracting key information, and then generating a mapping table (correspondence table) of "threat level - target device - linkage command" based on the deployment logic of campus security equipment (such as "dispatch equipment to the nearest location in the incident area") to ensure that the command matches the threat level and avoid resource waste or insufficient response.
[0093] First, the typical structure of a graded early warning signal is in JSON format. For example, a Level 3 (Medium Threat) signal is: {"Warning Level":"Level 3 (Medium Threat)","Incident Information":{"Location":"Main corridor of the teaching building (camera IP: 192.168.1.102)","Time":"20251001143010000","Person Involved":"STU-2025001 (Zhang San, student, 95% confidence level)","Behavior Description":"Running at 2.84 m / s during class"},"Generation Time":"20251001143010000"}. When analyzing the data, it is important to extract the "warning level" and "incident location": the warning level determines the intensity of the instruction, and the incident location determines the range of the target equipment (e.g., the camera 192.168.1.102 corresponds to the main corridor of the teaching building, and the nearest equipment is the SEC-003 security terminal on the first floor, the ALARM-002 sound and light alarm device in the corridor, and the ACCESS-005 and ACCESS-006 access control systems at both ends of the corridor).
[0094] The generation of the mapping table must follow preset rules and combine the instructions defined by the four threat levels on campus (Level 1: No threat; Level 2: Low threat; Level 3: Medium threat; Level 4: High threat). Level 1 (No threat, probability 0-0.2): No device linkage command, only the event is recorded on the security platform, and the mapping table entry is "Level 1: No target device, no command"; Level 2 (low threat, probability 0.2-0.4): Only push alarm information to the nearest security terminal (do not activate sound and light / access control), for example "Level 2: Target device = nearest security terminal (such as SEC-003), instruction = push text alarm ('Minor abnormal behavior found in the main corridor of the teaching building, involved person: Zhang San, behavior: jogging')"; Level 3 (Medium threat, probability 0.4-0.7): Push alarm to the nearest security terminal + activate area voice reminder (without locking access control), for example, "Level 3: Target device = SEC-003, ALARM-002, instruction = SEC-003 pushes an alarm containing video clips (with nearly 5 seconds of abnormal behavior video), ALARM-002 activates voice reminder ('Please keep quiet, do not run, security personnel have arrived', volume 60dB, loop playback)"; Level 4 (High Threat, Probability 0.7-1.0): Push alarm to all security terminals + activate audible and visual alarms + lock area access control. For example, "Level 4: Target devices = all security terminals (SEC-001 to SEC-010), ALARM-002, ACCESS-005 / 006, Command = push emergency alarm to all terminals ('High-threat behavior detected in the main corridor of the teaching building, the identity of the person involved is pending verification, behavior: climbing'), ALARM-002 activates audible and visual alarms (red light flashes 2 times / second, buzzer volume 80dB), ACCESS-005 / 006 locks for 10 minutes (prohibiting personnel from entering and exiting)."
[0095] In this example, the warning level is level three. The generated mapping table entries are: "Level three: Target device = SEC-003 (1st floor security terminal), ALARM-002 (corridor voice module), Command = SEC-003: Push an alarm containing a 5-second video ('14:30 teaching building main corridor, Zhang San (student) is running during class time, threat level is medium'); ALARM-002: Activate 60dB voice reminder, looping 'Please keep quiet, running is prohibited'". The mapping table needs to be marked with "generation time and associated warning signal ID" to ensure traceability.
[0096] Device linkage commands are sent to target devices via the MQTT IoT protocol. The target devices include at least a security terminal, an audible and visual alarm device, and an access control system. MQTT (Message Queuing Telemetry Transport) is a lightweight IoT protocol that uses a publish-subscribe model. It is suitable for low-bandwidth, high-latency scenarios (such as campus LANs) and ensures efficient and reliable command delivery (supporting message retransmission and quality level control). First, the MQTT communication environment needs to be configured: Deploy MQTTBroker (message server): The address is the campus security server IP (192.168.1.200), port 1883 (default port), and it supports QoS (Quality of Service) level 1 (at least one delivery, ensuring that the command is not lost). Target device registration: Each device (such as SEC-003, ALARM-002) needs to register a unique client ID (such as "SEC-003-CLIENT") with the Broker in advance and subscribe to the corresponding topic (such as security terminal subscribing to campus / security / command / terminal, audio-visual equipment subscribing to campus / security / command / alarm, access control subscribing to campus / security / command / access). Command format definition: Uses JSON format, including "cmd_id (unique command ID), device_id (target device ID), cmd_type (command type), cmd_content (command content), timestamp (send time), expire_time (expire time)" to ensure that the device can parse it.
[0097] Taking sending instructions to SEC-003 (security terminal) as an example, the specific process is as follows: Generate command content: cmd_id="CMD-20251001143012", device_id="SEC-003", cmd_type="video_alarm", cmd_content="{'Location': 'Main Corridor of Teaching Building', 'Time': '14:30:10', 'Person Involved': 'Zhang San (STU-2025001)', 'Behavior': 'Running (2.84m / s)', 'Video Link': 'http: / / 192.168.1.200 / video / 20251001143010.mp4'}", timestamp="20251001143012000", expire_time="20251001143512000" (Valid for 5 minutes); Publish command to Broker: The security platform, as the MQTT publisher, publishes the command to the topic "campus / security / command / terminal", specifying QoS level 1 (the Broker needs to return a "publish confirmation" to the publisher; if no confirmation is received, it will be resent, up to 3 times, with a 3-second interval between resentments). Device receives instructions: SEC-003, as a subscriber, listens to the topic in real time. After receiving the instruction, it verifies the uniqueness of cmd_id (to avoid duplicate execution) and the validity of expire_time (if it has not expired, it will be executed). The execution result is "displaying an alarm pop-up window + automatically playing a video link".
[0098] The command flow sent to ALARM-002 (voice module) is similar: command cmd_type="voice_remind", cmd_content="{'volume': 60, 'content': 'Please keep quiet, do not run, security personnel have arrived', 'loop count': unlimited (until stop command is received)}", published to the topic "campus / security / command / alarm". After receiving it, ALARM-002 activates the voice chip (model ISD1820), plays the voice according to the parameters, and returns "Executing" feedback.
[0099] If the access control system (such as ACCESS-005) issues a level 4 warning command, cmd_type="lock", cmd_content="{'lock duration': 600 (seconds), 'unlock condition': security terminal SEC-003 sends unlock command}", and publishes it to the topic "campus / security / command / access", the access controller will receive it and drive the electromagnetic lock to shut down and lock, and the LED indicator will turn red.
[0100] The system monitors the execution status of equipment linkage commands in real time and generates command execution status reports based on equipment feedback signals.
[0101] After the command is sent, the execution result needs to be monitored in real time to avoid response failure due to equipment failure (such as power failure or network disconnection). The core of monitoring is "receiving device feedback signals + judging the execution status + handling anomalies", and finally generating a status report containing "command execution status, anomaly cause, and handling suggestions" to provide a basis for subsequent event handling.
[0102] First, the format of the device feedback signal corresponds to the command format, including "cmd_id (associated command ID), device_id (device ID), status (execution status: success / fail / timeout), feedback_content (feedback content), timestamp (feedback time)". The feedback mechanism follows MQTT QoS1: after the device executes the command, it needs to publish a feedback signal to the topic "campus / security / feedback" to the Broker. The security platform subscribes to this topic to receive feedback.
[0103] Monitoring process and anomaly handling examples (using SEC-003 and ALARM-002 as examples): Normal monitoring execution: Within 1 second of receiving the command, SEC-003 will issue a feedback signal: "{'cmd_id': 'CMD-20251001143012', 'device_id': 'SEC-003', 'status': 'success', 'feedback_content': 'Alarm pop-up has been displayed, video has been loaded', 'timestamp': '20251001143012100'}"; ALARM Within 0.5 seconds of receiving the command, -002 will publish feedback "{'cmd_id':'CMD-20251001143013','device_id':'ALARM-002','status':'success','feedback_content':'Voice reminder has been started, volume 60dB','timestamp':'20251001143012500'}", and the platform will determine that the two devices have executed successfully.
[0104] Anomaly Handling: If ALARM-002 fails to respond within 3 seconds due to network interruption (preset timeout is 3 seconds), the platform triggers a retransmission mechanism: the first retransmission interval is 3 seconds, the second interval is 5 seconds, and the third interval is 10 seconds. If there is still no response after 3 retransmissions, it is determined that "execution failed" and the following message is issued: "status: fail, feedback_content: 'ALARM-002 has no feedback, suspected network interruption'". At the same time, the backup plan is triggered - a supplementary instruction of "ALARM-002 failure, manual inspection recommended" is pushed to the nearest security terminal SEC-003.
[0105] The instruction execution status report needs to integrate feedback information from all target devices. The format includes "Report ID, Associated Warning Signal ID, Instruction List, Execution Status of Each Device, Anomaly Summary, and Handling Suggestions". The report excerpt in this example is: "Report ID: REP-20251001143015, Associated Warning Signal ID: ALERT-20251001143010, Instruction List: 1. CMD-20251001143012 (SEC-003); 2. CMD-20251001143013 (ALARM-002); Execution Status: SEC-003 (success, feedback at 14:30:12.1), ALARM-002 (success, feedback at 14:30:12.5); Anomaly Summary: No anomalies; Handling Suggestions: Continuously monitor the behavior of the personnel involved. Security personnel SEC-003 have gone to the scene and are expected to arrive at 14:30:30."
[0106] After the report is generated, it is uploaded to the campus security management platform in real time and stored in the MySQL database (table name "command_status_report"). At the same time, the real-time status of each device is updated on the "Device Status" page of the platform (e.g., SEC-003 is marked "in normal execution", ALARM-002 is marked "in normal playback"), so that managers can keep track of the response progress in real time.
[0107] S205, based on the execution status of the device linkage command and the subsequent video analysis results, generate a complete security incident handling report including the location of the incident, images of the personnel involved, and handling suggestions.
[0108] Specifically, it can collect data on the execution status of device linkage commands and information on changes in human behavior in subsequent video streams to generate an event process dataset; The event process dataset needs to fully record the entire process information of "command execution - behavior change", covering two types of core data: equipment linkage command execution status data (reflecting whether the response link is effective) and subsequent personnel behavior change information (reflecting the development trend of the event). The collection must ensure the timeliness (sampling frequency ≥ 1 time / second) and correlation (linking the two types of data through event ID) of the data, so as to provide a foundation for subsequent map construction and disposal suggestion generation.
[0109] 1. Data acquisition of equipment linkage command execution status This type of data originates from the instruction execution status report generated in the preceding steps. Key fields need to be extracted and supplemented with real-time update information. Core fields include: A unique event ID (e.g., “EVENT-202510011430”, used throughout the entire process and associated with all data); Target device information: Device ID (e.g., “SEC-003”, “ALARM-002”), Device type (security terminal / audio-visual alarm / access control), Command ID (e.g., “CMD-20251001143012”); Execution status details: initial execution status (success / fail / timeout), status update time (e.g., "20251001143012", feedback 1 second after command is sent), exception handling record (if it fails, the reason must be recorded, such as "ALARM-002 network interruption, execution successful after reconnection at 14:30:15"), execution duration (e.g., ALARM-002 voice reminder continues to play until 14:35:00, duration 4 minutes and 48 seconds).
[0110] Example data snippet: "Event ID: EVENT-202510011430, Device ID: SEC-003, Device Type: Security Terminal, Command ID: CMD-20251001143012, Initial State: success, Update Time: 20251001143012, Abnormal Records: None, Duration: None (One-time command); Device ID: ALARM-002, Device Type: Audible and Visual Alarm, Command ID: CMD-20251001143013, Initial State: success, Update Time: 20251001143012, Abnormal Records: None, Duration: 288 seconds (14:30:12-14:35:00)".
[0111] 2. Collection of information on changes in human behavior in subsequent video streams Subsequent video streams refer to standardized video streams (1920×1080, 25fps) within 30 minutes of the command being sent. The YOLOv4 behavior detection algorithm should be used, extracting behavioral features of the individuals involved (associated through the target tracking ID "PER-001") every 5 frames (0.2 seconds), focusing on whether their behavior changes from abnormal to normal or whether a new abnormality occurs. Core data collection fields include: Timestamp (e.g., "20251001143015", accurate to the second); Behavioral characteristics: actual speed (e.g., 2.5 m / s at 14:30:15, decreasing to 1.2 m / s at 14:30:20), abnormal posture score (0.7 at 14:30:15, decreasing to 0.1 at 14:30:20), and behavior type label ("running → normal walking"). Environmental related information: Changes in the current scene's crowd density (e.g., 0.05 people / ㎡ at 14:30:15, decreasing to 0.03 people / ㎡ at 14:30:30), and whether other personnel have intervened (e.g., security personnel "SEC-003-USER" enter the screen at 14:30:25 and interact with the person involved).
[0112] Example data snippet: "Event ID: EVENT-202510011430, Target ID: PER-001, Timestamp: 20251001143015, Speed: 2.5m / s, Abnormal Posture Score: 0.7, Behavior Tag: Running, Crowd Density: 0.05 people / m², Intervening Personnel: None; Timestamp: 20251001143020, Speed: 1.2m / s, Abnormal Posture Score: 0.1, Behavior Tag: Normal Walking, Crowd Density: 0.04 people / m², Intervening Personnel: None; Timestamp: 20251001143025, Speed: 0.8m / s, Abnormal Posture Score: 0.0, Behavior Tag: Still, Crowd Density: 0.03 people / m², Intervening Personnel: SEC-003-USER (Security Personnel)."
[0113] 3. Generation of event process dataset The two types of data are sorted and integrated by "Event ID-Timestamp" and stored in JSON format. Each data entry contains "Timestamp, Device Execution Status Snapshot, Personnel Behavior Snapshot, and Environment Snapshot". Example dataset fragment: {"Event ID":"EVENT-202510011430","Timestamp":"20251001143015","Device Execution Status":{"SEC-003":"success (alarm received)","ALARM-002":"success (voice playing)"},"Personnel Behavior":{"Target..."}} ID":"PER-001","Speed":"2.5m / s","Abnormal Posture Score":"0.7","Behavior Label":"Running"},"Environmental Snapshot":{"Crowd Density":"0.05 people / ㎡","Scene Type":"Teaching Building Main Corridor"}} The dataset is stored in a MongoDB database on the campus security server (supports unstructured data storage, query speed ≤100ms), and is labeled with "data integrity identifier" (e.g., "Complete: 30 minutes of data collected" "Incomplete: Data from 14:35 to 14:40 is missing due to camera disconnection").
[0114] By integrating the location coordinates of the incident, facial images of the individuals involved, and identification results, a spatiotemporal trajectory map of the incident is constructed; The spatiotemporal trajectory map of an event needs to intuitively present the relationship between "time-space-personnel-behavior". The core is to transform discrete data into a visualized trajectory chain and integrate three types of key information: the precise coordinates of the incident location (spatial dimension), the identity and images of the people involved (personnel dimension), and the behavioral changes on the timeline (time dimension), to ensure that the map can clearly trace the entire process of the event.
[0115] 1. Key Information Integration Location coordinates of the incident: obtained through pre-calibrated geographic information of the cameras—each camera IP (e.g., 192.168.1.102) corresponds to unique latitude and longitude coordinates (e.g., "120.654321°E, 30.123456°N"), and the boundary coordinates of the camera's coverage area are marked (e.g., "Northeast boundary: 120.6545°E, 30.1236°N; Southwest boundary: 120.6541°E, 30.1233°N"); the pixel coordinates of the persons involved in the incident in the video frames. The marker (e.g., "(300,225)") is converted into the actual position relative to the camera (e.g., "6 meters in front of the camera, 0.5 meters to the left") through the camera calibration parameters (1 pixel corresponds to 0.02 meters), and finally superimposed into global latitude and longitude coordinates (e.g., "120.654321°E+0.5 meters×cos(0°)→120.654321°E,30.123456°N+6 meters×sin(0°)→30.123456°N", where 0° is the camera facing due north).
[0116] Face image of the person involved: Select the standardized face image (112×112 pixels, BRISQUE score 25) output from the above steps, and extract the frame image of the "most obvious moment of abnormal behavior" in the video stream (such as the frame image at 14:30:00, where the person involved is running, resolution 1920×1080), and label the image source ("camera IP: 192.168.1.102, frame timestamp: 20251001143000").
[0117] Identity verification results: Extract the final results, including personnel ID (“STU-2025001”), name (“Zhang San”), identity type (“student”), confidence score (“95 points”), result label (“reliable”), and supplement identity association information (such as “Class: 2025 Computer Science Class 1, Class Teacher: Teacher Li, Contact Number: 138XXXX1234”, synchronized from the campus student information system).
[0118] 2. Construction of event spatiotemporal trajectory map The map uses a structure of "timeline + spatial map + information card", arranging key nodes in chronological order (from 5 minutes before the event to 5 minutes after it ends). Each node includes "time point, spatial location, personnel status, behavioral description, and associated image": Timeline node example: 14:25:00 (5 minutes before the incident): Location "120.654321°E, 30.123456°N (west end of the main corridor of the teaching building)", Person's status "walking normally, speed 1.2m / s", Behavior description "no abnormality", Related image "14:25:00 frame image (person involved walking normally)"; 14:30:00 (Event Triggered): Location "120.654321°E, 30.123456°N (middle of the corridor)", Personnel Status "Running, speed 2.84m / s, abnormal posture score 0.9", Behavior Description "Triggered Level 3 warning", Associated Images "Standardized face image + 14:30:00 abnormal frame image"; 14:30:25 (Security intervention): Location "120.654321°E, 30.123456°N (middle of the corridor)", Personnel status "Still, speed 0m / s, abnormal posture score 0.0", Behavior description "Security personnel SEC-003-USER arrived, the person involved stopped running", Related image "Frame image at 14:30:25 (interaction between security and the person involved)"; 14:35:00 (Event Ended): Location "120.654321°E, 30.123456°N (Corridor Duty Room)", Personnel Status "Normal Walking, Speed 0.8m / s", Behavior Description "Taken to the duty room for verbal reprimand", Related Image "Frame image at 14:35:00 (Personnel involved walking towards the duty room)".
[0119] The map is visualized in SVG format (supporting zooming and detailed viewing), with an embedded campus map background (marking locations such as teaching buildings, corridors, and duty rooms). Clicking on each time point will pop up an information card, displaying complete personnel identity, behavioral data, and related images, making it easy for managers to quickly trace the sequence of events.
[0120] Based on the event process dataset and spatiotemporal trajectory map, a rule engine is used to generate disposal suggestions; The rule engine is the core of generating handling suggestions. It needs to be based on a pre-set rule library of "event type-scenario-behavioral result". By matching the event process data with key information in the spatiotemporal graph (such as the identity of the person involved, whether the behavior has been terminated, and whether personnel have intervened), it outputs targeted and actionable handling suggestions, avoiding generalities.
[0121] 1. Rule base design The rule base adopts a "condition-action" structure, with each rule containing "trigger conditions, priority, and suggested actions," covering common campus security scenarios. Example of a core rule: Rule 1 (Students running in the hallways during class time has been stopped and security has intervened): Triggering conditions: Person involved = "student", Scene type = "corridor", Time period = "class time", Behavior change = "abnormal (running) → normal (stationary / walking)", Intervening personnel type = "security personnel", No other injuries or property damage; Priority: Medium (No emergency response required, routine record keeping is necessary); Recommended actions: "1. Have the security personnel present give a verbal warning to the student involved (Zhang San, STU-2025001), emphasizing the importance of hallway discipline during class time; 2. Simultaneously report the incident to the student's homeroom teacher (Teacher Li, 138XXXX1234), and suggest that the class issue a disciplinary reminder; 3. Mark the student as a 'minor violation' in the campus security system, and remove the mark if there are no further violations within 3 months; 4. Check whether the sound and light alarm device (ALARM-002) in the area involved (main hallway of the teaching building) is functioning properly, and ensure that it can respond promptly to any future anomalies."
[0122] Rule 2 (Strangers loitering on campus at night without stopping or security intervention): Triggering conditions: Person involved identity type = "Not matched whitelist", Scene type = "Playground", Time period = "Night (after 21:30)", Behavioral change = "Continuous wandering (speed 0.5-1.0m / s)", Intervening personnel type = "None"; Priority: High (requires urgent action); Recommended actions: "1. Immediately dispatch security personnel from the nearest security terminal (SEC-005, located on the east side of the playground) to the scene to verify their identity; 2. Activate all audible and visual alarm devices (ALARM-004, ALARM-005) in the playground area to deter potential risks; 3. Lock the access control gates at the entrances and exits around the playground (ACCESS-008, ACCESS-009) to prevent personnel from escaping; 4. If the person involved is not found within 10 minutes, retrieve historical video from the surrounding cameras (192.168.1.105, 192.168.1.106) to track their movements."
[0123] The rule base is stored in the rule engine (using the Drools rule engine, which supports dynamic rule updates). Each rule is marked with "effective scenario" and "update time" (e.g., rule 1 was updated on 20250901 to adapt to the new semester's class time adjustment).
[0124] 2. Rule matching and suggestion generation Taking this case (student Zhang San running in the corridor during class, which has been stopped and security has intervened) as an example, the matching process is as follows: Extract the following information from the event process dataset: "Identity of Person Involved = Student", "Scene = Corridor", "Time Period = Class Time", "Behavioral Change = Running → Stillness", and "Intervening Personnel = Security Guard". The rule engine traverses the rule base and calculates the matching degree of each rule (e.g., the matching degree of rule 1 is 100%, and the matching degree of rule 2 is 0%). Select the rule with the highest matching degree and priority that matches the urgency of the current event (Rule 1, medium priority); Based on the details in the spatiotemporal trajectory map (such as the person involved being taken to the duty room and the sound and light devices operating normally), the suggestion of Rule 1 is refined and supplemented with "5. The duty room must keep the educational records of the student involved, including the educational time (14:35-14:40), the educator (SEC-003-USER), and the student's signature confirmation"; Generate final handling recommendations, and mark the "recommendation type" (routine handling / emergency handling / follow-up) and "execution time limit" (e.g., "verbal warnings must be completed within 10 minutes after the incident ends, and feedback from the homeroom teacher must be completed within 24 hours").
[0125] The incident location, images of those involved, and suggested handling data are formatted and output as a security incident handling report and uploaded to the campus security management platform.
[0126] Security incident handling reports must be "complete, standardized, and archiveable," using a standardized format to integrate all key information for easy retrieval, auditing, and review. They must also be uploaded to the campus security management platform via a security protocol to ensure data traceability and sharing.
[0127] 1. Formatted report output The report should be in PDF format (supporting electronic signatures and encryption), and its structure is divided into 6 core parts, each of which must contain specific data and visualization elements (such as images and graph screenshots): I. Basic Event Information: Event ID (EVENT-202510011430), Event Trigger Time (20251001143000), Event End Time (20251001144000), Location of Incident (Main Corridor of Teaching Building, Latitude and Longitude 120.654321°E, 30.123456°N), Warning Level (Level 3, Medium Threat), Event Status (Completed). II. Information on the person involved: Person ID (STU-2025001), Name (Zhang San), Identity Type (Student), Class (Class 1, Computer Science, 2025), Identity Recognition Confidence (95 points), Image of the person involved (Standardized face image + screenshot of abnormal frame at 14:30:00, with image source labeled); III. Event Process Review: Present the entire process of "abnormal triggering → command response → behavior termination → handling completion" in the form of a timeline, insert screenshots of the event's spatiotemporal trajectory map, and mark key time nodes (14:30:00 warning triggered, 14:30:12 command sent, 14:30:25 security intervention). IV. Equipment Interaction Status: The list displays the instruction content, execution status, and duration of the target equipment (SEC-003, ALARM-002), and marks abnormal equipment (none) and handling measures (none). V. Handling Suggestions and Implementation Status: List the handling suggestions generated by the rule engine in bullet points (5 in total), and mark the "Implementation Status" after each suggestion (e.g., "1. Verbal Warning: Completed, execution time 14:35:00" "2. Class Teacher Feedback: Pending execution, scheduled to be sent at 14:50"). VI. Attachments: Summary of the event process dataset (device and behavior data at key time points), camera calibration parameters (to ensure traceability of location coordinates), and signature column for the personnel handling the incident (SEC-003-USER signature area).
[0128] The report should be named in the format "Campus Security Incident Handling Report Event ID_Date.pdf" (e.g., "Campus Security Incident Handling Report_EVENT-202510011430_20251001.pdf"), and the size should be kept under 10MB (images should be compressed to a suitable resolution to ensure clarity while reducing file size).
[0129] 2. Upload the report to the campus security management platform. Uploads use the HTTPS protocol (to ensure secure data transmission and prevent leaks). Upload process: The campus security server acts as a client, sending an upload request to the platform, carrying the report file and metadata (event ID, report name, file size, and generation time). After receiving the request, the platform verifies the client's identity (using the API key "SEC-API-KEY-2025"). Once verification is successful, the platform receives the file. Files are stored in the platform's distributed file system (such as HDFS, which supports massive file storage and has an access latency of ≤500ms), and report indexes (event ID, storage path, upload time, uploader) are recorded in a relational database (MySQL). The platform returns an "upload successful" response, which includes the report access URL, allowing administrators to view and download the report via the platform's web interface or app. If the upload fails (e.g., due to network interruption), the client triggers a retransmission mechanism (up to 3 times, with a 5-second interval between retransmissions) and records a failure log ("202510011445, upload failed, reason: network timeout") for subsequent manual processing.
[0130] Another embodiment of the present invention provides a campus intelligent alarm linkage system based on facial recognition, see [link to relevant documentation]. Figure 3 The system may include: The extraction module 301 is used to collect video streams in real time through cameras deployed in key areas of the campus, and to extract face images from video frames using a face detection algorithm; The comparison module 302 is used to compare the face image with the whitelist database pre-stored on campus in real time to generate the face identification result and confidence score. The analysis module 303 is used to calculate the threat level of the current behavior scenario and generate a graded early warning signal based on the identity recognition result and confidence score, combined with real-time analysis of personnel behavior scenarios, through a dynamic early warning judgment model. The early warning module 304 is used to automatically trigger corresponding device linkage instructions based on the graded early warning signal, wherein the device linkage instructions include pushing alarm information to the security terminal, activating the audible and visual alarm device, or locking the access control of the relevant area; The generation module 305 is used to generate a complete security incident handling report, including the location of the incident, images of the people involved, and handling suggestions, based on the execution status of the device linkage command and the subsequent video analysis results.
[0131] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.
Claims
1. A campus intelligent alarm linkage method based on facial recognition, characterized in that, The method includes: Video streams are captured in real time by cameras deployed in key areas of the campus, and face images are extracted from video frames using face detection algorithms. The facial images are compared in real time with a pre-stored whitelist database on campus to generate facial recognition results and confidence scores. Based on the identity recognition results and confidence scores, combined with real-time analysis of personnel behavior scenarios, the threat level of the current behavior scenario is calculated and a graded early warning signal is generated through a dynamic early warning judgment model. The corresponding device linkage command is automatically triggered according to the graded early warning signal. The device linkage command includes pushing alarm information to the security terminal, activating the sound and light alarm device, or locking the access control of the relevant area. Based on the execution status of the device linkage command and the subsequent video analysis results, a complete security incident handling report is generated, including the location of the incident, images of the personnel involved, and handling suggestions.
2. The method according to claim 1, characterized in that, The method of acquiring video streams in real time through cameras deployed in key areas of the campus and extracting facial images from video frames using a face detection algorithm includes: Deploy network cameras in key areas of the campus, including at least entrances, corridors, and playgrounds, and generate a camera deployment plan; Based on the camera deployment scheme, video stream data is collected in real time via the RTSP protocol, and the video stream is preprocessed to generate standardized video stream data; Using standardized video stream data, a YOLOv5-based face detection algorithm is used to analyze video frames in real time, locate and extract face regions, and generate an initial set of face images. The initial set of face images is quality-assessed, and face images with a sharpness higher than a preset sharpness threshold are selected. Size normalization and illumination compensation are then performed to output standardized face images.
3. The method according to claim 2, characterized in that, The step of comparing the facial image with a pre-stored whitelist database on campus in real time to generate facial recognition results and confidence scores includes: For standardized face images, deep convolutional neural networks are used to extract face feature vectors and generate multidimensional feature representations for each face image; Load the facial feature vectors of registered individuals from the pre-stored whitelist database on campus to construct a feature retrieval database; The facial feature vectors extracted in real time are compared with the feature retrieval database to calculate similarity and generate a list of similarity scores. Based on the similarity score list, a dynamic similarity threshold is set to determine whether a face matches a whitelisted identity and generate a preliminary identity recognition result. Based on the preliminary identity recognition results, combined with the face image quality score and similarity score, the final identity recognition result and confidence score are output through the confidence calculation model.
4. The method according to claim 3, characterized in that, Based on the identity recognition results and confidence scores, and combined with real-time analysis of personnel behavior scenarios, the dynamic early warning judgment model calculates the threat level of the current behavior scenario and generates graded early warning signals, including: Based on the identity recognition results and confidence scores, extract the personnel identity credibility index and generate an identity credibility vector; The YOLOv4 behavior detection algorithm is used to extract human behavior features from the video stream, including at least motion trajectory, speed and abnormal posture, and generate behavior feature vectors. By combining the camera's geographic location information, the current scene type, such as a classroom or corridor, is identified, and the scene context, including time point and crowd density, is analyzed to generate scene context features; The identity credibility vector, behavioral feature vector, and scene context features are input into the dynamic early warning and judgment model based on XGBoost to calculate the threat probability score. Threat levels are classified based on threat probability scores, and corresponding graded early warning signals are generated and updated in real time.
5. The method according to claim 4, characterized in that, The automatic triggering of corresponding device linkage commands based on the graded early warning signals, wherein the device linkage commands include pushing alarm information to the security terminal, activating the audible and visual alarm device, or locking the access control of the relevant area, including: Analyze the graded early warning signals to generate a mapping table of device linkage instructions corresponding to different threat levels; Device linkage commands are sent to target devices via the MQTT IoT protocol. The target devices include at least a security terminal, an audible and visual alarm device, and an access control system. The system monitors the execution status of equipment linkage commands in real time and generates command execution status reports based on equipment feedback signals.
6. The method according to claim 5, characterized in that, Based on the execution status of the device linkage command and subsequent video analysis results, a complete security incident handling report is generated, including the incident location, images of the personnel involved, and handling suggestions, including: Collect data on the execution status of device linkage commands and information on changes in human behavior in subsequent video streams to generate an event process dataset; By integrating the location coordinates of the incident, facial images of the individuals involved, and identification results, a spatiotemporal trajectory map of the incident is constructed; Based on the event process dataset and spatiotemporal trajectory map, a rule engine is used to generate disposal suggestions; The incident location, images of those involved, and suggested handling data are formatted and output as a security incident handling report and uploaded to the campus security management platform.
7. The method according to claim 3, characterized in that, Based on the preliminary identity recognition results, combined with facial image quality scores and similarity scores, the final identity recognition result and confidence score are output through a confidence calculation model, including: Based on the preliminary identity recognition results, the highest similarity score in the similarity score list and the corresponding whitelist identity are extracted to generate a candidate identity set and associated similarity scores; The quality score of a face image is calculated using an image quality assessment algorithm, and a standardized quality score is generated. The association similarity scores and standardized quality scores of the candidate identity sets are input into a pre-trained confidence calculation model, which uses a gradient boosting decision tree algorithm to output an initial confidence score. The initial identity recognition results are adjusted based on the preliminary confidence score and the preset confidence threshold, and the confidence is smoothed by combining time series analysis to generate the final identity recognition results and the corresponding confidence score.
8. A campus intelligent alarm linkage system based on facial recognition, characterized in that, The system includes: The extraction module is used to collect video streams in real time through cameras deployed in key areas of the campus and to extract face images from video frames using a face detection algorithm; The comparison module is used to compare the face image with the whitelist database pre-stored on campus in real time to generate the face identification result and confidence score. The analysis module is used to calculate the threat level of the current behavior scenario and generate a graded early warning signal based on the identity recognition results and confidence scores, combined with real-time analysis of personnel behavior scenarios, through a dynamic early warning judgment model. The early warning module is used to automatically trigger corresponding device linkage instructions based on the graded early warning signals. The device linkage instructions include pushing alarm information to the security terminal, activating the audible and visual alarm device, or locking the access control of the relevant area. The generation module is used to generate a complete security incident handling report, including the location of the incident, images of the people involved, and handling suggestions, based on the execution status of the device linkage command and the subsequent video analysis results.
9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-7 when it is run.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-7.
Citation Information
Cited By
AI cloud access control identification method and device
CN122223821A
AI cloud access control identification method and device
CN122223821B