Security and protection monitoring comprehensive early warning analysis system and method based on machine learning

By aligning and calibrating multi-source data, spatiotemporal collaborative standardization, and dynamic scene description, combined with privacy protection and adaptive early warning decision-making, the misjudgment problems of implicit spatial dependence and progressive abnormal behavior identification in existing technologies are solved, achieving high-precision analysis of abnormal group behavior and stable early warning report generation.

CN121808549APending Publication Date: 2026-04-07SHAANXI RADIO & TELEVISION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to capture implicit spatial dependencies, leading to missed group conflicts, and the identification of progressive anomalous behavior is prone to delays or misjudgments.

Method used

By aligning and calibrating multi-source data, saliency detection and spatiotemporal collaborative standardization are performed. Combined with spatiotemporal graph convolutional networks and multi-head self-attention mechanisms, dynamic scene description and threat classification are carried out. Differential privacy and behavioral cloning are used for adaptive early warning decisions, generating interpretable early warning reports.

Benefits of technology

It improves the accuracy of identifying abnormal group behavior, reduces misjudgments and delays, enhances data quality and identification accuracy, and ensures the interpretability of early warning reports and the stability of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808549A_ABST
    Figure CN121808549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition and analysis, solves the technical problems that implicit spatial dependence is difficult to capture, missing report group conflicts are caused, and recognition of progressive abnormal behaviors is prone to delay or misjudgment in the prior art, and particularly relates to a security and protection monitoring comprehensive early warning analysis system and method based on machine learning. The method comprises the following steps: S1, collecting an original video stream monitored in a campus, obtaining campus Internet of Things sensor data, and preprocessing the original video stream and the Internet of Things sensor data to obtain time-space standardized data. According to the method, spatial configuration and time evolution can be explicitly coded, information splitting can be avoided, long-distance and many-to-many complex dependence can be captured in a cross-frame mode, the understanding accuracy of group abnormal behaviors is improved, robustness is greatly improved, and the joint analysis capacity of individual behavior and group interaction is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition analysis, and in particular to a security monitoring comprehensive early warning analysis system and method based on machine learning. BACKGROUND

[0002] Image recognition analysis is a process of extracting, understanding and classifying information such as targets, scenes and features in images using computer vision technology. A security monitoring comprehensive early warning analysis system combined with image recognition analysis can automatically add highlight labels and behavior tags by analyzing abnormal targets in video images in real time on the basis of traditional monitoring, visually presenting machine decision results to assist security personnel in quickly confirming or adjusting disposal methods.

[0003] In the prior art, the identification of campus security problems is often based on target-independent trajectory analysis, which relies on spatial distance threshold to determine the relationship, and is difficult to capture implicit spatial dependence, resulting in missed group conflicts. In addition, campus abnormal behaviors often evolve in multiple stages, and existing technologies easily lose cross-frame correlation features, leading to delayed or misjudged identification of progressive abnormal behaviors. SUMMARY

[0004] To overcome the shortcomings of the prior art, the present application provides a security monitoring comprehensive early warning analysis system and method based on machine learning, which solves the technical problems that the prior art is difficult to capture implicit spatial dependence, resulting in missed group conflicts, and the identification of progressive abnormal behaviors is prone to delay or misjudgment.

[0005] To solve the above technical problems, the present application provides the following technical scheme: a security monitoring comprehensive early warning analysis method based on machine learning, the method comprising the following steps: S1, collecting original video streams monitored in the campus, obtaining campus Internet of Things sensor data, and preprocessing the original video streams and the Internet of Things sensor data to obtain spatio-temporal standardized data; S2, obtaining single-frame images of video segment sequences in the spatio-temporal standardized data, and performing fine classification processing on the single-frame images to obtain a classification instance set; S3, obtaining a dynamic scene description according to the classification instance set through spatio-temporal joint modeling and dynamic global processing; S4, obtaining security threat classification according to the dynamic scene description through privacy protection processing and threat classification; S5, performing adaptive early warning decision according to the security threat classification to obtain early warning state, disposal strategy and updated advantage value; S6, generating a security monitoring early warning report according to the early warning state, the disposal strategy and the corresponding original video stream through early warning information integration, and sending the security monitoring early warning report to a processing center.

[0006] Preferably, in S1, the specific implementation steps are as follows: S11, obtaining integrated data by alignment and calibration according to the original video stream and the Internet of Things sensor data; S12, obtaining high-quality data by saliency detection according to the integrated data; S13, obtaining spatio-temporal standardized data by spatio-temporal collaborative standardization processing according to the high-quality data.

[0007] Preferably, in S2, the specific implementation steps are as follows: S21, obtaining target boundary, category and confidence by boundary detection according to the single-frame image of the spatio-temporal standardized data; S22, obtaining a standardized image set by cropping and scaling according to the target boundary and the single-frame image; S23, obtaining a fine-grained probability distribution by self-attention mechanism processing according to the standardized image set; S24, obtaining comprehensive confidence and final label by weighted average formula based on confidence and fine-grained probability distribution; S25, classifying the target boundary according to the comprehensive confidence and the final label to obtain a classification instance set.

[0008] Preferably, in S3, the specific implementation steps are as follows: S31, constructing a human key point model according to the target boundary in the classification instance set, and obtaining a spatio-temporal graph set; S32, constructing a spatio-temporal convolution layer according to the spatio-temporal graph set and the classification instance set, and obtaining associated node features by spatio-temporal convolution; S33, calculating posterior node features by multi-head self-attention mechanism according to the associated node features; S34, obtaining behavior probability distribution and interaction relationship matrix by processing the posterior node features through multi-layer perception and graph-level classifier respectively; S35, obtaining dynamic scene description by MLP neural network processing based on behavior probability distribution and interaction relationship matrix.

[0009] Preferably, in S4, the specific implementation steps are as follows: S41, obtaining global target trajectory by differential privacy processing according to the dynamic scene description; S42, obtaining desensitization trajectory features by information desensitization according to the global target trajectory; S43, obtaining security threat classification by clustering processing according to the desensitization trajectory features.

[0010] Preferably, in S5, the specific implementation steps are as follows: S51, obtaining a decision vector by encoding according to the security threat classification; S52. Construct a decision layer based on the decision vector, and generate an initial policy set through behavior cloning; S53. Construct a value layer based on the decision vector and the initial strategy set, and obtain the current strategy, actual reward and advantage value through strategy scoring; S54. Obtain the hyperparameters of the decision-making layer and the value layer, construct the parameter update layer, obtain the updated parameters through parameter updates, and obtain the optimized decision-making layer and the optimized value layer; S55. Replace the decision layer and value layer with the optimized decision layer and optimized value layer. The input decision vector is used to obtain the early warning status, disposal strategy and updated advantage value through the optimized decision layer and optimized value layer.

[0011] Preferably, in S6, the specific implementation steps are as follows: S61. Generate early warning video based on the early warning status, response strategy, original video stream, and target trajectory; S62. Based on the early warning video, design the early warning interactive interface; S63. Generate a security monitoring early warning report based on the early warning interactive interface, early warning status, handling strategy, original video stream and target trajectory.

[0012] Preferably, the early warning interactive interface includes a video display box, text description, and multiple quick processing buttons.

[0013] The technical solution also provides a system for the aforementioned machine learning-based comprehensive early warning analysis method for security monitoring, the system comprising: The standardization module is used to collect raw video streams from campus surveillance cameras, acquire campus IoT sensor data, and preprocess the raw video streams and IoT sensor data to obtain spatiotemporally standardized data. The fine classification module is used to acquire single-frame images of video segment sequences in spatiotemporally standardized data, and to perform fine classification processing on the single-frame images to obtain a classification instance set; The dynamic spatiotemporal module is used to obtain a dynamic scene description based on the classification instance set through spatiotemporal joint modeling and dynamic global processing. The threat classification module is used to classify security threats based on dynamic scenario descriptions through privacy protection processing and threat classification. The early warning decision module is used to make adaptive early warning decisions based on the classification of security threats, and to obtain the early warning status, handling strategy and updated advantage value. The information integration module is used to generate a security monitoring early warning report by integrating early warning information based on the early warning status, handling strategy and corresponding raw video stream, and then send the security monitoring early warning report to the processing center.

[0014] By employing the above technical solutions, this invention provides a comprehensive security monitoring early warning analysis system and method based on machine learning, which has at least the following beneficial effects: 1. This invention obtains integrated data by aligning and calibrating raw video streams and IoT sensor data, which can break down barriers between different data sources, fuse multi-source heterogeneous data, and use saliency detection to filter out high-quality data from the integrated data. This effectively removes noise and invalid information, improves data quality, reduces the burden of subsequent processing, and unifies the spatiotemporal dimensions of data through spatiotemporal collaborative standardization processing, thereby improving the efficiency of accurate decision-making and in-depth data value mining.

[0015] 2. This invention avoids information loss from multi-stage independent processing by using boundary detection, standardized image generation, self-attention mechanism and weighted fusion, thereby improving recognition accuracy, enhancing the detection capability for occlusion and small targets, capturing the correlation between local and global features of the target, generating a more accurate fine-grained probability distribution, and improving the overall confidence and reliability of the final label through dynamic weight adaptive balancing, while reducing post-processing costs.

[0016] 3. This invention combines spatiotemporal graph convolutional networks with multi-head self-attention mechanisms to explicitly encode spatial configuration and temporal evolution, avoiding information fragmentation. It can also capture complex dependencies over long distances and many-to-many relationships across frames, breaking through the limitations of local connectivity. Furthermore, it jointly optimizes spatial structure, temporal dynamics, and relational weights, thereby improving the accuracy of understanding abnormal group behavior, significantly enhancing robustness, and substantially improving the joint analysis capability of individual behavior and group interaction in complex dynamic scenarios.

[0017] 4. This invention transforms desensitized threat summaries into structured state vectors and rapidly forms initial policies through behavior cloning, significantly reducing initial exploration costs and accelerating convergence. It introduces online advantage estimation and PPO Clip pruning mechanisms to guide policy updates by dynamically calculating the advantage function, while limiting the update magnitude to prevent policy collapse and ensuring stable training. Based on the optimized decision layer and value layer, it outputs interpretable warning states and disposal strategies, balancing decision quality and practicality. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the machine learning-based comprehensive early warning analysis method for security monitoring according to the present invention; Figure 2 This is a structural block diagram of the security monitoring integrated early warning analysis system based on machine learning of the present invention. Detailed Implementation

[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0020] Example 1: Due to the difficulty of capturing implicit spatial dependencies in existing technologies, resulting in missed group conflicts, and the potential for delays or misjudgments in identifying progressively anomalous behavior, please refer to [the relevant documentation / reference]. Figure 1 This embodiment provides a comprehensive early warning analysis method for security monitoring based on machine learning. It can explicitly encode spatial configuration and temporal evolution, avoid information fragmentation, and capture complex dependencies over long distances and many-to-many relationships across frames. It breaks through the limitations of local connectivity, improves the accuracy of understanding abnormal group behavior, and significantly enhances robustness. The method includes the following steps: S1. Collect raw video streams from campus surveillance cameras and acquire campus IoT sensor data. Preprocess the raw video streams and IoT sensor data to obtain spatiotemporally standardized data. Existing technologies neglect the information gain brought by the fusion of multi-source data such as campus video and sensor data during multi-dimensional data preprocessing, resulting in biased analysis results. Furthermore, the integrated data is not effectively filtered, leading to low accuracy. To address these issues, the specific implementation steps are as follows: S11. Integrated data is obtained by aligning and calibrating the original video stream and IoT sensor data. The original video stream can be acquired through campus video surveillance equipment. The local calibration time point and a normally distributed synchronization residual are obtained. The video alignment time is obtained by subtracting the calibration time point from the time of each frame of the original video stream, multiplying the result by the clock frequency drift rate, adding the synchronization residual, and then adding the time of each frame. A fixed clock offset obtained through NTP calibration is also added. IoT sensor data includes data related to campus security such as door magnetic sensors, facial recognition records, smoke detectors, temperature sensors, vibration sensors, WiFi usage time, AP locations, device information, biosafety cabinet status, airflow velocity in chemical laboratories, power distribution box power data, vehicle recognition sensors, and vehicle speed detection radar. The timestamps of the IoT sensor data are obtained. The product of subtracting the calibration time point from the timestamp, multiplying it by the clock frequency drift rate, adding the synchronization residual, and then adding the timestamp again, plus a fixed clock offset obtained through NTP calibration, is then calculated. The fixed clock offset obtained by P calibration yields the sensor alignment time. The pixel coordinates and camera intrinsic matrix of the original video stream are acquired. The z-value of the pixel coordinates is set to 1 to generate 3D coordinates. The global processing coordinates within the campus are obtained by multiplying the transpose of the 3D coordinates by the -1 power of the intrinsic matrix. The camera's orientation angle relative to the campus is obtained. Angle coordinates with the orientation angle as the z-value are generated based on the pixel coordinates. The coarse calibration coordinates are obtained by multiplying the transpose of the processing coordinates, the transpose of the angle coordinates, and the camera's rotation matrix. The video alignment time, sensor alignment time, and calibration coordinates are then integrated using spatiotemporal anchor points. NTP calibration is a network protocol used for computer clock synchronization. The synchronization residual, which follows a normal distribution, is a small random error between the ideal synchronization time and the actual achievable synchronization time, obtainable through statistical analysis of data transmission time fluctuations. The camera intrinsic matrix is ​​a mathematical transformation matrix that projects 3D points in the camera coordinate system onto the image pixel coordinate system; it will not be elaborated upon here.

[0021] S12. High-quality data is obtained through saliency detection based on the integrated data; the absolute difference between the integrated data of the current frame and the integrated data of the previous frame at each pixel is calculated and summed to obtain the inter-frame difference saliency measure. Then, the difference saliency measure is dynamically calculated through the adjustment function to obtain the corresponding enhancement parameters. Finally, the integrated data of the current frame and these parameters are input into the enhancement function to generate a quality-optimized output frame. The adjustment function is a commonly used piecewise adjustment function, and the enhancement function is a set of parameterized image processing operators that receive the integrated data and enhancement parameters to generate a quality-optimized high-quality data.

[0022] S13. Spatiotemporally standardized data is obtained from high-quality data through spatiotemporal collaborative standardization. Spatiotemporal collaborative standardization is a multi-standardization method. High-quality data is first standardized for color and illumination, then for resolution, and finally for features. Color and illumination standardization, resolution standardization, and feature standardization are commonly used standardization methods, which will not be elaborated here. The standardization results are merged into a standardized vector, which is the spatiotemporally standardized data. This invention obtains integrated data by aligning and calibrating the original video stream and IoT sensor data. It can break down the barriers between different data sources, fuse multi-source heterogeneous data, and use saliency detection to filter out high-quality data from the integrated data. It effectively removes noise and invalid information, improves data quality, reduces the burden of subsequent processing, and unifies the spatiotemporal dimensions of the data through spatiotemporal collaborative standardization, thereby improving the efficiency of accurate decision-making and in-depth data value mining.

[0023] S2. Obtain single-frame images from the video segment sequence in the spatiotemporally standardized data, and perform fine classification processing on the single-frame images to obtain a classification instance set; Existing technologies, in detection and attribute classification, are trained independently, which can easily lead to bounding box deviations or misclassifications in campus security identification, and make it difficult to capture the complex relationships between local and global targets on campus, resulting in low recognition accuracy. To solve the above problems, the specific implementation steps are as follows: S21. Based on a single-frame image of spatiotemporally standardized data, the target boundary, category, and confidence level are obtained through boundary detection. Feature extraction is performed using a target detection model based on the single-frame image, typically implemented using the Detector General function. The Detector General function can directly obtain preliminary detection results, including target boundary, category, and confidence level. The Detector General function is the core processing unit of the target detection model. After inputting a single-frame image, features are extracted and decoded through a convolutional neural network, and preliminary detection results can be directly output. Both the Detector General function and the convolutional neural network are commonly used feature processing methods, which will not be elaborated here.

[0024] S22. Obtain a standardized image set by cropping and scaling based on the target boundary and single-frame images. A single-frame image is a standardized image, which can be an RGB image or an HSV image. The target in the single-frame image is cropped according to the target boundary. Cropping can be performed using the Crop function to obtain a cropped image. An adjusted image is obtained by resizing the cropped image. Resizing is generally achieved using the Resize function to adjust the image size to 224×224. Multiple single-frame images are then cropped and scaled to obtain a standardized image set. The Crop and Resize functions are commonly used methods for cropping and scaling images, and will not be elaborated here.

[0025] S23. Based on the standardized image set, a fine-grained probability distribution is obtained through self-attention mechanism. The images in the standardized image set are divided into N×N pixel blocks, converted into D-dimensional vectors through linear projection, and positional encoding is added to preserve spatial information. A cls token is inserted at the beginning of each sequence, and all D-dimensional vectors are converted to obtain the input vector. Based on the input vector, the correlation of each D-dimensional vector is calculated through multi-head self-attention mechanism, and the correlation is weighted and aggregated to obtain global features. The MLP layer maps the categories to probabilities. For example, the probability of the target being a student is 0.8, and the probability of the posture being walking is 0.9. Here, the cls token is a learnable embedding vector that provides a global representation for the sequence. Multi-head self-attention mechanism is a commonly used parallel variant of self-attention mechanism, widely used in deep learning, especially the Transformer model in natural language processing. Weighted aggregation is a commonly used calculation method, which will not be elaborated here.

[0026] S24. Based on the confidence level and fine-grained probability distribution, a weighted average formula is used to obtain the comprehensive confidence level and the final label. The confidence level reflects the confidence level of the existence of the target, and the fine-grained probability distribution reflects the confidence level of the classification result. A dynamic weight is obtained by cross-validation based on historical data of the confidence level and fine-grained probability distribution. Historical data can be obtained from data before the current time. The confidence level is obtained by multiplying the dynamic weight by the confidence level. The comprehensive confidence level is obtained by multiplying the maximum value of the fine-grained probability distribution by 1 and subtracting the difference of the dynamic weight, plus the confidence level. The final label is generated by combining the classification result of the comprehensive confidence level.

[0027] S25. Classify the target boundaries based on the comprehensive confidence score and the final label, and obtain a classification instance set. Based on the comprehensive confidence score and the final label, the target boundaries are processed by high-threshold filtering and non-maximum suppression to obtain the classification instance set. That is, some target boundaries with low comprehensive confidence scores are filtered out by a high threshold, and the remaining target boundaries are sorted by confidence score from high to low. The intersection-union ratio (IUR) is calculated using the target boundary with the highest confidence score. If the IUR is greater than 0.5, the overlapping boundaries are deleted. These overlapping boundaries are considered redundant. Finally, multiple classification instances are obtained and integrated into a classification instance set. The high threshold can be set in the range of 0.8 to 0.95 to facilitate strict filtering. This invention avoids information loss from multi-stage independent processing through boundary detection, standardized image generation, self-attention mechanism and weighted fusion, improves recognition accuracy, improves the detection capability of occlusion and small targets, can capture the correlation between local and global features of the target, generate more accurate fine-grained probability distribution, improves the reliability of comprehensive confidence score and final label through dynamic weight adaptive balancing, and reduces post-processing costs.

[0028] S3. Based on the classification instance set, a dynamic scene description is obtained through spatiotemporal joint modeling and dynamic global processing. Existing technologies for identifying campus security often rely on independent target trajectory analysis and spatial distance thresholds to determine relationships, making it difficult to capture such implicit spatial dependencies, leading to missed reports of group conflicts. Furthermore, abnormal campus behavior often involves multi-stage evolution, and existing technologies are prone to losing cross-frame correlation features, resulting in delays or misjudgments in the identification of progressive abnormal behavior. To solve the above problems, the specific implementation steps are as follows: S31. Construct a human keypoint model based on the target boundary in the classification instance set and obtain a spatiotemporal atlas. Based on the target boundary of each frame in the classification instance set, construct spatial relationships based on the skeletal connection relationship of human keypoints, and construct edges between nodes that conform to the skeletal relationship to form a spatial point edge set. For overlapping parts, interpolation can be used to fill in the gaps and obtain a complete spatial point edge set. For example, if the target is a human body, construct edges according to the keypoints, including the anatomical connection relationship of the head, shoulders, knees, etc., such as connecting the head node to the shoulder node. If there is overlap, add a spatial edge to indicate that they may interact. If two nodes are determined to be nodes at the same position on the same target boundary in different frames, add a temporal edge to connect them. If the target boundary reappears after disappearing in a certain frame, the temporal edge can be used to associate the preceding and following segments. The temporal edge represents the continuity of the target in the time dimension. Merge the spatial point edge sets and temporal edges in all classification instances to generate a complete spatiotemporal atlas. Merging does not mean merging into a single image, but rather establishing temporal association labels to generate an atlas that is related in time and space.

[0029] S32. Construct a spatiotemporal convolutional layer based on the spatiotemporal graph atlas and the classification instance set, and obtain the associated node features through spatiotemporal convolution. The spatiotemporal convolutional layer is a method for extracting spatiotemporal graph features based on the spatiotemporal graph convolutional network. First, generate weights based on the spatial and temporal edges in the spatiotemporal graph atlas. Based on the weights, generate processing groups by combining the corresponding spatiotemporal graph atlas and the classification instance set. For example, the spatiotemporal graph and classification instances corresponding to the time t-1 to t+1 are grouped together. Multiple fusion values ​​of spatiotemporal information are obtained by weighted summation. Adjacent fusion values ​​are nonlinearly processed through an activation function to generate feature values. All feature values ​​are combined into the associated node features. The spatiotemporal graph convolutional network is a common neural network for processing spatiotemporal graph features. The ReLU function can be selected as the activation function, which will not be elaborated here.

[0030] S33. Calculate the subsequent node features based on the associated node features using a multi-head self-attention mechanism; based on all associated node features, obtain three matrices—query, key, and value—through current projection, and obtain the dimension of the key. Multiply the query by the transpose of the key, divide by the quotient of the square root of the dimension, and obtain the weight coefficients through softmax. Multiply the weight coefficients by the value to obtain the weighted node features, which are the subsequent node features.

[0031] S34. Based on the features of the subsequent nodes, the behavior probability distribution and interaction matrix are obtained by processing them through a multilayer perceptron and a graph classifier, respectively. Based on the features of the subsequent nodes, the weighted features of a single node are nonlinearly transformed by a multilayer perceptron (MLP) to output the behavior probability distribution. The similarity matrix of the features of the subsequent nodes is calculated by multiplying all the features of the subsequent nodes by the transpose of the features of the subsequent nodes through matrix multiplication. Then, the similarity matrix of the features of the subsequent nodes is normalized to a probability matrix by Sigmoid, which is the interaction matrix.

[0032] S35. A dynamic scene description is obtained by processing the behavior probability distribution and interaction relationship matrix through an MLP neural network. The behavior probability distribution and interaction relationship matrix are first integrated through an MLP neural network to generate structured semantics. For example, if multiple individual behaviors are judged to be walking and talking, and two individuals have an interaction relationship, then after integration, the description is generated that the two individuals meet after walking and have a conversation. After processing all behavior probability distributions and interaction relationship matrices, a dynamic scene description with individual behavior labels and group interaction relationship graphs is obtained. The MLP neural network is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset. It is a commonly used method for processing and integrating multiple datasets, which will not be elaborated here. This invention combines a spatiotemporal graph convolutional network with a multi-head self-attention mechanism to explicitly encode spatial configuration and temporal evolution, avoid information fragmentation, capture complex dependencies of long distance and many-to-many relationships across frames, break through the limitations of local connectivity, and jointly optimize spatial structure, temporal dynamics, and relationship weights. This improves the accuracy of understanding abnormal group behavior, significantly enhances robustness, and significantly improves the joint analysis capability of individual behavior and group interaction in complex dynamic scenes.

[0033] S4. Based on the dynamic scene description, security threats are categorized through privacy protection processing and threat classification. Existing technologies often face the challenge of balancing privacy protection and data utility when analyzing campus security threats. During data analysis, network attacks can easily lead to the leakage of the target's appearance data and information. Existing technologies rely solely on simple anonymization or static desensitization, making them vulnerable to re-identification attacks. To address these issues, the specific implementation steps are as follows: S41. Based on the dynamic scene description, the global target trajectory is obtained through differential privacy processing. Local trajectories of the target in multiple camera scenes are collected and associated with the global trajectory. First, local appearance features of the target are extracted using MobileNetV3. Then, Laplacian noise is added using Differential Privacy to obtain privacy-preserving features. Similarity matching is performed through attention fusion. After similarity matching, trajectory features are generated based on the movement paths of the same target using the trajectory curvature formula. The trajectories of the same target in images from different cameras are associated using CRF to obtain the global target trajectory with hidden target information. Among these, MobileNetV3 is a commonly used method for feature recognition of target appearance, Differential Privacy is a method for differential privacy protection of information, Attention fusion is a commonly used method for feature similarity comparison based on attention, and CRF is a probabilistic graphical model used for modeling and predicting sequential or structured data. It is particularly good at handling annotations with contextual dependencies and is a commonly used method for constructing trajectory associations, which will not be elaborated here.

[0034] S42. Obtain desensitized trajectory features by desensitizing information based on the global target trajectory; perform privacy protection by k-anonymization based on the global target trajectory, and obtain the desensitized trajectory features after privacy protection. Among them, k-anonymization is a privacy protection technique. Its core idea is to ensure that each record is indistinguishable from at least k-1 other records in the dataset in terms of quasi-identifiers by generalizing or suppressing quasi-identifiers in the data. It is a commonly used method for protecting privacy and will not be elaborated here.

[0035] S43. Based on the de-identified trajectory features, clustering is used to obtain security threat classifications; based on the de-identified trajectory features, DBSCAN is used. DBSCAN is a density-based clustering algorithm. Its core idea is to divide the data into the largest set of density-connected points through the local density association of sample points, and automatically identify security threat features. All security threat features are integrated to obtain security threat classifications. DBSCAN is a commonly used data clustering method, which will not be elaborated here. Security threat classification can be understood as obtaining the identification results of multiple targets' behaviors, actions, interaction relationships, and multiple scenarios after security identification and privacy protection. For example, anonymous individuals 001, 002, and 003 are located in region B, and the time is t1-t3, which is the class time. The feature clustering result is a special cluster. The analysis is based on the fact that at time t1, individual 001 appears in region B, and individuals 002 and 003 simultaneously appear in region A at time t2, quickly merge with individual 001, and remain at time t3. The abnormal behavior is a special long-term clustering behavior during class time, which is defined as a special cluster. This invention uses differential privacy processing to handle global trajectories in dynamic scenarios, injects controllable noise during the data release stage, and mathematically ensures the indistinguishability of individuals while preserving the distribution characteristics of group behavior. After information desensitization, trajectory features are extracted, and sensitive information is further stripped away. Through cluster analysis, fine-grained threat classification is achieved, such as distinguishing between loitering, tailing, and clustering patterns. This not only meets privacy compliance requirements but also improves the accuracy and scenario adaptability of threat detection.

[0036] S5. Based on the classification of security threats, adaptive early warning decisions are made to obtain the early warning status, response strategy, and updated advantage value. In campus security early warning, existing technologies often rely on simple heuristic rules, which leads to low efficiency in the early training stage and a tendency to get stuck in local optima. Furthermore, the model may collapse during strategy updates due to large differences between the old and new strategies, making it difficult to guarantee the robustness of long-term decisions in dynamic threat environments. To solve the above problems, the specific implementation steps are as follows: S51. Encode the decision vector according to the security threat classification; Encode the security threat classification using an MLP encoder. The MLP neural network is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset. It is a commonly used method for processing and integrating multiple datasets, which will not be elaborated here. Here, the MLP neural network is used as the encoder of the decision vector to encode and convert it into a state vector, which is the decision vector.

[0037] S52. Construct a decision layer based on decision vectors and generate an initial strategy set through behavior cloning. Integrate countermeasures into a response strategy set based on decision vectors and security threat classification. Construct an imitation learning loss function based on decision vectors. Using the imitation learning loss function as a reference, the system can learn the behavioral patterns of the response strategy set by comparing the differences between the analyzed response strategies and the response strategy set. These initial strategies are then integrated into an initial strategy set. For example, when a student is seen walking in the corridor during class time, a security warning can be generated and sent to the teaching affairs office processing center. The imitation learning loss function is a commonly used behavior cloning method, which will not be elaborated here.

[0038] S53. Construct a value layer based on the decision vector and initial strategy set, and obtain the current strategy, actual reward, and advantage value through strategy scoring; construct an advantage estimation layer based on the decision vector and initial strategy set. First, give an initial strategy based on the current decision vector, such as a student appearing in the corridor during class, which belongs to the intermediate warning level, or the temperature on the second floor of the teaching building rising abnormally, which belongs to the high warning level, and generate a processing action, such as collecting video and generating evidence and warning reports for the intermediate warning level, including time, location, and people, and generating an alarm bell action for the high warning level, and retaining the data evidence, thus forming a state-action pair. After this processing action is implemented in practice, an initial score is given based on the result, which is the actual reward, and a new state is generated based on the result. In different states, the processing action will bring different scores. For example, if the current strategy is continued in the new state, a new total score can be obtained in the future. If the current strategy is maintained, the current total score can be obtained in the future. The advantage value is obtained by adding a discount factor to the initial score, multiplying the new total score by the current total score, and subtracting the current total score. The discount factor is a hyperparameter, and the initial value can be set to 0.8.

[0039] S54. Obtain the hyperparameters of the decision-making layer and value layer to construct a parameter update layer. Update the parameters through parameter updates to obtain the optimized decision-making layer and optimized value layer. Construct the PPO Clip algorithm model based on the hyperparameters of the decision-making layer and value layer, where PPO... The Clip algorithm is a major variant of the PPO algorithm. It improves the stability of reinforcement learning by introducing a pruning operation into the objective function to control the magnitude of policy updates. The process involves calculating the probability ratio of the current policy to the old policy for each state-action pair, and constructing a pruning surrogate objective function based on the advantage value. Specifically, the advantage value is multiplied by the probability ratio to obtain the previous ratio, and the pruned advantage value is multiplied by the pruned probability ratio to obtain the next ratio. The smaller of the two ratios is used as the update target, forming the pruning objective function. This limits the magnitude of policy updates and prevents training instability due to excessive differences between the old and new policies. The value layer updates synchronously by minimizing the mean squared error to more accurately estimate the actual reward and advantage value. Gradient descent is used to update the policy network parameters and value network parameters along the negative gradient direction of the objective function, resulting in updated parameters. The objective function is the loss function of the PPO Clip algorithm model, typically the pruning surrogate objective function. The pruning surrogate objective function, the probability ratio of the current policy to the old policy, and minimizing the mean squared error are commonly used PPO parameters. The calculation formulas in the Clip algorithm and statistics will not be elaborated here.

[0040] S55. By replacing the decision layer and value layer with an optimized decision layer and an optimized value layer, the input decision vector is processed through the optimized decision layer and optimized value layer to obtain the warning state, disposal strategy, and updated advantage value. This invention transforms the desensitized threat summary into a structured state vector and quickly forms an initial strategy through behavior cloning, which significantly reduces the initial exploration cost and accelerates convergence. It introduces an online advantage estimation and PPO Clip pruning mechanism to guide the strategy update by dynamically calculating the advantage function, while limiting the update magnitude to prevent strategy collapse and ensure the stability of the training process. Based on the optimized decision layer and value layer, it outputs interpretable warning states and disposal strategies, taking into account both decision quality and practicality.

[0041] S6. Based on the warning status, handling strategy, and corresponding original video stream, a security monitoring warning report is generated by integrating the warning information, and the report is sent to the processing center. Existing technologies for integrating warning information are prone to problems such as information dispersion and opaque decision-making processes, leading to difficulties in understanding and low trust among security personnel. To address these issues, the specific implementation method is as follows: S61. Generate an early warning video based on the early warning status, handling strategy, original video stream, and target trajectory; overlay the early warning status, handling strategy, original video stream, and target trajectory, such as highlighting the target box, behavior labels, early warning level icons, and text suggestions, to generate an early warning video.

[0042] S62. Based on the warning video, an interactive interface for warnings is designed. The interactive interface is designed using HTML5 and can simultaneously display the video display box of the warning video, input text descriptions, and add quick processing buttons, such as confirm, false alarm, save, upgrade handling, and downgrade handling. HTML5 is the fifth generation of hypertext markup language, which is the core technology standard for building web pages and web applications and is a commonly used method for browser programming. It will not be elaborated here.

[0043] S63. Generate a security monitoring early warning report based on the early warning interactive interface, early warning status, handling strategy, original video stream and target trajectory; This invention visualizes the machine decision-making results intuitively, enabling security personnel to quickly obtain key information and provide direct feedback, forming a closed loop of decision execution feedback, significantly improving response efficiency and decision accuracy.

[0044] Example 2: Due to the difficulty of existing technologies in capturing implicit spatial dependencies, leading to missed group conflicts, and the potential for delays or misjudgments in identifying progressive anomalous behavior, please refer to [link to relevant documentation]. Figure 2 The diagram shown is a structural block diagram of a security monitoring and early warning analysis system based on machine learning provided in this embodiment. The system includes a standardization module, a fine classification module, a dynamic spatiotemporal module, a threat classification module, an early warning decision module, and an information integration module.

[0045] The standardization module is used to collect raw video streams from campus surveillance cameras, acquire campus IoT sensor data, and preprocess the raw video streams and IoT sensor data to obtain spatiotemporally standardized data. The fine classification module is used to acquire single-frame images of video segment sequences in spatiotemporally standardized data, and to perform fine classification processing on the single-frame images to obtain a classification instance set; The dynamic spatiotemporal module is used to obtain a dynamic scene description based on the classification instance set through spatiotemporal joint modeling and dynamic global processing. The threat classification module is used to classify security threats based on dynamic scenario descriptions through privacy protection processing and threat classification. The early warning decision module is used to make adaptive early warning decisions based on the classification of security threats, and to obtain the early warning status, handling strategy and updated advantage value. The information integration module is used to generate a security monitoring early warning report by integrating early warning information based on the early warning status, handling strategy and corresponding raw video stream, and then send the security monitoring early warning report to the processing center.

[0046] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc.

[0047] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A comprehensive early warning analysis method for security monitoring based on machine learning, characterized in that, The method includes the following steps: S1. Collect raw video streams from campus surveillance cameras, acquire campus IoT sensor data, and preprocess the raw video streams and IoT sensor data to obtain spatiotemporally standardized data. S2. Obtain single-frame images from the video segment sequence in the spatiotemporal standardized data, and perform fine classification processing on the single-frame images to obtain a classification instance set; S3. Based on the classification instance set, a dynamic scene description is obtained through spatiotemporal joint modeling and dynamic global processing; S4. Based on the dynamic scenario description, security threats are categorized through privacy protection processing and threat classification. S5. Based on the classification of security threats, make adaptive early warning decisions to obtain the early warning status, handling strategy and updated advantage value; S6. Based on the warning status, handling strategy, and corresponding original video stream, the system integrates the warning information to generate a security monitoring warning report and sends the report to the processing center.

2. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S1, the specific implementation steps are as follows: S11. Integrated data is obtained by alignment and calibration based on the raw video stream and IoT sensor data; S12. Obtain high-quality data through significance testing based on the integrated data; S13. Spatiotemporally standardized data are obtained by spatiotemporally collaborative standardization processing based on high-quality data.

3. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S2, the specific implementation steps are as follows: S21. Based on a single-frame image of spatiotemporally standardized data, obtain the target boundary, category, and confidence level through boundary detection; S22. Obtain a standardized image set by cropping and scaling based on the target boundary and single-frame images; S23. Fine-grained probability distribution is obtained by processing the standardized image set through a self-attention mechanism. S24. Based on the confidence level and fine-grained probability distribution, the comprehensive confidence level and final label are obtained through a weighted average formula; S25. Classify the target boundary based on the comprehensive confidence level and the final label to obtain a set of classification instances.

4. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S3, the specific implementation steps are as follows: S31. Construct a human keypoint model based on the target boundary in the classification instance set, and obtain a spatiotemporal atlas; S32. Construct a spatiotemporal convolutional layer based on the spatiotemporal atlas and the classification instance set, and obtain the features of associated nodes through spatiotemporal convolution; S33. Calculate the post-node features based on the features of associated nodes using a multi-head self-attention mechanism; S34. Based on the features of the subsequent nodes, the behavior probability distribution and interaction matrix are obtained by processing them through a multilayer perceptron and a graph classifier, respectively. S35. Dynamic scene descriptions are obtained by processing the behavior probability distribution and interaction relationship matrix through an MLP neural network.

5. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S4, the specific implementation steps are as follows: S41. Obtain the global target trajectory through differential privacy processing based on the dynamic scene description; S42. Obtain desensitized trajectory features based on the global target trajectory through information desensitization; S43. Based on the desensitized trajectory characteristics, security threats are classified through clustering.

6. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S5, the specific implementation steps are as follows: S51. Obtain the decision vector by encoding according to the classification of security threats; S52. Construct a decision layer based on the decision vector, and generate an initial policy set through behavior cloning; S53. Construct a value layer based on the decision vector and the initial strategy set, and obtain the current strategy, actual reward and advantage value through strategy scoring; S54. Obtain the hyperparameters of the decision-making layer and the value layer, construct the parameter update layer, obtain the updated parameters through parameter updates, and obtain the optimized decision-making layer and the optimized value layer; S55. Replace the decision layer and value layer with the optimized decision layer and optimized value layer. The input decision vector is used to obtain the early warning status, disposal strategy and updated advantage value through the optimized decision layer and optimized value layer.

7. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 1, characterized in that, In S6, the specific implementation steps are as follows: S61. Generate early warning video based on the early warning status, response strategy, original video stream, and target trajectory; S62. Based on the early warning video, design the early warning interactive interface; S63. Generate a security monitoring early warning report based on the early warning interactive interface, early warning status, handling strategy, original video stream and target trajectory.

8. The comprehensive early warning analysis method for security monitoring based on machine learning according to claim 7, characterized in that, The warning interface includes a video display box, text description, and multiple quick processing buttons.

9. A system applied to the machine learning-based security monitoring comprehensive early warning analysis method according to any one of claims 1-8, characterized in that, The system includes: The standardization module is used to collect raw video streams from campus surveillance cameras, acquire campus IoT sensor data, and preprocess the raw video streams and IoT sensor data to obtain spatiotemporally standardized data. The fine classification module is used to acquire single-frame images of video segment sequences in spatiotemporally standardized data, and to perform fine classification processing on the single-frame images to obtain a classification instance set; The dynamic spatiotemporal module is used to obtain a dynamic scene description based on the classification instance set through spatiotemporal joint modeling and dynamic global processing. The threat classification module is used to classify security threats based on dynamic scenario descriptions through privacy protection processing and threat classification. The early warning decision module is used to make adaptive early warning decisions based on the classification of security threats, and to obtain the early warning status, handling strategy and updated advantage value. The information integration module is used to generate a security monitoring early warning report by integrating early warning information based on the early warning status, handling strategy and corresponding raw video stream, and then send the security monitoring early warning report to the processing center.