Multi-mode city security video abnormal behavior real-time detection system and method

Through a real-time detection system for abnormal behavior of multimodal urban security video, combined with video and sensor data, cross-modal feature fusion and dynamic resource regulation are achieved, which solves the problems of privacy leakage, bandwidth pressure and insufficient real-time in urban security systems, and is adapted to heterogeneous equipment to improve detection accuracy and efficiency.

CN120257092APending Publication Date: 2025-07-04浪潮智慧城市科技有限公司
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510372017.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

There are problems of privacy leakage, bandwidth pressure and insufficient real-time performance in existing urban security systems, especially in terms of multimodal data processing and heterogeneous equipment adaptation.

Method used

A real-time detection system for abnormal behavior of multimodal urban security video is adopted, including edge node module, central coordination server module, multimodal data fusion module and dynamic federal policy module. Combined with video and sensor data, cross-modal feature fusion is achieved through space-time alignment and attention mechanism, and the model parameter aggregation frequency is adaptively adjusted according to edge node calculation load and network state.

Benefits of technology

On the premise of ensuring data privacy, multi-modal collaboration, dynamic resource adaptation and lightweight are realized, monitoring accuracy is improved, computing overhead is reduced, heterogeneous devices are adapted to meet the real-time detection needs of milliseconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257092A_ABST
    Figure CN120257092A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode city security video abnormal behavior real-time detection system and method, and belongs to the technical field of artificial intelligence and city security, and the system comprises an edge node module which is used for collecting a local video stream and sensor data, and executing local model training and real-time abnormal detection; the central coordination server module is used for federal model parameter aggregation, global model distribution and dynamic resource coordination; the multi-modal data fusion module is used for realizing cross-modal feature fusion through space-time alignment and an attention mechanism; the dynamic federation strategy module is used for calculating a load and a network state according to an edge node and adaptively adjusting a model parameter aggregation frequency; and the alarm linkage module triggers a local sound-light alarm when an abnormal behavior is detected, and encrypts and transmits event information to a command center. According to the method, the problems of privacy leakage, bandwidth pressure and multi-modal data collaboration in the prior art can be solved, and on the premise that data privacy is guaranteed, calculation overhead is minimized, and heterogeneous equipment is adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and urban security, and specifically to a multi-modal real-time detection system and method for abnormal behaviors in urban security videos. Background Art

[0002] In traditional urban security systems, video analysis usually adopts a centralized architecture, that is, all cameras transmit high-definition video streams (such as 1080p / 30fps) in real time to a central server for processing. This mode has significant defects:

[0003] (1) Risk of privacy leakage: Video data may be intercepted or misused by malicious attackers during the upload, storage, and analysis processes. For example, sensitive scenes (such as face and license plate information) can be restored through reverse engineering. Although some systems adopt data desensitization or encrypted transmission (such as the TLS protocol), centralized storage still faces the risk of internal personnel leakage or database breach.

[0004] (2) Bandwidth pressure: A single camera may generate several GB of video data per hour (taking H.264 encoding as an example). When deployed on a large scale (such as thousands of cameras), the network bandwidth occupancy can reach the TB level, resulting in a sharp increase in transmission delay and cost. For example, a certain urban security project once caused the core switch to overload due to the concurrent transmission of tens of thousands of cameras, and was forced to reduce the resolution to 720p to relieve the pressure.

[0005] (3) Lack of real-time performance: The central server needs to queue up to process a large amount of video streams, resulting in analysis delays (usually in seconds), which are difficult to meet the millisecond-level response requirements in scenarios such as fires and violent incidents. Experiments show that under the centralized architecture, the average delay from the occurrence of an event to the triggering of an alarm exceeds 2 seconds, and the delay may further deteriorate to more than 5 seconds in crowded scenarios (such as stadiums).

[0006] Although existing security systems based on federated learning can alleviate privacy and bandwidth problems, there are still the following technical bottlenecks:

[0007] 1. Single-modal data processing: Most existing solutions only utilize video data (such as using CNN to extract visual features), ignoring the auxiliary information of sensor data such as temperature, smoke, and sound. For example, fire detection only relies on video flame recognition and cannot combine the sudden increase in temperature sensor data (such as a 10°C increase in 10 seconds), resulting in a relatively high false alarm rate (experiments show that the false alarm rate of a pure visual model reaches more than 15%).

[0008] 2. Static federated policy: Most systems adopt a fixed parameter aggregation frequency (e.g., once an hour) without considering the dynamic resource status of edge nodes. For example, during the idle network period at night, aggregation is still performed at a low frequency, wasting available bandwidth; or during peak traffic periods (such as the peak flow of people during holidays), forced high-frequency aggregation is carried out, causing edge devices to be overloaded (when the CPU utilization rate exceeds 90%, frequency scaling protection may be triggered). In addition, the traditional federated average algorithm (FedAvg) does not distinguish the data quality of nodes (such as annotation accuracy), and the global model may be polluted by parameters of low-quality nodes.

[0009] 3. Poor compatibility of heterogeneous devices: Existing solutions assume that all edge devices have the same computing power (e.g., all are equipped with GPUs), while in actual scenarios, old cameras (such as those only supporting the ARM Cortex-A53 CPU) are difficult to run complex models (such as ResNet-50), resulting in insufficient inference speed of the federated model on low-end devices (e.g., the single-frame processing time > 200 ms), which cannot meet the real-time requirements. Summary of the Invention

[0010] The technical task of the present invention is to address the above deficiencies and provide a multi-modal real-time detection system and method for abnormal behaviors in urban security videos, which can solve the problems of privacy leakage, bandwidth pressure, and multi-modal data collaboration in the prior art, minimize the computational overhead while ensuring data privacy, and adapt to heterogeneous devices.

[0011] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0012] A multi-modal real-time detection system for abnormal behaviors in urban security videos includes:

[0013] An edge node module, deployed at the camera terminal, for collecting local video streams and sensor data, and performing local model training and real-time abnormal detection;

[0014] A central coordination server module, for federated model parameter aggregation, global model distribution, and dynamic resource coordination;

[0015] A multi-modal data fusion module, integrating video data, sensor data, and historical alarm logs, and achieving cross-modal feature fusion through spatio-temporal alignment and attention mechanisms;

[0016] A dynamic federated policy module, adaptively adjusting the model parameter aggregation frequency according to the computing load and network status of edge nodes;

[0017] An alarm linkage module, triggering local audible and visual alarms when abnormal behaviors are detected, and encrypting and transmitting event information to the command center.

[0018] This system is used to detect abnormal behaviors (such as fires, violent incidents, break-ins, etc.) in urban security scenarios in real time, while ensuring data privacy and edge computing efficiency.

[0019] Furthermore, the edge node module specifically includes:

[0020] Video acquisition unit, a camera supporting the RTSP protocol, outputting a 1080p video stream encoded in H.264 / H.265;

[0021] Sensor interface unit, integrating temperature, smoke, and sound sensors, with a data sampling frequency ≥ 1Hz;

[0022] Local module unit, including a lightweight 3D convolutional neural network (3D MobileNetV3) and an LSTM network, used to extract video spatio-temporal features and sensor temporal features;

[0023] Privacy protection unit, used to add Laplace noise (ε = 0.5) to the fully connected layer before uploading the model gradient to prevent parameter inversion attacks.

[0024] Furthermore, the multi-modal data fusion module includes:

[0025] Hardware-level clock synchronization unit, aligning the camera and sensor clocks through the PTP protocol (error < 1ms);

[0026] Cross-modal attention fusion unit, using video features as Query, and sensor features as Key / Value, to calculate multi-head attention weights;

[0027] Abnormal event cross-validation unit, when an abnormality is detected in the video, verifying whether the sensor data conforms to the expected pattern (such as the temperature rising ≥ 10°C within 10 seconds in a fire scenario).

[0028] Furthermore, the weights of video and sensor features are dynamically allocated through the multi-head attention layer, and the formula is as follows:

[0029]

[0030] Among them, Q is the video feature, and K and V are the sensor features, realizing information complementarity between modalities.

[0031] Furthermore, the working logic of the dynamic federated policy module is:

[0032] When the average bandwidth utilization rate of the edge node < 30% and the CPU utilization rate < 60%, trigger high-frequency aggregation (once every 5 minutes);

[0033] When the bandwidth utilization rate > 70% or the CPU utilization rate > 80%, switch to low-frequency aggregation (once every 30 minutes);

[0034] When a high-confidence abnormal event is detected (probability > 90%), immediately trigger an emergency aggregation and preferentially update the global model.

[0035] Furthermore, when aggregating, weights are assigned according to the node data quality (annotation consistency score) and data volume for weighted averaging, and the weight formula is as follows:

[0036]

[0037] where N i is the data volume of node i, and Acc i is the accuracy of the local model of node i on the validation set.

[0038] Furthermore, the alarm linkage module includes:

[0039] A local alarm unit that controls the LED and buzzer through GPIO to trigger an audible and visual alarm signal;

[0040] An encryption transmission unit that encrypts the abnormal video segment using AES-256 and pushes it to the cloud through the MQTT protocol;

[0041] An emergency resource scheduling unit that automatically generates an optimal resource scheduling plan (such as dispatching the nearest fire truck).

[0042] The present invention also claims to protect a real-time detection method for abnormal behaviors in multi-modal urban security videos, and the implementation of this method includes the following steps:

[0043] Step S1, the edge node collects the local video stream and sensor data, and performs spatio-temporal alignment and preprocessing;

[0044] Step S2, train a lightweight 3D CNN model based on the local dataset, add noise and masks, and then upload the parameters to the central server;

[0045] Step S3, the central server dynamically aggregates the parameters according to the node resource status, generates a global model and distributes it to the edge nodes;

[0046] Step S4, the edge node loads the global model, analyzes the real-time video stream frame by frame, and detects abnormal behaviors;

[0047] Step S5, confirm the abnormal event through multi-modal cross-validation, trigger a local alarm and encrypt and report it to the command center.

[0048] Furthermore, the dynamic aggregation strategy in step S3 includes:

[0049] Assign aggregation weights according to the node data volume and annotation consistency score;

[0050] Adopt the asynchronous federated learning (FedAsync) protocol, allowing high-load nodes to delay uploading parameters;

[0051] Downweight the parameters of low-performance nodes to avoid degradation of the global model performance.

[0052] Furthermore, the real-time detection process in step S4 includes:

[0053] Continuously detect the anomaly probability using a sliding window (window size 10 frames); if the anomaly probability is > 0.8 for 10 consecutive frames, it is determined as a suspected anomaly event;

[0054] Automatically enable INT8 quantization for low-resolution devices to ensure that the inference latency < 120ms.

[0055] Compared with the prior art, a multi-modal urban security video anomaly behavior real-time detection system and method of the present invention have the following beneficial effects:

[0056] The present invention can simultaneously meet the requirements of a new type of security system for multi-modal collaboration, dynamic resource adaptation, and lightweight privacy protection. It deeply integrates video and sensor data, improves the monitoring accuracy through cross-modal alignment and attention mechanism; adjusts the federated learning strategy in real time according to the network bandwidth and edge computing load to balance efficiency and stability; minimizes the computational overhead and adapts to heterogeneous devices on the premise of ensuring data privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a system architecture diagram of the multi-modal urban security video anomaly behavior real-time detection system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] An embodiment of the present invention provides a multi-modal urban security video anomaly behavior real-time detection system, including:

[0059] An edge node module, deployed at the camera terminal, for collecting local video streams and sensor data, and performing local model training and real-time anomaly detection;

[0060] A central coordination server module, for federated model parameter aggregation, global model distribution, and dynamic resource coordination;

[0061] A multi-modal data fusion module, integrating video data, sensor data, and historical alarm logs, and realizing cross-modal feature fusion through spatio-temporal alignment and attention mechanism;

[0062] A dynamic federated policy module, adaptively adjusting the model parameter aggregation frequency according to the edge node computing load and network status;

[0063] The alarm linkage module triggers local audible and visual alarms when detecting abnormal behaviors, and encrypts and transmits event information to the command center.

[0064] Among them, the edge node module specifically includes:

[0065] The video acquisition unit, a camera supporting the RTSP protocol, outputs a 1080p video stream encoded in H.264 / H.265.

[0066] The sensor interface unit integrates temperature, smoke, and sound sensors, and the data sampling frequency ≥ 1Hz.

[0067] The local module unit includes a lightweight 3D convolutional neural network (3D MobileNetV3) and an LSTM network, which are used to extract video spatio-temporal features and sensor temporal features.

[0068] The privacy protection unit adds Laplace noise (ε = 0.5) to the fully connected layer.

[0069] The working logic of the dynamic federated policy module is as follows:

[0070] When the average bandwidth utilization rate of the edge node < 30% and the CPU utilization rate < 60%, high-frequency aggregation is triggered (once every 5 minutes);

[0071] When the bandwidth utilization rate > 70% or the CPU utilization rate > 80%, switch to low-frequency aggregation (once every 30 minutes);

[0072] When a high-confidence abnormal event (probability > 90%) is detected, immediately trigger an emergency aggregation and preferentially update the global model.

[0073] The multimodal data fusion module includes:

[0074] The hardware-level clock synchronization unit aligns the camera and sensor clocks through the PTP protocol (error < 1ms);

[0075] The cross-modal attention fusion unit uses video features as Query, and sensor features as Key / Value to calculate the multi-head attention weights;

[0076] The abnormal event cross-verification unit, when an abnormality is detected in the video, verifies whether the sensor data conforms to the expected pattern (such as the temperature rising ≥ 10℃ within 10 seconds in a fire scene).

[0077] The alarm linkage module includes:

[0078] The local alarm unit controls the LED and buzzer through GPIO to trigger audible and visual alarm signals;

[0079] Encryption transmission unit, which encrypts abnormal video clips using AES-256 and pushes them to the cloud through the MQTT protocol;

[0080] Emergency resource scheduling unit, which automatically generates the optimal resource scheduling plan (such as dispatching the nearest fire truck).

[0081] Combined with the attached Figure 1 As shown, the specific implementation of this system is as follows:

[0082] I. System architecture:

[0083] 1. Edge node layer.

[0084] The edge node layer is responsible for tasks such as local multi-modal data collection, real-time inference, and model training. The edge node layer consists of a data collection module, a local model module, and a privacy protection module.

[0085] Among them, the data collection module supports three types of data input:

[0086] (1) Video input: Cameras supporting the RTSP protocol, with a resolution of 1080P, a frame rate of 30fps, and H.264 / H.265 encoding;

[0087] (2) Sensor interface: Integrated temperature, smoke, and sound sensors, with a sampling frequency of 1Hz and a data format of JSON;

[0088] (3) Historical log synchronization: Loading labeled abnormal events (including information such as timestamps, types, and locations) from the local database.

[0089] The local model module adopts lightweight 3D MobileNetV3. When dealing with videos, feature optimization is required, and a temporal convolutional layer (kernel_size = 3*3*3) is added; the spatio-temporal features of video clips are extracted through 3D CNN; the sensor temporal data (temperature, smoke, sound) is processed through the LSTM network; the weights of video and sensor features are dynamically allocated through the multi-head attention layer, and the formula is as follows:

[0090]

[0091] Among them, Q is the video feature, and K and V are the sensor features, realizing information complementarity between modalities.

[0092] The privacy protection module adds Laplace noise to the fully connected layer before uploading the model gradient to prevent parameter inversion attacks.

[0093] 2. Central coordination layer.

[0094] The central coordination layer is responsible for global model aggregation, dynamic frequency regulation, and resource coordination. It includes:

[0095] (1) Federal aggregation unit,

[0096] When aggregating, weighted average is performed according to the node data quality (annotation consistency score) and data volume, and the weight formula is as follows:

[0097]

[0098] Where N i is the data volume of node i, and Acc i is the accuracy of the local model of node i on the validation set.

[0099] (2) Dynamic frequency regulation unit,

[0100] Monitor the resources of each node by collecting the CPU / GPU utilization, network bandwidth, and remaining power of the nodes in real time, and adopt different aggregation decisions by analyzing the resource status of different nodes: when the average bandwidth utilization of the node < 30% and the CPU utilization < 60%, trigger the high-frequency aggregation mode (aggregate once every five minutes); when the bandwidth > 70% or the CPU > 80%, trigger the low-frequency aggregation mode (once every 30 minutes) to reduce resource contention.

[0101] 3. Multimodal data fusion layer.

[0102] The multimodal data fusion layer is used to achieve spatio-temporal alignment and feature enhancement of cross-modal data.

[0103] Spatio-temporal alignment strategy:

[0104] Synchronize the clocks of the camera and the sensor through PTP (Precision Time Protocol) with an error < 1ms to achieve hardware-level synchronization; perform linear interpolation on the low-frequency sensor data to match the video frame rate.

[0105] Abnormal event enhancement:

[0106] Use the GAN network locally at the edge node to generate rare abnormal events (such as smoke diffusion simulation) to improve the generalization of the model; inject adversarial samples (such as occluding video frames) during the training phase to enhance the robustness of the model.

[0107] The specific process of this system to achieve real-time detection of abnormal behaviors in multimodal urban security videos is as follows:

[0108] 1. Model initialization and data preparation.

[0109] The central server initializes the local model parameters and pre-trains the base model based on the public dataset, which includes a temporal convolutional layer and a cross-modal attention module; it distributes the model weights to all edge nodes via the HTTPS protocol. Each edge node constructs a multi-modal dataset containing video clips, sensor time-series signals, and historical annotations based on the data collected by local cameras and sensors. The central server sends the pre-trained global initial model to each node to ensure that the model has the basic anomaly detection ability. The edge device automatically adapts the model calculation precision according to its own hardware performance (such as GPU computing power and memory size). For example, high-performance devices use floating-point operations, and low-end devices enable quantization compression to reduce the computing overhead.

[0110] 2. Local model training.

[0111] Local model training is the core link for the system to achieve distributed intelligent analysis. Its goal is to complete the feature extraction and model optimization of multi-modal data locally at the edge node, avoid the transmission of raw data, and at the same time achieve cross-node knowledge sharing through federated learning. The training process fully combines the characteristics of edge computing resources and privacy protection requirements.

[0112] The local model processes video and sensor data through a dual-branch structure:

[0113] (1) Video analysis branch: A lightweight 3D convolutional network is used to extract the spatio-temporal features of the video and capture the changing patterns of the relationships of dynamic behaviors (such as the spread of fire).

[0114] (2) Sensor analysis branch: A time-series model (such as LSTM) is used to encode the periodic laws of signals such as temperature and smoke to identify mutation events (such as a sudden rise in temperature and abnormal sounds).

[0115] (3) Cross-modal fusion: The weights of video and sensor features are dynamically allocated through the attention mechanism. For example, in fire detection, when the data of the temperature sensor is abnormal, the model automatically enhances the attention to the fire area in the video branch to achieve the complementarity of multi-source information.

[0116] Based on the local dataset, the edge node uses the focal loss function for training, focusing on optimizing the recognition ability of rare abnormal events and avoiding model bias caused by the overloading of normal event samples. During the training process, the model parameters are only updated locally, and the raw data is retained inside the device throughout the process, fundamentally eliminating the risk of privacy leakage.

[0117] Before uploading the model gradients to the central server, the edge node performs desensitization processing on the parameters through lightweight privacy protection technologies: adding controllable noise to the gradients so that attackers cannot reverse the original data content through the parameters; randomly masking some non-critical parameters to further reduce the possibility of privacy leakage.

[0118] During the local training process, a flexible resource adaptation mechanism is supported. When computing resources are insufficient, the model input resolution is automatically reduced or some training rounds are skipped to prioritize real-time inference performance. When a high-value abnormal event is detected, incremental learning is immediately triggered, adding the latest event data to the training set to quickly improve the model's sensitivity to similar scenarios.

[0119] 3. Federated parameter aggregation.

[0120] Federated parameter aggregation is the core mechanism for the system to achieve distributed intelligent collaboration. Its goal is to integrate the local knowledge of multiple edge nodes and construct a global model with high accuracy and strong generalization ability while protecting data privacy. This process realizes the dynamic fusion and optimization of model parameters through the efficient interaction between the central coordination server and the edge nodes. It includes:

[0121] (1) Parameter upload and privacy protection.

[0122] After the local model training is completed, each edge node does not directly upload the original data, but transmits the desensitized model parameters (such as gradients or weights) to the central server. To prevent the leakage of sensitive information, the node uses lightweight privacy protection technology to preprocess the parameters: adding random noise to the parameters to ensure that attackers cannot restore the original data features through reverse engineering.

[0123] (2) Dynamic aggregation strategy.

[0124] The central server intelligently adjusts the frequency and priority of parameter aggregation according to the real-time status of all edge nodes in the network to balance the model update efficiency and system stability. During idle network periods, the aggregation frequency is automatically increased to make full use of the limited bandwidth to accelerate model iteration; during resource-intensive periods, the aggregation frequency is reduced to avoid aggravating network congestion or device overload due to frequent communication; when a node detects a high-confidence abnormal event, the "urgent aggregation" mode is immediately triggered to prioritize integrating the parameters of this node to ensure that the global model can quickly respond to threats.

[0125] (3) Weighted fusion and quality assessment.

[0126] To improve the robustness of the global model, the central server does not simply average all node parameters, but introduces a data quality assessment mechanism to differentially allocate node weights. First, based on the annotation accuracy and data volume scale of the node's historical data, its contribution weight is dynamically calculated to increase the weight of nodes with high label consistency and large data volume in the aggregation and avoid contaminating the global model with low-quality data; second, for old devices, due to their first computing capabilities leading to model deviations, the central server automatically reduces their weight ratio and simultaneously calibrates them implicitly through the parameters of high-performance nodes to ensure the balanced performance of the model among different devices.

[0127] (4) Global model distribution and collaborative optimization.

[0128] After parameter aggregation, the central server distributes the updated global model to all edge nodes, initiating a new round of local training and inference. This process forms a closed-loop optimization. Through a version control mechanism, it ensures that all nodes synchronously load the latest model, avoiding inconsistent detection results caused by version fragmentation;

[0129] During local training, each node can combine newly collected abnormal event data to continuously optimize the model's detection ability for region-specific scenarios (such as special fire hazards in industrial parks), and these improvements will contribute to the global model in the next round of federated aggregation.

[0130] 4. Real-time inference stage.

[0131] The real-time inference stage is the core link for the implementation of this system. It realizes millisecond-level abnormal detection and linkage response through edge device localization processing. The design of this stage fully considers the complexity and real-time requirements of urban security scenarios, and combines multi-modal data fusion and dynamic resource adaptation technologies to ensure the stable operation of the system under high concurrency and multi-scenarios. The following details its workflow from aspects such as data preprocessing, multi-modal inference, and alarm linkage.

[0132] (1) Data preprocessing and model loading. Including:

[0133] Video decoding and caching: Video data comes from the camera RTSP video stream (1080p / 30fps, H.264 encoding). Use the FFmpeg decoder to extract YUV420 frames and convert them to RGB format; maintain a circular buffer to store the latest 10s video clip.

[0134] Sensor data synchronization: Achieve camera and sensor clock synchronization (error < 1ms) through the PTP protocol; perform linear interpolation on low-frequency sensor data (1Hz) to generate 30Hz timing data that matches the video frame rate.

[0135] Convert the global model into a format suitable for edge devices and accelerate it according to the hardware conditions of the edge devices.

[0136] (2) Multi-modal real-time inference. Including:

[0137] Single-frame abnormal probability calculation: Construct a 3D input tensor from the current frame and the 5 frames before and after it (a total of 11 frames) to obtain the abnormal probability p of the current frame (0 <= p <= 1);

[0138] Temporal continuity verification: Through a sliding window test, if p > the threshold (such as 0.8) for 10 consecutive frames, it is determined as a suspected abnormal event.

[0139] (3) Alarm linkage and emergency response.

[0140] It is achieved through various methods such as local device alarm, SMS alarm, and email alarm. At the same time, alarm-related information such as event type, timestamp, and location is written into the local database and pushed to the command center; at the same time, abnormal video clips are pushed to the cloud and the command center through the MQTT protocol for resource scheduling and emergency plans.

[0141] Figure 1 The figure shows the architecture diagram of a multi-modal urban security video abnormal behavior real-time detection system based on federated learning, demonstrating the processing flow of the system, which mainly includes the following stages:

[0142] Edge node data collection stage, where video, sensor, and historical log data are respectively input into their local data sets from different devices;

[0143] Multi-modal data fusion stage, where the timestamps of multi-modal data are aligned through hardware synchronization and software compensation, and the aligned data is enhanced to improve the model training effect. The processed data is stored locally for subsequent model training and real-time inference;

[0144] Local model training stage, where each edge node uses its local data set on the initial model parameters through its respective computing unit for model training;

[0145] Local model upload stage, where each edge node sends the trained local model to the central coordination server through the uplink;

[0146] Global model aggregation stage, where the central coordination server performs an aggregation operation after receiving the model parameter data of each edge node;

[0147] Global model distribution stage, where the central coordination server sends the aggregated global model parameters to each edge node through the downlink.

[0148] This system realizes edge node localization processing through federated learning, combines multi-modal data such as video, temperature, and smoke, and uses dynamic federated frequency adjustment and lightweight privacy protection technology to solve the problems of privacy leakage, bandwidth pressure, and multi-modal collaboration in traditional solutions. The system deploys the 3D MobileNetV3 model on edge devices, supports millisecond-level abnormal detection (such as fires and violent events), with a 12% improvement in detection accuracy and a 95% reduction in bandwidth occupancy, and is applicable to scenarios such as smart cities and campus security.

[0149] The embodiment of the present invention also provides a method for real-time detection of abnormal behaviors in multi-modal urban security videos. The implementation of this method includes the following steps:

[0150] Step S1, the edge node collects local video streams and sensor data, performs spatio-temporal alignment and preprocessing;

[0151] Step S2, train a lightweight 3D CNN model based on the local dataset, add noise and masks, and then upload the parameters to the central server;

[0152] Step S3, the central server dynamically aggregates the parameters according to the node resource status, generates a global model and distributes it to the edge nodes;

[0153] Step S4, the edge node loads the global model, performs frame-by-frame analysis on the real-time video stream, and detects abnormal behaviors;

[0154] Step S5, confirm the abnormal event through multi-modal cross-verification, trigger a local alarm and encrypt and report it to the command center.

[0155] Among them, the dynamic aggregation strategy in step S3 includes:

[0156] Allocate aggregation weights according to the node data volume and annotation consistency score;

[0157] Adopt the asynchronous federated learning (FedAsync) protocol to allow high-load nodes to delay uploading parameters;

[0158] Downweight the parameters of low-performance nodes to avoid degradation of the global model performance.

[0159] The real-time detection process in step S4 includes:

[0160] Continuously detect the abnormal probability using a sliding window (window size 10 frames); if the abnormal probability is > 0.8 for 10 consecutive frames, it is determined as a suspected abnormal event;

[0161] Automatically enable INT8 quantization for low-resolution devices to ensure that the inference latency < 120ms.

[0162] This method is implemented based on the real-time detection of abnormal behaviors in the modal urban security video described in the above embodiments, and the specific implementation is as follows:

[0163] 1. Implement tasks such as local multi-modal data collection, real-time inference, and model training based on the edge node layer. The edge node layer consists of a data collection module, a local model module, and a privacy protection module.

[0164] Among them, the data collection module supports three types of data input:

[0165] (1) Video input: cameras supporting the RTSP protocol, with a resolution of 1080P, a frame rate of 30fps, and H.264 / H.265 encoding;

[0166] (2) Sensor Interface: Integrates temperature, smoke, and sound sensors with a sampling frequency of 1 Hz and a data format of JSON;

[0167] (3) Historical Log Synchronization: Loads labeled abnormal events (including information such as timestamps, types, and locations) from the local database.

[0168] The local model module uses lightweight 3D MobileNetV3. When dealing with videos, feature optimization is required, and a temporal convolutional layer (kernel_size = 3*3*3) is added; spatial-temporal features of video segments are extracted through 3D CNN; sensor temporal data (temperature, smoke, sound) is processed through the LSTM network; the weights of video and sensor features are dynamically allocated through the multi-head attention layer, and the formula is as follows:

[0169]

[0170] Among them, Q is the video feature, and K and V are the sensor features, realizing information complementarity between modalities.

[0171] The privacy protection module adds Laplace noise to the fully connected layer before uploading the model gradient to prevent parameter inversion attacks.

[0172] 2. Based on the central coordination layer, global model aggregation, dynamic frequency regulation, and resource coordination are realized. It includes:

[0173] (1) Federated Aggregation Unit,

[0174] When aggregating, weighted averaging is performed according to the node data quality (labeling consistency score) and data volume, and the weight formula is as follows:

[0175]

[0176] Among them, N i is the data volume of node i, and Acc i is the accuracy of the local model of node i on the validation set.

[0177] (2) Dynamic Frequency Regulation Unit,

[0178] Monitors the resources of each node by collecting the CPU / GPU utilization rate, network bandwidth, and remaining power of the nodes in real time. Different aggregation decisions are adopted by analyzing the resource status of different nodes: when the average bandwidth utilization rate of the node < 30% and the CPU utilization rate < 60%, the high-frequency aggregation mode is triggered (aggregation once every five minutes); when the bandwidth > 70% or the CPU > 80%, the low-frequency aggregation mode is triggered (once every 30 minutes) to reduce resource contention.

[0179] 3. Based on the multi-modal data fusion layer, spatio-temporal alignment and feature enhancement of cross-modal data are realized.

[0180] Spatio-temporal alignment strategy:

[0181] Synchronize the clocks of the camera and the sensor through PTP (Precision Time Protocol) with an error < 1ms to achieve hardware-level synchronization; perform linear interpolation on the low-frequency sensor data to match the video frame rate.

[0182] Abnormal event enhancement:

[0183] Use the GAN network locally at the edge node to generate rare abnormal events (such as smoke diffusion simulation) to improve the generalization of the model; inject adversarial samples (such as occluding video frames) during the training phase to enhance the robustness of the model.

[0184] The specific process of this method to achieve real-time detection of abnormal behaviors in multi-modal urban security videos is as follows:

[0185] 1. Model initialization and data preparation.

[0186] The central server initializes the local model parameters and pre-trains the basic model based on the public dataset, including the temporal convolutional layer and the cross-modal attention module; distributes the model weights to all edge nodes through the HTTPS protocol. Each edge node constructs a multi-modal dataset containing video clips, sensor time series signals, and historical annotations based on the data collected by the local camera and sensor. The central server issues the pre-trained global initial model to each node to ensure that the model has the basic abnormal detection ability. The edge device automatically adapts the model calculation accuracy according to its own hardware performance (such as GPU computing power, memory size). For example, high-performance devices use floating-point operations, and low-end devices enable quantization compression to reduce the computing overhead.

[0187] 2. Local model training.

[0188] Local model training is the core link for the system to achieve distributed intelligent analysis. Its goal is to complete the feature extraction and model optimization of multi-modal data locally at the edge node, avoid the transmission of raw data, and at the same time achieve cross-node knowledge sharing through federated learning. The training process fully combines the characteristics of edge computing resources and privacy protection requirements.

[0189] The local model processes video and sensor data through a dual-branch structure:

[0190] (1) Video analysis branch: Use a lightweight 3D convolutional network to extract the spatio-temporal features of the video and capture the connection change patterns of dynamic behaviors (such as flame diffusion).

[0191] (2) Sensor analysis branch: Use a time series model (such as LSTM) to encode the periodic patterns of signals such as temperature and smoke, and identify mutation events (such as sudden temperature rise, abnormal sound).

[0192] (3) Cross-modal fusion: Dynamically allocate weights to video and sensor features through the attention mechanism. For example, in fire detection, when the temperature sensor data is abnormal, the model automatically enhances the attention to the flame area in the video branch to achieve complementary multi-source information.

[0193] Based on the local dataset, the edge nodes are trained using the Focal Loss function to focus on optimizing the recognition ability of rare abnormal events and avoid model bias caused by overloading of normal event samples. During the training process, the model parameters are only updated locally, and the original data is retained inside the device throughout, fundamentally eliminating the risk of privacy leakage.

[0194] Before uploading the model gradients to the central server, the edge nodes desensitize the parameters through lightweight privacy protection techniques: adding controllable noise to the gradients so that attackers cannot reverse-engineer the original data content; randomly masking some non-critical parameters to further reduce the possibility of privacy leakage.

[0195] During the local training process, a flexible resource adaptation mechanism is supported. When the computing resources are insufficient, the model input resolution is automatically reduced or some training rounds are skipped to prioritize real-time inference performance; when a high-value abnormal event is detected, incremental learning is immediately triggered, and the latest event data is added to the training set to quickly improve the model's sensitivity to similar scenarios.

[0196] 3. Federated parameter aggregation.

[0197] Federated parameter aggregation is the core mechanism for the system to achieve distributed intelligent collaboration. Its goal is to integrate the local knowledge of multiple edge nodes and build a global model with high accuracy and strong generalization ability while protecting data privacy. This process realizes the dynamic fusion and optimization of model parameters through the efficient interaction between the central coordination server and the edge nodes. It includes:

[0198] (1) Parameter upload and privacy protection.

[0199] After the local model training is completed, each edge node does not directly upload the original data, but transmits the desensitized model parameters (such as gradients or weights) to the central server. To prevent the leakage of sensitive information, the node preprocesses the parameters using lightweight privacy protection techniques: adding random noise to the parameters to ensure that attackers cannot reverse-engineer the original data features.

[0200] (2) Dynamic aggregation strategy.

[0201] Based on the real-time status of all edge nodes in the network, the central server intelligently adjusts the frequency and priority of parameter aggregation to balance model update efficiency and system stability. During idle network periods, it automatically increases the aggregation frequency to fully utilize the limited bandwidth to accelerate model iteration; during periods of resource tension, it reduces the aggregation frequency to avoid exacerbating network congestion or device overload due to frequent communication; when a node detects a high-confidence abnormal event, it immediately triggers the "emergency aggregation" mode to preferentially integrate the parameters of that node and ensure that the global model quickly responds to threats.

[0202] (3) Weighted fusion and quality assessment.

[0203] To enhance the robustness of the global model, the central server does not simply average the parameters of all nodes, but instead introduces a data quality assessment mechanism to differentially allocate node weights. First, based on the annotation accuracy and data volume scale of the node's historical data, it dynamically calculates its contribution weight, increasing the weight of nodes with high label consistency and large data volumes in the aggregation to avoid contaminating the global model with low-quality data; second, for old devices that cause model deviations due to their limited computing power, the central server automatically reduces their weight ratio and simultaneously performs implicit calibration on them through the parameters of high-performance nodes to ensure the balanced performance of the model across different devices.

[0204] (4) Global model distribution and collaborative optimization.

[0205] After completing parameter aggregation, the central server distributes the updated global model to all edge nodes, initiating a new round of local training and inference. This process forms a closed-loop optimization. Through a version control mechanism, it ensures that all nodes synchronously load the latest model to avoid inconsistent detection results caused by version fragmentation;

[0206] During local training, each node can combine newly collected abnormal event data to continuously optimize the model's detection ability for region-specific scenarios (such as special fire hazards in industrial parks), and these improvements will contribute to the global model in the next round of federated aggregation.

[0207] 4. Real-time inference stage.

[0208] The real-time inference stage is the core link for the implementation of this system, achieving millisecond-level abnormal detection and linkage response through edge device localization processing. The design of this stage fully considers the complexity and real-time requirements of urban security scenarios, combining multi-modal data fusion and dynamic resource adaptation technologies to ensure the stable operation of the system under high concurrency and multi-scenarios. The following details its workflow from aspects such as data preprocessing, multi-modal inference, and alarm linkage.

[0209] (1) Data preprocessing and model loading. Including:

[0210] Video Decoding and Caching: The video data is from the camera RTSP video stream (1080p / 30fps, H.264 encoded). The FFmpeg decoder is used to extract YUV420 frames and convert them to the RGB format. A circular buffer is maintained to store the video segments of the most recent 10s.

[0211] Sensor Data Synchronization: The camera and sensor clocks are synchronized through the PTP protocol (error < 1ms). Linear interpolation is performed on the low-frequency sensor data (1Hz) to generate 30Hz timing data that matches the video frame rate.

[0212] The global model is converted into a format suitable for edge devices and accelerated according to the hardware conditions of the edge devices.

[0213] (2) Multimodal Real-time Inference. It includes:

[0214] Calculation of Single-frame Anomaly Probability: A 3D input tensor is constructed from the current frame and the 5 frames before and after it (11 frames in total) to obtain the anomaly probability p of the current frame (0 <= p <= 1);

[0215] Temporal Continuity Verification: Through a sliding window check, if p > the threshold (such as 0.8) for 10 consecutive frames, it is determined as a suspected anomaly event.

[0216] (3) Alarm Linkage and Emergency Response.

[0217] It is achieved through various methods such as local device alarms, SMS alarms, and email alarms. At the same time, alarm-related information such as event type, timestamp, and location is written into the local database and pushed to the command center; at the same time, the abnormal video segments are pushed to the cloud and the command center through the MQTT protocol for resource scheduling and emergency response plans.

[0218] Through the above specific implementation manners, those skilled in the art of the technical field can easily implement the present invention. However, it should be understood that the present invention is not limited to the above specific implementation manners. Based on the disclosed implementation manners, those skilled in the art of the technical field can arbitrarily combine different technical features to implement different technical solutions.

[0219] Except for the technical features described in the specification, they are all known technologies to those skilled in the art.

Claims

1. A real-time detection system for abnormal behaviors in multimodal urban security videos, characterized in that, Including: Edge node module, deployed at the camera terminal, for collecting local video streams and sensor data, and performing local model training and real-time anomaly detection; Central coordination server module, for federated model parameter aggregation, global model distribution, and dynamic resource coordination; Multimodal data fusion module, integrating video data, sensor data, and historical alarm logs, and achieving cross-modal feature fusion through spatio-temporal alignment and attention mechanism; Dynamic federated policy module, adaptively adjusting the model parameter aggregation frequency according to the edge node computing load and network status; Alarm linkage module, triggering local audible and visual alarms when detecting abnormal behaviors, and encrypting and transmitting event information to the command center.

2. The real-time detection system for abnormal behaviors in multi-modal urban security videos according to claim 1, characterized in that, The edge node module specifically includes: Video acquisition unit, a camera supporting the RTSP protocol, outputting a 1080p video stream encoded by H.264 / H.265; Sensor interface unit, integrating temperature, smoke, and sound sensors, with a data sampling frequency ≥ 1Hz; Local module unit, including a lightweight 3D convolutional neural network (3D MobileNetV3) and an LSTM network, for extracting video spatio-temporal features and sensor temporal features; Privacy protection unit, for adding Laplace noise to the fully connected layer before uploading the model gradients.

3. A real-time detection system for abnormal behaviors in multimodal urban security videos according to claim 1, characterized in that, The multimodal data fusion module includes: Hardware-level clock synchronization unit, aligning the camera and sensor clocks through the PTP protocol; Cross-modal attention fusion unit, using video features as Query, and sensor features as Key / Value, and calculating multi-head attention weights; Abnormal event cross-validation unit, when an abnormality is detected in the video, verifying whether the sensor data conforms to the expected pattern.

4. A real-time detection system for abnormal behaviors in multimodal urban security videos according to claim 3, characterized in that, Dynamically allocate the weights of video and sensor features through the multi-head attention layer, and the formula is as follows: Where Q is the video feature, and K and V are the sensor features, realizing information complementarity between modalities.

5. A real-time abnormal behavior detection system for multi-modal urban security videos according to claim 1, characterized in that, The working logic of the dynamic federated policy module is: When the average bandwidth utilization rate of the edge node < 30% and the CPU utilization rate < 60%, trigger high-frequency aggregation; When the bandwidth utilization rate > 70% or the CPU utilization rate > 80%, switch to low-frequency aggregation; When a high-confidence abnormal event is detected, immediately trigger emergency aggregation and preferentially update the global model.

6. The real-time detection system for abnormal behaviors in multimodal urban security videos according to claim 5, wherein, When aggregating, allocate weights according to the node data quality and data volume for weighted average, and the weight formula is as follows: Among them, N i is the data volume of node i, and Acc i is the accuracy of the local model of node i on the validation set.

7. A real-time detection system for abnormal behaviors in multi-modal urban security videos according to claim 1, characterized in that, The alarm linkage module includes: Local alarm unit, controlling the LED and buzzer through GPIO to trigger audible and visual alarm signals; Encryption transmission unit, encrypting abnormal video segments using AES-256 and pushing them to the cloud through the MQTT protocol; Emergency resource scheduling unit, automatically generating the optimal resource scheduling plan.

8. A real-time detection method for abnormal behaviors in multimodal urban security videos, characterized in that, The implementation of this method includes the following steps: Step S1, the edge node collects local video streams and sensor data, and performs spatio-temporal alignment and preprocessing; Step S2, train a lightweight 3D CNN model based on the local dataset, add noise and masks, and then upload the parameters to the central server; Step S3, the central server dynamically aggregates the parameters according to the node resource status, generates a global model, and distributes it to the edge nodes; Step S4: The edge node loads the global model, analyzes the real-time video stream frame by frame, and detects abnormal behaviors. Step S5: Confirm the abnormal event through multimodal cross-verification, trigger a local alarm, and encrypt and report it to the command center.

9. A real-time detection method for abnormal behaviors in multi-modal urban security videos according to claim 8, characterized in that, The dynamic aggregation strategy in step S3 includes: Allocating aggregation weights according to the node data volume and the annotation consistency score; Adopting an asynchronous federated learning protocol to allow high-load nodes to delay uploading parameters; Reducing the weights of the parameters of low-performance nodes to avoid degradation of the global model performance.

10. A real-time detection method for abnormal behaviors in multimodal urban security videos according to claim 8, characterized in that, The real-time detection process in step S4 includes: Continuously detecting the abnormal probability using a sliding window; if the abnormal probability of 10 consecutive frames > 0.8, it is determined as a suspected abnormal event; Automatically enabling INT8 quantization for low-resolution devices to ensure that the inference latency < 120 ms.

Citation Information

Cited By

  • Fire-fighting detection alarm system and method for emergency management

    CN120977065A

  • Marine engine room monitoring alarm system based on artificial intelligence

    CN121096069A

  • An artificial intelligence-based ship engine room monitoring and alarm system

    CN121096069B

  • Federal learning system supporting dynamic node access

    CN121356923A

  • Multi-mode Internet of Things field computing AI gateway

    CN121356944A