Intelligent security and protection monitoring system and method fused with multi-modal data analysis

Through multimodal data analysis and dynamic privacy protection, the target identification and privacy protection problems of intelligent monitoring systems in complex scenarios are solved, efficient target tracking and behavior analysis are achieved, and the reliability and privacy protection capabilities of the monitoring system are improved.

CN120339959AInactive Publication Date: 2025-07-18菏泽泰康工贸有限公司
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510552167.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent monitoring system has insufficient target recognition rate in complex scenarios, and the privacy protection solution has led to the failure of behavioral analysis. Multi-device collaboration is difficult to cope with dynamic target tracking. Centralized processing brings high latency and bandwidth pressure, which cannot effectively solve the contradiction between multi-modal data fusion, real-time privacy desensitization and distributed intelligent collaboration.

Method used

The heterogeneous sensor array is used to collect multimodal data in real time, and the space-time alignment and feature-level fusion is performed through lightweight neural networks. Combined with dynamic privacy protection mechanisms and distributed consensus algorithms, abnormal detection and behavioral analysis are realized, cross-camera target tracking, output hierarchical alarm signals and update the cloud knowledge base.

Benefits of technology

Maintain high target recognition accuracy in scenarios such as extreme light and dense fog, realize irreversible face desensitization and behavioral characteristics retention, improve monitoring reliability, comply with international privacy regulations, and support multi-system linkage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339959A_ABST
    Figure CN120339959A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent security and protection monitoring system and method fusing multi-modal data analysis, and relates to the field of intelligent monitoring, and the method comprises the following steps: collecting a video stream, infrared thermal imaging data and an environment audio signal in real time through a heterogeneous sensor array; operating a lightweight neural network at an edge computing node to carry out space-time alignment and feature level fusion on the multi-source data, and generating an enhanced environment sensing matrix; a dual-channel anomaly detection mechanism is adopted, a first channel identifies short-term sudden abnormal events through a dynamic threshold self-adaptive model, and a second channel early warns potential risk behaviors through a time sequence prediction model; according to the system and the method, a dynamic privacy protection mechanism is adopted, behavior analysis is completed on the premise of ensuring biological feature safety, and transformation from passive monitoring to active early warning is realized in combination with a space-time prediction model; the problem of multi-view target tracking is solved through a distributed consensus algorithm, and the monitoring reliability in a complex scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent monitoring, and in particular to an intelligent security monitoring system and method integrating multi-modal data analysis. Background Art

[0002] Current intelligent monitoring systems mainly rely on visible light cameras and simple motion detection algorithms, and have significant defects in complex scenarios: traditional monocular vision systems are vulnerable to illumination changes, bad weather, and occlusion interference, resulting in a target recognition rate of less than 60% at night or in haze environments; mainstream privacy protection solutions mostly use global mosaics or full encryption, which, although meeting compliance requirements, lead to the failure of behavior analysis; in terms of multi-device collaboration, existing solutions rely on fixed rules to switch perspectives, making it difficult to meet the requirements of dynamic target tracking, and the centralized processing architecture brings high latency and bandwidth pressure. With the implementation of regulations such as the Personal Information Protection Law, how to achieve precise monitoring while protecting privacy has become a pain point in the industry, and the existing technologies have not effectively solved the contradictions among multi-modal data fusion, real-time privacy de-sensitization, and distributed intelligent collaboration. Summary of the Invention

[0003] To solve the above-mentioned problems of the prior art, the present invention provides an intelligent security monitoring system and method integrating multi-modal data analysis, which adopts a dynamic privacy protection mechanism to complete behavior analysis while ensuring the security of biometric features, and combines a spatio-temporal prediction model to realize the transformation from passive monitoring to active early warning; the problem of multi-perspective target tracking is solved through a distributed consensus algorithm, significantly improving the monitoring reliability in complex scenarios.

[0004] The technical solution adopted by the present invention to solve its technical problems is: an intelligent security monitoring method integrating multi-modal data analysis, including the following steps: (a) Real-time collect video streams, infrared thermal imaging data, and environmental audio signals through a heterogeneous sensor array; (b) Run a lightweight neural network at an edge computing node to perform spatio-temporal alignment and feature-level fusion on multi-source data, and generate an enhanced environmental perception matrix; (c) Adopt a dual-channel anomaly detection mechanism. The first channel identifies short-term sudden anomaly events through a dynamic threshold adaptive model, and the second channel warns of potential risk behaviors through a time series prediction model; (d) Perform differential privacy encryption on the face and biometric data in the video stream to generate an irreversible feature hash code for behavior analysis; (e) When a cross-camera target is detected, activate a distributed consensus algorithm to coordinate the perspectives of multiple devices and construct a three-dimensional trajectory prediction model; (f) Output a hierarchical alarm signal and a visualized decision map to a monitoring terminal, and synchronously update the adaptive learning parameters in a cloud knowledge base.

[0005] Furthermore, the feature-level fusion in step (b) specifically includes: Aligning the spatial resolution differences between the video and infrared data through a cross-modal attention mechanism; Using audio spectrum features to correct the false trigger areas of moving target detection; Establishing a multi-sensor confidence evaluation model to dynamically adjust the fusion weights.

[0006] Furthermore, the differential privacy encryption process in step (d) includes: Performing regional block dynamic blurring on the original face image; Injecting Gaussian noise matrix after extracting biometric features; Implementing audit traceability by storing hash values through blockchain nodes.

[0007] An intelligent security monitoring system integrating multi-modal data analysis, comprising: A multi-spectral imaging module for integrating a visible light camera and a millimeter-wave radar; An edge computing unit equipped with a heterogeneous computing architecture accelerated by FPGA; An adaptive learning module including an online incremental learning engine and a model distillation component; A secure storage unit for protecting the key system using a physically unclonable function PUF.

[0008] Compared with the prior art, the beneficial effects of the present invention are: The multi-modal data fusion mechanism enables the system to maintain a target recognition accuracy of >92% in scenarios such as extreme lighting (≤1 lux) and thick fog (visibility <10m), which is more than 35% higher than the traditional monocular vision scheme; The differential privacy encryption technology achieves irreversible desensitization of more than 97% of faces, while retaining more than 98% of the effective behavioral feature data; the blockchain-based audit system greatly improves the data operation traceability efficiency and complies with international privacy regulations such as GDPR / CCPA; The three-dimensional trajectory prediction model supports value-added functions such as crowd density analysis and abnormal aggregation warning; the open API interface is compatible with the smart city management platform to realize multi-system linkage of fire protection / security / emergency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the schematic diagram of the system of the present invention. Specific embodiments

[0011] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0012] Referring to Figure 1 , the present invention provides an intelligent security monitoring method that integrates multi-modal data analysis, including the following steps: (a) Real-time collect video streams, infrared thermal imaging data and environmental audio signals through a heterogeneous sensor array; (b) Run a lightweight neural network at the edge computing node to perform spatio-temporal alignment and feature-level fusion on multi-source data, and generate an enhanced environmental perception matrix; (c) Adopt a dual-channel anomaly detection mechanism. The first channel identifies short-term sudden anomaly events through a dynamic threshold adaptive model, and the second channel warns of potential risk behaviors through a time series prediction model; (d) Perform differential privacy encryption on the face and biometric data in the video stream, and generate an irreversible feature hash code for behavior analysis; (e) When a cross-camera target is detected, activate the distributed consensus algorithm to coordinate multi-device perspectives and build a three-dimensional trajectory prediction model; (f) Output a hierarchical alarm signal and a visual decision map to the monitoring terminal, and synchronously update the adaptive learning parameters in the cloud knowledge base.

[0013] Further, the feature-level fusion in step (b) specifically includes: Align the spatial resolution differences between video and infrared data through a cross-modal attention mechanism; Use audio spectrum features to correct the mis-triggered areas of moving target detection; Establish a multi-sensor confidence evaluation model to dynamically adjust the fusion weights.

[0014] Further, the differential privacy encryption process in step (d) includes: Perform regional block dynamic blurring on the original face image; Inject Gaussian noise matrix after extracting biometric features; Achieve audit traceability by storing the hash value through blockchain nodes.

[0015] Taking the deployment of an intelligent monitoring system for urban transportation hubs as an example, the above steps will be described in detail.

[0016] I. Hardware Deployment Configuration (1) Installation of Sensor Array: Intelligent monitoring terminals are deployed at intervals of 20 meters in the waiting hall. Each terminal includes: Visible light camera: SONY IMX585 sensor, supporting HDR mode, with a minimum illuminance of 0.001 lux; Thermal imaging module: FLIR A700, temperature measurement range -20°C to 1500°C, accuracy ±2°C; Ring microphone array: 6 digital MEMS microphones, frequency response 50Hz - 16kHz; The device is powered by PoE++, supports the 802.3bt standard, and has a maximum power consumption of 25W; (2) Edge Computing Node: Adopt the NVIDIA Jetson AGX Orin module, with built-in: 64GB DDR5 memory, 1TB NVMe solid-state storage; Dual 10Gbps fiber optic network interfaces; FPGA acceleration card: Xilinx Zynq UltraScale+, realizing hardware acceleration for feature fusion.

[0017] II. Multimodal Data Processing Flow (1) Spatiotemporal Alignment Operation: The video stream (30fps) and thermal imaging data (15fps) are synchronized to 30fps using the frame interpolation algorithm, and the audio signal is aligned with the video data through timestamps, with an error < 1ms; (2) Cross-modal Feature Fusion (Claim 2): Video-Infrared Alignment: Adopt the spatial attention mechanism, using the thermal imaging data as the query vector (Query) and the visible light features as the key-value (Key-Value), and automatically focusing on the temperature anomaly area; Audio-assisted Correction: When the microphone detects a sudden sound above 90dB, the following operations are triggered: Suspend the detection of moving targets within a radius of 2 meters from the corresponding spatial coordinates; Start the sound source localization algorithm to correct the movement trajectory of the video target.

[0018] III. Anomaly Detection and Privacy Protection (1) Dual-channel Detection Mechanism: Short-term Burst Detection Channel: Adopt the YOLOv7-tiny model and dynamically adjust the confidence threshold; Threshold = Base threshold (0.5) × (1 + Current frame motion intensity); Trigger an alarm when a target with a confidence > 0.8 is detected for 3 consecutive frames; Long-term risk warning channel: Use the Transformer architecture to build a time series model, and the input dimensions are: [Pedestrian flow, Average moving speed, Temperature gradient, Sound intensity variance]; Send a warning when the predicted abnormal probability within the next 5 minutes > 75%; (2) Differential privacy processing: Dynamic blur processing: Use Mask R-CNN to real-time segment the face area. For non-key personnel: Eyes / nose tip area: Apply Gaussian blur with a radius of 15 pixels (σ = 8); The rest of the face area: Apply block mosaic (8×8 pixel blocks); Feature encryption process: Extract the 128-dimensional face feature vector; Inject Gaussian noise with a mean of 0 and a variance of 0.1; Generate a hash value through SHA-256 and store it in the Hyperledger Fabric blockchain.

[0019] IV. Multi-camera collaborative tracking Target handover protocol: When the target leaves the field of view of camera A: A broadcasts the target feature vector and the predicted trajectory (in polar coordinates) The neighboring cameras start the scanning mode and preferentially search in the range of ±30° of the predicted trajectory direction, where the thermal imaging feature matching degree > 85%; The first camera B that discovers the target sends a confirmation signal to update the global trajectory library; 3D trajectory reconstruction: Based on multi-view geometric constraints, use the SFM algorithm to construct the target motion trajectory: 3D coordinates = ∑(w_i * P_i^+ * x_i) / ∑w_i, where P_i is the camera projection matrix and w_i is the confidence weight.

[0020] V. System output and optimization Alarm signal classification standard:

[0021] Model online update mechanism: Execute federated learning aggregation every 24 hours: Each edge node uploads the model gradient (gradient clipping + differential privacy processing); The cloud aggregator updates the global model using the FedAvg algorithm; The new model is sent to each node, and the version number is incremented for verification.

[0022] VI. Measured performance data Continuity of target tracking: The maximum continuous tracking time for a single target is increased from 4.7 min to 28.3 min; Privacy processing efficiency: The real-time desensitization frame rate of 1080P video stream ≥ 25 fps; Multi-machine cooperation latency: The average time taken for target handover is 123 ms (standard deviation ±18 ms); Abnormal detection accuracy: It still maintains an mAP of 89.7% in rainy and foggy weather.

[0023] Refer to Figure 2 , the present invention provides an intelligent security monitoring system integrating multi-modal data analysis, including a multi-spectral imaging module, an edge computing unit, an adaptive learning module, and a secure storage unit.

[0024] I. Multi-spectral imaging module (1) Hardware composition Visible light camera: Model: Sony STARVIS 2 IMX678 sensor; Resolution: 3840×2160 @ 30fps, supporting Starlight 2.0 technology (minimum illuminance of 0.0005 lux); Optical configuration: f / 1.2 large aperture, 82° wide-angle lens, with 6-layer anti-reflection coating; Microwave radar: Model: TI AWR6843AOP millimeter-wave radar; Operating frequency: 60 - 64 GHz, bandwidth 4 GHz; Detection ability: Up to 120 meters, speed resolution of 0.1 m / s, angle accuracy of ±1°; Antenna array: 3 transmit 4 receive MIMO architecture.

[0025] (2) Multi-spectral data fusion architecture Hardware synchronization mechanism: The GPS / PPS clock signal is used to synchronize the camera and radar timestamps, with an error < 100 ns; Spatial calibration: The mapping relationship between the camera pixel coordinates and the radar polar coordinates is established through a checkerboard calibration board.

[0026] Edge computing unit (1) Heterogeneous computing architecture Main Processor: NVIDIA Jetson Orin NX with 64 TOPS computing power; Running Ubuntu 22.04 LTS + ROS2 Humble; FPGA Acceleration Module: Chip: Xilinx Zynq UltraScale+ MPSoC; Acceleration Functions: Video Codec: H.265 8K@60fps hardware decoding; Neural Network Inference: Deploying the quantized YOLOv7 model with latency < 8ms; Sensor Data Preprocessing: Radar point cloud filtering and denoising.

[0027] III. Adaptive Learning Module (1) Online Incremental Learning Engine Abnormal Sample Collection: When a false alarm with a confidence level > 0.9 is detected, automatically save the multimodal data for the 30 seconds before and after; Storage Format: HDF5 compressed package (including video, radar point cloud, audio waveform); Model Update Strategy: Start incremental training at 2 am every day; Use knowledge distillation technology to compress the model: Teacher Model: ResNet152 (cloud); Student Model: MobileNetV3 (edge); Update Verification: Deploy the new model when the mAP decrease on the validation set is < 2%.

[0028] (2) Federated Learning Aggregation Mechanism Gradient Protection Scheme: Add Laplace noise (ε = 0.5, δ = 1e-5) to achieve differential privacy; Use Homomorphic Encryption for gradient encryption.

[0029] IV. Secure Storage Unit (1) Physical Unclonable Function (PUF) Key System Chip Selection: Main Controller: Microchip ATECC608B security component; PUF Source: SRAM startup state entropy source (generating a 256-bit root key); Key Derivation Process: Extract the original entropy value from the SRAM PUF when powering on; Generate the device-unique key K_dev through HMAC-SHA256; Encrypt and store using K_dev: User privacy data: AES-256-GCM mode; Model parameters: XChaCha20-Poly1305 mode.

[0030] Of course, the above description is not limited to the above examples. The technical features not described in the present invention can be realized by or adopted the prior art, which will not be elaborated here; the above embodiments and drawings are only used to illustrate the technical solutions of the present invention and are not limitations to the present invention. The present invention has been described in detail with reference to the preferred embodiments. Those of ordinary skill in the art should understand that the changes, modifications, additions or substitutions made by those of ordinary skill in the art within the scope of the essence of the present invention do not depart from the purpose of the present invention and should also fall within the protection scope of the claims of the present invention.

Claims

1. An intelligent security monitoring method integrating multi-modal data analysis, characterized in that, Including the following steps: (a) Collecting video streams, infrared thermal imaging data, and environmental audio signals in real time through a heterogeneous sensor array; (b) Running a lightweight neural network on an edge computing node to perform spatio-temporal alignment and feature-level fusion on multi-source data, generating an enhanced environmental perception matrix; (c) Adopting a dual-channel anomaly detection mechanism. The first channel identifies short-term sudden anomaly events through a dynamic threshold adaptive model, and the second channel warns of potential risk behaviors through a time series prediction model; (d) Performing differential privacy encryption on the face and biometric data in the video stream to generate irreversible feature hash codes for behavior analysis; (e) When a cross-camera target is detected, activating a distributed consensus algorithm to coordinate multi-device perspectives and constructing a three-dimensional trajectory prediction model; (f) Outputting a hierarchical alarm signal and a visual decision-making map to the monitoring terminal, and synchronously updating the adaptive learning parameters in the cloud knowledge base.

2. The intelligent security monitoring method integrating multi-modal data analysis according to claim 1, wherein The feature-level fusion in step (b) specifically includes: Aligning the spatial resolution differences between video and infrared data through a cross-modal attention mechanism; Using audio spectrum features to correct the mis-triggered areas of moving target detection; Establishing a multi-sensor confidence evaluation model to dynamically adjust the fusion weights.

3. The intelligent security monitoring method integrating multi-modal data analysis according to claim 1, wherein The differential privacy encryption process in step (d) includes: Performing regional block dynamic blurring on the original face image; Injecting a Gaussian noise matrix after extracting biometric features; Realizing audit traceability by storing hash values through blockchain nodes.

4. An intelligent security monitoring system integrating multi-modal data analysis, characterized in that, Including: A multi-spectral imaging module for integrating a visible light camera and a millimeter-wave radar; An edge computing unit equipped with a heterogeneous computing architecture accelerated by FPGA; An adaptive learning module containing an online incremental learning engine and a model distillation component; A secure storage unit for protecting the key system using a physically unclonable function PUF.

Citation Information

Cited By

  • Visual communication system based on weak current engineering and operation method thereof

    CN120935322A

  • Nuclear power plant production command center management system

    CN121094496A

  • Audio and video multi-mode identification method based on artificial intelligence

    CN121095997A

  • An audio and video multi-modal recognition method based on artificial intelligence

    CN121095997B

  • Subway fire safety assessment method and system based on artificial intelligence

    CN121581646A