Multi-data-source rapid grading and screening system and method for security and protection
Through multi-protocol unified encoding, space-time alignment and dynamic feedback mechanisms, combined with knowledge graph and multimodal joint inference, the problems of data compatibility and cross-modal inference in intelligent security systems are solved, and efficient and accurate multi-source data hierarchical screening is achieved, reducing false alarm rates and missed detection rates.
Patent Information
- Application Number
- CN202510515429.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent security systems have problems such as insufficient data compatibility, undynamic resource allocation, and weak cross-modal inference capabilities in multi-source data fusion and hierarchical processing, resulting in high false alarm rates and missed detection rates, making it difficult to achieve accurate alarms in complex environments.
The protocol conversion layer is used to realize unified coding for multi-protocols, combining space-time alignment engine and dynamic feedback mechanism, and optimizing feature weights through reinforcement learning algorithms, building a knowledge graph inference unit and multimodal joint inference engine, and using hardware accelerators and lightweight neural network models for hierarchical screening to form a closed-loop optimization mechanism.
It realizes efficient fusion and real-time response of heterogeneous data, reduces false alarm rates and missed detection rates, and improves the system's dynamic adaptability and accurate alarm capabilities in complex environments.
Smart Images

Figure CN120337154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent security, and particularly to a multi-data-source rapid grading and screening system and method for security protection. Background Technique
[0002] With the rapid development of intelligent security technology, multi-source heterogeneous data (such as video streams, sensor time-series data, audio streams, etc.) has become the core element of the security system. The traditional security system has accumulated a large amount of cross-modal data in long-term operation, covering multi-dimensional information such as real-time monitoring, environmental perception, and behavior recognition, and its application scenarios have covered key fields such as public security, smart cities, and industrial protection.
[0003] However, there are significant bottlenecks in the multi-source data fusion and grading processing of the existing technology:
[0004] First, the data fusion efficiency is low. Due to protocol differences (such as RTSP, MQTT, etc.) in the traditional system, the compatibility of heterogeneous data is insufficient, and the time synchronization accuracy is poor, making it difficult to realize the real-time correlation analysis of cross-modal data.
[0005] Second, the resource allocation and dynamic adaptability are insufficient. The proportion of redundant data processing is too high, there is a lack of an intelligent filtering mechanism for low-quality data, and the static threshold strategy cannot adapt to environmental mutations (such as light, noise interference), resulting in high false alarm rates and missed detection rates.
[0006] Third, the cross-modal reasoning ability is weak. The multi-source data relies on manual rule association, lacks an automated joint reasoning engine, and the knowledge graph and attention mechanism are not deeply integrated, making it difficult to support accurate alarm decisions in complex scenarios.
[0007] The current technology is difficult to balance the contradictions among real-time performance, accuracy, and resource efficiency. Therefore, we propose a multi-data-source rapid grading and screening system and method for security protection, which realizes the accurate grading and efficient response of massive data through protocol adaptation optimization, dynamic weight adjustment, and multi-modal joint reasoning. Summary of the Invention
[0008] The purpose of the present invention is to provide a multi-data-source rapid grading and screening system and method for security protection.
[0009] To achieve the above purpose, the present invention provides the following technical solution: A multi-data-source rapid grading and screening system for security protection, the grading and screening system includes a multi-source heterogeneous data fusion module, a grading and screening module, and a computing acceleration module;
[0010] The grading and screening module includes a preprocessing unit, a primary screening unit, a fine screening unit, and a dynamic feedback unit connected in sequence;
[0011] The multi-source heterogeneous data fusion module is configured with a spatio-temporal alignment engine for time synchronization of video stream data, sensor time-series data, and audio stream data, with a synchronization error less than 100 ms;
[0012] The dynamic feedback unit is built-in with a weight online update module to dynamically adjust the feature weight matrix of multi-source data based on the reinforcement learning algorithm.
[0013] As a further solution of the present invention: The multi-source heterogeneous data fusion module further includes:
[0014] (a) A protocol conversion layer that supports network protocol conversion including RTSP, ONVIF, and MQTT, and uniformly encodes heterogeneous data into the H.264 / H.265 format;
[0015] (b) A knowledge graph reasoning unit that stores an entity relationship library in the security field. The entity relationship library contains association rules between device entities, environmental entities, and behavior entities, and realizes cross-modal event reasoning through a graph database.
[0016] As a further solution of the present invention: The fine screening unit includes a multi-modal joint reasoning engine, and its data processing process satisfies:
[0017]
[0018] Among them, F fusion The fused cross-modal feature matrix, F represents the feature matrix, F m 、F n respectively represent the features of modality m and the features of modality n. M = {v, t, a} represents the set of video, thermal sensor, and audio modalities. n→m is the attention weight matrix from modality n to modality m, which is generated through the cross-modal attention mechanism. γ m is the learning weight coefficient of modality m, and Norm(·) represents the layer normalization operation.
[0019] As a further solution of the present invention: The preprocessing unit includes:
[0020] (a) A hardware accelerator that uses a programmable logic device to implement image denoising and outputs preprocessed data with a peak signal-to-noise ratio greater than 35 dB;
[0021] (b) A data quality evaluator that calculates the QoD index based on multi-dimensional indicators including signal-to-noise ratio, data integrity, and time continuity, and dynamically eliminates data packets with a QoD value lower than 0.75.
[0022] As a further solution of the present invention: The computing acceleration module includes:
[0023] (a) A graphics processing unit cluster for parallel processing of video stream data, with a single node supporting simultaneous processing of over 150 1080P video streams;
[0024] (b) A tensor processing unit array with a processing delay for sensor timing data of less than 100 ms;
[0025] (c) An edge computing node configured with a low-power embedded processor to implement front-end data preprocessing with a power consumption of less than 5W.
[0026] The present invention also provides a method for rapid hierarchical screening of multiple data sources for security applications. The hierarchical screening method includes the following steps:
[0027] S1. Collect security device data through a multi-protocol adaptation interface. The devices include cameras, infrared sensors, and voiceprint acquisition devices;
[0028] S2. Perform spatio-temporal alignment on multi-source data to generate a synchronized data stream with a timestamp deviation of less than 100 ms;
[0029] S3. Execute the hierarchical screening process:
[0030] (301) Preprocessing level: Dynamically filter low-quality data based on the QoD index, and the QoD threshold is configurable in the range of 0.6 - 0.9;
[0031] (302) Primary screening: Use a lightweight neural network model for rapid inference, with the number of model parameters less than 10M;
[0032] (303) Fine screening level: Analyze cross-device association features through a multi-modal joint inference engine;
[0033] (304) Dynamic feedback: Based on the data in the historical false alarm case library, dynamically update the feature weight allocation strategy through the policy gradient algorithm;
[0034] S4. Output a hierarchical alarm signal to the security execution terminal;
[0035] S5. Feed back the updated feature weight matrix to the hierarchical screening process in step S3 to form a closed-loop optimization mechanism.
[0036] As a further aspect of the present invention: The calculation of dynamically adjusting the detection threshold in step S3(302) satisfies:
[0037]
[0038] where T base is the reference threshold, calculated by sliding based on historical data, L c is the current ambient light intensity, L avg 、L stdThe average value and standard deviation of the light intensity in the past 30 minutes respectively, N c is the current noise decibel value, N max is the maximum noise range supported by the device, k1 = 0.1 - 0.2, k2 = 0.15 - 0.25 are adjustable environmental sensitivity coefficients.
[0039] As a further solution of the present invention: the multi-modal joint reasoning in the step S3(303) includes:
[0040] (a) Generating virtual abnormal samples through transfer learning, and its potential space perturbation formula is:
[0041] Z abn = Z nor + λ·C 1 / 2 ·ε
[0042] where, Z nor is the latent vector of the normal sample, C is the normal data covariance matrix, ε is the standard normal distribution random vector, λ is the abnormal intensity coefficient, and the value range is 0.2 - 1.5;
[0043] (b) Adding the generated virtual abnormal samples to the training data set to improve the generalization ability of the model.
[0044] As a further solution of the present invention: the weight update strategy in the step S3(304) is:
[0045] Collecting the data of the false alarm case library every 10 - 60 minutes, and updating the weight matrix through the policy gradient algorithm, and its objective function is:
[0046]
[0047] where, E s~D represents the expectation of the state s sampled from the experience pool D, R(s,a) represents the expected reward for taking the action a in the environmental state s, KL(·) is the KL divergence constraint term of the new and old policies, β = 0.3 - 0.7 is the policy update conservative coefficient, and D is the experience pool containing historical decision data.
[0048] Adopting the above technical solutions, compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] 1. The present invention realizes the unified encoding and standardization of multiple protocols through a protocol conversion layer, eliminates the differences in heterogeneous data formats, and significantly reduces the data loss rate. At the same time, the spatio-temporal alignment engine combines with a dynamic feedback mechanism to optimize the feature weight allocation of multi-modal data in real time by means of a reinforcement learning algorithm, ensuring high-precision time synchronization of video streams, sensor time-series data, and audio streams, effectively capturing instantaneous abnormal events. Compared with traditional static strategies, this solution realizes the coordinated improvement of data fusion efficiency and real-time response ability through a cross-modal attention mechanism and online weight update, providing a reliable guarantee for accurate alarm in complex environments;
[0050] 2. The hierarchical screening module of the present invention adopts a hierarchical processing architecture: the preprocessing unit quickly filters out data packets with low signal-to-noise ratio and discontinuous time series at the edge through a hardware accelerator and a dynamic quality assessment mechanism, reducing the consumption of invalid computing resources. The primary screening unit deploys a lightweight neural network model to achieve fast inference with extremely low parameter quantities, balancing real-time performance and accuracy. The dynamic feedback unit continuously optimizes the threshold strategy based on historical false alarm data, adapts to dynamic changes such as ambient light and noise, and through the closed-loop process of "preprocessing - primary screening - fine screening - feedback", the system significantly improves the dynamic adaptability to complex scenarios while reducing resource occupancy, avoiding false alarms and missed detections caused by traditional fixed thresholds;
[0051] 3. The present invention constructs an entity relationship library in the field of security through the collaborative design of a knowledge graph reasoning unit and a multi-modal joint reasoning engine, deeply excavating the potential correlation features of video, thermal sensor, and audio data. The knowledge graph is based on a graph database to realize the regularized expression of device entities, environmental entities, and behavior entities. The multi-modal joint reasoning engine automatically learns the semantic association weights between different modalities through a cross-modal attention mechanism, and combines virtual abnormal samples generated by transfer learning to enhance the model's generalization and recognition ability for rare threats. This dual mechanism of "rule-driven + data-driven" breaks through the limitations of traditional single-modal analysis, provides multi-dimensional data support for intelligent alarm decision-making across devices and scenarios, and significantly reduces the risk of threat missed detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of the multi-source heterogeneous data fusion module in an embodiment of the present invention;
[0053] Figure 2 It is a flowchart of the hierarchical screening module and the computing acceleration module in an embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of the dynamic weight update mechanism in an embodiment of the present invention;
[0055] Figure 4 It is a schematic diagram of the multi-modal data fusion algorithm in an embodiment of the present invention;
[0056] Figure 5 This is the architecture diagram of the computing acceleration module in the embodiment of the present invention;
[0057] Figure 6 This is the inference logic diagram of the knowledge graph in the embodiment of the present invention;
[0058] Figure 7 This is the deployment scenario diagram of the embodiment in the embodiment of the present invention. Detailed implementation manners
[0059] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings. It should be noted here that the description of these implementation manners is used to help understand the present invention, but does not limit the present invention.
[0060] In addition, the technical features involved in the various implementation manners of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0061] Please refer to the attached Figure 1 - attached Figure 7 For a multi-data-source fast hierarchical screening system and method for security protection of the present invention, the hierarchical screening system includes a multi-source heterogeneous data fusion module, a hierarchical screening module, and a computing acceleration module;
[0062] The hierarchical screening module includes a preprocessing unit, a primary screening unit, a fine screening unit, and a dynamic feedback unit connected in sequence;
[0063] The multi-source heterogeneous data fusion module is configured with a spatio-temporal alignment engine for synchronizing video stream data, sensor time-series data, and audio stream data, and the synchronization error is less than 100 ms;
[0064] The dynamic feedback unit is built-in with a weight online update module to dynamically adjust the feature weight matrix of multi-source data based on the reinforcement learning algorithm.
[0065] In an embodiment of the present invention: The multi-source heterogeneous data fusion module further includes:
[0066] (a) A protocol conversion layer that supports network protocol conversion including RTSP, ONVIF, MQTT, etc., and uniformly encodes heterogeneous data into the H.264 / H.265 format;
[0067] (b) A knowledge graph inference unit that stores an entity relationship library in the security field. The entity relationship library contains association rules between device entities, environmental entities, and behavior entities, and realizes cross-modal event inference through a graph database.
[0068] In an embodiment of the present invention: The fine screening unit includes a multi-modal joint inference engine, and its data processing process satisfies:
[0069]
[0070] Among them, F fusion The fused cross-modal feature matrix, where F represents the feature matrix, F m 、F n respectively represent the features of modality m and the features of modality n. M = {v, t, a} represents the set of video, thermal sensor, and audio modalities. n→m is the attention weight matrix from modality n to modality m, generated by the cross-modal attention mechanism, and γ m is the learning weight coefficient of modality m, and Norm(·) represents the layer normalization operation.
[0071] In an embodiment of the present invention: The preprocessing unit includes:
[0072] (a) A hardware accelerator that uses a programmable logic device to implement image denoising and outputs preprocessed data with a peak signal-to-noise ratio greater than 35 dB;
[0073] (b) A data quality evaluator that calculates the QoD index based on multi-dimensional metrics including signal-to-noise ratio, data integrity, and time continuity, and dynamically eliminates data packets with a QoD value lower than 0.75.
[0074] In an embodiment of the present invention: The computing acceleration module includes:
[0075] (a) A graphics processing unit cluster for parallel processing of video stream data, with a single node supporting simultaneous processing of more than 150 1080P video streams;
[0076] (b) A tensor processor array with a processing delay of sensor timing data lower than 100 ms;
[0077] (c) An edge computing node configured with a low-power embedded processor to implement front-end data preprocessing with a power consumption lower than 5W.
[0078] The present invention also provides a method for fast hierarchical screening of multi-data sources for security protection. The hierarchical screening method includes the following steps:
[0079] S1. Collect security device data through a multi-protocol adaptation interface. The devices include cameras, infrared sensors, and voiceprint acquisition devices;
[0080] S2. Perform spatio-temporal alignment on the multi-source data to generate a synchronized data stream with a timestamp deviation less than 100 ms;
[0081] S3. Execute the hierarchical screening process:
[0082] (301) Preprocessing level: Dynamically filter low-quality data based on the QoD index. The QoD (Quality of Data) threshold is configurable in the range of 0.6 - 0.9;
[0083] (302) Primary screening: A lightweight neural network model is used for fast inference with the number of model parameters less than 10M.
[0084] (303) Fine screening level: Analyze cross-device association features through a multimodal joint inference engine.
[0085] (304) Dynamic feedback: Based on the data in the historical false alarm case library, dynamically update the feature weight allocation strategy through the policy gradient algorithm.
[0086] S4. Output a graded alarm signal to the security execution terminal.
[0087] S5. Feed back the updated feature weight matrix to the graded screening process in step S3 to form a closed-loop optimization mechanism.
[0088] In an embodiment of the present invention: The calculation of dynamically adjusting the detection threshold in step S3(302) satisfies:
[0089]
[0090] where Tbase is the reference threshold, calculated by sliding according to historical data, Lc is the current environmental light intensity, L avg 、L std are respectively the average value and standard deviation of the light intensity in the past 30 minutes, N c is the current noise decibel value, N max is the maximum noise range supported by the device, and k1 = 0.1 - 0.2, k2 = 0.15 - 0.25 are adjustable environmental sensitivity coefficients.
[0091] In an embodiment of the present invention: The multimodal joint inference in step S3(303) includes:
[0092] (a) Generate virtual abnormal samples through transfer learning, and its potential space perturbation formula is:
[0093] Z abn =Z nor +λ·C 1 / 2 ·ε
[0094] where Z nor is the latent vector of the normal sample, C is the normal data covariance matrix, ε is a standard normal distribution random vector, and λ is the abnormal intensity coefficient with a value range of 0.2 - 1.5;
[0095] (b) Add the generated virtual abnormal samples to the training data set to improve the model generalization ability.
[0096] In an embodiment of the present invention: The weight update strategy in step S3(304) is:
[0097] Collect the false alarm case library data every 10 - 60 minutes, and update the weight matrix through the policy gradient algorithm. Its objective function is:
[0098] L = E s~D [R(s,a)] + β·KL(π old ||π new )
[0099] Among them, E s~D represents the expectation of the state s sampled from the experience pool D, R(s,a) represents the expected reward for taking the action a in the environmental state s, KL(·) is the KL divergence constraint term of the new and old policies, β = 0.3 - 0.7 is the policy update conservative coefficient, and D is the experience pool containing historical decision-making data.
[0100] Example 1. Real-time hierarchical screening of multi-source security data in a smart park:
[0101] Application scenario: A certain smart park deploys multiple types of security devices (high-definition cameras, infrared thermal sensors, voiceprint acquisition devices), and needs to process video streams, thermal sensor time-series data, and audio streams in real time to achieve accurate identification of intrusion behaviors;
[0102] Implementation steps:
[0103] 1. Data collection and protocol conversion:
[0104] The camera transmits the H.264 video stream through the RTSP protocol, the infrared sensor sends the JSON format thermal data through the MQTT protocol, and the voiceprint device transmits the audio stream through the ONVIF protocol.
[0105] The protocol conversion layer uniformly encodes the heterogeneous data into the H.265 format, eliminates data loss caused by protocol differences, and improves the compatibility to 98%.
[0106] 2. Space-time alignment and preprocessing:
[0107] The space-time alignment engine realizes hardware clock synchronization through PTP (Precision Time Protocol), and combines the dynamic interpolation compensation algorithm to ensure that the timestamp deviation of the video stream, sensor time-series data, and audio stream is less than 80ms.
[0108] The hardware accelerator (FPGA) denoises the video stream and outputs a clear image with PSNR > 38dB. The data quality evaluator dynamically eliminates low-quality audio packets with QoD < 0.7 (such as invalid data with background noise > 65dB).
[0109] 3. Hierarchical screening process:
[0110] Primary screening: A lightweight CNN model (with 8.5M parameters) runs on an edge computing node (NVIDIA Jetson Nano) to detect moving targets in videos with a latency of <50ms.
[0111] Fine screening level: The multimodal joint inference engine correlates moving targets in the video, abnormal temperatures (>40°C) of thermal sensors, and abnormal voiceprints (such as the sound of broken glass), calculates weights through a cross-modal attention mechanism, and outputs potential intrusion events.
[0112] Dynamic feedback: When the ambient light changes suddenly (such as a 60% sudden drop in light intensity due to heavy rain), the weight online update module, based on historical false alarm data, uses the policy gradient algorithm to increase the weight of thermal sensor data from 0.3 to 0.6 within 20 minutes, reducing the false detection rate of the video.
[0113] 4. Alarm output:
[0114] Hierarchical alarm signals (Level 1: moving targets, Level 2: temperature + voiceprint anomalies) are pushed to the security control system, triggering the corresponding area's light alarm and drone patrol.
[0115] Technical effects:
[0116] The efficiency of multi-source data fusion is increased by 40%, and the false alarm rate is reduced from 35% of the traditional system to 8%.
[0117] The power consumption of the edge node is <4W, supporting single-node parallel processing of 180 1080P video streams.
[0118] Technical effects:
[0119] The efficiency of multi-source data fusion is increased by 40%, and the false alarm rate is reduced from 35% of the traditional system to 8%. The power consumption of the edge node is <4W, supporting single-node parallel processing of 180 1080P video streams. Example 2: Dynamic environment adaptation in industrial protection scenarios:
[0120] Application scenario: Deploying a security system in a chemical plant, which needs to cope with environmental challenges such as high noise (equipment operation noise >75dB) and complex heat source interference (steam pipe temperature fluctuations).
[0121] Implementation steps:
[0122] 1. Enhancement of data quality:
[0123] The preprocessing unit dynamically configures the QoD threshold to 0.8 through the data quality evaluator, filtering out invalid data with a sensor signal interruption duration >200ms (such as sensor packet loss caused by electromagnetic interference).
[0124] The tensor processor array (Google TPU v4) processes 128 channels of thermal data in parallel with a latency of less than 90ms, identifying temperature anomalies (>100°C) caused by steam leaks.
[0125] 2. Cross-modal reasoning optimization:
[0126] The knowledge graph reasoning unit associates the equipment entity (pipeline valve ID: V-203), the environmental entity (regional temperature), and the behavioral entity (valve abnormal opening), and generates the rule: "If the V-203 valve status is open and the regional temperature is greater than 100°C for 10 seconds, a steam leak alarm is triggered."
[0127] The multimodal joint reasoning engine generates virtual abnormal samples (such as voiceprint features caused by simulated valve corrosion) through transfer learning and adds them to the training data set, which improves the model's recognition accuracy for low-frequency anomalies (occurring less than 5 times a year) by 25%.
[0128] 3. Dynamic threshold adjustment:
[0129] The noise sensitivity coefficient k2 is adjusted from 0.2 to 0.25. When the instantaneous peak value of the equipment noise is greater than 85dB, the detection weight of the audio data is dynamically reduced to avoid false triggering of electrical fire alarms.
[0130] The edge computing node (Raspberry Pi 4B) calculates the ambient light intensity L in real time c With noise N c , dynamically adjust the detection threshold T according to the formula adj , which reduces the missed detection rate of targets under sudden changes in lighting (such as turning on searchlights at night) by 30%.
[0131] 4. Resource collaborative management:
[0132] The computing acceleration module dynamically allocates resources according to the load: during peak video streaming hours (such as shift changes), the GPU cluster prioritizes processing 150 channels of video; during non-peak hours, 80% of the computing power is used for thermal data analysis.
[0133] Technical effect:
[0134] The false alarm rate is reduced from 42% of traditional solutions to 12%, and the threat response speed is less than 3 seconds.
[0135] Virtual anomaly samples reduced the model’s missed detection rate for rare events, such as pipeline corrosion leaks, from 28% to 9%.
[0136] Example 3: Multimodal real-time security monitoring of urban transportation hubs:
[0137] Application scenario: A multi-source security system is deployed in a subway station, which needs to monitor crowd density, abnormal behaviors (such as going against the flow, staying), and sudden accidents (such as falling, items left behind) in real time.
[0138] Implementation steps:
[0139] 1. Heterogeneous data fusion and spatio-temporal alignment:
[0140] The protocol conversion layer uniformly encodes 8 RTSP video streams (4K resolution), 12 MQTT infrared sensors (monitoring the heat map of the flow of people), and 6 ONVIF audio streams (collecting platform broadcasts and ambient sounds) into the H.265 format, and improves the data compatibility to 97%.
[0141] The spatio-temporal alignment engine synchronizes the clocks of multiple devices through the PTP protocol, and the time stamp deviation of video frames, sensor sampling points, and audio packets < 70 ms, ensuring the time consistency of cross-modal data.
[0142] 2. Hierarchical screening and dynamic resource allocation:
[0143] Preprocessing level: The hardware accelerator (Xilinx Zynq UltraScale+) denoises the video stream (PSNR > 40 dB), and the data quality evaluator dynamically eliminates blurred images with QoD < 0.65 (such as frames with a flow of people occlusion rate > 30%).
[0144] Primary screening: The lightweight YOLOv5s model (with 7.2M parameters) runs on the edge node (Intel NUC), and detects abnormal areas of the flow of people density in the video with a latency < 40 ms (such as the number of people in a certain area > 100 people / 10㎡).
[0145] Fine screening level: The multi-modal joint inference engine correlates reverse behaviors in the video, abnormal infrared heat maps (temperature gradient mutation > 5℃ / m 2 ) and screams in the audio, and generates a comprehensive alarm signal through cross-modal attention weights (video: 0.6, thermal sense: 0.3, audio: 0.1).
[0146] 3. Dynamic environment adaptation:
[0147] When the platform illumination drops to 50 lux due to a power outage, the dynamic feedback unit automatically adjusts the detection threshold T based on historical data (the average illumination in the past 30 minutes is 85 lux) adj , reduces the video weight from 0.6 to 0.4, and increases the infrared weight from 0.3 to 0.5, avoiding missed detections caused by insufficient illumination.
[0148] The computing acceleration module dynamically allocates resources: During peak hours (morning and evening commuting), the GPU cluster (NVIDIA A100) preferentially processes video streams (200 1080P), and during off-peak hours, 80% of the computing power is used for audio anomaly analysis.
[0149] Technical effects:
[0150] Cross-modal reasoning improves the accuracy of abnormal behavior recognition from 78% of traditional single-modal to 94%.
[0151] Dynamic resource allocation reduces the system power consumption by 22%, and a single node supports the processing of 200 video streams.
[0152] Example 4. Fire warning linkage in the complex environment of commercial complexes:
[0153] Application scenario: Large shopping malls need to monitor fire risks (smoke, high temperature, abnormal sound sources) in real time and link fire-fighting equipment.
[0154] Implementation steps:
[0155] Data quality optimization and protocol adaptation:
[0156] The data quality evaluator configures the QoD threshold to 0.85 to filter out the time-series interruption data (packet loss rate > 15%) of the smoke sensor caused by the interference of the ventilation system.
[0157] The protocol conversion layer uniformly encodes the data of 12 ONVIF smoke detectors, the video streams of 8 RTSP thermal cameras, and the data of 5 MQTT voiceprint sensors, and the protocol compatibility reaches 96%.
[0158] Knowledge graph-driven cross-device reasoning:
[0159] The knowledge graph reasoning unit constructs an entity relationship rule: "If the smoke concentration in area A > 5% LEL and the thermal sense temperature > 60°C for 10 seconds, then trigger a first-level fire warning."
[0160] The multi-modal joint reasoning engine generates virtual smoke samples (simulating the RGB characteristics of different combustible substances) through transfer learning, and adding them to the training set improves the recognition rate of the model for new types of fires (such as lithium battery combustion) by 18%.
[0161] Dynamic threshold and resource coordination:
[0162] When the environmental noise > 80 dB due to promotional activities, dynamically adjust the k2 coefficient in the formula to 0.22 to reduce the weight of audio data and avoid false alarms of fire alarms.
[0163] The edge node (Rockchip RK3588) processes thermal sense data in real time (delay < 90 ms) and links the fire sprinkler system, with a response speed < 2 seconds.
[0164] Reinforcement learning strategy optimization:
[0165] The weight online update module collects false alarm cases (such as misjudging steam pipes as smoke) every 30 minutes and updates the feature weights through the policy gradient algorithm, reducing the false alarm rate of smoke detection from 20% to 6%.
[0166] The dynamic feedback unit dynamically increases the weight of thermal sensation data from 0.5 to 0.7 during the refined screening stage based on historical fire data (such as the duration of high temperature).
[0167] Technical effects:
[0168] The accuracy rate of fire warning is increased from 82% of the traditional solution to 96%, and the linkage response speed is < 3 seconds.
[0169] The virtual sample training reduces the false negative rate of the new type of fire from 35% to 12%.
[0170] Although the present invention is disclosed above in a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modifications, equivalent changes and decorations made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the protection scope defined by the claims of the present invention.
Claims
1. A multi-data-source fast hierarchical screening system for security protection, characterized in that: The hierarchical screening system includes a multi-source heterogeneous data fusion module, a hierarchical screening module, and a computing acceleration module; The hierarchical screening module includes a preprocessing unit, a primary screening unit, a fine screening unit, and a dynamic feedback unit connected in sequence; The multi-source heterogeneous data fusion module is configured with a spatio-temporal alignment engine for time synchronization of video stream data, sensor time-series data, and audio stream data; The dynamic feedback unit is built with a weight online update module to dynamically adjust the feature weight matrix of multi-source data based on the reinforcement learning algorithm.
2. The multi-data-source fast hierarchical screening system for security protection according to claim 1, wherein, The multi-source heterogeneous data fusion module further includes: (a) A protocol conversion layer that supports network protocol conversion including RTSP, ONVIF, and MQTT, and uniformly encodes heterogeneous data into the H.264 / H.265 format; (b) A knowledge graph reasoning unit that stores an entity relationship library in the security field. The entity relationship library contains association rules between device entities, environmental entities, and behavior entities, and realizes cross-modal event reasoning through a graph database.
3. The multi-data-source rapid classification and screening system for security protection according to claim 1, characterized in that, The fine screening unit includes a multi-modal joint reasoning engine, and its data processing process satisfies: Among them, F fusion is the fused cross-modal feature matrix, F represents the feature matrix, F m , F n respectively represent the features of modality m and modality n. M = {v, t, a} represents the set of video, thermal sensor, and audio modalities. n→m is the attention weight matrix from modality n to modality m, generated by the cross-modal attention mechanism. γ m is the learning weight coefficient of modality m, and Norm(·) represents the layer normalization operation.
4. The multi-data-source fast hierarchical screening system for security protection according to claim 1, wherein, The preprocessing unit includes: (a) A hardware accelerator that uses a programmable logic device to implement image denoising and outputs preprocessed data with a peak signal-to-noise ratio greater than 35 dB; (b) A data quality evaluator that calculates the QoD index based on multi-dimensional metrics including signal-to-noise ratio, data integrity, and time continuity, and dynamically eliminates data packets with a QoD value lower than 0.
75.
5. The multi-data-source rapid classification and screening system for security protection according to claim 1, wherein The computing acceleration module includes: (a) A graphics processing unit cluster for parallel processing of video stream data, with a single node supporting simultaneous processing of more than 150 1080P video streams; (b) A tensor processing unit array with a processing delay of sensor time-series data less than 100 ms; (c) An edge computing node configured with a low-power embedded processor to implement front-end data preprocessing with a power consumption lower than 5W.
6. A method for rapid hierarchical screening of multiple data sources for security protection, characterized in that, The hierarchical screening method includes the following steps: S1. Collect security device data through a multi-protocol adaptation interface. The devices include cameras, infrared sensors, and voiceprint acquisition devices; S2. Perform spatio-temporal alignment on multi-source data to generate a synchronized data stream with a timestamp deviation less than 100 ms; S3. Execute the hierarchical screening process: (301) Preprocessing level: Dynamically filter low-quality data based on the QoD index, and the QoD threshold is configurable in the range of 0.6 - 0.9; (302) Primary screening: Use a lightweight neural network model for fast reasoning, with the number of model parameters less than 10M; (303) Fine screening level: Analyze cross-device association features through a multi-modal joint reasoning engine; (304) Dynamic feedback: Dynamically update the feature weight allocation strategy based on the data in the historical false alarm case library through the policy gradient algorithm; S4. Output a hierarchical alarm signal to the security execution terminal; S5. Feedback the updated feature weight matrix to the hierarchical screening process in step S3 to form a closed-loop optimization mechanism.
7. A multi-data-source fast hierarchical screening method for security protection according to claim 6, characterized in that The calculation of dynamically adjusting the detection threshold in step S3(302) satisfies: Among them, T base is the reference threshold, calculated by sliding according to historical data, L c is the current ambient light intensity, L avg , L std are respectively the average value and standard deviation of the light intensity in the past 30 minutes, N c is the current noise decibel value, N max is the maximum noise range supported by the device, and k1 = 0.1 - 0.2, k2 = 0.15 - 0.25 are adjustable environmental sensitivity coefficients.
8. A method for rapid hierarchical screening of multiple data sources for security protection according to claim 6, characterized in that, The multi-modal joint reasoning in step S3(303) includes: (a) Generate virtual abnormal samples through transfer learning, and its potential space perturbation formula is: Z abn = Z nor + λ·C 1 / 2 ·ε Among them, Z nor is the latent vector of the normal sample, C is the normal data covariance matrix, ε is a standard normal distribution random vector, and λ is the anomaly intensity coefficient, and the value range is 0.2 - 1.5; (b) Add the generated virtual anomaly samples to the training dataset to improve the generalization ability of the model.
9. A multi-data-source fast hierarchical screening method for security protection according to claim 6, characterized in that, The weight update strategy in step S3(304) is as follows: Collect the data in the false alarm case library every 10 - 60 minutes, and update the weight matrix through the policy gradient algorithm. Its objective function is: Among them, E s~D denotes the expectation of the state s sampled from the experience pool D, R(s, a) represents the expected reward for taking action a in the environmental state s, KL(·) is the KL divergence constraint term between the old and new policies, β = 0.3 - 0.7 is the policy update conservative coefficient, and D is the experience pool containing historical decision-making data.
Citation Information
Cited By
Safety monitoring management method and system based on Internet of Things
CN120932176A
Intelligent management and control system for real-time monitoring and early warning of construction site safety risk
CN121052655A
An intelligent management and control system for real-time monitoring and early warning of safety risks at construction sites
CN121052655B
Multi-modal edge computing security method and system
CN121167583A