A security system linkage and intrusion behavior recognition method fusing multi-modal data

By using multimodal data fusion and a dynamic strategy engine, the problem of insufficient perception capability of intelligent security systems under single-modality conditions is solved, enabling the identification and adaptive response to complex intrusion behaviors, and possessing self-optimization and cross-site knowledge sharing capabilities.

CN122286388APending Publication Date: 2026-06-26SHANDONG HAILIANXUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG HAILIANXUN INFORMATION TECH CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing intelligent security systems rely on single-modal data, resulting in decreased perception capabilities in low-light, inclement weather, or obstructed or camouflaged scenarios. They also suffer from poor robustness, high false alarm and false negative rates, and lack self-learning and optimization capabilities, leading to rigid response strategies that cannot adapt to new threat patterns and environmental changes.

Method used

By fusing multimodal data and employing a multimodal collaborative embedding and cognitive fusion network, a dynamic semantic scene graph is constructed. Combined with a dynamic policy engine and a multi-scale memory bank, the system can identify and self-optimize intrusion behaviors, and has the ability to continuously improve itself.

Benefits of technology

It achieves deep integration of multimodal data at the hardware, feature, and semantic levels, enabling it to understand complex intrusion behaviors, dynamically optimize response strategies, reduce false alarm and false negative rates, and possess self-learning and cross-site knowledge sharing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122286388A_ABST
    Figure CN122286388A_ABST
Patent Text Reader

Abstract

This invention discloses a method for security system linkage and intrusion behavior recognition that integrates multimodal data, specifically relating to the field of intelligent security technology. The method includes: First, performing three-level spatiotemporal alignment and semantic-level collaborative embedding fusion on multi-source data to generate unified fusion features; second, constructing and reasoning a dynamic semantic scene graph, identifying high-order intrusion behaviors through graph pattern matching and logical rules, and outputting structured behavioral events; then, based on behavioral events and multi-dimensional context, performing multi-objective optimization through a dynamic policy engine to generate and execute collaborative linkage strategies, and dynamically replanning based on feedback; finally, achieving continuous self-evolution and cross-scene knowledge transfer through multi-scale memory and metacognitive optimization mechanisms. This invention solves the problems of traditional systems such as single perception, rigid linkage, and inability to adapt, achieving highly accurate, interpretable, and self-evolving intelligent security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent security technology, and more specifically, to a method for security system linkage and intrusion behavior recognition that integrates multimodal data. Background Technology

[0002] Existing intelligent security systems typically employ a single or limited number of sensors (such as visible light cameras) for monitoring and trigger alarms through pre-defined video analysis algorithms (such as moving target detection and boundary crossing detection) or deep learning-based behavior recognition models (such as action classification based on convolutional neural networks). To achieve system linkage, a simple "event-triggered - pre-defined action" response mode is commonly adopted. That is, when a specific alarm event is detected (such as area intrusion), a set of predefined and fixed control commands are automatically executed (such as turning on lights, activating audible and visual alarms, or rotating the pan-tilt unit).

[0003] However, in practical use, it still has some drawbacks. For example, relying on single-modal data leads to a sharp decline in the system's perception capabilities under conditions of insufficient light, severe weather, obstruction, or camouflage, resulting in poor robustness and high false alarm and false negative rates. Simple rule linkages lack comprehensive consideration of intrusion intent, scene context, and system resource status, resulting in rigid response strategies that cannot achieve optimal collaborative handling and resource scheduling, and may cause interference or security vulnerabilities due to improper linkages. The system generally lacks the ability to learn and optimize from historical data and handling results. Its perception model, behavior rule base, and linkage strategies are fixed once deployed, making it unable to adapt to new threat patterns or dynamically changing environments. Operation, maintenance, and upgrades are highly dependent on manual intervention, and the level of intelligence is limited. Summary of the Invention

[0004] To overcome the aforementioned deficiencies of the prior art, this invention provides a method for security system linkage and intrusion behavior identification that integrates multimodal data, and solves the problems mentioned in the background art through the following scheme.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for security system linkage and intrusion behavior identification that integrates multimodal data, comprising:

[0006] S1: Hardware-level spatiotemporal synchronization and feature-level adaptive alignment are performed on data from multiple sensors, including vision, hearing, and radar. The aligned multimodal features are then input into a multimodal collaborative embedding and cognitive fusion network. This network first uses visual features as a guide to perform semantic filtering and enhancement on other modal features. Subsequently, a multi-head cross-modal Transformer encoder performs bidirectional collaborative encoding on all modal features, and finally outputs a unified collaborative embedding tensor.

[0007] S2: Based on the aforementioned collaborative embedding tensor, atomic actions, entities, and relationships between entities are detected in parallel, and a dynamic semantic scene graph is constructed and updated in real time. The dynamic semantic scene graph is matched with a predefined threat graph pattern template, and combined with temporal logic rules for reasoning and verification, thereby identifying high-order intrusion behaviors and outputting structured behavioral events containing behavior type, confidence level, threat score, and evidence chain.

[0008] S3: Construct a dynamic strategy engine that receives the structured behavioral events and integrates the current multi-dimensional system context information; based on a preset strategy knowledge base and utility evaluation model, the dynamic strategy engine generates a device linkage plan sequence that maximizes collaborative utility by solving a constrained temporal action sequence optimization problem, distributes and executes the plan, and performs dynamic replanning based on execution feedback and environmental changes.

[0009] S4: Establish a multi-scale memory bank to record contextual data, extracted semantic patterns, and policy rules during system operation; optimize perception, cognition, and decision-making models through offline learning based on memory replay; discover boundary cases and introduce expert knowledge through online active learning mechanisms; and achieve federated evolution of the system under the premise of protecting privacy through cross-scenario knowledge transfer, thereby enabling the system to have the ability to continuously improve itself.

[0010] The technical effects and advantages of this invention are as follows:

[0011] 1. Through the innovative "three-level alignment and collaborative embedding" framework, this invention achieves deep fusion of multimodal data such as video, audio, and radar at the hardware clock, feature space, and semantic levels; effectively overcoming the limitations of a single sensor (such as a camera) under adverse lighting, weather, or occlusion conditions.

[0012] 2. By constructing and calculating "dynamic semantic scene graphs" and combining graph pattern matching and symbolic logic reasoning, it is possible to understand complex intrusion behaviors constituted by "who, when, where, what, and what sequence of actions";

[0013] 3. The Dynamic Strategy Engine (DSGE), based on utility theory and context awareness, changes the rigid response pattern of the traditional "event-fixed action";

[0014] 4. The multi-scale memory and metacognitive optimization mechanism of this invention enables the system to "accumulate experience" and "reflect on strategies" like human experts; the system learns from each action, automatically optimizes the perception model, enriches the behavioral pattern library, corrects linkage strategies, and achieves cross-site knowledge sharing through federated learning while protecting privacy. Attached Figure Description

[0015] Figure 1This is a schematic diagram of the overall structure of the present invention.

[0016] Figure 2 This is a schematic diagram of the multimodal data fusion structure of the present invention.

[0017] Figure 3 This is a schematic diagram of the intrusion behavior recognition structure of the present invention.

[0018] Figure 4 This is a schematic diagram of the linkage structure of the security system of the present invention.

[0019] Figure 5 This is a schematic diagram of the adaptive learning and evolution structure of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] refer to Figures 1-5 The method for security system linkage and intrusion behavior recognition that integrates multimodal data, as shown, includes:

[0022] S1: Hardware-level spatiotemporal synchronization and feature-level adaptive alignment are performed on data from multiple sensors, including vision, hearing, and radar. The aligned multimodal features are then input into a multimodal collaborative embedding and cognitive fusion network. This network first uses visual features as a guide to perform semantic filtering and enhancement on other modal features. Subsequently, a multi-head cross-modal Transformer encoder performs bidirectional collaborative encoding on all modal features, and finally outputs a unified collaborative embedding tensor.

[0023] S2: Based on the aforementioned collaborative embedding tensor, atomic actions, entities, and relationships between entities are detected in parallel, and a dynamic semantic scene graph is constructed and updated in real time. The dynamic semantic scene graph is matched with a predefined threat graph pattern template, and combined with temporal logic rules for reasoning and verification, thereby identifying high-order intrusion behaviors and outputting structured behavioral events containing behavior type, confidence level, threat score, and evidence chain.

[0024] S3: Construct a dynamic strategy engine that receives the structured behavioral events and integrates the current multi-dimensional system context information; based on a preset strategy knowledge base and utility evaluation model, the dynamic strategy engine generates a device linkage plan sequence that maximizes collaborative utility by solving a constrained temporal action sequence optimization problem, distributes and executes the plan, and performs dynamic replanning based on execution feedback and environmental changes.

[0025] S4: Establish a multi-scale memory bank to record contextual data, extracted semantic patterns, and policy rules during system operation; optimize perception, cognition, and decision-making models through offline learning based on memory replay; discover boundary cases and introduce expert knowledge through online active learning mechanisms; and achieve federated evolution of the system under the premise of protecting privacy through cross-scenario knowledge transfer, thereby enabling the system to have the ability to continuously improve itself.

[0026] S1: Fusion of multimodal features: spatiotemporal alignment and collaborative embedding fusion: In order to overcome the heterogeneity of multimodal features at the temporal, spatial and semantic levels, a three-level alignment and collaborative embedding fusion framework is used to transform low-level heterogeneous sensor data into high-level, unified joint semantic representation.

[0027] S101: Level 1: Hardware-level spacetime synchronization

[0028] At the sensing layer, a synchronization module based on the IEEE 1588 Precise Time Protocol (PTP) is used to assign microsecond-level unified timestamps to all sensor data streams (video frames, audio frames, radar scan cycles, and event logs). Simultaneously, through joint calibration technology, a unified world coordinate system with the optical center of the panoramic camera as the origin is established. For each sensor Obtain its correlation through offline calibration rigid transformation matrix between ,in For rotation matrix, It is a translation vector.

[0029] For example, by using a calibration plate or a specific target, thermal imaging pixels are... Millimeter-wave radar point cloud Laser beam alarm line endpoint Precise mapping to this coordinate system enables hardware-level alignment of spatial positions: .

[0030] S102: Level 2: Feature-level Adaptive Spatiotemporal Alignment

[0031] Considering real-world factors such as network transmission latency, sensor internal clock drift, and trigger jitter, hardware synchronization cannot be absolutely perfect. Therefore, in the front end of the converged network, we introduce a lightweight learnable adaptive spatiotemporal alignment module (ASTA).

[0032] Adaptive temporal alignment: for non-visual modality feature sequences For applications such as audio and radar, learn a time offset function. This function uses visual features Dynamically predict modes using the context within the time window as a reference. The fine-tuning offset relative to the visual master clock. The aligned features are: The training objective is to minimize the lower bound of mutual information (InfoNCE loss) between cross-modal features, thereby forcing the aligned features to have maximum consistency at the event level.

[0033] Spatial attention refocusing: For non-spatial modalities (such as audio), by learning a spatial attention map This image, based on visual scene and audio features, predicts the probability distribution of sound sources in the image coordinate system. Then, this attention map is used as weights to spatially contextualize and embed global audio features. .

[0034] S103: Level 3: Semantic-level Collaborative Embedding and Fusion

[0035] The core of the Multimodal Collaborative Embedding and Cognitive Fusion Network (MCECF-Net) is a hierarchical cross-attention fusion module.

[0036] Modal-specific coding:

[0037] Vision: The video stream extracts appearance and motion features through a lightweight spatiotemporal segregated convolutional network (STS-ConvNet). ;

[0038] Audio: After Mel-frequency conversion, time-frequency features are extracted using a one-dimensional residual network. ;

[0039] Radar: 4D point cloud of millimeter-wave radar Target structure and micro-motion features were extracted using a customized PointPillar network. ;

[0040] Switching quantity: The sensor state is encoded as a time-series state vector. .

[0041] Layered cross-attention fusion: This process consists of two steps, forming a funnel-shaped fusion structure of "guidance-coordination":

[0042] Guided alignment phase: with high information density visual features As an "anchor," it provides semantic filtering and enhancement for features from other modalities through a guided attention mechanism. For non-visual modalities... :

[0043]

[0044]

[0045] For example, when the system learns to detect high-frequency micro-movements in a "glass area" in a video, it will automatically give high weight to the "crackling sound" characteristics in a specific frequency band of the audio. and radar point clouds Micro-Doppler diffusion patterns in the corresponding directions are used to generate features guided by visual semantics. , , .

[0046] Bidirectional collaborative embedding phase: and , , A multi-head cross-modal Transformer encoder (CM-Transformer) is used as the common input. At this stage, features from all modalities serve as context within a unified semantic space, enabling fully connected, bidirectional information exchange and complementarity. For a given modality... The query, whose key and value are derived from the concatenation of all modalities:

[0047]

[0048] Ultimately, this module outputs a unified, deeply intertwined cooperative embedding tensor. .

[0049] Engineering integration of the fusion framework: collaborative embedding tensors It serves as a unified starting point for advanced cognitive tasks. To support subsequent parallel tasks such as atomic action detection and entity relation detection, The system is structured along the channel dimension into a shared feature layer and a task-specific feature layer. The shared layer carries general scene semantics across modalities, while the task-specific layer extracts action-sensitive or entity-relationship-sensitive feature components from the fused features using lightweight task-specific attention heads.

[0050] S2: Intrusion Behavior Recognition: Hierarchical Recognition and Reasoning Based on Semantic Scene Graph: In order to leap from low-level perception to high-level behavior understanding, a dynamically evolving semantic scene graph of intrusion behavior is constructed, and a hierarchical recognition and reasoning mechanism based on graph computing and symbolic logic is designed.

[0051] S201: Detection of Atomic Entities and Relationships

[0052] Based on fusion features The system executes two detection tasks in parallel, forming the atomic elements of the scene graph:

[0053] Atomic Action Detection: Using a spatiotemporal Transformer-based decoder, it identifies and spatiotemporally locates sets of atomic actions such as "walking slowly," "running quickly," "squatting," "throwing," "climbing," "prying," "waving a tool," and "leaving an object." ;

[0054] Entity and Relationship Detection: Using a separate decoding head, detect entities (people) in the scene. ,vehicle , wall ,door ,pack And the binary relationships between them, including static spatial relationships and dynamic interactive relationships;

[0055] Static spatial relationships, such as: "close to ( , "On top ( , )”;

[0056] Dynamic interactive relationships, such as: "looking towards ( , "Handheld ( , )".

[0057] S202: Real-time Construction and Evolution of Dynamic Semantic Scene Graphs

[0058] The system maintains a time-slice Animated graph that scrolls in units :

[0059] node This includes entity nodes (with attributes such as type, location, and trajectory history) and action nodes (bound to the initiating entity). Each node... Each has a feature vector that updates over time. .

[0060] side This includes relationship edges between entities and "execution" edges.

[0061] The evolution of the graph is achieved through a temporal graphical neural network (T-GNN):

[0062]

[0063] in, It is a node Neighbors It is a node from arrive The process enables temporal smoothing of node features, cross-frame entity association (Re-ID), and dynamic updating of the graph structure, such as the addition of new nodes and the disappearance of old nodes.

[0064] S203: High-order behavior recognition based on graph pattern matching and logical reasoning

[0065] The system has a built-in scalable hierarchical behavioral knowledge base. High-order complex behaviors are defined as a composite of "graph pattern templates" and "temporal logic rules".

[0066] Graph pattern matching: A predefined set of threat graph pattern templates The system will display real-time images. Subgraph similarity matching is performed with these templates, and a matching score is calculated. Its specific implementation is accomplished through a trainable graph encoder-matcher network. This network processes graphs in real time. and pattern template Encode them as graph-level embedding vectors respectively and Similarity is measured by the cosine distance between vectors:

[0067]

[0068] During training, triplet loss is used to make the embeddings of samples with the same behavior closer together and those with different behaviors farther apart.

[0069] Temporal Logic Reasoning: For high-scoring matching candidate behaviors, linear temporal logic (LTL) rules are used for fine-grained verification and behavioral characterization. These rules are compiled into executable state machines or temporal queries within the system. (Inference engine monitoring dynamic graph) The evolutionary sequence is used to check whether a specific sequence of events satisfies the preorder, postorder, persistence, and negation conditions defined in the rules. For example, for the act of "premeditated vandalism and theft," a rule can be formally defined as:

[0070]

[0071]

[0072]

[0073]

[0074] Threat scoring and evidence chain generation: final output of structured events :

[0075]

[0076] Among them, threat score It is a composite function: Inherent risk level of the behavior (value of the target asset). Chain of evidence The core graph nodes that triggered the behavior, the matching templates, and the reasoning rules were recorded to ensure that the decision is traceable and explainable.

[0077] S3: Security System Linkage: Dynamic Strategy Engine Based on Semantics and Context: To achieve an intelligent and adaptive "perception-decision-action" closed loop, a dynamic strategy engine (DSGE) based on utility theory and context awareness is designed. It transforms abstract behavioral events into specific, collaborative, and optimal security device linkage sequences.

[0078] S301: Multidimensional Context-Aware Model

[0079] The engine maintains a dynamically updated multidimensional context vector. ,include:

[0080] Contextual information: absolute time (weekday / weekend / night), weather (sunny / rainy / foggy), lighting;

[0081] Business context: Building interior status (occupied / unoccupied), presence status of critical assets, number of currently active alerts;

[0082] Resource context: real-time location and status of security personnel, robot battery level and task queue, camera view of busy / idle status, and current open / closed status of access control.

[0083] S302: Strategy Knowledge Base and Candidate Action Set

[0084] The policy library is stored in the form of production rules and utility tables. Each rule is associated with a "condition-candidate action set". Each executable action... Examples such as "Turn the spotlight L1 to 70% brightness", "Send an AR alarm layer to the patrol PAD", and "Dispatch the robot R2 to coordinates (x,y)" all have predefined attribute vectors, including the expected effect. Execution costs Risk coefficient .

[0085] S303: Dynamic Strategy Optimization and Generation

[0086] Candidate strategy retrieval: based on the input behavioral events and context Retrieve all candidate actions that meet the conditions from the strategy library to form the initial set. .

[0087] Multi-objective utility evaluation and collaborative modeling: Designing a multi-objective utility function Each candidate action is quantitatively evaluated. This function is a weighted sum across multiple dimensions. More importantly, to model the synergistic and antagonistic effects between actions, we introduce an action-effect interaction matrix. .plan At any moment Overall synergistic effect The calculation is as follows:

[0088]

[0089] in, Quantified the actions and Simultaneous interaction effects during execution (positive values ​​indicate synergy, negative values ​​indicate antagonism). For interaction weights, This is an indicator function.

[0090] Cooperative strategy synthesis: The engine solves a temporal action sequence optimization problem. It searches for the optimal action sequence within a time window. Action plan within This maximizes overall synergistic effectiveness while simultaneously satisfying resource constraints and temporal logic constraints.

[0091]

[0092]

[0093] in, It is the discount factor, and the solution is obtained using a constrained heuristic look-ahead search algorithm.

[0094] Instruction distribution, closed-loop feedback, and replanning: The optimized plan is decomposed into instructions with precise spatiotemporal parameters and distributed to each execution terminal via a highly reliable message queue. A monitoring-evaluation-replanning loop is established; replanning is triggered not only by contextual mutations but also by deviations between internal predictions and feedback. The system utilizes current behavioral events. Based on the executed plan, short-term predictions of intruder behavior are made. By continuously comparing the predictions with the actual perceived state, if the deviation exceeds a threshold, it is immediately determined that the current strategy may be ineffective, triggering replanning and enabling a rapid response to adaptive intrusion strategies.

[0095] The S4: System Adaptive Learning and Evolution: Multi-scale Memory and Metacognitive Optimization Mechanism: In order to break through the limitations of traditional security systems' static configuration and inability to learn from experience, the "multi-scale memory" and "metacognitive optimization" mechanisms are introduced to enable the system to have the ability to continuously improve itself.

[0096] S401: Construction and Organization of Multi-Scale Memory

[0097] The system establishes a three-tiered, interconnected memory structure:

[0098] Episodic memory: Stores raw data and decision-making processes in the form of event chains. Each entry... For example, "memory snapshot":

[0099] .

[0100] Semantic pattern memory: Extracted from episodic memory through clustering and abstraction. It stores two key patterns: successfully intercepted patterns. Such as "standard procedures for a single person climbing over the southeast corner barbed wire fence at night"; failure / false alarm mode. For example, "visual-radar feature combination of false alarms caused by tree branches swaying in strong winds." Each pattern is a graph pattern template with statistical information.

[0101] Procedural Memory: This is the core knowledge base of the DSGE engine, storing optimized policy rules. Enhanced representation is used:

[0102] .

[0103] S402: Model Optimization Based on Memory Replay and Comparative Learning

[0104] During periods of low load, initiate the "offline reflection" process, replaying memories through a priority experience replay mechanism. Each scenario memory... Each element is assigned a "learning priority" score, determined by the outcome, the uncertainty of the system's decision, and the rarity of the scenario. Reflecting on the learning process prioritizes sampling and replaying high-priority memories, ensuring learning resources focus on "key experiences," achieving three levels of optimization:

[0105] Perception Layer Optimization (MCECF-Net): Difficult samples, such as high-uncertainty events or those ultimately proven to be false positives or false negatives, are sampled from the memory to fine-tune the fusion network through contrastive learning. The loss function encourages the system to better distinguish between cross-modal feature combinations of "real threats" and "security interference."

[0106]

[0107] in, It is a characteristic of the fusion of a certain event. These are characteristics of events of the same category. These are characteristics of different categories or false alarm events.

[0108] Cognitive layer optimization (behavioral knowledge base): automatically learn from successes This involves summarizing new graph pattern templates or revising the matching thresholds of existing templates. The analysis generates "counterexample" patterns, which are used to add exclusionary checks during reasoning and reduce false positives.

[0109] Decision-level optimization (strategy engine): For each strategy rule Perform posterior performance analysis. If a rule is most recently... Average utility in each call If it remains below the threshold, or the success rate A decrease in the rule threshold triggers a rule revision process. Revisions may include relaxing / tightening the conditions. Or adjust the recommendations The combination of actions in the text.

[0110] S403: Boundary Case Discovery and Expert Collaboration Based on Online Active Learning

[0111] The system can recognize its "unknown":

[0112] Uncertainty Quantification and Proactive Inquiry: For perceived and cognitive results, the system outputs not only confidence scores but also cognitive uncertainty estimates. When dealing with an event with persistently high uncertainty, the system marks it as a "boundary case." While executing a conservative default strategy, it proactively extracts key data fragments and generates a structured inquiry request to be submitted to the control center.

[0113] Expert knowledge injection and model update: The logic of experts' interpretation, annotation and handling of "boundary cases" is transformed into structured training samples or rules by the system; through interpretable AI technology, the decision-making basis of experts is transformed into a direct supplement to the perception model or cognitive rule base, realizing precise knowledge injection of humans in the loop.

[0114] S404: Cross-Scenario Knowledge Transfer and Federated Evolution

[0115] To achieve secure collaborative evolution among systems deployed at different locations, a hybrid paradigm of "knowledge distillation + federated learning" is adopted:

[0116] Local knowledge distillation: Each local system distills its learned "experiences" into a lightweight knowledge package that does not contain raw privacy data. It mainly includes: the weight increments of some layers of the updated MCECF-Net, the newly discovered graph pattern templates, the validated policy rules and their metadata;

[0117] Centralized knowledge fusion and distribution: A central server aggregates knowledge from multiple sites. By fusing federated averaging and knowledge distillation techniques, a more general and robust global prior knowledge package is generated. Then Securely distribute to each node;

[0118] Local adaptive fine-tuning: Each node introduces... Then, local data is used for short-term adaptive fine-tuning to better adapt general knowledge to the specific environment of this site.

[0119] S405: Security and Version Control

[0120] While bringing about intelligent improvements, the self-evolution mechanism also establishes a strict evolution sandbox and version rollback mechanism. All new models and rules generated through offline learning must first be evaluated in a "sandbox" environment using a regression test set containing a large number of historical boundary cases to confirm that their performance has not degraded before they can be deployed to the production environment in stages. At the same time, the system retains snapshots of all historical versions of models and rules. If unforeseen negative effects occur after a new version is launched, it can be automatically rolled back to a stable version immediately, ensuring the absolute reliability of core security functions.

[0121] Through this complete "perception-cognition-decision-memory-optimization" closed loop, the system described in this solution achieves a leap from static automation to dynamic intelligence, ultimately becoming an autonomously evolving security expert capable of continuously learning from practice, accumulating experience, optimizing strategies, and adapting to new threats.

[0122] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.

[0123] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for linkage and intrusion behavior recognition in a security system that integrates multimodal data, characterized in that, include: S1: Hardware-level spatiotemporal synchronization and feature-level adaptive alignment are performed on data from multiple sensors, including vision, hearing, and radar. The aligned multimodal features are then input into a multimodal collaborative embedding and cognitive fusion network. This network first uses visual features as a guide to perform semantic filtering and enhancement on other modal features. Subsequently, a multi-head cross-modal Transformer encoder performs bidirectional collaborative encoding on all modal features, and finally outputs a unified collaborative embedding tensor. S2: Based on the aforementioned collaborative embedding tensor, atomic actions, entities, and relationships between entities are detected in parallel, and a dynamic semantic scene graph is constructed and updated in real time. The dynamic semantic scene graph is matched with a predefined threat graph pattern template, and combined with temporal logic rules for reasoning and verification, thereby identifying high-order intrusion behaviors and outputting structured behavioral events containing behavior type, confidence level, threat score, and evidence chain. S3: Construct a dynamic strategy engine that receives the structured behavioral events and integrates the current multi-dimensional system context information; based on a preset strategy knowledge base and utility evaluation model, the dynamic strategy engine generates a device linkage plan sequence that maximizes collaborative utility by solving a constrained temporal action sequence optimization problem, distributes and executes the plan, and performs dynamic replanning based on execution feedback and environmental changes. S4: Establish a multi-scale memory bank to record contextual data, extracted semantic patterns, and policy rules during system operation; optimize perception, cognition, and decision-making models through offline learning based on memory replay; discover boundary cases and introduce expert knowledge through online active learning mechanisms; and achieve federated evolution of the system under the premise of protecting privacy through cross-scenario knowledge transfer, thereby enabling the system to have the ability to continuously improve itself.

2. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The multimodal features include: Video stream data is used to extract its appearance and motion features through a spatiotemporal separation convolutional network. The audio stream data, after Mel-frequency conversion, has its time-frequency features extracted using a one-dimensional residual network; Millimeter-wave radar data, including 4D point clouds of the target. The structure and micro-motion features were extracted using a customized PointPillar network; The data from the digital sensor is encoded into a time-series state vector.

3. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The collaborative embedding tensor includes: A unified deep interleaved feature representation generated through the semantic-level collaborative embedding fusion step; Its tensor structure is , where dimensions T, H, W, and D represent time, spatial height, spatial width, and fusion feature channels, respectively; It internally encodes joint semantic information formed through cross-modal bidirectional information exchange and complementarity, which transcends the independent representation of any single modality; Its channel dimension is structured into a shared feature layer and a task-specific feature layer, serving as a common and distinguishable input for the two parallel tasks of subsequent atomic action detection and entity relationship detection.

4. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The dynamic semantic scene graph includes: A graph structure that evolves over time based on the detection results of the atomic entities and relationships. ; Among them, the node set This includes entity nodes that represent various targets in the scene, as well as atomic action nodes that are bound to the initiating entity; Among them, the edge set This includes relation edges that represent static spatial relationships or dynamic interactive relationships between entities, and "execution" edges that connect entities to the actions they perform. The dynamic semantic scene graph uses a temporal graph neural network to perform temporal smoothing and updating of the features of nodes and edges, thereby realizing the dynamic evolution of the graph and cross-frame entity association.

5. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The threat graph pattern template includes: A standardized graph structure pattern, predefined and stored in a behavioral knowledge base, used to describe the core characteristics of typical intrusion behaviors; Its structure includes specific combinations of node types, definitions of relational edges between nodes, and optional node attribute constraints, used to characterize the semantic and structural essence of threatening behaviors such as "climbing", "prying", and "approaching after loitering". This template serves as a matching benchmark. Through graph encoding and similarity calculation, it performs subgraph matching with the dynamically constructed semantic scene graph in real time to generate preliminary behavioral hypotheses. This template is associated with one or more temporal logic rules, which together constitute the definition of a complete high-order intrusion behavior.

6. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The structured behavioral events include: Behavior type identifies the specific high-level intrusion behavior category that has been identified. Participating entity list, which enumerates all entities directly related to this action; Spatiotemporal scope, defining the start and end times and spatial area where the behavior occurs; Confidence level, characterizing the degree of certainty with which the system identifies the behavior; Threat score is a quantitative threat level calculated by combining factors such as fit, inherent risk level of behavior, and target value. The evidence chain records the key nodes in the real-time semantic scene graph on which the behavior is based, the matching threat graph pattern template, and the triggered temporal logic rules, so as to achieve traceability and explainability of the decision-making process.

7. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The strategy knowledge base and utility evaluation model include: A policy knowledge base that stores the triple rule of "condition-candidate action set-utility attribute", where the condition is based on the behavioral event attribute and the system context, and each candidate action is associated with a predefined expected effect vector, execution cost vector and risk coefficient; A multi-objective utility function Used for candidate actions A quantitative assessment is conducted, and the function is a weighted sum of multiple dimensions, including containment effectiveness, operational costs, evidence collection value, risks, and interference effects. An action-effect interaction matrix This matrix parameter is used to quantify the synergistic or antagonistic effects between different combinations of coordinated actions, and it is used to calculate the overall utility of a coordinated action plan. .

8. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The dynamic replanning includes: a closed-loop decision optimization mechanism based on monitoring and feedback; During the execution of the linkage strategy, the mechanism continuously compares the expected effect of the action with the actual effect of the environmental feedback, and automatically triggers a new round of strategy generation process when the deviation between the two exceeds a preset threshold. The triggering conditions specifically include: the intruder's actual behavioral trajectory deviating from the prediction based on the current strategy, or an unforeseen change occurring in the system's operating context; Once triggered, the dynamic strategy engine will immediately generate and execute a new, optimized collaborative action plan based on the latest perception results and context state, replacing the current plan.

9. The method for linkage and intrusion behavior recognition of a security system integrating multimodal data according to claim 1, characterized in that, The federated evolution includes: a mechanism for collaborative knowledge updates across multiple independently deployed security system nodes; Each local node extracts the experience and knowledge gained from its own operation and learning into a lightweight knowledge package that does not contain original privacy data. The central server aggregates knowledge packages from various nodes and generates a more universal global prior knowledge package through federated averaging and knowledge distillation fusion technology. After the global prior knowledge package is securely distributed to each node, each node then performs adaptive fine-tuning based on its own local data, thereby achieving secure knowledge sharing and collaborative evolution of system performance while protecting data privacy.