A security monitoring situation awareness method and device based on big data, and a medium
By performing spatiotemporal alignment and fusion of multi-source heterogeneous monitoring data, generating entity association graphs and conducting causal analysis, the technical problems in existing security monitoring technologies are solved. This enables a deep understanding of correlations and causal situation assessment that cannot be solved in existing technologies, and achieves a deep understanding and proactive early warning of security monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing security monitoring methods have limitations in deep correlation understanding and causal situation assessment. They are unable to perform in-depth modeling and mining of complex interaction networks between entities across modalities and time and space, resulting in fragmented situation understanding, delayed risk warnings, and insufficient interpretability of decision support.
By receiving multi-source heterogeneous monitoring data and monitoring and positioning data, spatiotemporal alignment and fusion are performed. A multimodal entity detection method is used to generate an entity association graph. A graph embedding operation is performed to generate low-dimensional feature vectors. Association patterns are identified and the temporal sequence between entities is analyzed through a causal discovery algorithm to generate a causal knowledge graph. Dynamic matching and risk inference are performed, and a situational awareness report is output.
It has achieved a deep understanding of the security monitoring situation and proactive early warning, and constructed a closed loop from deep correlation cognition to forward-looking risk intervention, thereby improving the interpretability of risk warning and decision support capabilities.
Smart Images

Figure CN121561358B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of security monitoring technology, and in particular to a security monitoring situational awareness method, device and medium based on big data. Background Technology
[0002] In recent years, security monitoring methods have been evolving from traditional passive video surveillance to proactive intelligent sensing. In the field of security monitoring, existing methods are mainly based on computer vision and IoT approaches. Through structured video analysis, multi-sensor data association, and deep learning models, they detect and identify entities such as people, vehicles, and events. They also utilize big data to aggregate multi-source information and combine it with rule engines or classification models to trigger alarms for abnormal events, forming an intelligent security framework centered on multimodal perception and event detection. Further, the introduction of graph computing for shallow analysis of entity relationships or the use of time-series models to predict specific risks enhances the automation level of security monitoring.
[0003] However, existing methods have limitations in terms of deep correlation understanding and causal situation assessment. Current methods mostly focus on the identification and simple correlation of isolated events, making it difficult to deeply model and explore complex interaction networks between entities across modalities and time and space. This leads to fragmented situation understanding, and the analytical methods lack the ability to evolve from statistical correlation to causal mechanisms. They are unable to analyze the root causes and derivative paths of event chains, resulting in delayed risk warnings and insufficient interpretability of decision support. This restricts the qualitative improvement of security monitoring from post-event tracing to pre-event warning and in-event intervention. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a security monitoring situational awareness method based on big data to solve the limitations in deep correlation understanding and causal situational judgment.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, this invention provides a security monitoring situational awareness method based on big data, comprising: receiving multi-source heterogeneous monitoring data and monitoring location data; unifying the multi-source heterogeneous monitoring data and monitoring location data to the same spatiotemporal dimension to form security fusion data; extracting entity categories from the security fusion data using a multimodal entity detection method, outputting a security entity dataset; performing spatiotemporal logical relationship analysis and entity integration on the security entity dataset to generate an entity association graph; performing graph embedding operations on the entity association graph to generate low-dimensional feature vectors; identifying recurring association patterns in the low-dimensional feature vectors and outputting a set of association patterns; analyzing the temporal order between entities within the set of association patterns using a causal discovery algorithm to generate a preliminary causal hypothesis set; performing statistical independence tests on the preliminary causal hypothesis set and outputting a causal knowledge graph; dynamically matching the multi-source heterogeneous monitoring data with the causal knowledge graph; tracing the root causes and extrapolating derived risks of security monitoring events according to the dynamic matching results; and integrating and outputting a security monitoring situational awareness report.
[0008] As a preferred embodiment of the big data-based security monitoring situational awareness method described in this invention, the specific steps for unifying multi-source heterogeneous monitoring data and monitoring location data to the same spatiotemporal dimension to form security fusion data are as follows:
[0009] Multi-source heterogeneous monitoring data and monitoring and positioning data are aligned in time reference and transformed in spatial coordinates to form spatiotemporally aligned data pairs.
[0010] The spatiotemporally aligned data pairs are standardized in format and correlated and fused to generate security fusion data.
[0011] As a preferred embodiment of the big data-based security monitoring situational awareness method described in this invention, the step of using a multimodal entity detection method to extract entity categories from the security fusion data and outputting a security entity dataset involves the following steps:
[0012] A multimodal entity detection method is used to identify and extract categorized entities from security fusion data, and integrate them to form a preliminary entity set;
[0013] The initial entity set is categorized and its attributes are assigned to generate a category attribute entity set.
[0014] Perform entity parsing and uniqueness determination on the set of category attribute entities, and output a security entity dataset.
[0015] As a preferred embodiment of the big data-based security monitoring situational awareness method of the present invention, the specific steps of performing spatiotemporal logical relationship analysis and entity integration on the security entity dataset to generate an entity association graph are as follows:
[0016] Analyze the spatiotemporal logical relationships between entities in the security entity dataset and output a set of spatiotemporal relationships.
[0017] Perform relationship strength statistics and filtering on the spatiotemporal correlation set, and output the security correlation set;
[0018] The entity association graph is constructed by taking the entities of each category in the security entity dataset as security nodes and the associations in the security association set as security relationship edges.
[0019] As a preferred embodiment of the big data-based security monitoring situational awareness method of the present invention, the specific steps of performing graph embedding operation on the entity association graph to generate a low-dimensional feature vector are as follows:
[0020] Extract the direct adjacent node information and indirect connection path information of each security node in the entity association graph, and integrate them to generate graph structure data;
[0021] The Node2Vec algorithm is used to map graph-structured data to vector representations, and an initial set of vector representations is output.
[0022] The initial vector representation set is scaled and reorganized into a matrix to generate low-dimensional feature vectors.
[0023] As a preferred embodiment of the big data-based security monitoring situational awareness method of the present invention, the specific steps for identifying recurring association patterns in low-dimensional feature vectors and outputting a set of association patterns are as follows:
[0024] Similarity statistics are performed on entities of each category in the low-dimensional feature vector, and pattern clustering is performed on entities of each category based on the similarity statistics results to generate a candidate association pattern set;
[0025] Evaluate the repetition frequency and statistical significance of each pattern in the candidate association pattern set, and then filter and output the association pattern set.
[0026] As a preferred embodiment of the big data-based security monitoring situational awareness method of the present invention, the steps include: analyzing the temporal order of entities within a set of association patterns using a causal discovery algorithm to generate a preliminary causal hypothesis set; performing a statistical independence test on the preliminary causal hypothesis set; and outputting a causal knowledge graph.
[0027] Analyze the temporal order of entities within the statistical association pattern set to generate a temporal relationship set.
[0028] Based on the set of temporal relationships, the causal direction between entities is deduced, and a preliminary set of causal hypotheses is output.
[0029] Perform statistical independence tests on the initial set of causal hypotheses and output the set of tested causal hypotheses;
[0030] The set of causal hypotheses is structured and integrated to output a causal knowledge graph.
[0031] As a preferred embodiment of the big data-based security monitoring situational awareness method described in this invention, the steps include: dynamically matching multi-source heterogeneous monitoring data with a causal knowledge graph, tracing the root causes of security monitoring events and extrapolating derived risks based on the dynamic matching results, and integrating and outputting a security monitoring situational awareness report.
[0032] Multi-source heterogeneous monitoring data is matched with causal knowledge graphs to generate event matching data.
[0033] Traverse the matching events in the event matching data, locate the node position corresponding to the matching event in the causal knowledge graph, and perform reverse causal tracing and forward causal deduction of the matching event in the causal knowledge graph along the node position, and output the root entity set, the derivative risk event set, and the event tracing path set.
[0034] The system integrates the set of root entities, the set of derived risk events, and the set of event tracing paths into a report, and outputs a security monitoring situational awareness report.
[0035] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements any step of the security monitoring situational awareness method based on big data as described in the first aspect of the present invention.
[0036] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the security monitoring situational awareness method based on big data as described in the first aspect of the present invention.
[0037] The beneficial effects of this invention are as follows: By synergistically employing graph embedding pattern mining and causal discovery inference mechanisms, a deep understanding of security monitoring situation and proactive early warning are achieved. Low-dimensional feature vectors are generated by performing graph embedding operations on entity association graphs. Then, pattern clustering of entities within these low-dimensional feature vectors outputs a set of association patterns. Stable and computable potential association patterns are extracted from complex interactive networks, laying a structured knowledge foundation for high-level situation understanding. A causal discovery algorithm analyzes the temporal sequence of entities within the set of association patterns to generate a preliminary set of causal hypotheses. Statistical independence tests on this preliminary set of causal hypotheses output a causal knowledge graph, realizing the evolution from statistical association to causal mechanisms. This constructs interpretable reasoning that characterizes the root causes and derivative chains of events. By dynamically matching and bidirectionally inferring multi-source heterogeneous monitoring data with this causal knowledge graph, a situation awareness report integrating root causes, paths, and derivative risks is output, achieving a closed loop from deep association cognition to proactive risk intervention. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a security monitoring situational awareness method based on big data.
[0040] Figure 2 This is a flowchart for outputting the security entity dataset.
[0041] Figure 3 The flowchart for generating low-dimensional feature vectors.
[0042] Figure 4 This is a flowchart for generating a security monitoring situational awareness report. Detailed Implementation
[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0044] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0045] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0046] Reference Figures 1-4 This is one embodiment of the present invention, which provides a security monitoring situational awareness method based on big data, including the following steps:
[0047] S1. Receive multi-source heterogeneous monitoring data and monitoring location data, and unify the multi-source heterogeneous monitoring data and monitoring location data into the same spatiotemporal dimension to form security fusion data.
[0048] It receives multi-source heterogeneous monitoring data and monitoring location data, and performs time base alignment and spatial coordinate transformation on the multi-source heterogeneous monitoring data and monitoring location data to form a spatiotemporally aligned data pair.
[0049] Specifically, receiving multi-source heterogeneous monitoring data and monitoring location data refers to the real-time or near-real-time collection of raw monitoring information and corresponding spatial location information from different sources and in different formats through various data acquisition interfaces deployed in the security monitoring network. Multi-source heterogeneous monitoring data mainly includes real-time video streams and historical video clips from cameras, images captured by mobile devices, access control card swipe logs, vehicle checkpoint capture records, perimeter protection sensor alarm signals, and status reporting data from other IoT devices. Monitoring location data specifically refers to the geographical location information associated with the multi-source heterogeneous monitoring data, such as the geographical coordinates of cameras, the installation location of access control systems, the latitude and longitude of vehicle capture points, and the physical location coordinates of various sensors. Multi-source heterogeneous monitoring data and monitoring location data are continuously received through network transmission protocols.
[0050] Using a unified time standard (such as UTC time) as a benchmark, each record in the multi-source heterogeneous monitoring data is calibrated based on its timestamp. That is, each record in the multi-source heterogeneous monitoring data is calibrated based on its timestamp, and offset statistics and compensation are performed between the timestamp and the unified time standard (such as UTC time). This converts the timestamps of each data source into standard time values under the same time benchmark, so that the timestamps of the corresponding records in the multi-source heterogeneous monitoring data and the monitoring and positioning data are on the same time axis. Based on the geographical coordinates of the cameras, access control installation points, vehicle capture points, latitude and longitude, and physical location coordinates of various sensors provided in the monitoring and positioning data, the spatial information involved in the multi-source heterogeneous monitoring data is uniformly converted to the same spatial coordinate system (such as WGS-84 or the geodetic coordinate system). This ensures that each piece of multi-source heterogeneous monitoring data and the corresponding monitoring and positioning data are accurately matched in time and space, thereby forming a spatiotemporally aligned data pair.
[0051] The spatiotemporally aligned data pairs are standardized in format and correlated and fused to generate security fusion data.
[0052] Specifically, spatiotemporal aligned data is generated from real-time video streams and historical video clips from cameras, images captured by mobile devices, access control card swipe logs, vehicle checkpoint capture records, perimeter protection sensor alarm signals, and status reports from other IoT devices. Field alignment, encoding format conversion, and semantic tag mapping are performed according to a unified data structure. This unified data structure includes five fields: event time, spatial location, entity type, data source, and content carrier. Event time is represented using the ISO 8601 standard time string, unifying card swipe time, capture time, and alarm reporting time from multi-source heterogeneous monitoring data to ISO 8601 standard. The 8601 standard time string represents spatial location in decimal degrees of WGS-84 latitude and longitude. Entity types include four categories: personnel, vehicles, equipment, and behavioral events. Captured images are classified as personnel, license plate recognition results as vehicles, and sensor boundary crossing signals as behavioral events. Data sources include six categories: cameras, access control systems, checkpoints, sensors, mobile devices, and IoT devices, labeled according to the actual source of the record. The content carrier is filled with video stream address, image storage path, text log, or status code according to the data format, so that data from different sources have the same field names, data types, and value ranges. Based on the timestamp and spatial coordinates aligned in the spatiotemporal alignment data pair for each record, multi-source records belonging to the same spatiotemporal event are associated and bound to form a fused record indexed by time and space, i.e., security fused data.
[0053] S2. Employ multimodal entity detection methods to extract entity categories from security fusion data, output a security entity dataset, perform spatiotemporal logical relationship analysis and entity integration on the security entity dataset, and generate an entity association graph.
[0054] A multimodal entity detection method is used to identify and extract categorized entities from security fusion data, and integrate them to form a preliminary entity set.
[0055] Specifically, the system performs frame-by-frame image analysis on the real-time or historical video content pointed to by the video stream addresses in the security fusion data. Multimodal entity detection methods are used to detect face regions and overall vehicle outlines in the images, recording each detected face region as a person entity and each identified vehicle as a vehicle entity. Face and vehicle detection operations are performed on the static images pointed to by the image storage paths in the security fusion data to generate corresponding person and vehicle entities. The system also scans the text logs in the security fusion data line by line, identifies access control format records, extracts personnel identification and access timestamps to generate person entities, and identifies keyword combinations describing abnormal behavior in the text, such as prolonged stays and unauthorized entry. For each set of valid behavioral descriptions (valid behavioral descriptions refer to a complete description in the text log consisting of clearly listed words and context for prolonged loitering, unauthorized entry, and tailgating, which can be directly used to create and define a specific behavioral event entity), a behavioral event entity is generated. The status codes in the security fusion data are parsed value by value. When the status code indicates the device's operating status, such as online, offline, or fault signal, a device entity is generated. When the status code indicates a perimeter sensor triggered event, such as infrared obstruction, microwave boundary crossing, or vibration alarm, a behavioral event entity is generated. All generated personnel entities, vehicle entities, device entities, and behavioral event entities are aggregated to form a preliminary entity set.
[0056] The initial entity set is labeled with categories and assigned attribute values to generate a category attribute entity set.
[0057] Specifically, for each entity in the initial entity set, if the current entity is a face region detected through video content pointed to by a video stream address, then the current entity is categorized as a person, and its attributes include person identification, face feature vector, time of appearance, and spatial location. If the current entity is a vehicle outline obtained through image recognition pointed to by a video stream address or image storage path, then the current entity is categorized as a vehicle, and its attributes include license plate recognition result, vehicle color, vehicle type, time of appearance, and spatial location. If the current entity is an access control record extracted from a text log, then the current entity is categorized as a person, and its attributes include person identification, access time, and access control installation location, where the spatial location is determined by the access control system. The installation location is determined, and the occurrence time is determined by the passage time. If the current entity is a combination of abnormal behavior keywords identified from the text log, the current entity is categorized as a behavior event, and the current entity is assigned attributes including behavior type, occurrence time, and associated location. If the current entity is a device operating status parsed from a status code, the current entity is categorized as a device, and the current entity is assigned attributes including device type, device number, status value, reporting time, and device physical location coordinates. If the current entity is a sensor trigger signal parsed from a status code, the current entity is categorized as a behavior event, and the current entity is assigned attributes including event type, trigger time, and sensor physical location coordinates. After completing the categorization and attribute assignment of all entities, a category attribute entity set is formed.
[0058] Perform entity parsing and uniqueness determination on the set of category attribute entities, and output a security entity dataset.
[0059] Specifically, the process iterates through each entity in the set of category attribute entities. For entities categorized as "personnel," it determines whether they belong to the same real individual as previously processed personnel entities based on personnel identification or facial feature vectors. If they belong to the same real individual, the attributes are merged, retaining the earliest appearance time and complete attributes. For entities categorized as "vehicles," it determines whether they are duplicates based on license plate recognition results. If the license plates are the same and their appearance time and spatial location are consecutive, they are considered the same vehicle entity and their records are merged. For entities categorized as "equipment," uniqueness is determined based on the equipment number. For entities categorized as "behavioral events," since each trigger is independent, they are not merged. After parsing and deduplicating all entities, the security entity dataset is output.
[0060] Analyze the spatiotemporal logical relationships between entities in the security entity dataset and output a set of spatiotemporal association relationships.
[0061] Specifically, the system iterates through all entity pairs in the security entity dataset. For each entity pair, it compares their occurrence times. If the occurrence times of one entity and another are the same, adjacent, or within the same time interval, then the entity pair is considered to be related in the time dimension. It also compares the spatial locations of each entity pair. If their spatial coordinates are the same or they are located in the same monitoring area (such as a single access control point, checkpoint, or sensor coverage area), then the entity pair is considered to be related in the spatial dimension. When two entities simultaneously satisfy both temporal continuity or overlap and spatial similarity or adjacency, the current entity pair is determined to constitute a spatiotemporal logical relationship. A structured record is generated for the spatiotemporal logical relationship, containing the first entity identifier, the second entity identifier, the time relationship type (such as synchronous and sequential), and the spatial relationship type (such as same point or neighbor). All such structured records are then aggregated to form a set of spatiotemporal relationship records.
[0062] Perform relationship strength statistics and filtering on the spatiotemporal correlation set, and output the security correlation set.
[0063] Specifically, each structured record in the spatiotemporal relationship set is traversed, and the two entity identifiers within it are used as unique keys. The total number of times the current entity pair appears in all structured records is counted, and this total number is used as the relationship strength of the current entity pair. For structured records with a relationship strength of 1 (1 is naturally obtained by counting the total number of times the current entity pair appears in all structured records; when the current entity pair appears only once in all records, the relationship strength is 1), it is checked whether they share the same time interval or spatial region with other structured records. If there is no sharing, they are determined to be isolated relationship records and are removed. All structured records with a relationship strength greater than 1, as well as structured records with a relationship strength of 1 but supported by spatiotemporal context, are retained. The retained structured records are then aggregated to form a security relationship set.
[0064] The entity association graph is constructed by taking the entities of each category in the security entity dataset as security nodes and the associations in the security association set as security relationship edges.
[0065] Specifically, iterate through each entity in the security entity dataset, create a corresponding security node for each entity, with the node identifier using the unique identifier of the current entity, and the node attributes including the category and assigned attributes of the current entity; iterate through each structured record in the security association set, locate the two corresponding security nodes in the created security nodes based on the first and second entity identifiers contained in the record, and establish a security relationship edge between the two security nodes, with the edge attributes including the time relationship type and spatial relationship type in the structured record; after completing the creation of all security nodes and security relationship edges, an entity association graph is formed.
[0066] S3. Perform graph embedding operation on the entity association graph to generate low-dimensional feature vectors, identify recurring association patterns in the low-dimensional feature vectors, and output a set of association patterns.
[0067] Extract the direct adjacent node information and indirect connection path information of each security node in the entity association graph, and integrate them to generate graph structure data.
[0068] Specifically, each security node in the entity association graph is traversed. For the current security node, all other security nodes directly connected by a security relationship edge are collected to form direct adjacent node information. Starting from the current security node, a breadth-first search or depth-first search is performed along the security relationship edge to traverse paths of length 2 (2 is determined according to the definition of path length in graph theory, indicating that an indirect connection path contains at least one intermediate node) to the maximum number of hops, recording all reachable security nodes and path sequences to form indirect connection path information. The direct adjacent node information and indirect connection path information corresponding to each security node are organized into a structured form according to node identifiers. After collecting the information of all security nodes, the graph structure data is integrated to generate the graph structure data.
[0069] The Node2Vec algorithm is used to map graph-structured data to vector representations, outputting an initial set of vector representations.
[0070] Specifically, based on each security node and its connections in the graph structure data, the Node2Vec algorithm is used to perform a biased random walk on the entity association graph. Starting from the current security node, at each step, when selecting the next security node, nodes that can form a local closure or expand outwards from the directly connected neighbors of the current node are chosen, thus generating a node sequence that balances depth-first and breadth-first search characteristics. When generating the node sequence, neighboring nodes that are structurally related to the current security node are retained, while also considering the exploration of local neighborhoods and global structures. Based on the contextual co-occurrence relationships of each security node in all node sequences, a numerical vector is generated for each security node. This is determined by the frequency with which each security node co-occurs with other security nodes within the preceding and following time windows across all node sequences, ensuring that security nodes with higher co-occurrence frequencies have larger inner product values in the vector space. The numerical vectors corresponding to each security node are then aggregated to form an initial vector representation set.
[0071] The initial vector representation set is scaled and reorganized into a matrix to generate low-dimensional feature vectors.
[0072] Specifically, the numerical vectors corresponding to each security node in the initial vector representation set are arranged sequentially according to the unique identifiers of the security nodes in the entity association graph, forming a vector matrix. Each row corresponds to a security node, and each column corresponds to a dimension of the initial vector representation set. Each column of the vector matrix is linearly scaled so that all elements in the current column are mapped to a uniform numerical range (e.g., 0 to 1), completing the scaling adjustment. Multiple adjacent dimensions of the vector matrix are merged into one dimension by averaging the values of each row of the dimension, thereby reducing the total number of dimensions. Each row of the new matrix obtained after merging is the low-dimensional feature vector of the corresponding security node, and all rows together constitute the low-dimensional feature vector.
[0073] Similarity statistics are performed on entities of each category in the low-dimensional feature vector, and pattern clustering is performed on entities of each category based on the similarity statistics results to generate a candidate association pattern set.
[0074] Specifically, the process iterates through the low-dimensional feature vectors corresponding to all security nodes, organizing the low-dimensional feature vectors of four types of entities—personnel, vehicles, equipment, and behavioral events—by category. For each category, the cosine similarity of all low-dimensional feature vectors is calculated pairwise to obtain the similarity value of all entity pairs within that category. Entity pairs with similarity values greater than or equal to a set similarity clustering threshold are grouped into the same group, and all entity pairs with similarity values greater than or equal to the set similarity clustering threshold are considered connected, forming multiple entity clusters. Each entity cluster represents a recurring structural pattern, and all entity clusters are considered as candidate association patterns, which are then aggregated to form a candidate association pattern set.
[0075] Furthermore, the similarity clustering threshold is set based on the inflection point of the similarity distribution of low-dimensional feature vectors in each category of entities. By analyzing the cumulative distribution curve of cosine similarity, the similarity value corresponding to the maximum curvature is selected as the similarity clustering threshold; specifically, the value is 0.75. A value of 0.75 ensures that entities with the same structural pattern are grouped into the same cluster, while avoiding the erroneous merging of entities with different patterns due to accidental similarity. If the similarity clustering threshold is greater than 0.75, the clustering standard will be too strict, and entities that should belong to the same structural pattern will be separated due to slight differences, resulting in pattern fragmentation. If the similarity clustering threshold is less than 0.75, the clustering standard will be too lenient, and entities with different structural patterns may be erroneously merged, leading to pattern mixing and distortion. The similarity clustering threshold of 0.75 is a balance point selected based on the inflection point of the distribution curve, aiming to achieve the optimal trade-off between intra-pattern aggregation and inter-pattern distinguishability.
[0076] The formula for calculating the similarity value is:
[0077] ;
[0078] in, This represents the similarity value between entity pairs in a low-dimensional feature vector. The dimension index representing the dimension of the low-dimensional feature vector. This represents the number of dimensions of a low-dimensional feature vector. Representing low-dimensional feature vectors In the The values of each dimension Representing low-dimensional feature vectors In the The numerical values of each dimension.
[0079] Evaluate the repetition frequency and statistical significance of each pattern in the candidate association pattern set, and then filter and output the association pattern set.
[0080] Specifically, each candidate association pattern in the candidate association pattern set is traversed, and the number of security nodes covered by the current candidate association pattern in the low-dimensional feature vector is counted as the repetition frequency of the current candidate association pattern. Based on the distribution of all security nodes in the low-dimensional feature vector, a reference pattern set is generated multiple times by randomly shuffling the entity category labels, and the proportion of each repetition frequency in the reference pattern set is counted as the statistical significance index of the current candidate association pattern. Candidate association patterns with a repetition frequency greater than the median of the corresponding frequency in the reference pattern set and a statistical significance index less than the median of the corresponding index in the reference distribution are retained. The retained candidate association patterns are aggregated to form an association pattern set.
[0081] S4. Analyze the temporal order between entities in the association pattern set through the causal discovery algorithm, generate a preliminary causal hypothesis set, perform statistical independence tests on the preliminary causal hypothesis set, and output a causal knowledge graph.
[0082] The temporal order of entities within the statistical association pattern set is used to generate a temporal relationship set.
[0083] Specifically, each association pattern in the association pattern set is traversed, and all security nodes contained in the current association pattern are extracted. Based on the occurrence time of each security node recorded in the security entity dataset, the security nodes are sorted in ascending order by time value. For two adjacent security nodes after sorting, if the occurrence time of the former security node is earlier than that of the latter security node, a time-series relationship record from the former security node to the latter security node is generated. The time-series relationship records generated in all association patterns are collected to form a time-series relationship set.
[0084] Based on the set of temporal relationships, the causal direction between entities is deduced, and a preliminary set of causal hypotheses is output.
[0085] Specifically, each temporal relationship record in the temporal relationship set is traversed, and the temporal relationship from the previous security node to the next security node in the record is regarded as a potential causal direction, that is, the previous security node is the cause entity and the next security node is the result entity. Combining the context of the association pattern to which the two security nodes belong in the association pattern set, it is confirmed that the cause entity and the result entity have structural association support, that is, check whether the cause entity and the result entity co-occur in at least one of the same association patterns. If there is at least one association pattern that contains both the cause entity and the result entity, it is considered that there is a stable structural co-occurrence relationship in the entity association graph, which constitutes association support. If the cause entity and the result entity never appear in the same association pattern in all association patterns, it is considered that there is no structural co-occurrence basis between the cause entity and the result entity, and no association support is constituted. If association support exists, a causal hypothesis record containing the cause entity, the result entity and the corresponding association pattern identifier is generated. If there is no association support, no causal hypothesis record is generated. All generated causal hypothesis records are collected to form a preliminary causal hypothesis set.
[0086] Perform statistical independence tests on the initial set of causal hypotheses and output the set of tested causal hypotheses.
[0087] Specifically, the process iterates through each causal hypothesis record in the initial causal hypothesis set, extracting the causal entity and the result entity. Based on all historical observation instances in the security entity dataset, the number of times the causal entity and result entity appear simultaneously, only the causal entity appears, only the result entity appears, and neither the causal entity nor the result entity appears is counted, forming four types of observation counts. The mutual information value is calculated based on these four types of observation counts. The mutual information value is determined by comparing the difference between the actual frequency of the causal entity and the result entity co-occurring and the frequency of their co-occurrence when they are independent. The greater the difference, the higher the mutual information value. If the mutual information value is greater than zero, the causal entity and the result entity are considered to have a dependency relationship, and the current causal hypothesis record is retained. If the mutual information value is equal to zero, the causal entity and the result entity are considered to be independent, and the current causal hypothesis record is removed. All retained causal hypothesis records are then aggregated to form a set for testing causal hypotheses.
[0088] The set of causal hypotheses is structured and integrated to output a causal knowledge graph.
[0089] Specifically, the process iterates through each causal hypothesis record in the set of causal hypotheses, treating the causal entities and result entities as causal nodes, and the causal relationships between them as causal directed edges. Each causal node is assigned its category and attribute information from the security entity dataset, and each causal directed edge is labeled with its corresponding association pattern identifier and causal direction. All causal nodes and causal directed edges are organized according to a unified graph data format. This unified graph data format uses triples (head node, relation type, and tail node) to represent each causal directed edge, where the head and tail nodes are uniquely identified causal nodes. Each causal node records its category and all attribute information from the security entity dataset in a key-value pair structure, and each causal directed edge records its association pattern identifier and causal direction in a key-value pair structure. All causal nodes are stored in a node list, and all causal directed edges are stored in an edge list. The node list and edge list together constitute a structured representation of the causal knowledge graph, forming a causal knowledge graph with node attributes, edge types, and causal semantics.
[0090] S5. Dynamically match multi-source heterogeneous monitoring data with causal knowledge graphs, trace the root causes of security monitoring events and extrapolate derivative risks based on the dynamic matching results, and integrate and output a security monitoring situational awareness report.
[0091] Multi-source heterogeneous monitoring data is matched with causal knowledge graphs to generate event matching data.
[0092] Specifically, entity information and timestamps are extracted from real-time multi-source heterogeneous monitoring data using data query methods. The extracted entity information is compared with the categories, attributes, and time ranges of all causal nodes in the causal knowledge graph to determine if there are any causal nodes that meet the spatiotemporal consistency criteria. Spatiotemporal consistency means that the entity information and timestamps extracted from the real-time monitoring data match the entity category, attributes, and associated time and spatial ranges represented by the causal nodes in the causal knowledge graph. Meeting spatiotemporal consistency means that the entity extracted from the real-time monitoring data falls within the corresponding category, attribute, time range, and spatial range defined by the corresponding causal node in the causal knowledge graph in terms of category, attribute, timestamp, and spatial location. Not meeting spatiotemporal consistency means that the entity extracted from the real-time monitoring data does not fall within the corresponding category, attribute, time range, and spatial range defined by the corresponding causal node in the causal knowledge graph in terms of category, attribute, timestamp, and spatial location. If causal nodes that meet the criteria exist, a matching relationship is established between the current monitoring record and the corresponding causal node, and the causal node identifier, monitoring record content, and matching time are recorded. All real-time established matching relationships are aggregated in chronological order to form event matching data.
[0093] Traverse the matching events in the event matching data, locate the node position corresponding to the matching event in the causal knowledge graph, and perform reverse causal tracing and forward causal deduction of the matching event in the causal knowledge graph along the node position, outputting the root entity set, the derivative risk event set, and the event tracing path set.
[0094] Specifically, each matching event in the event matching data is read sequentially. Based on the causal node identifier recorded in the matching event, the corresponding causal node is precisely located in the causal knowledge graph and used as the starting node for the current analysis. Starting from the starting node, a reverse causal tracing operation is performed, that is, traversing all reachable upstream causal nodes layer by layer in the opposite direction of the causal directed edge until the maximum tracing depth (the maximum tracing depth is an integer parameter, ranging from 1 to 5 layers, and the value is determined based on the average length of the causal chain of typical events in security scenarios; in practical applications, if not explicitly specified, the default maximum tracing depth is 3 layers to balance tracing completeness and computational efficiency). Root cause nodes belonging to the initial triggering factor among all upstream causal nodes traversed during the tracing process are classified as root causes. The source entity set records each complete node sequence from the root cause node to the starting node, forming the component of the event tracing path set. A forward causal deduction operation is performed on the starting node, which involves traversing all reachable downstream causal nodes layer by layer along the positive direction of the causal directed edge until no subsequent causal nodes remain. Downstream causal nodes with risk attributes or marked as abnormal categories identified during the deduction process are added to the derived risk event set, and the complete path from the starting node to each derived risk event node is added to the event tracing path set. For each matching event, reverse causal tracing and forward causal deduction are repeated, summarizing the root entities, derived risk events, and corresponding paths generated by all matching events, forming the root entity set, the derived risk event set, and the event tracing path set, respectively.
[0095] The system integrates the set of root entities, the set of derived risk events, and the set of event tracing paths into a report, and outputs a security monitoring situational awareness report.
[0096] Specifically, for each root entity in the root entity set, the category, attributes, and occurrence time information in the security entity dataset are extracted; for each derivative risk event in the derivative risk event set, the category, risk level, occurrence time, and associated causal node identifiers are extracted; for each event tracing path in the event tracing path set, the causal nodes and causal relationship directions contained in the path are organized in chronological order to form a complete causal chain description from the root entity to the derivative risk event; the information extracted from the root entity set, derivative risk event set, and event tracing path set is arranged according to a four-part structure: event summary—root source analysis—risk warning—causal path. The event summary summarizes the basic situation of the currently matched event, the root source analysis lists all root entities and their characteristics, the risk warning presents derivative risk events and their potential impacts, and the causal path displays the corresponding event tracing path one by one; the complete causal chain description is output in text or document form to form a security monitoring situational awareness report.
[0097] This embodiment also provides a computer device applicable to the security monitoring situational awareness method based on big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the security monitoring situational awareness method based on big data proposed in the above embodiment.
[0098] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0099] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the big data-based security monitoring situational awareness method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0100] In summary, this invention achieves deep understanding and proactive early warning of security monitoring situations through the synergistic effect of graph embedding pattern mining and causal discovery inference. By performing graph embedding operations on the entity association graph to generate low-dimensional feature vectors, and then performing pattern clustering on the entities of each category within these low-dimensional feature vectors to output a set of association patterns, stable and computable potential association rules are extracted from complex interactive networks, laying a structured knowledge foundation for high-level situation understanding. The causal discovery algorithm analyzes the temporal sequence of entities within the set of association patterns to generate a preliminary set of causal hypotheses, and then performs statistical independence tests on these preliminary causal hypotheses to output a causal knowledge graph. This realizes the evolution from statistical association to causal mechanisms, constructing interpretable reasoning that characterizes the root causes and derivative chains of events. By dynamically matching and bidirectionally inferring multi-source heterogeneous monitoring data with this causal knowledge graph, a situation awareness report integrating root causes, paths, and derivative risks is output, achieving a closed loop from deep association cognition to proactive risk intervention.
[0101] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A big data based security monitoring situation awareness method, characterized in that: include, It receives multi-source heterogeneous monitoring data and monitoring location data, unifies them into the same spatiotemporal dimension, and forms security fusion data. The multi-source heterogeneous monitoring data includes real-time video streams and historical video clips, images captured by mobile devices, access control card swipe logs, vehicle checkpoint capture records, perimeter protection sensor alarm signals, and status reporting data from other IoT devices. The monitoring location data includes the geographic coordinates of cameras, the installation locations of access control systems, the latitude and longitude of vehicle capture points, and the physical location coordinates of various sensors. A multimodal entity detection method is used to extract entity categories from security fusion data, outputting a security entity dataset. Spatiotemporal logical relationship analysis and entity integration are performed on the security entity dataset to generate an entity association graph. Perform graph embedding operations on the entity association graph to generate low-dimensional feature vectors, identify recurring association patterns in the low-dimensional feature vectors, and output a set of association patterns. By analyzing the temporal order of entities within the set of association patterns using a causal discovery algorithm, a preliminary set of causal hypotheses is generated. A statistical independence test is then performed on the preliminary set of causal hypotheses, and a causal knowledge graph is output. Multi-source heterogeneous monitoring data is dynamically matched with causal knowledge graphs. Based on the dynamic matching results, the root causes of security monitoring events are traced and derivative risks are inferred. The results are then integrated and output as a security monitoring situational awareness report. The specific steps for performing spatiotemporal logical relationship analysis and entity integration on the security entity dataset to generate an entity association graph are as follows: Analyze the spatiotemporal logical relationships between entities in the security entity dataset and output a set of spatiotemporal relationships. Perform relationship strength statistics and filtering on the spatiotemporal correlation set, and output the security correlation set; The entity association graph is constructed by taking the entities of each category in the security entity dataset as security nodes and the associations in the security association set as security relationship edges. The process involves dynamically matching multi-source heterogeneous monitoring data with a causal knowledge graph, tracing the root causes of security monitoring events and extrapolating derived risks based on the dynamic matching results, and then integrating and outputting a security monitoring situational awareness report. The specific steps are as follows: Multi-source heterogeneous monitoring data is matched with causal knowledge graphs to generate event matching data. Traverse the matching events in the event matching data, locate the node position corresponding to the matching event in the causal knowledge graph, and perform reverse causal tracing and forward causal deduction of the matching event in the causal knowledge graph along the node position, and output the root entity set, the derivative risk event set, and the event tracing path set. The system integrates the set of root entities, the set of derived risk events, and the set of event tracing paths into a report, and outputs a security monitoring situational awareness report.
2. The big data based security surveillance situation awareness method of claim 1, wherein: The specific steps for unifying multi-source heterogeneous monitoring data and monitoring location data to the same spatiotemporal dimension to form security fusion data are as follows: Multi-source heterogeneous monitoring data and monitoring and positioning data are aligned in time reference and transformed in spatial coordinates to form spatiotemporally aligned data pairs. The spatiotemporally aligned data pairs are standardized in format and correlated and fused to generate security fusion data.
3. The big data based security surveillance situation awareness method of claim 1, wherein: The multi-modal entity detection method is used to extract entity categories from the security and protection fusion data, and output security and protection entity data sets, and the specific steps are, The multi-modal entity detection method is used to identify and extract category entities from the security and protection fusion data, and to integrate to form a preliminary entity set; The preliminary entity set is labeled and attributed, and a category attribute entity set is generated; The category attribute entity set is parsed and uniquely determined, and the security and protection entity data set is output.
4. The big data based security surveillance situation awareness method of claim 1, wherein: The graph embedding operation is performed on the entity association graph to generate a low-dimensional feature vector, and the specific steps are, Extract the direct adjacent node information and indirect connection path information of each security node in the entity association graph, and integrate to generate graph structure data; The graph structure data is mapped to a vector representation by the Node2Vec algorithm, and an initial vector representation set is output; The initial vector representation set is scaled and matrixed to generate a low-dimensional feature vector.
5. The big data based security surveillance situation awareness method as claimed in claim 1, wherein: The repeated association patterns in the low-dimensional feature vector are identified, and an association pattern set is output, and the specific steps are, The similarity of each category entity in the low-dimensional feature vector is counted, and the mode clustering of each category entity is performed according to the similarity statistical result, and a candidate association pattern set is generated; The repetition frequency and statistical significance of each pattern in the candidate association pattern set are evaluated, and the association pattern set is output.
6. The big data based security surveillance situation awareness method of claim 1, wherein: The time sequence between entities in the association pattern set is analyzed by the causal discovery algorithm to generate a preliminary causal hypothesis set, and the statistical independence of the preliminary causal hypothesis set is tested to output a causal knowledge graph, and the specific steps are, The time sequence between entities in the association pattern set is counted to generate a time sequence relationship set; The causal direction between entities is derived according to the time sequence relationship set, and a preliminary causal hypothesis set is integrated and output; The statistical independence of the preliminary causal hypothesis set is tested, and a test causal hypothesis set is output; The test causal hypothesis set is structured and integrated to output a causal knowledge graph. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that: The processor executes the computer program to realize the steps of the security and protection monitoring situation awareness method based on big data according to any one of claims 1-6.
8. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to realize the steps of the security and protection monitoring situation awareness method based on big data according to any one of claims 1-6.
Citation Information
Patent Citations
Knowledge graph construction method and device for network security situation awareness
CN120196790A
Safety control method based on knowledge graph
CN120671683A