Metadata-based abnormal data intelligent monitoring method and system
Through intelligent monitoring methods based on metadata, node blood ties graphs are constructed and abnormal detection is performed in combination with graph attention network and GraphSAGE model, which solves the detection accuracy and efficiency problems of the existing technology when processing complex and large-scale data, and achieves efficient and accurate abnormal detection and scalability.
Patent Information
- Application Number
- CN202510621756.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing anomaly detection technology is difficult to capture the potential relationships and structural characteristics between data when processing complex and large-scale data, resulting in reduced detection accuracy and efficiency, lack of flexibility and scalability, and cannot handle dynamic data changes in real time.
Using an intelligent monitoring method based on metadata, node tables and edge tables are extracted by collecting data logs, nested list structures are constructed for nested compression storage, and node blood relationship diagrams are constructed using SQL parser and RelMetadataQuery API, mixed anomaly detection model is constructed in combination with graph attention network and GraphSAGE model, and exception propagation chains are constructed based on abnormal nodes.
It improves data storage efficiency, reduces computing resource consumption, enhances the graph representation ability, provides richer context information for abnormal detection, and significantly improves the accuracy and scalability of abnormal detection.
Smart Images

Figure CN120123960A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data detection, and particularly to an intelligent monitoring method and system for abnormal data based on metadata. Background Art
[0002] With the rapid development of big data and Internet of Things (IoT) technologies, the scale of data generation, storage, and processing has shown an explosive growth. In this context, how to effectively manage and monitor anomalies in large-scale data has become an important topic in the fields of data analysis and intelligent monitoring. Existing anomaly detection technologies mainly rely on traditional statistical methods and machine learning algorithms, usually requiring a large amount of labeled data or pre-set rules for analysis. This makes these methods face many challenges when dealing with complex and large-scale data. For example, when facing complex data structures, traditional methods may be difficult to capture the potential relationships and structural features between data, resulting in a significant reduction in the accuracy and efficiency of anomaly detection. In addition, with the increase in data scale and complexity, traditional anomaly detection methods often lack flexibility and scalability and cannot process and respond to a large number of dynamic data changes in real time.
[0003] In recent years, methods based on graph neural networks (GNNs) have begun to make some progress in the field of anomaly detection. Graph neural networks can effectively capture the complex relationships between nodes and perform feature extraction in large-scale data through mechanisms such as graph convolution and graph attention. However, current graph neural network technologies still have certain limitations when facing large-scale, high-dimensional, and complex data. Especially when dealing with multi-level and multi-dimensional data structures, how to efficiently perform data storage, graph structure construction, and anomaly detection has become a technical problem to be solved urgently. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an intelligent monitoring method and system for abnormal data based on metadata, which solves the problem of how to efficiently perform data storage, graph structure construction, and anomaly detection.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides an intelligent monitoring method for abnormal data based on metadata, which includes: Collecting data logs, extracting a node table and an edge table from the logs, constructing an in-edge aggregation subgraph based on each node in the node table, and constructing a nested list structure for nested compression storage of each node; Generate the original RelNode tree based on nodes through the SQL parser and perform logical optimization. Use the RelMetadataQuery API to extract the node lineage from the optimized RelNode tree to construct the node lineage graph; Adopt a graph attention network and a GraphSAGE model to construct a hybrid anomaly detection model to detect anomalies in the node lineage graph, and construct an anomaly propagation chain based on the abnormal nodes.
[0007] As a preferred solution of the metadata-based intelligent monitoring method for abnormal data described in the present invention, wherein: extracting the node table and the edge table from the log, constructing an incoming edge aggregation subgraph based on each node in the node table, and constructing a nested list structure to perform nested compression storage on each node refers to extracting the identification name and type of each type of data from the collected data event log to construct nodes and form a node table, assigning an index value to each node in the node table, and recording the node, timestamp, and operation type when the event occurs in the event log as edges to form the original edge table; For each node v in the node table, record the relationship between each edge and the node v to construct a mapping table, and connect the edge with the node v to construct an incoming edge aggregation subgraph. Extract all the nodes connected to the node v in the incoming edge aggregation subgraph as adjacent nodes to form an adjacent node set, and synchronously record the earliest timestamp among all the edges connected to the node v. Construct a nested list structure for all the connected edges of each node ; The nested list structure , the node index, and the adjacent node set of the node v constitute the final storage node format :
[0008] Calculate the storage space of the original edge table ; Synchronously calculate the storage space after constructing the nested list structure ; ; Calculate the ratio of the number of nodes after constructing the nested list structure to the total number of connected edges between nodes as the average number of connected edges, and compare and . When the average number of connected edges is greater than or equal to the set threshold and is greater than , it is judged that the compression is effective, and the nodes are stored according to the final storage node format .
[0009] As a preferred solution of the intelligent monitoring method for abnormal data based on metadata according to the present invention, wherein: generating an original RelNode tree based on nodes through an SQL parser and performing logical optimization, and extracting the node lineage relationship from the optimized RelNode tree using the RelMetadataQueryAPI to construct a node lineage graph means parsing an SQL query through the SQL parser to generate an abstract syntax tree AST; Generating an original RelNode tree from the AST, each tree node including an operation type, a source node, and a target node, and performing logical optimization on the original RelNode tree, including inferring the field source, merging adjacent operations, and merging grouping operations; Using the RelMetadataQuery API method to extract the data source nodes and flow nodes of the operations between the nodes in the optimized RelNode tree, defining the data source nodes as source nodes, the flow nodes as target nodes, and adding a directed lineage marker to the connection edge between the source nodes and the target nodes, and traversing all nodes to gradually construct a node lineage graph.
[0010] As a preferred solution of the intelligent monitoring method for abnormal data based on metadata according to the present invention, wherein: using a graph attention network and a GraphSAGE model to construct a hybrid anomaly detection model to perform anomaly detection on the node lineage graph means, based on the constructed node lineage graph, concatenating the identification names and types of the data in the nodes to form node vectors, and using the graph attention network to calculate the attention coefficients between the nodes And performing a normalization operation on the attention coefficients; Updating the node vectors according to the normalized attention coefficients to obtain node attention feature vectors ; Using the GraphSAGE model to perform neighbor aggregation on the nodes in the node lineage graph; Updating the node vectors through neighbor aggregation to obtain node aggregation feature vectors ; Combining the graph attention network and the GraphSAGE model to form a hybrid anomaly detection model, including an input layer, an embedding layer, an aggregation layer, and an output layer; Optimizing the hybrid anomaly detection model through a cross-entropy loss function; Inputting the node lineage graph into the trained hybrid anomaly detection model, obtaining the node anomaly probabilities in the node lineage graph through the output layer, and comparing the node anomaly probabilities with an anomaly threshold to determine the node anomaly situation.
[0011] As a preferred solution of the intelligent monitoring method for abnormal data based on metadata according to the present invention, wherein: constructing an abnormal propagation chain based on abnormal nodes means extracting adjacent nodes from the node blood relationship graph according to the abnormal nodes, defining the propagation chain direction according to the directed blood relationship marks of the node connection edges, connecting the nodes according to the propagation chain direction to form an abnormal propagation chain, and traversing all adjacent nodes to form a set of abnormal propagation chains, and calculating an abnormal score for each abnormal propagation chain in the set of abnormal propagation chains. ; Select the abnormal propagation chain with the highest abnormal score from the set of abnormal propagation chains as the main propagation chain.
[0012] As a preferred solution of the intelligent monitoring method for abnormal data based on metadata according to the present invention, wherein: collecting data logs means using a system audit framework to collect various data event logs in real time and perform preprocessing.
[0013] As a preferred solution of the intelligent monitoring method for abnormal data based on metadata according to the present invention, wherein: after collecting the data logs, store the data logs in a database, and synchronously store the node blood relationship graph and the abnormal propagation chain. The database is classified and stored according to the timestamp and the node data type and is regularly backed up.
[0014] In a second aspect, the present invention provides an intelligent monitoring system for abnormal data based on metadata, including A data collection and compression module, configured to collect data logs, extract a node table and an edge table from the data logs, and construct a nested list structure to perform nested compression storage on each node; A node blood relationship extraction module, configured to generate an original RelNode tree based on nodes through an SQL parser and perform logical optimization, and extract node blood relationship from the optimized RelNode tree using the RelMetadataQuery API to construct a node blood relationship graph; An abnormal detection module, configured to use a graph attention network and a GraphSAGE model to construct a hybrid abnormal detection model to perform abnormal detection on the node blood relationship graph, and construct an abnormal propagation chain based on abnormal nodes.
[0015] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the intelligent monitoring method for abnormal data based on metadata as described in the first aspect of the present invention is implemented.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, any step of the intelligent monitoring method for abnormal data based on metadata as described in the first aspect of the present invention is implemented.
[0017] The beneficial effects of the present invention are as follows: By collecting data logs, extracting node tables and edge tables, and constructing a nested list structure to perform nested compression storage on each node, the present invention effectively improves the data storage efficiency and reduces the consumption of computing resources. By using the RelMetadataQuery API to extract node lineage from the optimized RelNode tree and constructing a node lineage graph, the representation ability of the graph is improved, providing richer context information for anomaly detection. At the same time, by combining the Graph Attention Network (GAT) and the GraphSAGE model to perform anomaly detection on the node lineage graph and construct an anomaly propagation chain, the accuracy and scalability of anomaly detection are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of the intelligent monitoring method for abnormal data based on metadata in Embodiment 1; Figure 2 It is a structural diagram of the intelligent monitoring system for abnormal data based on metadata in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0021] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0022] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0023] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an intelligent monitoring method for abnormal data based on metadata, including the following steps: S1. Collect data logs, extract the node table and edge table from the logs, construct an incoming edge aggregation subgraph based on each node in the node table, and construct a nested list structure for nested compression storage of each node. Specifically, collecting data logs means using the system audit framework to collect various data event logs in real time and perform preprocessing. The data sources include but are not limited to process operations, file interactions, network communications, etc.
[0024] Furthermore, extracting the node table and edge table from the logs, constructing an incoming edge aggregation subgraph based on each node in the node table, and constructing a nested list structure for nested compression storage of each node means extracting the identification names and types of each type of data from the collected data event logs to construct nodes and form a node table, assigning an index value to each node in the node table, and recording the nodes, timestamps, and operation types of events occurring in the event logs as edges to form an original edge table. For each node v in the node table, record the relationship between each edge and node v to construct a mapping table, connect the edge with node v to construct an incoming edge aggregation subgraph, extract all the nodes connected to node v in the incoming edge aggregation subgraph as adjacent nodes to form an adjacent node set, and synchronously record the earliest timestamp among all the edges connected to node v. Construct a nested list structure for all the connected edges of each node :
[0025] Where is the index value of the node in the node table, is the relative time offset, obtained by the difference between the event timestamp in the edge and the earliest timestamp, is the operation type; Combine the nested list structure , the node index, and the adjacent node set of node v to form the final storage node format :
[0026] Calculate the storage space of the original edge table :
[0027] Where m is the total number of nodes, is the number of bytes of the identification of node i, is the number of bytes of the identification of the adjacent nodes of node i, obtained by summing the number of bytes of the identification of the adjacent nodes, is the number of bytes of the timestamp of the edge connected to the node, is the number of bytes of the operation type of the edge connected to the node, obtained by summing the number of bytes of the timestamp of the edge connected to the node and the number of bytes of the operation type; Synchronous computing constructs a nested list structure The storage space after :
[0028] Wherein is the number of connected nodes of node i, is the number of connected edges of node i, is the total number of operation types in the connected edges of node i, is the relative time offset of the connected edges of node i under operation o; Calculate the ratio of the number of nodes to the total number of connected edges between nodes after constructing the nested list structure as the average number of connected edges, and and are compared. When the average number of connected edges is greater than or equal to the set threshold and is greater than it is judged that the compression is effective, and the nodes are stored according to the final storage node format Store the nodes
[0029] By extracting the node table and edge table from the data log and assigning an index value to each node, it is possible to ensure clear organization of large-scale data and facilitate subsequent data analysis and processing. The structured storage of the node table and edge table enables the system to efficiently process a large amount of dynamic event data and provides strong data support for subsequent anomaly detection. Especially in the big data environment, through this efficient storage method, large-scale data can reduce storage overhead and calculation time while ensuring accuracy. By aggregating all the incoming edge information (timestamp, operation type, etc.) related to node v, the model can analyze the behavior of the node in a larger context. This aggregation method improves the global perception ability of the model. Especially when dealing with complex relationships between nodes, it can more effectively discover potential abnormal behaviors. The design of the nested list structure not only solves the problem of redundant node and edge information, but also significantly reduces the memory overhead of the system through compressed storage. This structure is particularly suitable for storing a large number of graph data with complex connection relationships, while improving storage efficiency and maintaining the flexibility of data access. By compressing and storing the relevant events and their timing information of each node, the memory occupancy is reduced, and the data can be quickly extracted for processing when needed. When calculating the original edge table and the compressed storage space, by comparing the storage occupancy of the two, the data compression effect can be intuitively understood. By calculating the average number of connected edges and its threshold judgment, the system can dynamically identify which data has been effectively compressed
[0030] S2. Generate an original RelNode tree based on the nodes through an SQL parser and perform logical optimization. Use the RelMetadataQuery API to extract the node lineage relationships from the optimized RelNode tree to construct a node lineage graph. Specifically, generating an original RelNode tree based on the nodes through an SQL parser and performing logical optimization, and using the RelMetadataQuery API to extract the node lineage relationships from the optimized RelNode tree to construct a node lineage graph means parsing the SQL query through the SQL parser to generate an Abstract Syntax Tree (AST). The AST is a tree-like representation that describes the structure of the SQL statement and contains symbols for operations in all connection edges, such as selection, filtering, joining, etc. Generate an original RelNode tree from the AST. Each tree node includes an operation type, source nodes, and target nodes. Perform logical optimization on the original RelNode tree, including inferring the field sources, merging adjacent operations, and merging grouping operations. The so-called inferring the field sources means analyzing how each step of the operation affects the input fields, inferring the sources of the fields, and forming a simplified lineage graph structure. The so-called merging adjacent operations means that if there is no data change or filtering between consecutive tree nodes, they can be merged to reduce the generation of intermediate results, thereby reducing the computational amount. The so-called merging grouping operations means that if multiple operations belong to the same data grouping, these operations can be merged into one processing step. Use the RelMetadataQuery API method to extract the data source nodes and flow nodes of the operations between the nodes in the optimized RelNode tree. Define the data source nodes as source nodes, the flow nodes as target nodes, and add directed lineage markers to the connection edges between the source nodes and target nodes. Traverse all nodes to gradually construct a node lineage graph.
[0031] Compared with traditional database query processing methods, the advantage of using this abstract tree - like representation method is that it can efficiently capture the operation order and field dependencies in the query, thereby laying a foundation for subsequent optimization and data - flow analysis. Through the clear representation of the tree - like structure, the system can perform logical optimization more quickly and effectively reduce unnecessary operations during the query process. The logical optimization steps greatly improve the query execution efficiency. Inferring the field sources can help the system simplify the lineage graph, enabling faster identification of data dependencies during large - scale data processing. Merging adjacent operations can avoid duplicate calculations and intermediate result storage, thus reducing memory occupancy and calculation time. The merging of grouping operations ensures the optimal processing of data grouping and avoids unnecessary duplicate operations. Extracting node lineage relationships from the optimized RelNode tree through the RelMetadataQuery API provides a clear description of the dependencies between the entire data flow and operations, which can provide a reliable basis for subsequent anomaly detection and data problem tracing. The construction of the node lineage graph can provide a visual structure for data flow, facilitating subsequent tracking and management of data changes. When dealing with complex queries, the node lineage graph provides an intuitive way to help understand and analyze the order of query operations and data flow directions. Through the lineage graph, the system can effectively identify the propagation paths of data changes, thereby optimizing query strategies or diagnosing data problems. In addition, the node lineage graph plays an important role in anomaly detection and can help the system locate possible anomaly sources in the data flow, thus improving the accuracy of anomaly detection.
[0032] S3. Construct a hybrid anomaly detection model using a graph attention network and the GraphSAGE model to detect anomalies in the node lineage graph, and construct an anomaly propagation chain based on the abnormal nodes; Specifically, constructing a hybrid anomaly detection model using a graph attention network and the GraphSAGE model to detect anomalies in the node lineage graph means that based on the constructed node lineage graph, the identification name and type of the data in the nodes are concatenated to form node vectors, and the graph attention network is used to calculate the attention coefficients between nodes And perform a normalization operation on the attention coefficients:
[0033] where is the parameter of the shared attention mechanism, W is the weight matrix, and are the vectors of node i and node j respectively; Update the node vectors according to the normalized attention coefficients to obtain the node attention feature vectors :
[0034] where is the set of adjacent nodes of node i, is the attention coefficient of node i and node j after normalization; Use the GraphSAGE model to perform neighbor aggregation on the nodes in the node blood relationship graph:
[0035] where is the result of neighbor aggregation, is the degree of node i, that is, the number of neighbor nodes, is the vector of node j in the (k - 1)-th layer of the GraphSAGE model; Update the node vector through neighbor aggregation to obtain the node aggregation feature vector :
[0036] where is the node vector of node i in the (k - 1)-th layer; Combine the graph attention network and the GraphSAGE model to form a hybrid anomaly detection model, including an input layer, an embedding layer, an aggregation layer, and an output layer; Among them, the input layer is used to input the node blood relationship graph, the embedding layer embeds the graph attention network and the GraphSAGE model respectively to extract the node attention feature vector and the node aggregation feature vector, the aggregation layer is used to splice the node attention feature vector and the node aggregation feature vector to form a node hybrid vector, and the output layer is used to classify the node hybrid vector through a linear layer; Optimize the hybrid anomaly detection model through the cross-entropy loss function; Input the node blood relationship graph into the trained hybrid anomaly detection model, obtain the node anomaly probability in the node blood relationship graph through the output layer, and compare the node anomaly probability with the anomaly threshold to judge the node anomaly situation.
[0037] The graph attention network can automatically learn the most important relationships in the data by assigning different attention coefficients to the edges between nodes. In the anomaly detection task, the correlation between nodes often determines the propagation path of anomaly events. Through this dynamic weighting mechanism, the graph attention network can effectively capture the potential and weak correlations between nodes, which may be crucial for anomaly detection. Compared with the traditional graph convolutional network (GCN), GAT can significantly improve the model's recognition ability when dealing with complex graph data, especially in the case of sparse data and complex connection relationships. The neighbor aggregation ability of GraphSAGE gives it an obvious advantage when dealing with large-scale graph data. In traditional graph neural networks, the representation of each node usually depends on the information of the entire graph, which may bring computational bottlenecks in large-scale graph structures. However, GraphSAGE only depends on the local graph structure by aggregating the neighbor information of nodes, greatly improving the computational efficiency. The hybrid model formed by combining the graph attention network and the GraphSAGE model can give full play to the advantages of both. The graph attention network can focus on the neighbors that have important relationships with the target node, while GraphSAGE can effectively aggregate the neighborhood information of nodes. Through the collaborative work of the two, the hybrid model can not only handle the complex dependence relationships between nodes, but also capture the local features in large-scale graph data. By modeling the blood relationship of nodes, it can more accurately capture the flow and change paths of data. The anomaly of each node may not only be the change of its own features, but may also be affected by the behaviors of its surrounding nodes. Therefore, anomaly detection based on blood relationship can make judgments in a more comprehensive context, significantly improving the effect of anomaly recognition.
[0038] Furthermore, constructing an anomaly propagation chain based on anomaly nodes means extracting adjacent nodes from the node blood relationship graph according to the anomaly nodes, and defining the propagation chain direction according to the directed blood relationship marks of the node connection edges. Connect the nodes in the propagation chain direction to form an anomaly propagation chain, and traverse all adjacent nodes to form an anomaly propagation chain set. Calculate the anomaly score for each anomaly propagation chain in the anomaly propagation chain set. :
[0039] where is the node anomaly probability of the i-th node in the anomaly propagation chain, and n is the number of nodes in the anomaly propagation chain; Select the anomaly propagation chain with the highest anomaly score from the anomaly propagation chain set as the main propagation chain.
[0040] The construction of the exception propagation chain is based on exception nodes and their blood relationships. By tracing the scope of influence of exception nodes, it helps to deeply understand how an exception event spreads from one node to other nodes. Simply detecting exception nodes cannot reflect the diffusibility of exceptions. However, by constructing a propagation chain, the propagation path and potential impact of exception behaviors can be revealed. For complex systems, especially large-scale distributed networks or multi-level systems, the propagation of exceptions usually expands step by step. Locating the exception propagation chain can more accurately identify key issues globally. When constructing the propagation chain, directed blood relationship markers are used to define the direction of the propagation chain between nodes. Through this directed relationship, the flow path of data or exception status from the source node to the target node can be clearly identified. In this way, not only can the source node of the exception be traced, but also the impact path of the exception on other nodes can be accurately determined, providing a clear basis for subsequent decision-making and handling. Different from the simple adjacency relationship of an undirected graph, directed blood relationship markers can more accurately express the direction of dependence and influence between nodes. Especially in large-scale data systems, it can prevent misjudgment and omission of key exception sources. Traversing all nodes adjacent to the exception node and forming a set of propagation chains can comprehensively evaluate the scope of influence of the exception node on the entire system. Calculating the exception score for each exception propagation chain can quantitatively evaluate different exception propagation chains based on the exception probability of the node and the number of nodes in the chain. The score not only considers the exception degree of a single node but also combines the complexity and breadth of the entire propagation chain. This multi-dimensional evaluation method can significantly improve the accuracy and processing efficiency of exception detection. Especially when facing large-scale data and complex relationships, it can quickly locate the exception chains that most need attention.
[0041] Furthermore, after collecting the data logs, store the data logs in a database, and synchronously store the node blood relationship graph and the exception propagation chain. The database stores and classifies the data according to the timestamp and node data type and performs regular backup processing.
[0042] This embodiment also provides an intelligent exception data monitoring system based on metadata, including: A data collection and compression module, which is used to collect data logs, extract node tables and edge tables from the data logs, and construct a nested list structure to store each node in a nested and compressed manner; A node blood relationship extraction module, which is used to generate an original RelNode tree based on nodes through an SQL parser and perform logical optimization, and extract the node blood relationship from the optimized RelNode tree using the RelMetadataQuery API to construct a node blood relationship graph; An exception detection module, which is used to construct a hybrid exception detection model using a graph attention network and a GraphSAGE model to detect exceptions in the node blood relationship graph, and construct an exception propagation chain based on the exception nodes.
[0043] This embodiment also provides a computer device, which is applicable to the situation of the intelligent monitoring method for abnormal data based on metadata, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent monitoring method for abnormal data based on metadata as proposed in the above embodiment.
[0044] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0045] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the intelligent monitoring method for abnormal data based on metadata as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0046] In summary, the present invention collects data logs to extract node tables and edge tables and constructs a nested list structure to perform nested compression storage on each node, effectively improving data storage efficiency and reducing computational resource consumption. By using the RelMetadataQuery API to extract node lineage from the optimized RelNode tree and constructing a node lineage graph, the graph representation ability is improved, providing richer context information for anomaly detection. At the same time, by combining the Graph Attention Network (GAT) with the GraphSAGE model to perform anomaly detection on the node lineage graph and construct an anomaly propagation chain, the accuracy and scalability of anomaly detection are significantly improved.
[0047] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for intelligent monitoring of abnormal data based on metadata, characterized in that: include, Collect data logs, extract node tables and edge tables from the logs, build an in-edge aggregation subgraph based on each node in the node table, and build a nested list structure to perform nested compression storage on each node; Generate the original RelNode tree based on the nodes through the SQL parser and perform logic optimization. Use the RelMetadataQueryAPI to extract the node blood relationship from the optimized RelNode tree to build a node blood relationship graph. A hybrid anomaly detection model is constructed using graph attention network and GraphSAGE model to detect anomalies in node lineage graphs, and anomaly propagation chains are constructed based on abnormal nodes.
2. The method for intelligently monitoring abnormal data based on metadata according to claim 1, characterized in that: The extracting of the node table and edge table from the log, constructing an in-edge aggregation subgraph based on each node in the node table, and constructing a nested list structure to perform nested compression storage on each node refers to extracting the identification name and type of each type of data from the collected data event log to construct a node and form a node table, assigning an index value to each node in the node table, and recording the node, timestamp, and operation type where the event occurred in the event log as an edge to form an original edge table; For each node v in the node table, record the relationship between each edge and node v to build a mapping table, connect the edge with node v to build an in-edge aggregation subgraph, extract all nodes connected to node v in the in-edge aggregation subgraph as adjacent nodes to form an adjacent node set, and synchronously record the earliest timestamp of all edges connected to node v; Construct a nested list structure for all connected edges of each node ; The nested list structure , node index and the set of adjacent nodes of node v Composition of the final storage node format : ; Calculate the storage space of the original edge table ; Synchronous calculation to construct nested list structure Storage space after ; Calculate the ratio of the number of nodes after constructing the nested list structure to the total number of connecting edges between nodes as the average number of connecting edges, and and For comparison, when the average number of connected edges is greater than or equal to the set threshold and Greater than When the compression is judged to be effective, the final storage node format is Store the nodes.
3. The method for intelligently monitoring abnormal data based on metadata according to claim 2, characterized in that: The SQL parser generates the original RelNode tree based on the node and performs logic optimization, and uses the RelMetadataQuery API to extract the node lineage relationship from the optimized RelNode tree to build a node lineage graph, which refers to parsing the SQL query through the SQL parser to generate an abstract syntax tree AST; Generate the original RelNode tree from the AST. Each tree node includes the operation type, source node, and target node. Perform logical optimization on the original RelNode tree, including inferring the source of fields, merging adjacent operations, and grouping operations. Use the RelMetadataQuery API method to extract the data source nodes and flow nodes of the operations between the nodes in the optimized RelNode tree, define the data source nodes as source nodes, define the flow nodes as target nodes, add directed lineage markers to the connecting edges between the source nodes and the target nodes, traverse all nodes and gradually build a node lineage graph.
4. The method for intelligently monitoring abnormal data based on metadata according to claim 3, characterized in that: The hybrid anomaly detection model constructed by using the graph attention network and the GraphSAGE model to perform anomaly detection on the node lineage graph refers to concatenating the identification name and type of the data in the node to form a node vector based on the constructed node lineage graph, and using the graph attention network to calculate the attention coefficient between the nodes. And normalize the attention coefficient; Update the node vector according to the normalized attention coefficient to obtain the node attention feature vector ; The GraphSAGE model is used to aggregate neighbors of nodes in the node lineage graph; Update the node vector through neighbor aggregation to obtain the node aggregation feature vector ; The graph attention network and the GraphSAGE model are combined to form a hybrid anomaly detection model, including an input layer, an embedding layer, an aggregation layer, and an output layer; The hybrid anomaly detection model is optimized through the cross entropy loss function; The node lineage graph is input into the trained hybrid anomaly detection model, the node anomaly probability in the node lineage graph is obtained through the output layer, and the node anomaly probability is compared with the anomaly threshold to determine the node anomaly.
5. The method for intelligently monitoring abnormal data based on metadata according to claim 4, characterized in that: The construction of abnormal propagation chain based on abnormal nodes refers to extracting adjacent nodes from the node lineage graph according to the abnormal nodes, defining the propagation chain direction according to the directed lineage mark of the node connection edge, connecting the nodes according to the propagation chain direction to form an abnormal propagation chain, and traversing all adjacent nodes to form an abnormal propagation chain set, and calculating the abnormal score for each abnormal propagation chain in the abnormal propagation chain set. ; The anomaly propagation chain with the highest anomaly score is selected from the anomaly propagation chain set as the main propagation chain.
6. The method for intelligently monitoring abnormal data based on metadata according to claim 5, characterized in that: The data log collection refers to the use of a system audit framework to collect various data event logs in real time and perform pre-processing.
7. The method for intelligently monitoring abnormal data based on metadata according to claim 6, characterized in that: After collecting data logs, they are stored in the database, and the node lineage graph and abnormal propagation chain are stored synchronously. The database is classified and stored according to timestamps and node data types and backed up regularly.
8. A metadata-based abnormal data intelligent monitoring system, based on the metadata-based abnormal data intelligent monitoring method according to any one of claims 1 to 7, characterized in that: include, The data collection and compression module is used to collect data logs, extract node tables and edge tables from the data logs, and construct a nested list structure to perform nested compression storage on each node; The node lineage extraction module is used to generate the original RelNode tree based on the node through the SQL parser and perform logic optimization, and use the RelMetadataQuery API to extract the node lineage relationship from the optimized RelNode tree to build a node lineage graph; The anomaly detection module is used to construct a hybrid anomaly detection model using the graph attention network and the GraphSAGE model to perform anomaly detection on the node lineage graph and build an anomaly propagation chain based on abnormal nodes.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the abnormal data intelligent monitoring method based on metadata described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the abnormal data intelligent monitoring method based on metadata described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
APT detection method and system based on multi-branch graph neural network
CN120378226A