An intelligent data governance system based on behavioral path analysis

Through an intelligent data governance system based on behavioral path analysis, graph neural networks and graph diffusion models are used to generate behavioral path maps, which solves the problem of lack of global correlation analysis in existing technologies, realizes early identification and real-time blocking of complex attacks, and improves the security and accuracy of data access.

CN120524484BActive Publication Date: 2025-09-16GUANGZHOU PRINCIPAL DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511016468.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-09-16
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing data governance technologies lack global correlation analysis of data operation behaviors, cannot effectively detect complex attacks, and cannot quantify the impact scope of abnormal operations, resulting in insufficient accuracy in data access governance.

Method used

An intelligent data governance system based on behavioral path analysis is adopted, which generates behavioral path maps through graph neural networks, uses graph diffusion models to simulate the potential impact range of abnormal operations, extracts behavioral path characteristic parameters, marks high-risk nodes and related edges, and combines multi-dimensional risk assessment to perform data access governance.

Benefits of technology

It achieves early identification and real-time blocking of complex attacks, improves the accuracy of risk node identification and the security of data access, and dynamically adjusts confidence levels to cope with changing data operation patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524484B_ABST
    Figure CN120524484B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field related to data governance analysis, and specifically includes an intelligent data governance system based on behavioral path analysis, the system comprising: a path generation module that generates data flow paths using a graph neural network according to data behavior logs, and draws up a behavioral path map; a first and a second risk propagation path set determination module that respectively uses graph diffusion and feature parameter marking to determine risk paths, and uses an operation blocking mechanism to perform data access governance. This solves the technical problems of relying on single rule matching and single-point anomaly detection in data access, lacking global correlation analysis of data operation behaviors, and being unable to effectively detect complex attacks. It achieves the technical effect of generating behavioral path maps through graph neural networks, quantifying the impact range of abnormal operations using graph diffusion models, predicting potential paths for sensitive data leakage in advance, dynamically adjusting confidence levels, improving the accuracy of risk node identification, and performing real-time blocking governance to ensure data access security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to data governance analysis, and in particular to an intelligent data governance system based on behavior path analysis. Background Art

[0002] At a time when digital transformation and data security demands are surging, the scale of enterprise data assets is growing exponentially, and data flow scenarios are becoming increasingly complex, such as cross-system access scenarios and multi-agent collaboration scenarios. Conventional data governance methods lack the ability to dynamically model data behavior and are therefore unable to cope with the risk identification needs in massive operation logs.

[0003] Existing data governance technologies rely on rule matching or single-point anomaly detection, and are unable to analyze the correlation of behavioral paths from a global perspective, making it difficult to identify advanced persistent threats. In addition, due to the lack of quantitative simulation capabilities for risk propagation paths, they can only locate isolated risk points and are unable to assess the impact of abnormal operations through graph diffusion models. The behavioral feature extraction dimension is single, and the accuracy of data access governance is insufficient.

[0004] In summary, existing technologies rely on single rule matching and single-point anomaly detection in data access, lack global correlation analysis of data operation behaviors, and are unable to effectively detect complex attacks. Summary of the Invention

[0005] This application provides an intelligent data governance system based on behavioral path analysis, aiming to solve the technical problems in the existing technology that data access relies on single rule matching and single-point anomaly detection, lacks global correlation analysis of data operation behavior, and cannot effectively detect complex attacks.

[0006] In view of the above problems, the technical solution to implement this application is:

[0007] The present application provides an intelligent data governance system based on behavioral path analysis, wherein the system includes: generating data flow paths using graph neural networks based on data behavior logs, and formulating a behavioral path graph containing nodes and edges, wherein the nodes in the behavioral path graph are used to represent operating entities, and the edges in the behavioral path graph are used to represent operating relationships; based on the operation sequence pattern of the abnormal behavior path, using graph diffusion in the behavioral path graph to simulate the potential impact range of abnormal operations, and determining a first risk propagation path set; based on the behavioral path graph, extracting behavioral path characteristic parameters with path length, node type distribution, and time distribution; based on the behavioral path characteristic parameters of the abnormal behavior path, marking high-risk nodes and associated edges in the behavioral path graph, and determining a second risk propagation path set; based on the first risk propagation path set and the second risk propagation path set, determining key risk diffusion nodes and sensitive data access paths, and performing data access governance with an operation blocking mechanism.

[0008] Preferably, a path propagation aggregation analysis is performed on the first risk propagation path set and the second risk propagation path set to determine the risk propagation domain; based on the behavioral path map, operation context feature labels are added, and a multi-dimensional risk confidence assessment is performed in combination with the first risk propagation path set and the second risk propagation path set to determine the key risk diffusion nodes and sensitive data access paths in the risk propagation domain.

[0009] Preferably, the collection dimensions of the data behavior log include: operation subject, operation timestamp, operation object, operation type, operation permission and associated metadata; wherein, the operation subject includes user ID and system account, the operation object includes data table name, field name, file path, the operation type includes query, modification, deletion, and export; the associated metadata includes data sensitivity level and business scenario label.

[0010] Preferably, databases, data services, file systems and business systems are used as data sources; based on the data sources, embedded probes, log collection tools and API interfaces are used to synchronously collect the data behavior logs.

[0011] Preferably, the operation relationship is defined as a directed edge, and the edge attributes of the directed edge include the operation timestamp, operation object, and operation authority; based on the directed edge, the dynamic association pattern between nodes is learned to generate a behavior path graph containing attribute characteristics.

[0012] Preferably, the continuous operation sequence is aggregated with the operation subject as the index; at the same time, the edge weight value is dynamically assigned according to the operation type and the associated metadata; incremental graph calculation is performed based on the continuous operation sequence and the edge weight value, and the data flow path is injected into the graph neural network in real time to update the node attributes and edge relationships of the behavior path graph.

[0013] Preferably, a reference database under a predefined normal behavior mode is determined; based on the reference database, outlier paths are extracted, and the transmission risk probability is evaluated to perform risk determination on the outlier paths; and according to the risk determination result, the abnormal behavior path is determined.

[0014] Preferably, the transmission risk assessment factors include node conduction factors and edge conduction factors; wherein, the node conduction factor is obtained by weighted fusion of the historical risk frequency of the operating subject and the current operating authority level, and the edge conduction factor is dynamically generated based on the comprehensive operation type risk coefficient, data sensitivity level, and time series density.

[0015] Preferably, the number of path operation nodes corresponding to the sensitive data access path is recorded; a hierarchical blocking decision tree is constructed through the number of path operation nodes, node conduction factors, and edge conduction factors to determine the operation blocking mechanism.

[0016] In summary, one or more technical solutions provided in this application achieve the technical effect of generating behavioral path maps through graph neural networks, quantifying the impact range of abnormal operations using graph diffusion models, predicting potential paths for sensitive data leakage in advance, dynamically adjusting confidence levels, improving the accuracy of risk node identification and performing real-time blocking and governance, and ensuring data access security. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A structural diagram of an intelligent data governance system based on behavioral path analysis is provided for this application.

[0018] Figure 2 This application provides a flow chart of the node attributes and edge relationships of the updated behavior path graph of an intelligent data governance system based on behavior path analysis.

[0019] Explanation of the reference numerals: path generation module M100, first risk propagation path set determination module M200, second risk propagation path set determination module M300, access governance module M400. DETAILED DESCRIPTION

[0020] Embodiment: The present application will be described in detail below with reference to the accompanying drawings. Figure 1 As shown, the present application provides an intelligent data governance system based on behavioral path analysis, wherein the system includes:

[0021] The path generation module M100 is used to generate a data flow path using a graph neural network based on the data behavior log, and to formulate a behavior path graph containing nodes and edges, wherein the nodes in the behavior path graph are used to represent the operation subjects, and the edges in the behavior path graph are used to represent the operation relationships; the first risk propagation path set determination module M200 is used to determine the first risk propagation path set based on the operation sequence pattern of the abnormal behavior path, using graph diffusion in the behavior path graph to simulate the potential impact range of abnormal operations.

[0022] Specifically, graph neural networks are algorithms that use graph-structured data for deep learning and can capture the complex relationships between nodes and edges. Behavior path graphs are models that abstract data operation behaviors into graph structures, where nodes represent operation subjects, including users and system accounts, and edges represent operation relationships, including queries and modifications. Graph diffusion refers to the process of simulating the spread of information or risks in a graph structure, and operation sequence patterns refer to sequences of operation behaviors arranged in chronological order.

[0023] Execution steps: Model the data behavior log through the graph neural network, generate the data flow path, and build a behavior path map containing nodes and edges. Specifically, by collecting the user's query, modification and other operation logs on the data table, use the graph neural network to learn the association pattern between nodes to form a comparative map. For the operation logs generated by the enterprise every day, single-point detection needs to be analyzed one by one. Preferably, these logs are integrated into a map to greatly reduce the amount of data and retain related information.

[0024] Based on the operation sequence pattern of the abnormal behavior path, graph diffusion is used in the graph to simulate the potential impact range of abnormal operations and determine the first risk propagation path set. Furthermore, if an abnormal query operation sequence pattern is detected for a certain user, such as frequent queries of highly sensitive data tables in a short period of time, starting from the user node, graph diffusion is used to simulate other nodes and edges that may be affected, determine the risk propagation path, form the first risk propagation path set, accurately identify possible risk propagation paths, and provide a basis for subsequent governance.

[0025] The second risk propagation path set determination module M300 is used to extract behavioral path characteristic parameters with path length, node type distribution, and time distribution based on the behavioral path graph; mark high-risk nodes and associated edges in the behavioral path graph according to the behavioral path characteristic parameters of abnormal behavioral paths, and determine the second risk propagation path set; the access governance module M400 is used to determine key risk diffusion nodes and sensitive data access paths based on the first risk propagation path set and the second risk propagation path set, and perform data access governance with an operation blocking mechanism.

[0026] Specifically, path length refers to the number of nodes and edges in the behavior path, indicating the complexity of the operation; node type distribution refers to the proportion of different types of nodes in the path, reflecting the diversity of the operation subjects; time distribution refers to the distribution of operation behaviors in the time dimension, reflecting the timing characteristics of the operation; behavior path characteristic parameters are a set of indicators used to describe and quantify the characteristics of behavior paths; marking high-risk nodes and associated edges refers to identifying high-risk nodes and edges in the behavior path map based on the characteristic parameters of abnormal behavior paths; key risk diffusion nodes refer to nodes that play a key role in the risk propagation process; sensitive data access paths refer to paths that may lead to sensitive data leakage; and operation blocking mechanisms refer to measures to intercept or restrict high-risk data access behaviors through technical means.

[0027] Implementation steps: Based on the behavioral path graph, we extract behavioral path characteristic parameters, including path length, node type distribution, and time distribution. By analyzing the user's access behavior path to the data table, for example, if the average path length is 5, the node type distribution shows 30% user nodes and 70% data table nodes, and the time distribution shows 10% nighttime access and 90% workday access, the behavioral path characteristic parameters of path length, node type distribution, and time distribution can comprehensively describe the characteristics of the behavioral path and provide a basis for subsequent risk identification.

[0028] Based on the behavioral path characteristic parameters of the abnormal behavior path, high-risk nodes and associated edges are marked in the behavioral path graph to determine the second risk propagation path set. For example, the path length of an abnormal behavior path is 10, far exceeding the average value of 5, and the node type distribution contains multiple highly sensitive data table nodes. The time distribution shows frequent access during non-working hours. These nodes and edges can be marked as high-risk to form the second risk propagation path set, further enriching the identification results of the risk propagation path.

[0029] Based on the first risk propagation path set and the second risk propagation path set, key risk diffusion nodes and sensitive data access paths are determined, and data access governance is performed with an operation blocking mechanism. For example, by integrating the risk propagation paths in the first risk propagation path set and the second risk propagation path set, node A, which is involved multiple times, is determined to be a key risk diffusion node, and its access path is a sensitive data access path; data access governance is achieved through an operation blocking mechanism, such as limiting user access rights to node A, or performing secondary identity authentication when accessing. Specifically, the method of determining key nodes and paths and blocking them effectively improves the accuracy of risk identification and the targeted governance, thereby enhancing data security protection capabilities.

[0030] Furthermore, the system is also used to perform the following method:

[0031] Perform path propagation aggregation analysis on the first risk propagation path set and the second risk propagation path set to determine the risk propagation domain; based on the behavioral path map, add operation context feature labels, and perform multi-dimensional risk confidence assessment in combination with the first risk propagation path set and the second risk propagation path set to determine the key risk diffusion nodes and sensitive data access paths in the risk propagation domain.

[0032] Specifically, path propagation aggregation analysis is a technique that comprehensively analyzes multiple risk propagation paths to determine the overall scope of risk propagation, known as the risk propagation domain. Operational context feature labels refer to various background information tags related to operational behavior, such as operation time, location, and equipment. Multidimensional risk confidence assessment quantitatively evaluates the likelihood and credibility of a risk by integrating characteristic parameters from multiple dimensions.

[0033] Execution steps: Perform path propagation aggregation analysis on the first risk propagation path set and the second risk propagation path set to determine the risk propagation domain. For example, the first risk propagation path set has 5 paths and the second risk propagation path set has 3 paths. Through path propagation aggregation analysis, these paths are integrated into a risk propagation domain, which contains the nodes and edges of 8 paths, covering the possible risk propagation range, and then comprehensively identifying the overall scope of risk propagation to avoid missing potential risks.

[0034] Based on the behavior path map, add operation context feature labels. For example, in the behavior path map, add labels such as operation time, operation location, and operation device to each operation node. For example, if the operation time of a user node is late at night, the operation location is a non-office area, and the operation device is an unauthorized device, these labels enrich the feature information of the behavior path and provide a more comprehensive basis for subsequent risk assessment.

[0035] A multi-dimensional risk confidence assessment is performed in combination with the first risk propagation path set and the second risk propagation path set. For example, by comprehensively considering characteristic parameters such as path length, node type distribution, time distribution, and operation context feature labels, a multi-dimensional risk confidence is performed on the nodes and paths in the risk propagation domain, and the key risk diffusion nodes and sensitive data access paths in the risk propagation domain are determined. Preferably, the accuracy and credibility of risk identification are improved through multi-dimensional assessment, and the data security protection capability is enhanced.

[0036] Furthermore, the path generation module M100 is further configured to execute the following method:

[0037] The collection dimensions of the data behavior log include: operation subject, operation timestamp, operation object, operation type, operation permission and associated metadata; among which, the operation subject includes user ID and system account, the operation object includes data table name, field name, file path, the operation type includes query, modification, deletion, and export; the associated metadata includes data sensitivity level and business scenario label.

[0038] Specifically, a data behavior log refers to a detailed log file that records data operation behaviors, containing information in multiple dimensions such as the operation subject, operation timestamp, operation object, operation type, operation permissions, and associated metadata; the operation subject refers to the entity that initiates the data operation, such as a user ID or system account; the operation timestamp refers to the specific time when the operation occurs; the operation object refers to the data entity being operated on, such as a data table name, field name, or file path; the operation type refers to the type of operation behavior, such as query, modification, deletion, or export; the operation permission refers to the permission level of the operation subject to the operation object; and the associated metadata refers to other metadata related to the operation behavior, such as data sensitivity level and business scenario label.

[0039] Execution steps: The collection dimensions of data behavior logs include the operation subject, operation timestamp, operation object, operation type, operation permission, and associated metadata. For example, the system will record the operation of user A (operation subject) querying data table T1 (operation object) (operation type) at the operation timestamp, and at the same time record that the user has read permission (operation permission) on data table T1 and the sensitivity level of data table T1 is high (associated metadata). These detailed operation information provides a rich data foundation for the subsequent construction of the behavior path map.

[0040] By collecting data from these dimensions, we can comprehensively record the details of data operation behaviors and provide rich feature information for subsequent risk identification and assessment. For example, by analyzing the user ID and system account of the operating subject, we can understand the identity of the initiator of the operation; by analyzing the operation timestamp, we can understand the time pattern of the operation; by analyzing the operation object, we can understand the specific data entity of the operation; by analyzing the operation type, we can understand the behavioral nature of the operation; by analyzing the operation permissions, we can understand the rationality of the operation; by analyzing the associated metadata, we can understand the data background of the operation and build a more accurate behavior path map, thereby better identifying potential risk behaviors.

[0041] The use of multi-dimensional data collection methods provides data governance with more comprehensive and detailed operation records, which helps to improve the accuracy and reliability of risk identification. Through this multi-dimensional collection method, abnormal operation behaviors can be identified more accurately, such as a user frequently querying highly sensitive data tables during non-working hours, thereby timely discovering potential data leakage risks, providing a data foundation for subsequent risk propagation path analysis and data access governance, and helping to achieve more effective data security management.

[0042] Furthermore, the system is also used to perform the following method:

[0043] Use databases, data services, file systems, and business systems as data sources; based on the data sources, use embedded probes, log collection tools, and API interfaces to synchronously collect the data behavior logs.

[0044] Specifically, a database is a collection of stored data that organizes, stores, and manages data according to data structures; data services are services that provide functions such as data access, processing, and transmission; a file system is a system responsible for managing the storage, reading, writing, and organization of files; a business system is a system that supports the operation of business processes of an enterprise or organization; a point-of-sale probe is a technical tool used to collect system runtime data, usually embedded in an application in the form of code; a log collection tool is a tool specifically used to collect and manage log files; API interface docking refers to the use of an application programming interface to achieve connection and data interaction between different systems or components.

[0045] Execution steps: Use databases, data services, file systems, and business systems as data sources. For example, MySQL databases store user information, Hadoop data services process big data, NTFS file systems store files, and ERP business systems manage enterprise resources. The above data sources cover the main storage and processing locations of enterprise data, providing a comprehensive source for the collection of data behavior logs.

[0046] Based on the data source, we use tracking probes, log collection tools and API interface docking to synchronously collect data behavior logs. Specifically, we embed tracking probes in the MySQL database to collect user query operations in real time; use the Flume log collection tool to collect data processing logs from the Hadoop data service; obtain business operation logs through the API interface of the ERP business system to achieve real-time synchronous data collection and ensure the integrity and timeliness of data behavior logs.

[0047] Multi-source data collection, combined with a variety of collection tools, enables comprehensive, real-time collection of data behavior logs, providing a rich data foundation for subsequent behavior path mapping and risk identification. In large enterprises encompassing multiple business systems, synchronized data collection allows for real-time monitoring of all data operations, enabling timely detection of abnormal operations and improving the efficiency and effectiveness of data governance. Comprehensive and real-time data collection is crucial for accurately constructing behavior path mapping, ensuring timely and accurate risk identification and enabling more effective data security management.

[0048] Furthermore, the system is also used to perform the following method:

[0049] The operation relationship is defined as a directed edge, and the edge attributes of the directed edge include the operation timestamp, operation object, and operation permission; based on the directed edge, the dynamic association pattern between nodes is learned to generate a behavior path graph containing attribute features.

[0050] Specifically, a directed edge refers to a line pointing from one node to another, representing a specific operation behavior between the operation subject and the operation object; the operation object is the data entity being operated on. Edge attributes include the operation timestamp, operation object, and operation permissions. These attributes can provide rich contextual information and describe the details of the operation behavior. At the same time, the dynamic association pattern refers to the interaction pattern between the operation subject and the operation object that changes over time, reflecting the dynamic characteristics of data operation behavior.

[0051] Execution steps: Define the operation relationship as a directed edge, where edge attributes include the operation timestamp, operation object, and operation permission. For example, in an enterprise database, when user A (node) performs a query operation on data table T1 (node), a directed edge is generated from user A to data table T1. The edge attributes include the current timestamp, the operation object is data table T1, and the operation permission is read permission. The representation of the directed edge can clearly reflect the relationship between the operation subject and the operation object and the operation details.

[0052] Based on these directed edges, the system learns the dynamic association patterns between nodes and generates a behavioral path graph containing attribute features. For example, by analyzing user A's query operations on data tables T1 and T2 at different times, as well as the operational relationship between user A and system account B, the system can learn user A's typical operation patterns and preferences. The behavioral path graph is dynamically updated to include new operation records and update node and edge attributes. For example, during a one-hour observation period, the system recorded 100 operations by user A and generated 100 directed edges. By analyzing the timestamps, operation objects, and permissions of these edges, it can be found that user A's operation pattern is to access work-related data tables, mainly during working hours. This directed edge-based graph generation method can effectively record and reflect the dynamic characteristics of data operation behavior. By analyzing the dynamic association patterns between nodes, anomalous operation behaviors can be detected, such as user A suddenly frequently accessing highly sensitive data tables during non-working hours. Identifying such anomalous behavior helps improve the efficiency of data security governance.

[0053] Furthermore, if Figure 2 As shown, the system is also used to perform the following method:

[0054] Taking the operation subject as the index, the continuous operation sequence is aggregated; at the same time, the edge weight value is dynamically assigned according to the operation type and the associated metadata; incremental graph calculation is performed based on the continuous operation sequence and the edge weight value, the data flow path is injected into the graph neural network in real time, and the node attributes and edge relationships of the behavior path graph are updated.

[0055] Specifically, the operation subject index refers to the identifier of the operation subject, which is used to associate and integrate all its operation behaviors; the continuous operation sequence refers to a series of operation behaviors continuously performed by the operation subject within a certain time range, arranged in chronological order to form a sequence; the edge weight value refers to the weight dynamically assigned to the edge in the behavior path graph according to the operation type and associated metadata, which is used to measure the importance or risk level of the edge; incremental graph calculation refers to the calculation method of dynamically adjusting the graph structure by updating the information of nodes and edges in real time based on the existing graph; real-time injection of data flow path refers to the timely input of the data flow path collected in real time into the graph neural network to update the behavior path graph.

[0056] Execution steps: Use the operation subject as the index to aggregate continuous operation sequences. For example, after user A's operation records are collected, the system uses user A's ID as the index to arrange the operations (such as query, modification, deletion, etc.) performed continuously by user A within a period of time in chronological order to form a continuous operation sequence. User A performed 5 query operations and 2 modification operations in succession within 10 minutes. The system integrates these operations into a sequence to facilitate subsequent analysis of user A's operation pattern.

[0057] Edge weights are dynamically assigned based on the operation type and associated metadata. For example, a query operation might have a lower weight, while a delete or modify operation might have a higher weight. Operations on highly sensitive data tables have higher weights than those on less sensitive data tables. Accordingly, a query operation has a weight of 1, a modify operation has a weight of 3, and a delete operation has a weight of 5. Edges with higher data sensitivity levels receive an additional weight of 2. In this way, edge weights can be dynamically adjusted to reflect the risk level of different operations.

[0058] Incremental graph calculation is performed based on the continuous operation sequence and edge weight values, and the data flow path is injected into the graph neural network in real time to update the node attributes and edge relationships of the behavior path graph. For example, after user A's continuous operation sequence and dynamic weight value are calculated, the system inputs this information into the graph neural network in real time to update the behavior path graph. For example, if user A's operation causes the access frequency of the node corresponding to the highly sensitive data table to increase, the node attributes will be updated, and the edge weight will also be adjusted according to the operation type and data sensitivity level; the graph neural network dynamically updates the graph through incremental calculation to ensure the timeliness and accuracy of the graph.

[0059] The aggregation and dynamic edge weight assignment method based on the operation subject index can more accurately reflect the behavior patterns and operation risks of the operation subject, and provide more accurate data support for subsequent risk identification and data governance; by updating the behavior path map in real time, the system can promptly detect abnormal operation behaviors and improve the efficiency and effectiveness of data security governance.

[0060] Furthermore, the first risk propagation path set determination module M200 is configured to execute the following method:

[0061] Determine a benchmark database under a predefined normal behavior mode; extract outlier paths based on the benchmark database, evaluate the transmission risk probability, and perform risk determination on the outlier paths; and determine the abnormal behavior path based on the risk determination result.

[0062] Specifically, predefined normal behavior patterns refer to operational behavior patterns set in advance based on historical data and business rules, which are used as a benchmark for normal behavior; the benchmark database is a database that stores these normal behavior patterns and is used to compare and identify abnormal behavior; outlier paths refer to behavior paths that are significantly different from normal behavior patterns, which may usually contain abnormal or high-risk operations; transmission risk probability refers to the possibility of path risk occurring based on multiple factors; risk judgment refers to the process of determining whether a path is an abnormal behavior path based on the transmission risk probability.

[0063] Execution steps: Determine a baseline database under predefined normal behavior patterns. For example, by analyzing the normal operation behavior logs of the enterprise system over the past year, extract common operation sequences and patterns, such as user querying data tables and modifying field values ​​during working hours, and store these patterns in the baseline database to form a baseline for normal behavior.

[0064] Based on the benchmark database, outlier paths are extracted, and the transmission risk probability is evaluated, and risk assessment is performed on the outlier paths. For example, in real-time monitoring, if a user is found to frequently query a highly sensitive data table during non-working hours, which is inconsistent with the normal pattern in the benchmark database, the system will identify it as an outlier path; by evaluating the transmission risk probability, such as based on factors such as the historical risk frequency of the operating subject, the level of operating authority, and the sensitivity level of the operating object, the risk probability of the outlier path is determined. If the risk probability is higher than the set threshold, it is determined to be an abnormal behavior path; based on the risk assessment results, the abnormal behavior path is determined. For example, the system identifies an outlier path with a risk probability higher than 70% as an abnormal behavior path and records it in the abnormal behavior database; by identifying and recording abnormal behavior paths, the system can promptly detect potential data security threats and improve the efficiency and accuracy of data security governance.

[0065] In these steps, we effectively distinguish between normal and abnormal behavior, providing strong support for data security governance. Through real-time monitoring and risk assessment, the system can promptly detect and address abnormal behavior paths, reducing the risk of data leakage and ensuring the security of enterprise data assets.

[0066] Furthermore, the first risk propagation path set determination module M200 is further configured to execute the following method:

[0067] The transmission risk assessment factors include node conduction factors and edge conduction factors; wherein, the node conduction factor is obtained by weighted fusion of the historical risk frequency of the operating subject and the current operation authority level, and the edge conduction factor is dynamically generated by comprehensive operation type risk coefficient, data sensitivity level, and time series density.

[0068] Specifically, transmission risk assessment factors refer to various indicators used to evaluate the potential risks of data during transmission. The node transmission factor and edge transmission factor are two key transmission risk assessment factors. The node transmission factor comprehensively considers the historical risk frequency and current operation authority level of the operator, and is calculated through a weighted fusion method to reflect the risk transmission capability of the operator. The edge transmission factor is dynamically generated by combining factors such as the risk coefficient of the operation type, the data sensitivity level, and the time series density, and reflects the risk transmission capability of the operation relationship.

[0069] Execution steps: Determine the transmission risk assessment factors, including the node transmission factor and the edge transmission factor. For example, for a certain operation subject (such as user A), its historical risk frequency is 0.2 (that is, 20 out of every 100 operations are identified as high risk) and the current operation permission level is 3 (the permission level ranges from 1 to 5, with 5 being the highest). The node transmission factor can be calculated by weighted fusion of these two factors. Assuming the weights are 0.6 and 0.4 respectively, the node transmission factor is 0.6×0.2+0.4×(3 / 5)=0.12+0.24=0.36.

[0070] The edge conduction factor is dynamically generated by integrating the operation type risk coefficient, data sensitivity level, and time series density. For example, for user A's operation of querying a highly sensitive data table, the operation type risk coefficient is 0.7 (the risk coefficient of query is low, and the risk coefficient of modification or deletion is high), the data sensitivity level is 0.8 (highly sensitive data), and the time series density is 0.6 (frequent operations in a short period of time). The edge conduction factor can be calculated by integrating these three factors according to certain weights (such as 0.4, 0.4, and 0.2), that is, 0.4×0.7+0.4×0.8+0.2×0.6=0.28+0.32+0.12=0.72.

[0071] Through comprehensive analysis of node conduction factors and edge conduction factors, we can more comprehensively evaluate the transmission risk probability and further determine whether it is an abnormal behavior path. This dynamic assessment method can more accurately identify potential data security threats and improve the efficiency and effectiveness of data security governance.

[0072] Furthermore, the access management module M400 is further configured to execute the following method:

[0073] Record the number of path operation nodes corresponding to the sensitive data access path; construct a hierarchical blocking decision tree based on the number of path operation nodes, node conduction factors, and edge conduction factors to determine the operation blocking mechanism.

[0074] Specifically, the number of path operation nodes refers to the number of operation nodes involved in the sensitive data access path, reflecting the complexity of the access path; the hierarchical blocking decision tree is a rule-based decision model that classifies data access behaviors according to characteristics such as the number of path operation nodes, node conduction factors, and edge conduction factors, and determines whether blocking is required; the operation blocking mechanism refers to measures to intercept or restrict high-risk data access behaviors based on the results of the hierarchical blocking decision tree.

[0075] Execution steps: Record the number of path operation nodes corresponding to the sensitive data access path. For example, in an enterprise system, a sensitive data access path involves 5 operation nodes, such as user A, data table T1, system account B, data table T2, and user C. The number of path operation nodes is 5. The number of path operation nodes reflects the complexity of the access path. The more nodes there are, the higher the potential risk.

[0076] A hierarchical blocking decision tree is constructed based on the number of path operation nodes, node conduction factors, and edge conduction factors to determine the operation blocking mechanism. For example, the decision tree rules are set as follows: if the number of path operation nodes is ≥5, and the node conduction factor is ≥0.4, or the edge conduction factor is ≥0.6, it is judged as high risk, triggering the highest level of blocking, directly blocking access and issuing a reminder; if the number of path operation nodes is between 3-4, and the node conduction factor is between 0.2-0.4 or the edge conduction factor is between 0.4-0.6, it is judged as medium risk, triggering intermediate blocking, limiting access speed and issuing a reminder; if the number of path operation nodes is ≤2, and the node conduction factor is ≤0.2, and the edge conduction factor is ≤0.4, it is judged as low risk, and normal access is allowed.

[0077] The operational blocking mechanism based on the hierarchical blocking decision tree can take different blocking measures according to the degree of risk, which not only ensures data security but also avoids the impact of excessive blocking on normal business. In this way, the blocking strategy can be dynamically adjusted to adapt to the ever-changing data access patterns, improve the flexibility and effectiveness of data security governance, and effectively reduce the risk of data leakage.

[0078] In summary, the beneficial effects of the embodiments of the present application are:

[0079] Due to the adoption of a path generation module, which is used to generate data flow paths using a graph neural network based on data behavior logs, and to formulate a behavior path graph containing nodes and edges, wherein the nodes in the behavior path graph are used to represent the operating subjects, and the edges in the behavior path graph are used to represent the operation relationships; a first risk propagation path set determination module is used to use graph diffusion to simulate the potential impact range of abnormal operations in the behavior path graph based on the operation sequence pattern of the abnormal behavior path, and determine the first risk propagation path set; a second risk propagation path set determination module is used to extract behavioral path characteristic parameters with path length, node type distribution, and time distribution based on the behavior path graph; based on the behavioral path characteristic parameters of the abnormal behavior path, high-risk nodes and associated edges are marked in the behavior path graph to determine the second risk propagation path set; an access governance module is used to determine key risk diffusion nodes and sensitive data access paths based on the first risk propagation path set and the second risk propagation path set, and perform data access governance with an operation blocking mechanism. The present application provides an intelligent data governance system based on behavioral path analysis. It has achieved the technical effect of generating behavioral path maps through graph neural networks, quantifying the impact range of abnormal operations using graph diffusion models, predicting potential paths for sensitive data leakage in advance, dynamically adjusting confidence levels, improving the accuracy of risk node identification and conducting real-time blocking governance to ensure data access security.

[0080] In summary, any step can be stored as a computer instruction or program in an unlimited computer memory and can be called and recognized by an unlimited computer processor, without any unnecessary restrictions.

[0081] Furthermore, the above technical solution only reflects the preferred technical solution of the technical solution of the embodiment of the present application. Some changes that may be made to certain parts thereof by technical personnel in this technical field all reflect the novel principles of the embodiment of the present application. Obviously, technical personnel in this field can make various changes and modifications to the present application without departing from the scope of the present application.

Claims

1. An intelligent data governance system based on behavioral path analysis, characterized in that: include: A path generation module is used to generate data flow paths using a graph neural network based on data behavior logs, and to formulate a behavior path graph containing nodes and edges, wherein the nodes in the behavior path graph are used to represent operation entities, and the edges in the behavior path graph are used to represent operation relationships; A first risk propagation path set determination module is configured to determine a first risk propagation path set based on an operation sequence pattern of an abnormal behavior path and using graph diffusion to simulate a potential impact range of an abnormal operation in the behavior path graph; A second risk propagation path set determination module is configured to extract behavioral path characteristic parameters including path length, node type distribution, and time distribution based on the behavioral path graph; mark high-risk nodes and associated edges in the behavioral path graph based on the behavioral path characteristic parameters of abnormal behavioral paths to determine a second risk propagation path set; An access governance module, configured to determine key risk diffusion nodes and sensitive data access paths based on the first and second risk diffusion path sets, and to perform data access governance using an operation blocking mechanism; The system is further configured to perform the following method: Performing path propagation aggregation analysis on the first risk propagation path set and the second risk propagation path set to determine a risk propagation domain; Based on the behavior path map, add operation context feature tags, and combine the first risk propagation path set and the second risk propagation path set to perform a multi-dimensional risk confidence assessment to determine the key risk diffusion nodes and sensitive data access paths in the risk propagation domain; According to the operation sequence pattern of the abnormal behavior path, graph diffusion is used in the behavior path graph to simulate the potential impact range of the abnormal operation, and a first risk propagation path set is determined, including: Determine a baseline database under predefined normal behavior patterns; Extracting outlier paths based on the benchmark database, evaluating transmission risk probabilities, and performing risk assessment on the outlier paths; Determine the abnormal behavior path according to the risk determination result; Among them, the assessment of transmission risk probability includes: The transmission risk assessment factors include node conduction factor and edge conduction factor; The node transmission factor is obtained by weighted fusion of the historical risk frequency of the operating subject and the current operation authority level, and the edge transmission factor is dynamically generated by integrating the risk coefficient of the operation type, the data sensitivity level, and the time series density; Among them, data access governance is carried out with an operation blocking mechanism, including: Record the number of path operation nodes corresponding to the sensitive data access path; A hierarchical blocking decision tree is constructed based on the number of path operation nodes, node conduction factors, and edge conduction factors to determine the operation blocking mechanism.

2. The intelligent data governance system based on behavior path analysis according to claim 1, characterized in that: Based on the data behavior log, the graph neural network is used to generate the data flow path, which also includes: The collection dimensions of the data behavior log include: operation subject, operation timestamp, operation object, operation type, operation permission and associated metadata; Among them, the operation subject includes user ID and system account, the operation object includes data table name, field name, file path, the operation type includes query, modification, deletion, and export; the associated metadata includes data sensitivity level and business scenario label.

3. The intelligent data governance system based on behavior path analysis according to claim 2, characterized in that: include: Use databases, data services, file systems, and business systems as data sources; Based on the data source, the data behavior logs are collected synchronously using embedded probes, log collection tools and API interface docking.

4. The intelligent data governance system based on behavior path analysis according to claim 3, characterized in that: include: The operation relationship is defined as a directed edge, wherein the edge attributes of the directed edge include an operation timestamp, an operation object, and an operation permission; Based on the directed edges, the dynamic association patterns between nodes are learned to generate a behavior path graph containing attribute features.

5. The intelligent data governance system based on behavior path analysis according to claim 4, characterized in that: include: Aggregate continuous operation sequences using the operation subject as an index; At the same time, according to the operation type and the associated metadata, edge weight values ​​are dynamically assigned; Incremental graph calculation is performed based on the continuous operation sequence and edge weight values, the data flow path is injected into the graph neural network in real time, and the node attributes and edge relationships of the behavior path graph are updated.

Citation Information

Patent Citations

  • Network security monitoring and early warning management method and system based on knowledge graph

    CN116232745A

  • Management software security maintenance method and system based on Internet information technology

    CN117349843A