Data leakage judgment method and system
By constructing a seasonal data breach monitoring map and combining multiple algorithm models, real-time monitoring and dynamic defense against data breach risks in the cloud environment have been achieved. This solves the problem of difficulty in accurately identifying and defending against data breaches in existing technologies and enhances the proactive security capabilities of highly sensitive data assets.
Patent Information
- Application Number
- CN202610002704.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies are insufficient for real-time monitoring and location of data leakage risks in cloud environments, especially in terms of multi-source data fusion, dynamic correlation analysis, and real-time early warning capabilities, making it difficult to accurately identify leakage sources and transmission paths.
A seasonal data breach monitoring map is constructed. By combining the Hidden Markov Model, the forward infectious disease model, and the reverse path reasoning algorithm, a data breach prevention strategy is generated through simulation, enabling accurate judgment and dynamic defense against data breach risks.
It achieves accurate identification and proactive defense of data leakage risks across the entire chain, breaks through the limitations of traditional static defense, adapts to seasonal attack characteristics, and supports dynamic data protection for real-time risk evolution.
Smart Images

Figure CN121462318A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data security, and in particular relates to a method and system for determining data leakage. Background Technology
[0002] The current risks of enterprise data breaches mainly exhibit three characteristics: data assets are deeply intertwined with business systems, access permissions, supply chains, and other dimensions, making it difficult for traditional models to capture hidden leakage paths such as privilege abuse and cross-system vulnerabilities; data flows extremely quickly in the cloud environment, making static auditing methods unable to address real-time leakage risks; most SMEs rely on expert experience rule bases, resulting in a low rate of identification of new leakage methods and a lack of real-time monitoring capabilities for dark web data transactions; existing solutions, lacking multi-source data fusion, dynamic correlation analysis, and real-time early warning capabilities, struggle to accurately pinpoint the source and transmission path of leaks. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention proposes a data leakage assessment method and system. This method responds to a pre-defined seasonal data leakage monitoring map, collects access status data and interaction information of each entity in real time, and generates a first leakage risk logical path by combining it with a leakage risk logic graph. Based on a comprehensive leakage risk assessment model using a Hidden Markov Model, a forward contagion model, a reverse path reasoning algorithm, and a probability function for configuring leakage type distributions, it obtains forward and reverse visualized leakage discrimination path sequences. Through simulation combined with risk thresholds, it performs forward and reverse discrimination simulations to obtain leakage risk probability, risk level, and forward and reverse path overlap. Based on overlap, path, and time interval, it identifies core risk nodes, generates defense strategies by combining them with a pre-defined anti-leakage strategy execution library, and provides feedback for optimization until the risk meets security requirements. This method achieves accurate assessment and dynamic defense against data leakage risks through seasonal feature fusion and collaborative forward and reverse path discrimination.
[0004] To achieve the above objectives, the present invention provides the following technical solution: Data breach detection methods include: In response to the preset seasonal data leakage monitoring map, the access status data and interaction information of each entity in the dynamic data leakage monitoring map are collected in real time and combined with the preset leakage risk logic map to obtain the first leakage risk logic path. The first leakage risk logical path includes data leakage risk type, access subject and abnormal operation behavior score of access subject, access target entity and data sensitivity level and exposure risk of access target entity, access path and leakage risk logical path identification accuracy; Based on the data leakage risk type, access subject and abnormal operation behavior score of the access subject, access path and leakage risk logical path identification accuracy, combined with the hidden Markov algorithm and positive infectious disease model, a positive visual leakage discrimination path sequence is obtained. Based on the data leakage risk type, the target entity being accessed and the data sensitivity level and exposure risk of the target entity being accessed, the access path and the accuracy of the leakage risk logical path identification, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, a reverse visualization leakage discrimination path sequence is obtained. The leakage type distribution probability function is constructed by combining the time distribution of access subject type, leakage risk logical path type and corresponding occurrence frequency with a continuous-time Bayesian model, and is used to measure the seasonal distribution of different types of leakage time.
[0005] Specifically, the methods for determining leakage also include: The positive visualization leakage detection path sequence includes the positive predicted risk leakage probability of each entity node in the positive visualization leakage detection path for accessing the target entity, as well as the leakage risk path length and the positive comprehensive predicted risk probability. The reverse visualization leakage detection path sequence includes the reverse predicted risk leakage probability and the reverse comprehensive prediction risk probability of each entity node in the reverse visualization leakage detection path for accessing the target entity; Based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, a simulation algorithm is used to simulate forward and reverse data leakage discrimination using a leakage risk threshold. This yields the data leakage risk discrimination results for the same access subject to the target entity under different leakage risk logical paths, as well as the overlap between the simulated forward and reverse visualization leakage discrimination paths and the forward and reverse paths. The data leakage risk discrimination results include the leakage risk probability and risk level.
[0006] Specifically, the methods for determining leakage also include: Based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type, the core risk nodes of data leakage are determined. Based on the core risk nodes of data leakage, combined with the data leakage risk judgment results and the preset anti-leakage strategy execution library, the corresponding defense strategy is generated, and the generated defense strategy is fed back to the simulation algorithm for adjustment. The change status of leakage risk probability of the corresponding positive and negative visualized leakage judgment path after adjustment is monitored. Based on the changes in the leakage risk probability of the corresponding positive and negative visualization leakage identification paths after the adjustment, a secondary data leakage defense strategy is determined until the leakage risk probability and risk level of the accessed target entity meet the data leakage security requirements.
[0007] Specifically, the construction of the seasonal data breach monitoring map includes: Based on data breach scenarios, a pre-defined text decomposition algorithm is used to determine the entity composition and data transmission logic of the corresponding scenario. This results in a set of entity nodes including data assets, access subjects, threat sources, and path nodes, as well as relationship connection types including access permissions, data flow, and threat propagation. Basic attributes and seasonal extended attributes are also pre-defined for each entity node. The relationship connection types include access permission relationships, data flow relationships, threat propagation relationships, and seasonal association sensitive relationships; the preset basic attributes include the sensitivity level of data assets and the access scope of the access subject; the seasonal extended attributes include the quarterly access peak of data assets and the quarterly abnormal operation frequency of the access subject; the seasonal association sensitive relationship is used to adjust the quarterly authorization strength coefficient of the access target entity according to the access time attribute of the access subject.
[0008] Specifically, the construction of the seasonal data breach monitoring map also includes: Obtain operation logs, permission change records, and vulnerability intelligence logs corresponding to historical leakage cases within a preset time period under different data leakage scenarios. Separate trend items and seasonal items through time series decomposition algorithm, extract seasonal characteristics of different entity nodes and seasonal correlation leakage risk characteristics between them and the remaining entity nodes, and obtain the seasonal risk coefficient of different entity nodes under the corresponding scenario. Based on entity nodes, the corresponding relationship connection types between different entity nodes, and the seasonal risk coefficient in the current scenario, an initial seasonal data leakage monitoring map is constructed using graph algorithms. Based on the initial seasonal data breach monitoring map, a seasonal data breach monitoring map is obtained by combining a breach risk logic graph obtained based on historical breach paths and preset trigger rules through a community algorithm.
[0009] Specifically, the process of obtaining the positive visualization leakage discrimination path sequence includes: Based on the data leakage risk type, abnormal score of the access subject's operation behavior, access path, and leakage risk logical path identification accuracy in the first leakage risk logical path, the product of the abnormal score of the access subject's operation behavior and the leakage risk logical path identification accuracy is mapped to the corrected observation state. The data leakage risk type, seasonal extension attribute, and associated time-sensitive weight of the entity node are set as hidden states. A hidden Markov model is trained to obtain the leakage behavior state transition probability matrix. Based on the state transition probability matrix of leakage behavior combined with the positive infectious disease model, and combined with the real-time access paths marked as susceptible, infected or recovered, the corrected risk transmission probability between entity nodes is obtained.
[0010] Specifically, the process of obtaining the positive visualization leakage discrimination path sequence also includes: Based on the preset basic attributes of each entity node in the current leakage risk logic graph in the seasonal data leakage monitoring map, the graph traversal algorithm is used to obtain the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node, starting from the preset initial risk node in the current leakage risk logic graph. At the same time, the number of entity nodes contained in the path and the risk propagation probability between every two entity nodes in each path and the original length of the corresponding associated connection are counted to obtain the leakage risk path length. Based on the positive predicted risk leakage probability of each entity node in each path and the risk propagation probability between entity nodes, a weighted average algorithm is used to obtain the positive comprehensive predicted risk probability. Based on the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node, and the corrected risk propagation probability between the corresponding entity nodes of each path, combined with the mapping rules constructed by the preset HSV color space and the leakage risk path length, a positive visual leakage discrimination path sequence is obtained.
[0011] Specifically, the process of obtaining the reverse prediction risk leakage probability and the reverse comprehensive prediction risk probability for each entity node for accessing the target entity includes: Based on the first leakage risk logical path, extract the data leakage risk type, access target entity and the data sensitivity level and exposure surface risk of the target entity, access path and leakage risk logical path identification accuracy. Based on the temporal distribution of access subject type, leakage risk logical path type, and the frequency of occurrence of corresponding entity nodes, a leakage type distribution probability function is obtained through a continuous-time Bayesian model. The Bayesian model, with seasonal time intervals as conditions, outputs the probability of occurrence of the corresponding leakage type within a preset time interval, which is used to describe the seasonal distribution of data leakage. Based on the data sensitivity level and exposure risk of the target entity and the abnormal operation score of the current access subject, the basic risk value of the target entity is calculated by quantifying the sensitivity level into an access sensitivity weight value and the exposure risk into a vulnerability coefficient. Based on the product of the target entity's basic risk value, the probability of occurrence at the current time output by the probability distribution function of the leakage type, and the correctness of the leakage risk logical path identification, combined with the comprehensive leakage risk assessment model, the reverse comprehensive predicted risk probability of the target entity is obtained. Starting from the target entity, traverse the entity nodes in the path in reverse, and filter the effective paths between the accessing main node and the target entity node by combining the preset time constraints. Based on the leakage type distribution probability function of each entity in the effective path, and combining the reverse comprehensive prediction risk probability of the target entity with the access risk probability between the two nodes in the corresponding effective path, the Monte Carlo inference algorithm is used to perform forward risk inference to obtain the reverse prediction risk leakage probability of each entity node in the effective path between the target entity node and any access subject.
[0012] The data breach detection system includes: a monitoring and classification module, a forward discrimination module, and a reverse discrimination module; The monitoring and classification module is used to respond to the preset seasonal data leakage monitoring map, collect the access status data and interaction information of each entity in the dynamic data leakage monitoring map in real time, and combine them with the preset leakage risk logic map to obtain the first leakage risk logic path. The positive discrimination module, based on the data leakage risk type, the access subject and the abnormal score of the access subject's operation behavior, the access path and the recognition accuracy of the leakage risk logical path, combined with the hidden Markov algorithm and the positive infectious disease model, obtains a positive visual leakage discrimination path sequence. The reverse discrimination module, based on the data leakage risk type, the access target entity and the data sensitivity level and exposure risk of the access target entity, the access path and the accuracy of the leakage risk logical path identification, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, obtains the reverse visualized leakage discrimination path sequence.
[0013] Specifically, the leakage detection system also includes: a simulation module, a risk assessment module, an adjustment monitoring module, and an adjustment feedback module; The simulation module, based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, uses a simulation algorithm combined with a leakage risk threshold to perform forward and reverse data leakage discrimination simulation, and obtains the data leakage risk discrimination result of the same access subject to the target entity under different leakage risk logical paths, as well as the corresponding simulated forward and reverse visualization leakage discrimination paths and the overlap of forward and reverse paths; The risk determination module determines the core risk nodes of data leakage based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type. The adjustment monitoring module generates corresponding defense strategies based on the core risk nodes of data leakage, the data leakage risk judgment results, and the preset anti-leakage strategy execution library. It then feeds the generated defense strategies back to the simulation algorithm for adjustment and monitors the change in leakage risk probability of the corresponding positive and negative visualized leakage judgment paths after adjustment. The adjustment feedback module determines a secondary data leakage defense strategy based on the change in leakage risk probability of the corresponding positive and negative visualized leakage discrimination paths after adjustment, until the leakage risk probability and risk level of the accessed target entity meet the data leakage security requirements.
[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention addresses the shortcomings of existing technologies by constructing a dynamic seasonal data leakage monitoring map and combining it with a dual-path risk analysis engine to achieve accurate identification and proactive defense against data leakage risks across the entire data chain. Specifically, it generates forward leakage detection paths based on Hidden Markov Models and forward infectious disease models, and accurately quantifies the risk propagation probability of entity nodes through dynamic adjustments based on behavioral anomaly scores and path identification accuracy. Simultaneously, it utilizes a Bayesian seasonal distribution function to drive the reverse leakage assessment path, integrating the target entity's data sensitivity level, exposure risk, and temporal distribution characteristics to achieve reverse risk tracing from the access source to the data target. Furthermore, through forward and reverse path overlap analysis and simulation, it dynamically locates core risk nodes and links with the anti-leakage strategy execution library to generate targeted defense mechanisms, forming a closed-loop optimization system of monitoring, location, defense, and verification. This overcomes the limitations of traditional static defense, constructing a dynamic data protection architecture that adapts to seasonal attack characteristics and supports real-time risk evolution, significantly improving the proactive security capabilities of highly sensitive data assets. Attached Figure Description
[0015] Figure 1 This is a flowchart of the data leakage judgment method of the present invention; Figure 2 This is a block diagram of the data leakage detection system of the present invention. Detailed Implementation
[0016] Example 1: Please see Figure 1 The present invention provides an embodiment of a data leakage detection method, the steps of which include: S1. Respond to the preset seasonal data leakage monitoring map, collect the access status data and interaction information of each entity in the dynamic data leakage monitoring map in real time, and combine them with the preset leakage risk logic map to obtain the first leakage risk logic path. It should be further explained that the leakage risk logic diagram in this embodiment includes types such as internal privilege escalation, API abuse, supply chain attack, social engineering phishing, and device loss. Each subgraph has a specific entity-relationship chain. Internal privilege escalation is identified by low-privilege user → privilege escalation → access to core library → USB drive copying, and K-means clustering is used to extract abnormal privilege escalation patterns to block illegal exports. API abuse is characterized by external IP → calling unauthorized APIs → batch data acquisition → dark web transactions. Graph editing distance is used to compare the similarity between real-time behavior and historical patterns, and multi-node association is used to reduce false alarms. Supply chain attack is based on the path of third-party system vulnerability → lateral movement → intrusion into internal network → database data breach, and time series decomposition is used to extract time series features to warn of lateral movement. Social engineering phishing is based on phishing email → obtaining account → logging into OA → downloading sensitive documents. A Bayesian model is used to construct a leakage type distribution probability function to cover phishing vector variant attacks. Device loss is based on the causal relationship of laptop theft → unencrypted hard drive → data leakage, and device status and encrypted data are integrated to trigger warnings. These subgraphs extract common patterns from historical cases using K-means clustering and combine them with graph edit distance to achieve real-time graph similarity matching. This transforms abstract scenarios into computable graph structure templates, solving the challenge of identifying attacks with varied methods but repetitive core logic, and improving the accuracy of early warnings and the efficiency of real-time response in high-concurrency scenarios. Specifically, data leakage risk types are categorized by attack patterns, and the nature of the risk is determined by matching real-time paths with historical subgraphs. The access subject refers to the entity initiating the access; its abnormal behavior score is quantified by comparing it with historical baselines to assess risk. The access target entity is the data carrier being accessed; its sensitivity level is marked with confidentiality and exposure surface risk assessment to evaluate the likelihood of attack. The access path presents an entity connection sequence generated from interaction logs. The accuracy of identifying leakage risk logic paths is calculated by real-time overlap with historical path features, providing prior logic for risk judgment.
[0017] It should be further explained that this embodiment is based on the repetitive patterns of attack paths in historical data breach cases. It uses the K-means clustering algorithm to mine common patterns in public data breach events, extracts typical paths with standardized entity-relationship chains, and transforms abstract attack scenarios into computable graph structures. At the same time, it uses a graph edit distance algorithm to achieve fast similarity matching between subgraphs and real-time graphs, and identifies risk associations in a quantitative way. It should be further explained that the first leakage risk logic path in this embodiment includes the data leakage risk type, the access subject and the abnormal operation behavior score of the access subject, the access target entity and the data sensitivity level and exposure risk of the access target entity, the access path and the identification accuracy of the leakage risk logic path; It should be further explained that the data leakage risk type in this embodiment refers to the risk category classified according to the attack mode, such as internal privilege escalation, API abuse, etc., which is used to locate the nature of the risk. It is obtained by matching the real-time access path with the graph edit distance extracted by K-means clustering of historical typical subgraphs. Because the core attack logic is repetitive, it can quickly associate known risk patterns. The access subject refers to the entity that initiates the access, such as user account or external IP. Its abnormal operation behavior score reflects the degree of deviation from the baseline. Its function is to assess the risk level of the subject. The score is obtained by comparing real-time operations, such as access frequency and data volume, with the historical normal behavior baseline and calculating the deviation using the standard deviation or entropy method, providing a quantitative basis for identifying abnormal behavior. The access target entity is the data carrier that is accessed, such as database or document, whose data is sensitive, etc. The levels reflect the importance of data confidentiality, with level 5 being the highest. These levels measure the severity of risk consequences and are categorized using preset sensitive data classification standards, such as personal information and trade secrets. Exposure risk refers to the likelihood of unauthorized access to the target, assessing its vulnerability. This is achieved by analyzing storage location (public / internal network) and access control policies. The access path is the sequence of nodes connecting the accessing entity to the target entity, such as user → API → database. This presents the risk transmission chain and is generated by real-time tracking of entity interaction logs, such as access logs and data flow records. The accuracy rate of leakage risk logical path identification refers to the degree of matching between the current risk path and the corresponding subgraph in the leakage risk logical graph. This is obtained by calculating the feature overlap between the real-time path and the corresponding path in historical cases, such as node type and relationship attribute matching rate. It should be further explained that the implementation process for obtaining the first leakage risk logic path in this embodiment includes: Based on a pre-defined seasonal data leakage monitoring map, access status data and interaction information of each entity in the dynamic data leakage monitoring map are collected through real-time data acquisition technology to obtain a real-time behavior dataset of the entity. It should be further noted that the access status data in this embodiment includes, but is not limited to, login time and operation type; the interaction information includes, but is not limited to, data flow records and node connection relationships. Based on the access path information in the real-time behavior dataset of entities, the graph editing distance algorithm is used to calculate the similarity between the real-time access path and each subgraph in the leakage risk logic graph, and the attack mode corresponding to the subgraph with the highest matching degree is selected to obtain the data leakage risk type. This process is used to locate the specific attack mode to which the risk belongs. Based on the operation records of the accessing subject in the real-time behavior dataset, the real-time operation data is compared with the historical normal behavior baseline of the subject by means of standard deviation or entropy value method. The degree of deviation of the behavior from the baseline is calculated to obtain the abnormal operation score of the accessing subject. By quantifying the degree of abnormality of the accessing subject's behavior, its risk level is assessed. It should be further noted that the accessing subject in this embodiment includes, but is not limited to, user accounts, external IPs, etc.; the operation records include, but are not limited to, access frequency, amount of data processed, operation time period, etc. Based on the target entity information accessed in the real-time behavior dataset, the target entities are labeled according to the preset sensitive data classification rules to obtain their data sensitivity level. The sensitive data classification rules include, but are not limited to, classifying data into personal information, trade secrets, etc., according to information confidentiality. The corresponding data sensitivity level is marked by a person skilled in the art according to the corresponding data security protection level. The target entities include, but are not limited to, databases and documents containing sensitive data. Simultaneously, network topology analysis algorithms are used to analyze the storage location and access control policies of the corresponding target entities, thereby identifying exposure surface risks. This process is used to assess the confidentiality importance of the target entities and their vulnerability to unauthorized access. Storage locations include, but are not limited to, public networks and intranets; access control policies include, but are not limited to, permission configuration and encryption status. Based on the interaction logs in the real-time behavior dataset of entities, the node connection relationship from the access subject initiating the operation to the target entity being accessed is extracted through depth-first search to obtain the access path. This process is used to present the complete link of risk transmission from the access subject to the target entity. The interaction logs include, but are not limited to, access logs, data flow records, etc. The node connection relationship, for example, is user account → API interface → database. Based on the path features of the obtained access path and the corresponding subgraph in the leakage risk logic graph, the cosine similarity algorithm is used to compare the degree of matching between the two to obtain the accuracy of leakage risk logic path identification. This process is used to verify the degree of consistency between the current path and the known risk pattern, and to ensure the reliability of the initial risk judgment. Path features include, but are not limited to, node type, relationship attributes, connection order, etc. Based on the data breach risk type, the abnormal scores of the access subject and its operational behavior, the access target entity and its data sensitivity level and exposure risk, the access path, and the accuracy of identifying the leakage risk logical path, the first leakage risk logical path is formed. This process is used to construct an initial risk profile by integrating multi-dimensional information, providing comprehensive input for subsequent positive and negative risk assessments.
[0018] It should be further noted that the seasonal data leakage monitoring map preset in this embodiment includes: Based on data breach scenarios, a pre-defined text decomposition algorithm is used to determine the entity composition and data transmission logic of the corresponding scenario. This results in a set of entity nodes including data assets, access subjects, threat sources, and path nodes, as well as relationship connection types including access permissions, data flow, and threat propagation. Basic attributes and seasonal extended attributes are also pre-defined for each entity node. The relationship connection types include access permission relationships, data flow relationships and threat transmission relationships, and seasonal association sensitive relationships; the preset basic attributes include, but are not limited to, the sensitivity level of data assets, the access subject's permission scope, and the exposure status of path nodes; the seasonal extended attributes include, but are not limited to, the quarterly access peak of data assets and the quarterly abnormal operation frequency of access subjects; the seasonal association sensitive relationship is used to adjust the quarterly authorization strength coefficient of the access target entity according to the access time attribute of the access subject. The seasonal association sensitive relationship uses a time series analysis model to mine the association rules between access time and sensitive operations, constructs a quarterly authorization strength coefficient matrix, and combines a sliding window algorithm to calculate the access time weight in real time, driving the permission management system to automatically adjust the access threshold.
[0019] It should be further explained that, in this embodiment, the sensitivity level of data assets is based on a preset data classification and labeling system, which serves to define the baseline weight for risk assessment. The scope of access permissions for the accessing entity is determined through the job authorization records of the enterprise's permission management platform. For example, financial personnel are granted permission to view salary data, while interns are only granted permission to browse publicly available information. This system identifies unauthorized behavior by comparing the deviation between actual operations and the scope of permissions. When an intern attempts to access salary data, the deviation triggers a risk warning. The quarterly authorization intensity coefficient is calculated by analyzing the permission approval logs for each quarter over the past three years combined with time series algorithm analysis. The seasonal correlation sensitivity relationship is determined by seasonal correlation. Sensitivity coefficients are constructed; in this embodiment, the seasonal correlation sensitivity coefficient is obtained by statistically analyzing data flow logs from different historical quarters and combining them with different data leakage risk types under the corresponding business cycle. The attributes of the threat sources in this embodiment are obtained from the vulnerability database of the threat intelligence platform, such as the attack tool characteristics of known ransomware vulnerabilities. Their function is to provide early warning of threat propagation risks. When the system is detected to have such a vulnerability, it immediately correlates and warns of ransomware attack risks. The exposure status of path nodes is obtained through regular detection using network port scanning tools. The quarterly access peak of data assets is extracted by analyzing the access logs of the corresponding quarters over the past three years using time series decomposition. It should be further explained that one implementation method for determining the logical relationship between the corresponding scene entity composition and data transmission in this embodiment is as follows: The original descriptive text of data breach scenarios, such as requirement documents, security specifications, and historical case reports, is cleaned using text cleaning algorithms, such as regular expressions, to remove redundant information, such as punctuation marks and repeated sentences, thus obtaining standardized scenario text; providing a clean data foundation for subsequent entity and relation extraction.
[0020] Based on standardized scene text, named entity recognition is performed by combining bidirectional long short-term memory network with conditional random field algorithm to label entity words such as database, employee account, hacker IP, network port, etc. in the text and obtain a preliminary entity candidate set; this process is used to accurately locate the core entity objects involved in the scene.
[0021] Based on the initial candidate entity set, the support vector machine classification algorithm is used to classify the candidate entities according to the pre-set entity category feature library, such as "objects storing data" corresponding to data assets and "objects initiating access" corresponding to access subjects, to obtain a set of entity nodes containing data assets, access subjects, threat sources, and path nodes. This process is used to clarify the functional category of each entity and provide a premise for relationship analysis.
[0022] Based on standardized scene text and entity node sets, the sentence components are parsed using dependency parsing algorithms such as the Stanford parser, and verb phrases such as "access", "transmit", and "attack" and prepositional phrases such as "through" and "from" are extracted between entities to obtain relational description phrases between entities. This process is used to capture the action features of entity interactions and provide a basis for relation type determination.
[0023] Based on relational description phrases, the cosine similarity calculated by semantic matching algorithms such as Word2Vec is compared with a preset relation type dictionary. For example, "access permission" corresponds to "authorized access" and "permission allocation", and "data flow" corresponds to "transfer to" and "synchronize to". This determines the relational connection type between entities, specifically including access permission, data flow, and threat propagation. Based on the entity node set and relationship connection type, the Graphviz directed graph construction algorithm is used to draw a data transmission logic graph with entities as nodes and relationships as directed edges, such as access subject → [access permissions] → data assets, to obtain the data transmission logic relationship in the scenario; this process is used to visualize the interaction path between entities and intuitively reflect the data flow and risk transmission link.
[0024] Obtain operation logs, permission change records, and vulnerability intelligence logs corresponding to historical leakage cases within a preset time period under different data leakage scenarios. Separate trend items and seasonal items through time series decomposition algorithm, extract seasonal characteristics of different entity nodes and seasonal correlation leakage risk characteristics between them and the remaining entity nodes, and obtain the seasonal risk coefficient of different entity nodes under the corresponding scenario. It should be further explained that the process of obtaining the seasonal risk coefficient of different entity nodes in the corresponding scenario in this embodiment includes: Based on the operation logs, permission change records, and vulnerability intelligence logs corresponding to historical leakage cases within a preset time period under different data leakage scenarios, the original log data is preprocessed by data cleaning algorithms such as interpolation to handle missing values and the 3σ rule to remove outliers, resulting in a multi-dimensional time series dataset organized by timestamps such as hours or days. Specifically, it includes indicators such as the operation frequency of entity nodes, the number of permission changes, and the number of vulnerabilities. Based on a regular time series dataset, the Loess decomposition algorithm is used to decompose the time series indicators of each entity node, such as the quarterly access count of a data asset, to separate the trend items that reflect long-term trends, such as the year-on-year increase in access volume, the seasonal items that reflect periodic fluctuations, such as the access peak at the end of each quarter, and the random fluctuations of the residual items. Based on the seasonal terms obtained from the decomposition, the seasonal terms of each entity node are quantitatively analyzed by calculating the quarterly average, the difference between the peak and the trough, and the fluctuation cycle length through feature extraction algorithms. This yields the seasonal characteristics of a single entity node, such as the average frequency of abnormal operations in Q4 for a certain access subject and the fluctuation range of quarterly access volume for a certain path node. Based on the seasonal items of multiple entity nodes, the synchronization degree of seasonal items between entities is calculated by association rule mining algorithm. For example, the Pearson correlation coefficient between the quarterly abnormal operation frequency of a certain access subject and the quarterly access peak of a certain data asset is used to extract the seasonal association leakage risk characteristics between entity nodes. For example, the association pattern of "when the abnormal operation of access subject A increases in Q4, the leakage risk of data asset B in Q4 increases". Based on the seasonal characteristics of individual entities, such as fluctuation amplitude and peak frequency, and the seasonal correlation characteristics between entities, such as correlation strength and synchronization frequency, a risk quantification model is used to sum the feature values according to the impact weights in historical leakage cases. The weights are determined by training a logistic regression model and then comprehensively calculated to obtain the seasonal risk coefficients of different entity nodes in the corresponding scenarios.
[0025] Based on entity nodes, the corresponding relationship connection types between different entity nodes, and the seasonal risk coefficient in the current scenario, an initial seasonal data leakage monitoring map is constructed using graph algorithms. It should be further explained that the process of obtaining the initial seasonal data leakage monitoring map in this embodiment includes: Based on the entity node set and the preset basic attributes and seasonal risk coefficients of each entity node, an attribute labeling algorithm is used to add attribute tags to each entity node, resulting in an attributed entity node set. This is used to clarify the basic characteristics and seasonal risk level of each entity node, providing a quantitative basis at the node level for subsequent risk assessment. It should be further noted that the entity node set includes, but is not limited to, data assets, access subjects, threat sources, and path nodes. Attribute tags are added to each entity node as follows: data assets are labeled with sensitivity level combined with seasonal risk coefficient; access subjects are labeled with permission scope combined with quarterly abnormal operation frequency; threat sources are labeled with vulnerability type combined with quarterly activity coefficient; and path nodes are labeled with exposure status combined with quarterly transmission risk. Based on the relationship connection types between different entity nodes and the corresponding seasonal correlation sensitivity coefficients, a weight value is assigned to each relationship connection through a relationship weight calculation algorithm. For example, the access permission relationship weight of the access subject-data asset is calculated as b1 × quarterly authorization strength coefficient + b2 × basic access frequency strength, resulting in a weighted set of relationship connections. This is used to quantify the seasonal risk transmission strength of relationships between entities, enabling the relationship connections to reflect seasonal fluctuations. The seasonal correlation sensitivity coefficient includes, but is not limited to, the quarterly authorization strength coefficient of access permissions and the quarterly transmission sensitivity coefficient of data flow. Here, b1 represents the contribution ratio of the quarterly authorization strength coefficient in the access permission relationship weight, reflecting the weight of the impact of seasonal permission adjustments on data leakage risk; b2 represents the contribution ratio of the basic access frequency strength in the access permission relationship weight, reflecting the baseline impact weight of normalized access frequency on data leakage risk.
[0026] Based on the set of attributed entity nodes and the set of weighted relational connections, a directed graph construction algorithm is used to generate an initial graph structure containing vertex attributes and edge weights, with data assets, access subjects, threat sources, and path nodes as vertices and weighted relational connections as directed edges. This process is used to build the basic framework of the graph and intuitively present the relationship between entity nodes and relational connections. Based on the initial graph structure and the seasonal risk coefficients of each entity node, a graph with fused node seasonal risk and edge weights is obtained through a graph feature fusion algorithm. This process is used to make the graph simultaneously reflect the seasonal risk of the node itself and the seasonal transmission risk of the relationship between nodes, thereby strengthening the graph's ability to capture seasonal features. It should be further explained that the graph feature fusion algorithm in this embodiment is specifically as follows: performing matrix operations on the seasonal risk coefficient of the entity node and the weight of the connected edge to update the comprehensive risk value of the edge, such as the comprehensive risk value of the edge from path node to data asset = seasonal risk coefficient of path node × weight of data flow relationship. Based on a graph incorporating seasonal features, an initial seasonal data leakage monitoring graph is obtained by correcting isolated nodes or invalid edges using a graph structure verification algorithm. This ensures the integrity and rationality of the graph structure, providing a reliable foundation for subsequent dynamic updates and risk analysis. The graph structure verification algorithm checks the integrity of the association between nodes and edges, ensuring that each data asset has at least one access permission edge connecting the accessing entity, and that each path node has at least one data flow edge connecting upstream and downstream nodes. It should be further explained that, in this embodiment, the data asset node is used to mark the core protected object, and its sensitivity level and seasonal risk coefficient are directly related to the severity of the leakage consequences; the access subject node is used to record the access initiator's permissions and quarterly abnormal characteristics, which is the core basis for identifying unauthorized operations; the threat source node is used to mark the type and seasonal activity of potential threats, providing source information for external attack warnings; the path node is used to record the intermediate links of data transmission and their exposed risks, which is the key hub for tracking the risk diffusion path from the threat source to the data asset. Based on the initial seasonal data breach monitoring map, a seasonal data breach monitoring map is obtained by combining a breach risk logic graph obtained based on historical breach paths and preset trigger rules through a community algorithm.
[0027] It should be further explained that the process of obtaining the seasonal data leakage monitoring map in this embodiment includes: Based on the initial seasonal data leakage monitoring map, the entity nodes in the map are divided into communities using the Louvain community discovery algorithm. The modularity between nodes is calculated. This index measures the tightness of connections within a community and the sparsity of connections between communities, thus obtaining a community cluster composed of multiple highly correlated nodes. Based on leakage path data from historical leakage cases, such as the internal privilege escalation path "privileged user → privilege elevation → core library", a standardized leakage risk logic graph is extracted from historical data using a depth-first search algorithm. This leakage risk logic graph includes subgraphs such as internal privilege escalation type and API abuse type, and the core node sequence is marked for each subgraph. Based on the community cluster and leakage risk logic graph, the similarity between the node sequence of each community cluster and the core node sequence of each subgraph is calculated using the graph editing distance algorithm. For example, the similarity between the financial community node sequence and the core node sequence of the internal unauthorized subgraph is calculated. Community clusters with similarity ≥ a preset threshold are selected and marked as risk-associated communities. This is used to associate community clusters with known leakage patterns and locate high-risk communities.
[0028] Based on preset trigger rules, Rule 1: If a community cluster contains core nodes in historical leakage paths, such as core libraries, and the seasonal risk coefficient is ≥1.2, then the weight of threat transmission relationships within the community is strengthened; Rule 2: If the similarity between a community cluster and a certain subgraph is ≥0.8, then corresponding leakage type tags are added to nodes within the community, such as internal unauthorized access tags; Rule 3: If the frequency of abnormal operations of nodes within the community exceeds twice the historical average in a quarter, then a community risk warning is triggered, and the attributes of risk-related communities are adjusted and marked; This is used to transform historical risk patterns and seasonal characteristics into dynamic attributes of the graph through rules, enhancing the graph's sensitivity to risks.
[0029] Based on the adjusted community cluster, including enhanced relationship weights, leakage type labels, and early warning markers, the seasonal risk coefficients and label attributes of nodes within the community are integrated into the corresponding positions of the initial seasonal data leakage monitoring graph through a graph fusion algorithm. The attribute values of nodes and relationships are updated, and redundant nodes and weakly related edges outside the community cluster are deleted to obtain the seasonal data leakage monitoring graph. Based on the seasonal data breach monitoring map, a rule-based verification algorithm is used to verify whether all marked risk communities contain at least one path that matches the breach risk logic graph, and whether all enhanced relationship weights are associated with nodes with high seasonal risk coefficients. Community attributes that do not conform to the rules are corrected, such as removing mislabeled breach type tags and lowering unreasonable relationship weights. Finally, the seasonal data breach monitoring map is determined to ensure the accuracy of the map and the consistency with the rules, and to ensure the reliability of subsequent risk monitoring.
[0030] This process aggregates highly correlated nodes through community algorithms, uses graph matching to associate historical leakage patterns, and then transforms risk characteristics into graph attributes through triggering rules. The three steps are progressive: community division provides group boundaries for risk positioning, historical path templates provide risk comparison benchmarks, and triggering rules realize the transformation of features into attributes, ultimately forming a monitoring graph that integrates community associations, historical risks, and seasonal characteristics.
[0031] S2. Based on the data leakage risk type, access subject and abnormal operation behavior score of the access subject, access path and leakage risk logical path identification accuracy, combined with the hidden Markov algorithm and positive infectious disease model, a positive visual leakage discrimination path sequence is obtained. S3. Based on the data leakage risk type, the access target entity and the data sensitivity level and exposure risk of the access target entity, the access path and the identification accuracy of the leakage risk logical path, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, a reverse visualization leakage discrimination path sequence is obtained. The leakage type distribution probability function is constructed by combining the time distribution of access subject type, leakage risk logical path type and corresponding occurrence frequency with a continuous-time Bayesian model, and is used to measure the seasonal distribution of different types of leakage time. The positive visualization leakage detection path sequence includes the positive predicted risk leakage probability of each entity node in the positive visualization leakage detection path for accessing the target entity, as well as the leakage risk path length and the positive comprehensive predicted risk probability. The reverse visualization leakage detection path sequence includes the reverse predicted risk leakage probability and the reverse comprehensive prediction risk probability of each entity node in the reverse visualization leakage detection path for accessing the target entity; Furthermore, the process of obtaining the positive visualization leakage discrimination path sequence in this embodiment includes: Based on the data leakage risk type, abnormal score of the access subject's operation behavior, access path, and leakage risk logical path identification accuracy in the first leakage risk logical path, the product of the abnormal score of the access subject's operation behavior and the leakage risk logical path identification accuracy is mapped to the corrected observation state. The data leakage risk type, seasonal extension attribute, and associated time-sensitive weight of the entity node are set as hidden states. A hidden Markov model is trained to obtain the leakage behavior state transition probability matrix. It should be further explained that the process of obtaining the state transition probability matrix of the leakage behavior in this embodiment includes: Based on the entity node data in the first leakage risk logic path, the data leakage risk types are extracted, including internal unauthorized access, API abuse, etc., seasonal extended attributes, including the quarterly access peak of data assets, the quarterly abnormal operation frequency of access subjects, etc., and time-sensitive weights, including the quarterly authorization strength coefficient of access permissions, etc., and are transformed into discretized state values through classification and coding algorithms to obtain a set of hidden states. Based on the abnormal behavior score of the access subject and the accuracy of the logical path identification of leakage risk, the original observation value is obtained through multiplication. Then, the original observation value is divided into three levels of low, medium and high through the equal width discretization algorithm to obtain the corrected observation state set. This is used to transform the continuous abnormal behavior quantification results into discrete observation states that the model can process, reflecting the comprehensive impact of the abnormality of the access subject's operation and the credibility of the path. Based on the frequency of occurrence of hidden states in historical leak cases, the initial probability of occurrence of each hidden state is calculated using a frequency statistics algorithm to obtain the initial probability distribution. Based on the transition records between hidden states in historical data, including the number of transitions from low-season-risk states to high-season-risk states, a state transition probability matrix is initialized using a transition frequency calculation algorithm, i.e., the transition probability equals the number of transitions divided by the total number of transitions. The matrix elements represent the probability of transitioning from state i to state j. This is used to initially characterize the transition patterns between hidden states and provide an initial matrix for model iterative optimization.
[0032] Based on the correspondence between hidden states and observed states in historical data, including the number of times that internal overriding and high-quarterly abnormal frequency states correspond to high observation levels, a conditional probability calculation algorithm is used, i.e., the emission probability equals the corresponding occurrence count divided by the total occurrence count of the hidden state, to initialize the emission probability matrix. The matrix elements represent the probability that the hidden state p generates the observed state q. This is used to establish the correlation between the hidden state and the observed state, so that the model can infer the hidden state from the observed state.
[0033] Based on the observed state sequence in the first leakage risk logic path, i.e. the corrected observed states arranged in chronological order, the initial probability distribution, state transition probability matrix and emission probability matrix are iteratively optimized using the Baum-Welch algorithm. The expectation is calculated using the forward-backward algorithm, and then each probability parameter is updated until convergence is obtained to obtain the trained leakage behavior state transition probability matrix.
[0034] Based on the state transition probability matrix of leakage behavior combined with the positive infectious disease model, and combined with the real-time access paths marked as susceptible, infected or recovered, the corrected risk transmission probability between entity nodes is obtained. It should be further explained that the process of obtaining the corrected risk propagation probability between entity nodes in this embodiment includes: Based on real-time access path data, each entity node in the real-time access path is marked as susceptible, infected, or recovered state according to the state determination rules. Susceptible state refers to a node that is not affected by risk, infected state refers to a node that has been affected by risk, and recovered state refers to a node that was affected by risk but has been relieved. The set of entity nodes after being marked with state is obtained. Based on the state transition probability matrix of the leakage behavior, the probability values of state transition between corresponding entity nodes in the matrix are extracted by the matrix extraction algorithm and used as the basic propagation probability between entity nodes. The basic propagation probability reflects the initial possibility of risk transfer between different state nodes. Based on the set of entity nodes after marking the state and the basic transmission probability, combined with the state transition rules in the positive infectious disease model, namely that the probability of infected nodes spreading to susceptible nodes is higher than that of other state combinations, the basic transmission probability is adjusted through a probability adjustment algorithm to obtain the initially adjusted transmission probability. Based on the time correlation of the state transition probabilities in the state transition probability matrix of the leakage behavior, a time decay factor is introduced through a time decay algorithm. The time decay factor decreases as the duration of the node state increases, which corrects the initially adjusted propagation probability and obtains a propagation probability that takes time factors into account. This is used to reflect the dynamic changes of the risk propagation probability over time and enhance the timeliness of the propagation probability.
[0035] Based on the inhibitory effect of the recovery state node on risk transmission in the positive infectious disease model, an inhibition coefficient algorithm is used to introduce an inhibition coefficient for the transmission probability between the recovery state node and other state nodes. The inhibition coefficient ranges from 0 to 0.5. The transmission probability considering the time factor is then corrected again to obtain the second-corrected transmission probability. Based on the propagation probability after secondary correction, a normalization algorithm is used to map the propagation probability value to the 0-1 range to ensure the validity and consistency of the probability value, thereby obtaining the corrected risk propagation probability between entity nodes. This is used to ensure the rationality of the propagation probability and to provide a reliable quantitative indicator for subsequent risk path analysis.
[0036] This process integrates entity node state marking, initialization of the leakage behavior state transition probability matrix, adjustment of the state rules of the positive infectious disease model, correction of time factors, and suppression of recovery states. It achieves the combination of the leakage behavior state transition probability matrix and the positive infectious disease model. The final corrected risk transmission probability not only reflects the state transition law but also conforms to the transmission characteristics of the infectious disease model, and can accurately reflect the actual possibility of risk transmission between entity nodes.
[0037] Based on the preset basic attributes of each entity node in the current leakage risk logic graph in the seasonal data leakage monitoring map, the graph traversal algorithm is used to obtain the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node, starting from the preset initial risk node in the current leakage risk logic graph. At the same time, the number of entity nodes contained in the path and the risk propagation probability between every two entity nodes in each path and the original length of the corresponding associated connection are counted to obtain the leakage risk path length. This embodiment constructs a dynamically evolving data leakage risk defense system and extracts the core logic from historical attack cases into standardized graph structure templates, making stable intrusion paths explicit amidst diverse attack methods. It utilizes a graph edit distance algorithm to achieve similarity matching between real-time behavioral paths and risk subgraphs, identifying logical consistency threats early in the attack chain. Simultaneously, it introduces a seasonal risk coefficient to dynamically calibrate the weights of entity nodes and relationship connections, accurately responding to changes in risk patterns caused by business cycle fluctuations. A triple defense mechanism is constructed, integrating historical experience-based structuring, real-time multi-node correlation analysis, and time-aware capabilities: historical leakage paths are transformed into internal privilege escalation, application programming interface abuse, etc., through clustering algorithms. The computable subgraph focuses on the essence of the logical chain by stripping away specific attack vectors, covering similar attack variants; it effectively distinguishes between normal business operations and malicious behavior by calling unauthorized application programming interfaces through external Internet Protocol addresses and verifying evidence chains from multiple nodes such as dark web transactions; it extracts seasonal features such as quarterly access peaks and abnormal operation frequencies based on time series decomposition, automatically increasing the risk transmission intensity of core data assets during peak business periods; and finally, it constructs a complete cognitive closed loop from risk diffusion prediction to attack root cause tracing by combining forward leakage probability extrapolation of Hidden Markov Models with reverse tracing driven by Bayesian probability, forming a proactive defense architecture that combines logical stability, time adaptability, and behavioral accuracy.
[0038] It should be further explained that the process of obtaining the positive prediction risk leakage probability and leakage risk path length in this embodiment includes: Based on the seasonal data breach monitoring map, a node extraction algorithm is used to extract preset initial risk nodes and target entity nodes from the current breach risk logic graph. Initial risk nodes include known threat source nodes or abnormal access subject nodes, and target entity nodes include core data asset nodes. The identification information of the initial risk nodes and target entity nodes is obtained to clarify the starting and ending points of the path traversal and to define the scope for subsequent path search. Based on the identification information of the initial risk node and the target entity node, a graph traversal algorithm using a breadth-first search algorithm is employed. Starting from the initial risk node, the algorithm visits the directly associated entity nodes in a hierarchical manner, recording the access status of each entity node to avoid repeated traversals, until the target entity node is reached. This yields a complete set of all paths from the initial risk node to the target entity node, which is used to comprehensively capture all possible paths from the starting point to the target, ensuring the completeness of path coverage. Based on the entity nodes in each complete path, preset basic attributes are extracted for each entity node. These preset basic attributes include the sensitivity level of the data asset, the access scope of the access subject, and the exposure status of the path node. The basic attributes are converted into risk weight values through an attribute quantification algorithm. The higher the sensitivity level, the greater the risk weight value; the lower the matching degree between the access scope and the operation, the greater the risk weight value; and the more open the exposure status, the greater the risk weight value. This yields the basic risk weight for each entity node. This weight is used to convert the inherent attributes of the node into calculable risk quantification indicators, providing basic parameters for calculating the node risk probability. Based on the corrected risk propagation probability between adjacent entity nodes in each path, and combined with the basic risk weight of each entity node, a forward accumulation algorithm is used to calculate the forward predicted risk leakage probability of the current node by multiplying the forward predicted risk leakage probability of the preceding node with the basic risk weight of the current node and the risk propagation probability between adjacent nodes, starting from the initial risk node. This calculation is iterated until the target entity node, thus obtaining the forward predicted risk leakage probability of each entity node in each path. This is used to quantify the risk contribution of each node in the path during the risk propagation process, reflecting the impact of the node on the leakage risk of the target entity. For each complete path, the number of entity nodes contained in the path is counted by a node counting algorithm; the risk propagation probability between two adjacent entity nodes in each path and the original length of the corresponding associated connection are extracted by a relationship extraction algorithm. The original length of the associated connection includes the physical transmission distance or the number of logical jumps. The risk propagation probability between adjacent nodes is multiplied and summed by a weighted summation algorithm to obtain the leakage risk path length of each path. This process, by clearly defining the starting and ending points of traversal, comprehensively searching the path, quantifying the risk probability of nodes, and statistically calculating the path length, enables the system to obtain the probability of positive risk leakage and the path length of leakage risk. It not only accurately depicts the specific risk contribution of each node in risk transmission, but also comprehensively reflects the complexity of risk transmission through the path length, providing core quantitative basis for the construction of positively visualized leakage judgment path sequences.
[0039] Based on the positive predicted risk leakage probability of each entity node in each path and the risk propagation probability between entity nodes, a weighted average algorithm is used to obtain the positive comprehensive predicted risk probability. Based on the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node and the corrected risk propagation probability between the corresponding entity nodes of each path, combined with the mapping rules constructed by the preset HSV color space and the leakage risk path length, a positive visual leakage discrimination path sequence is obtained. It should be further explained that the specific implementation process of the positive visualization leakage identification path sequence in this embodiment includes: Based on the positive prediction risk leakage probability of each entity node in each path between the access subject node and the target entity node, a threshold classification algorithm is used to divide the risk leakage probability value into multiple risk levels, and obtain the risk level classification result of each entity node; this is used to transform continuous probability values into discrete risk levels, providing a classification basis for subsequent color mapping.
[0040] Based on the corrected risk propagation probability between entity nodes, a normalization algorithm is used to map the propagation probability value to a preset range. Then, a mapping algorithm is used to map the normalized propagation probability value to a saturation value in the HSV color space, obtaining the saturation value of the connection between each entity node. This is used to convert the risk propagation intensity into the saturation feature of the color, so that the connection color can intuitively reflect the possibility of risk transmission.
[0041] Based on the path length of the leakage risk, a data transformation algorithm is used to transform the path length value to reduce the skewness of the data. Then, a normalization algorithm is used to map the transformed path length value to a preset range. Finally, a mapping algorithm is used to map the normalized path length value to a luminance value in the HSV color space to obtain the luminance value of each path. This is used to transform the path length feature into the luminance feature of the color, so that the path luminance can reflect the complexity of risk transmission.
[0042] Based on the risk level classification results of entity nodes, different risk levels are mapped to different color values through preset HSV color space mapping rules to obtain the color value of each entity node; the color feature is used to intuitively convert the risk level into color, so that the node color can quickly convey the degree of risk.
[0043] Based on node color values, connection saturation values, and path brightness values, a visualization rendering algorithm is used. With nodes as vertices and connections as edges, a layout algorithm is employed to determine node positions. Node colors are applied to node graphics, and connection saturation and path brightness are blended and applied to connection lines to obtain a preliminary visualized path map. This is used to transform abstract risk data into an intuitive graphical representation and construct the basic visual form of the path.
[0044] Based on the initial visualized path map, a sorting algorithm is used to sort all paths according to the positive comprehensive prediction risk probability. Then, a hierarchical layout algorithm is used to arrange the sorted paths in order of levels. The brightness value of each level of path remains consistent, but the color tone is gradually processed according to the positive comprehensive prediction risk probability to obtain a hierarchical visualized path map. This is used to highlight high-risk paths, so that the visualization results can be displayed in an orderly manner according to risk level, and the readability is enhanced.
[0045] Based on a hierarchical visualization path map, an animation generation algorithm is used to add color transition animations when the risk level changes for each node, arrow animations indicating the risk propagation direction for each connection, and hierarchical expansion animations based on risk probability from high to low for the entire path map, resulting in a dynamic and positively visualized leakage identification path sequence. This enhances the interactivity and dynamism of the visualization results, making the risk transmission process more vividly displayed.
[0046] It should be further explained that the process of obtaining the reverse prediction risk leakage probability and the reverse comprehensive prediction risk probability for each entity node in this embodiment includes: Based on the first leakage risk logical path, extract the data leakage risk type, access target entity and the data sensitivity level and exposure surface risk of the target entity, access path and leakage risk logical path identification accuracy. Based on the temporal distribution of access subject type, leakage risk logical path type, and the frequency of occurrence of corresponding entity nodes, a leakage type distribution probability function is obtained through a continuous-time Bayesian model. The Bayesian model, with seasonal time intervals as conditions, outputs the probability of occurrence of the corresponding leakage type within a preset time interval, which is used to describe the seasonal distribution of data leakage. It should be further explained that the process of obtaining the leakage type distribution probability function in this embodiment includes: Based on the first leakage risk logical path, a feature extraction algorithm is used to extract data leakage risk types from the path data, including internal unauthorized access and API abuse, access target entities, including core data assets and key system interfaces, and the data sensitivity level and exposure risk of the target entities. Data sensitivity levels include public, internal, and confidential, and exposure risk includes network exposure and interface vulnerability. At the same time, the access path and the identification accuracy of the leakage risk logical path are extracted to obtain a feature set containing risk type, target entity, data sensitivity level, exposure risk, access path, and identification accuracy. Based on the access subject type, including internal users, external partners, anonymous attackers, etc., the leakage risk logical path type, including direct access path, indirect detour path, etc., and the time distribution of the frequency of occurrence of corresponding entity nodes, the historical data is divided into preset seasonal time intervals through a time series segmentation algorithm. The seasonal time intervals include quarters, months, etc., to obtain statistical results of the occurrence frequency of different access subject types, leakage risk logical path types, and entity nodes in each time interval. Based on the frequency statistics of occurrence in each time interval, the conditional probability of different access subject types is calculated under a given seasonal time interval using a conditional probability calculation algorithm. The calculation formula is: P(access subject type|seasonal time interval) = number of times the access subject type occurs in the seasonal time interval / total number of times all access subject types occur in the seasonal time interval, thus obtaining the conditional probability distribution of access subject types. Based on the frequency statistics of occurrence in each time interval, the conditional probability of different leakage risk logical path types under a given seasonal time interval is calculated using a conditional probability calculation algorithm. The calculation formula is: P(Leakage risk logical path type|Seasonal time interval) = Number of times the leakage risk logical path type occurs in the seasonal time interval / Total number of times all leakage risk logical path types occur in the seasonal time interval, thus obtaining the conditional probability distribution of leakage risk logical path types. Based on the conditional probability distributions of access subject type and leakage risk logical path type, a joint probability algorithm is used to calculate the joint conditional probability of different combinations of access subject type and leakage risk logical path type under a given seasonal time interval. The calculation formula is: P(Access Subject Type ∩ Leakage Risk Logical Path Type | Seasonal Time Interval) = P(Access Subject Type | Seasonal Time Interval) × P(Leakage Risk Logical Path Type | Access Subject Type, Seasonal Time Interval), thus obtaining the joint conditional probability distribution. In this embodiment, P(Access Subject Type | Seasonal Time Interval) is implemented as follows: Access log data within historical seasonal time intervals is extracted based on the sliding window algorithm, and the conditional probability distribution of various access subjects in a specific seasonal time interval is calculated through frequency statistics and maximum likelihood estimation. P(Leakage Risk Logical Path Type | Access Subject Type, Seasonal Time Interval) is implemented as follows: Association rule mining, such as the Apriori algorithm, is used to analyze the co-occurrence patterns of access subject type, seasonal time interval, and risk logical path type in historical data, and the conditional probability of path type under a given subject type and seasonal conditions is dynamically updated using a Bayesian estimation model.
[0047] Based on the joint conditional probability distribution, and using Bayes' theorem, the posterior probability of different data breach types occurring under given seasonal time intervals and combinations of access subject type and leakage risk logical path type is calculated. The formula is: P(data breach type|access subject type, leakage risk logical path type, seasonal time interval) = [P(access subject type, leakage risk logical path type|data breach type, seasonal time interval) × P(data breach type|seasonal time interval)] / P(access subject type, leakage risk logical path type|seasonal time interval), thus obtaining the posterior probability distribution of each data breach type within different seasonal time intervals. In this embodiment, P(data breach type|access subject type, leakage risk logical path type, seasonal time interval) represents the probability of the target data breach type occurring when a specific access subject, leakage risk logical path, and seasonal characteristics are observed; P(access subject type, leakage risk logical path type|data breach type) is the likelihood probability, describing the probability of a specific data breach occurring within a given seasonal time interval. Under seasonal and leakage type conditions, the probability of the access subject experiencing the current leakage risk logical path type; P(data leakage type|seasonal time interval) represents the basic probability of the leakage type occurring in a specific season; P(access subject type, leakage risk logical path type|seasonal time interval) is the probability that the data leakage type corresponding to the current access subject is the current leakage risk logical path type under a certain seasonal time interval; it is used to calculate the posterior probability of various data leakage types occurring under a specific seasonal time interval, given access subject type and leakage risk logical path type. The calculation logic is as follows: the numerator is the product of the prior probability of the data leakage type in the seasonal time interval and the conditional probability of the combination of the data leakage type and the leakage risk logical path type in the seasonal time interval, and the denominator is the joint conditional probability of the combination of the access subject type and the leakage risk logical path type in the seasonal time interval. The final posterior probability is obtained by the ratio of the two, thereby quantifying the possibility of various data leakages occurring under different scenario combinations.
[0048] Based on the posterior probability distribution of each data leakage type within different seasonal time intervals, a Gaussian mixture model is used to fit the posterior probability distribution through a function fitting algorithm to obtain a continuous function with the seasonal time interval as the independent variable and the probability of data leakage type as the dependent variable, namely the leakage type distribution probability function. Based on the leakage type distribution probability function, the function parameters are verified using a parameter verification algorithm and cross-validation. The goodness of fit of the function is evaluated by calculating the mean square error between the predicted probability and the actual occurrence frequency. If the mean square error exceeds a preset threshold, the Gaussian mixture model parameters are readjusted until the mean square error meets the requirements. Finally, the leakage type distribution probability function is determined to ensure the accuracy and reliability of the function, so that the function can truly reflect the seasonal pattern of data leakage.
[0049] Based on the data sensitivity level and exposure risk of the target entity and the abnormal operation score of the current access subject, the basic risk value of the target entity is calculated by quantifying the sensitivity level into an access sensitivity weight value and the exposure risk into a vulnerability coefficient. Based on the product of the target entity's basic risk value, the probability of occurrence at the current time output by the probability distribution function of the leakage type, and the correctness of the leakage risk logical path identification, combined with the comprehensive leakage risk assessment model, the reverse comprehensive predicted risk probability of the target entity is obtained. It should be further explained that the process of obtaining the reverse comprehensive prediction risk probability of the target entity in this embodiment includes: Based on the access target entity information in the first leakage risk logic path, the data sensitivity level and exposure risk of the access target entity are obtained through attribute extraction algorithm, and the basic attribute set of the target entity is obtained; this is used to clarify the inherent risk characteristics of the target entity itself and provide raw attribute data for the calculation of the risk base value; Based on the basic attribute set of the target entity, a quantitative scoring algorithm is used to assign values to the sensitivity level and exposure risk according to the preset scoring criteria. The higher the sensitivity level, the higher the score; the greater the exposure risk, the higher the score. Then, a weighted summation algorithm is used to calculate the comprehensive score to obtain the basic risk value of the target entity. This is used to transform the qualitative attribute description into a quantitative risk benchmark value, which serves as a basic reference for reverse risk assessment. Based on the probability distribution function of leakage types, the algorithm inputs the current time information into the function through the time parameter input, calculates the probability of leakage types related to the target entity occurring within the current time interval, and obtains the probability of leakage types occurring at the current time; this is used to introduce the impact of seasonal time factors on risk and reflect the risk fluctuation characteristics at different time points.
[0050] Based on the first leakage risk logical path, the accuracy rate of leakage risk logical path identification corresponding to the path is extracted through a parameter extraction algorithm. This accuracy rate reflects the degree of matching between the current path and the historical real leakage path; it is used to measure the reliability of path risk judgment and provide a basis for risk probability correction.
[0051] Based on the target entity's basic risk value, the probability of leakage type occurring at the current time, and the accuracy of leakage risk logical path identification, the product of the three is calculated using a multiplication algorithm to obtain a preliminary risk product value. This value is then used to integrate the target entity's inherent risk, time-related risk, and path reliability to form a basic risk quantification result.
[0052] Based on the reverse path reasoning algorithm, the set of reverse paths to the target entity is obtained. The node contribution analysis algorithm is used to calculate the risk impact weight of each entity node in the reverse path on the target entity. The closer the association between the node and the target entity, the greater the weight. The risk contribution weight of each node is obtained. This is used to quantify the degree of impact of different nodes in the reverse path on the risk of the target entity, and to provide node-level risk input for the comprehensive leakage risk assessment model.
[0053] It should be further explained that this embodiment constructs a comprehensive leakage risk assessment model, which includes the following sub-steps: Based on the initial risk product value and the risk contribution weight of each node, a weighted fusion algorithm is used to sum the initial risk product value and the risk contribution weight of each node to obtain the intermediate risk value of the fused node's influence; this is used to integrate the overall path risk and the individual node risk to improve the comprehensiveness of risk assessment.
[0054] Based on the seasonal extended attributes of the target entity, including quarterly peak visits and quarterly abnormal visit frequency, the deviation rate between the current attribute value and the historical average value is calculated through a deviation correction algorithm. The deviation rate is then converted into a correction coefficient (if the deviation rate is positive, the correction coefficient is >1; if the deviation rate is negative, the correction coefficient is <1) to obtain the seasonal correction coefficient. This coefficient is used to introduce the impact of seasonal dynamic characteristics on risk, so that the assessment results conform to the seasonal risk pattern.
[0055] Based on the intermediate risk value and the seasonality correction coefficient, the product of the two is calculated using a product correction algorithm to obtain the corrected risk value; this is used to dynamically adjust the risk value to reflect the risk changes caused by seasonal fluctuations.
[0056] Based on the corrected risk value output by the comprehensive leakage risk assessment model, a normalization algorithm is used to map it to the 0-1 range to ensure the validity and comparability of the risk probability, thereby obtaining the reverse comprehensive prediction risk probability of the target entity. This is used to transform the comprehensive assessment results into standardized probability values, providing core risk indicators for reverse visualization of leakage identification path sequences.
[0057] Starting from the target entity, traverse the entity nodes in the path in reverse, and filter the effective paths between the accessing main node and the target entity node by combining the preset time constraints. It should be further explained that the process of obtaining the effective path in this embodiment includes: Based on the seasonal data leakage monitoring map, the location of the target entity node to be accessed is determined from the map by the node localization algorithm, and the unique identification information of the node is obtained; this is used to determine the starting point of the reverse traversal and to define the starting point for the path search.
[0058] Based on the unique identifier of the target entity node, a reverse traversal algorithm with a depth-first search strategy is used. Starting from the target entity node, the algorithm visits the upstream entity nodes connected to it along the relation edges, recording the nodes visited each time and their relationships, until the main entity node is reached or backtracking can no longer continue, thus obtaining a set of all possible reverse paths. This is used to comprehensively explore all potential paths from the target to the main entity and construct a candidate set for path filtering.
[0059] Based on each path in the reverse path set, an access timestamp for each entity node on the path is extracted from the graph using a timestamp extraction algorithm. The access timestamp includes the start time and end time of data access, thus obtaining the timestamp sequence of each node in each path. This is used to obtain the temporal information of node access on the path, providing a data foundation for temporal constraint filtering.
[0060] Based on preset time-series constraints, including access time order constraints and access time interval constraints, a time sequence verification algorithm is used to verify the timestamp sequence of each path to determine whether it meets the time-series constraints. For example, it determines whether the access end time of each upstream node in the path is earlier than the access start time of the downstream node, thereby obtaining a subset of paths that meet the time sequence constraints. This is used to ensure that the selected paths are logically reasonable in terms of time and to exclude invalid paths with time inconsistencies.
[0061] Based on a subset of paths that satisfy time sequence constraints, an interval calculation algorithm is used to calculate the access interval between adjacent nodes in the path. Then, a threshold comparison algorithm is used to compare the calculated interval with a preset interval threshold to filter out paths where the interval between all adjacent nodes meets the threshold requirement, thus obtaining a subset of paths that satisfy the interval constraints. This is used to ensure that the access interval between nodes in the path is within a reasonable range and to exclude abnormal paths with excessively long or short intervals.
[0062] Based on a subset of paths that satisfy the time interval constraint, a path integrity verification algorithm is used to check whether each path contains complete access subject nodes, intermediate transmission nodes, and target entity nodes, eliminating incomplete path segments and obtaining a set of structurally complete and valid paths. This ensures that the selected paths contain complete access links and avoids risk misjudgment due to incomplete paths.
[0063] Based on a complete set of valid paths, a duplicate path identification algorithm is used to compare the node sequences and relation sequences of each path, identify and eliminate duplicate paths, and obtain a set of valid paths without duplicates. This is used to reduce the interference of redundant paths on subsequent analysis and improve analysis efficiency.
[0064] Based on a set of valid paths without repetition, a path quality assessment algorithm is used to comprehensively consider factors such as path length, risk level of nodes facing target entity nodes (e.g., access authorization level of the current entity node facing the target entity node at the current time), and relationship strength. Each valid path is scored for data leakage risk, and the paths are sorted from high to low according to their data leakage risk scores to obtain the final valid path sequence. This sequence is used to assess the quality and prioritize the selected valid paths, providing ordered path data for subsequent risk analysis.
[0065] This process constructs a candidate set of paths through reverse traversal, and then performs multi-dimensional screening in conjunction with temporal constraint rules to finally obtain a valid path sequence. This process not only ensures the rationality of the path in terms of temporal logic, but also ensures the integrity and quality of the path structure. It provides an accurate and reliable path data foundation for constructing the reverse visualization leakage judgment path sequence, enabling subsequent risk analysis to be carried out based on real and valid access paths.
[0066] Based on the leakage type distribution probability function of each entity in the effective path, and combining the reverse comprehensive prediction risk probability of the target entity with the access risk probability between the two nodes in the corresponding effective path, the Monte Carlo inference algorithm is used to perform forward risk inference to obtain the reverse prediction risk leakage probability of each entity node in the effective path between the target entity node and any access subject.
[0067] It should be further explained that one specific implementation of forward risk reasoning using the Monte Carlo inference algorithm in this embodiment is as follows: Based on the set of effective paths, the identification information and connection relationships of each entity node in each effective path are extracted through a path parsing algorithm to obtain the path topology. This is used to clarify the path structure for risk propagation and provide a basic framework for subsequent Monte Carlo simulations.
[0068] Based on the leakage type distribution probability function, a parameter sampling algorithm is used to randomly sample each entity node in each valid path under the current time condition, generating the possible leakage types of the node in this simulation and their corresponding probability values, thus obtaining simulated leakage type samples for each node; this is used to transform the continuous probability distribution into a specific simulation scenario, providing random input for risk inference.
[0069] Based on the reverse comprehensive prediction of risk probability of the target entity, the probability value is transformed into the initial risk value in the Monte Carlo simulation through a probability mapping algorithm, which serves as the risk benchmark of the target entity in the simulation; it is used to determine the starting point strength of risk propagation and set the initial conditions for forward inference.
[0070] Based on the access risk probability between two nodes in an effective path, a weight allocation algorithm is used to assign the access risk probability of each connection as the risk propagation weight of that connection, thereby obtaining a risk propagation weight matrix for all connections in the path. This matrix is used to quantify the possibility of risk propagation between nodes and to provide a transfer probability for risk propagation.
[0071] Initialize Monte Carlo simulation parameters, including the number of simulations and convergence threshold. Using a random number generation algorithm, for each simulation, starting from the target entity node in the effective path, randomly select the next propagation node according to the risk propagation weight matrix until the visiting entity node is reached or the path terminates. Record the propagation path and node status for each simulation to obtain the risk propagation trajectory of a single simulation. This is used to simulate the propagation process of risk in the path through random sampling, generating a large number of possible propagation scenarios.
[0072] Based on the risk propagation trajectory of a single simulation, a risk accumulation algorithm is used to calculate the risk value of each entity node in the path during this simulation, combining the simulated leakage type samples of nodes and the initial risk value of the target entity. The risk value calculation method is: current node risk value = previous node risk value × current node leakage type probability × connection risk propagation weight. This yields the risk value of each node in a single simulation, which is used to quantify the change in the intensity of risk propagation along the path in a single simulation and to provide data points for statistical analysis.
[0073] Repeat the above steps until the preset number of simulations is reached. Use a statistical analysis algorithm to statistically analyze all simulation results, calculate the average risk value of each entity node in all simulations, and obtain the reverse prediction risk leakage probability of each entity node. This is used to eliminate the influence of randomness through a large number of simulations and obtain a stable and reliable risk probability estimate.
[0074] Based on the reverse prediction risk leakage probability of each entity node, a confidence interval calculation algorithm is used to calculate the confidence interval of each probability value, evaluate the uncertainty of the simulation results, and if the confidence interval width exceeds the preset threshold, the number of simulations is increased until the requirements are met. Finally, the reverse prediction risk leakage probability corresponding to each entity node in the effective path between the target entity node and any access subject is determined; this is used to evaluate the reliability of the risk probability estimation and ensure the statistical significance of the results.
[0075] S4. Based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, a simulation algorithm is used to simulate forward and reverse data leakage discrimination using a leakage risk threshold. This yields the data leakage risk discrimination results for the same access subject to the target entity under different leakage risk logical paths, as well as the corresponding simulated forward and reverse visualization leakage discrimination paths and the overlap between the forward and reverse paths. The data leakage risk discrimination results include the leakage risk probability and risk level. It should be further explained that the process of obtaining the overlap between the forward and reverse paths in this embodiment includes: Based on the forward visualization leakage discrimination path sequence, the entity node sequence and connection relationship in each forward path are extracted through the path parsing algorithm to obtain the forward path topology; Based on the reverse visualization leakage discrimination path sequence, the entity node sequence and connection relationship in each reverse path are extracted through the path parsing algorithm to obtain the reverse path topology; Based on the forward and reverse path topology, a path matching algorithm is used to match the forward and reverse paths from the same access subject to the target entity, identify all pairings of forward and reverse paths, and obtain a path pairing set. For each path pair in the path pairing set, the number of common entity nodes between the forward and reverse paths is calculated using a node intersection calculation algorithm to obtain the node intersection count. This count is used to quantify the degree of overlap between the forward and reverse paths at the node level, providing a node-dimensional indicator for overlap calculation.
[0076] For each pair of paths in the path pairing set, the number of common connections between the forward and reverse paths is calculated using a relation intersection algorithm to obtain the relation intersection count. This count is used to quantify the degree of overlap between the forward and reverse paths at the connection level, providing a relation dimension indicator for overlap calculation.
[0077] Based on the number of node intersections and the number of relationship intersections, the overlap of each pair of paths is calculated using a weighted summation algorithm. The calculation formula is: overlap = number of node intersections × node weight coefficient + number of relationship intersections × relationship weight coefficient, where the sum of the node weight coefficient and the relationship weight coefficient is 1, thus obtaining the overlap value of each pair of paths. Based on the overlap value of each path pair, a clustering analysis algorithm is used to cluster all path pairs according to their overlap values, dividing them into different categories such as high overlap, medium overlap, and low overlap, thus obtaining the overlap classification results. Based on the dynamic data leakage monitoring map, through simulation algorithm, for different leakage risk logical paths, the positive comprehensive prediction risk probability in the forward path and the reverse comprehensive prediction risk probability in the reverse path are fused and calculated. The fusion method is weighted average, and the weight is determined according to the path reliability to obtain the fused risk probability under different leakage risk logical paths. Based on the fusion risk probability, the risk level classification algorithm is used to map the fusion risk probability to a preset risk level range to obtain the risk level under different leakage risk logic paths; Based on the overlap classification results and the fusion of risk probability and risk level, the correlation analysis algorithm is used to analyze the correlation between different overlap categories and risk probability and risk level, and obtain the correlation model between overlap and risk indicators. Based on the association model, a visualization rendering algorithm is used to overlay and display the positive and negative visualization leakage discrimination paths under different leakage risk logic paths. Different colors are used to mark the overlapping and non-overlapping parts, and the corresponding overlap value, fusion risk probability and risk level are marked to obtain a visual display of the positive and negative path overlap effect.
[0078] S5. Based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type, determine the core risk nodes of data leakage; It should be further explained that one implementation process for determining the core risk nodes of data leakage in this embodiment includes: Based on the positive and negative visualization leakage discrimination path, the entity node sequence and connection relationship in each path are extracted through the path parsing algorithm to obtain a node set containing node identifier, node type and node risk value; Based on the overlap of forward and reverse paths, each node in the node set is marked using a node marking algorithm. If a node exists in both the forward and reverse paths, it is marked as an overlapping node; otherwise, it is marked as a non-overlapping node, thus obtaining the marked node set. Based on the time interval of the corresponding leakage risk type, the time matching algorithm is used to filter out the nodes with access records or status changes within the time interval from the marked node set to obtain the time matching node set; Based on the time-matched node set, the nodes are sorted from high to low according to their risk values using a risk value sorting algorithm to obtain a risk-sorted node list. Based on the risk-ranked node list, a threshold filtering algorithm is used to set a risk value threshold and filter out nodes with risk values higher than the risk value threshold to obtain a set of high-risk nodes. Based on the set of high-risk nodes, a weight is assigned to each node using an overlap-weighted algorithm. The weight calculation formula is: Node weight = Node risk value × (1 + Overlap coefficient), where the overlap coefficient is determined based on whether the node is an overlapping node (the coefficient for overlapping nodes is 0.5, and the coefficient for non-overlapping nodes is 0), thus obtaining the weighted set of nodes. Based on the weighted set of nodes, clustering analysis algorithms are used to cluster the nodes according to features such as location, type, and risk value to obtain multiple node clusters; Based on node clusters, a centrality calculation algorithm is used to calculate the centrality index of nodes within each node cluster. The centrality index includes degree centrality, betweenness centrality, etc. The node with the highest centrality index is selected as the cluster center node, and the set of center nodes of each node cluster is obtained. This is used to determine the core nodes of each risk cluster area. These nodes usually play a key role in risk propagation. Based on the central node set of each node cluster, the central nodes are further screened using an expert rule matching algorithm, combined with preset core node judgment rules, such as whether a node is a critical data asset or a high-privilege access point, to obtain the final set of core risk nodes for data leakage. Based on the final set of core risk nodes for data leakage, the contribution of each core node to the overall risk is evaluated using an impact assessment algorithm. The contribution calculation formula is: Node contribution = Node weight × (1 + Percentage of nodes in the cluster), thus obtaining the risk contribution ranking of each core node.
[0079] This embodiment employs the overlap of forward and reverse paths and the visualization of leakage identification paths, combined with the time interval of the corresponding leakage risk type, to determine core nodes. The motivation for this approach is that the forward path infers possible leakage paths from the source, while the reverse path traces back possible access paths from the target. The overlap often represents the most likely true leakage path. Overlapping nodes are identified as key nodes in risk propagation from both forward and reverse reasoning perspectives, thus increasing their credibility. Simultaneously, filtering based on the time interval of the leakage risk type ensures that the selected nodes are consistent with the current risk event in terms of time, eliminating interference from historically irrelevant nodes. The principle is based on the bidirectional consistency and temporal correlation of the data leakage process. The true leakage path conforms to both the forward propagation logic from source to target and the reverse tracing logic from target to source, and it matches the time interval of the risk event. Therefore, by comprehensively considering the overlap of forward and reverse paths, visualization path characteristics, and time interval constraints, core risk nodes in the data leakage process can be effectively identified, providing precise target positioning for risk prevention and control.
[0080] S6. Based on the core risk nodes of data leakage, combined with the data leakage risk judgment results and the preset anti-leakage strategy execution library, a corresponding defense strategy is generated, and the generated defense strategy is fed back to the simulation algorithm for adjustment. The change in leakage risk probability of the corresponding positive and negative visualized leakage judgment path after adjustment is monitored. It should be further noted that the anti-leakage strategy execution library in this embodiment is obtained by those skilled in the art based on historical leakage risk types and the constructed defense strategy, combined with knowledge graph algorithms and text generation algorithms. S7. Based on the changes in the leakage risk probability of the corresponding positive and negative visualization leakage judgment paths after adjustment, determine the secondary data leakage defense strategy until the leakage risk probability and risk level of the access target entity meet the data leakage security requirements.
[0081] From a risk identification perspective, this embodiment first constructs a standardized entity-relationship model based on historical data breach cases and real-time data through algorithms such as text cleaning, entity recognition, and relationship extraction. Combined with time series decomposition to extract seasonal features, an initial seasonal data breach monitoring map is formed. This process solves the problem of traditional monitoring relying solely on real-time data while ignoring historical patterns and seasonal fluctuations. By quantifying the seasonal risk coefficients of entity nodes, the map can capture periodic risk patterns such as peak access at the end of the quarter and abnormal operations during holidays, providing a more accurate basis for subsequent risk assessments that aligns with actual business cycles. Building upon this foundation, a community algorithm aggregates highly correlated nodes, and combined with a risk logic graph and triggering rules formed from historical leakage paths, further optimization yields a seasonal data leakage monitoring map. This not only achieves precise positioning of risk groups but also solves the identification challenge of diverse attack methods with repetitive core logic through standardized subgraph matching, significantly improving the comprehensiveness of risk scenario coverage. In the risk assessment phase, a dual-path reasoning mechanism is employed. The forward path, based on a Hidden Markov Model and a forward infectious disease model, combines entity state transition probabilities with risk propagation patterns to quantify the forward comprehensive prediction risk probability of each node. The reverse path constructs a leakage type distribution probability function using a Bayesian model and calculates the reverse comprehensive prediction risk probability using a comprehensive leakage risk assessment model. This two-way reasoning not only verifies the risk transmission chain from source to target and from target to source, but also handles the uncertainty of risk propagation through Monte Carlo simulation, making the assessment results more robust. Simultaneously, the forward and reverse path overlap analysis identifies the most likely true leakage path through the intersection calculation of nodes and relationships, solving the problem of single-direction reasoning being easily interfered with by isolated anomalies. Higher overlap indicates stronger path credibility, providing an objective basis for risk priority ranking. Furthermore, this embodiment combines overlapping nodes of forward and reverse paths, high-risk values, and time matching constraints, using cluster analysis and centrality calculation to locate key nodes. These nodes are both risk hubs recognized by both forward and reverse reasoning and are highly correlated with the current time interval, avoiding the resource waste caused by blind coverage in traditional defenses. For example, in internal privilege escalation scenarios, overlapping nodes may be key accounts or core databases with elevated privileges, whose risk contribution is far higher than other nodes, enabling defense measures to accurately target key links in risk transmission. This embodiment also calls matching strategies from the anti-leakage strategy execution library based on core nodes and risk discrimination results, and verifies the adjustment effect through simulation until the risk meets security requirements. This closed-loop optimization not only connects historical experience with real-time defense, but also dynamically updates strategies based on changes in risk probability to cope with the iterative evolution of attack methods. For example, in response to API abuse risks, the initial strategy might be to restrict unauthorized API calls. If simulations reveal that risks still exist, dark web interaction monitoring can be added to form a multi-layered defense, significantly improving the flexibility and effectiveness of the defense.
[0082] In summary, this solution comprehensively addresses the problems of incomplete coverage, high false alarm rate, and delayed response in traditional data breach monitoring by integrating seasonal characteristics, two-way reasoning verification, precise node positioning, and dynamic strategy optimization. It achieves a transformation from passive defense to proactive prediction, and from extensive response to precise prevention and control. Ultimately, while ensuring data security, it improves the utilization efficiency of defense resources and provides a systematic solution for data breach governance in complex business scenarios.
[0083] Example 2: Please see Figure 2 Another embodiment of the present invention provides a data leakage judgment system, comprising: a monitoring and classification module, a forward discrimination module, a reverse discrimination module, a simulation module, a risk determination module, an adjustment monitoring module, and an adjustment feedback module; The monitoring and classification module is used to respond to the preset seasonal data leakage monitoring map, collect the access status data and interaction information of each entity in the dynamic data leakage monitoring map in real time, and combine them with the preset leakage risk logic map to obtain the first leakage risk logic path. The positive discrimination module, based on the data leakage risk type, the access subject and the abnormal score of the access subject's operation behavior, the access path and the recognition accuracy of the leakage risk logical path, combined with the hidden Markov algorithm and the positive infectious disease model, obtains a positive visual leakage discrimination path sequence. The reverse discrimination module, based on the data leakage risk type, the access target entity and the data sensitivity level and exposure risk of the access target entity, the access path and the accuracy of the leakage risk logical path identification, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, obtains the reverse visualized leakage discrimination path sequence.
[0084] The simulation module, based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, uses a simulation algorithm combined with a leakage risk threshold to perform forward and reverse data leakage discrimination simulation, and obtains the data leakage risk discrimination result of the same access subject to the target entity under different leakage risk logical paths, as well as the corresponding simulated forward and reverse visualization leakage discrimination paths and the overlap of forward and reverse paths; The risk determination module determines the core risk nodes of data leakage based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type. The adjustment monitoring module generates corresponding defense strategies based on the core risk nodes of data leakage, the data leakage risk judgment results, and the preset anti-leakage strategy execution library. It then feeds the generated defense strategies back to the simulation algorithm for adjustment and monitors the change in leakage risk probability of the corresponding positive and negative visualized leakage judgment paths after adjustment. The adjustment feedback module determines a secondary data leakage defense strategy based on the change in leakage risk probability of the corresponding positive and negative visualized leakage discrimination paths after adjustment, until the leakage risk probability and risk level of the accessed target entity meet the data leakage security requirements.
Claims
1. A data leakage detection method, characterized in that, include: In response to the preset seasonal data leakage monitoring map, the access status data and interaction information of each entity in the dynamic data leakage monitoring map are collected in real time and combined with the preset leakage risk logic map to obtain the first leakage risk logic path. The first leakage risk logical path includes data leakage risk type, access subject and abnormal operation behavior score of access subject, access target entity and data sensitivity level and exposure risk of access target entity, access path and leakage risk logical path identification accuracy; Based on the data leakage risk type, access subject and abnormal operation behavior score of the access subject, access path and leakage risk logical path identification accuracy, combined with the hidden Markov algorithm and positive infectious disease model, a positive visual leakage discrimination path sequence is obtained. Based on the data leakage risk type, the target entity being accessed and the data sensitivity level and exposure risk of the target entity being accessed, the access path and the accuracy of the leakage risk logical path identification, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, a reverse visualization leakage discrimination path sequence is obtained. The leakage type distribution probability function is constructed by combining the time distribution of access subject type, leakage risk logical path type and corresponding occurrence frequency with a continuous-time Bayesian model, and is used to measure the seasonal distribution of different types of leakage time.
2. The data leakage judgment method as described in claim 1, characterized in that, The leakage detection method also includes: The positive visualization leakage detection path sequence includes the positive predicted risk leakage probability of each entity node in the positive visualization leakage detection path for accessing the target entity, as well as the leakage risk path length and the positive comprehensive predicted risk probability. The reverse visualization leakage detection path sequence includes the reverse predicted risk leakage probability and the reverse comprehensive prediction risk probability of each entity node in the reverse visualization leakage detection path for accessing the target entity; Based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, a simulation algorithm is used to simulate forward and reverse data leakage discrimination using a leakage risk threshold. This yields the data leakage risk discrimination results for the same access subject to the target entity under different leakage risk logical paths, as well as the overlap between the simulated forward and reverse visualization leakage discrimination paths and the forward and reverse paths. The data leakage risk discrimination results include the leakage risk probability and risk level.
3. The data leakage judgment method as described in claim 2, characterized in that, The leakage detection method also includes: Based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type, the core risk nodes of data leakage are determined. Based on the core risk nodes of data leakage, combined with the data leakage risk judgment results and the preset anti-leakage strategy execution library, the corresponding defense strategy is generated, and the generated defense strategy is fed back to the simulation algorithm for adjustment. The change status of leakage risk probability of the corresponding positive and negative visualized leakage judgment path after adjustment is monitored. Based on the changes in the leakage risk probability of the corresponding positive and negative visualization leakage identification paths after the adjustment, a secondary data leakage defense strategy is determined until the leakage risk probability and risk level of the accessed target entity meet the data leakage security requirements.
4. The data leakage judgment method as described in claim 3, characterized in that, The construction of the seasonal data breach monitoring map includes: Based on data breach scenarios, a pre-defined text decomposition algorithm is used to determine the entity composition and data transmission logic of the corresponding scenario. This results in a set of entity nodes including data assets, access subjects, threat sources, and path nodes, as well as relationship connection types including access permissions, data flow, and threat propagation. Basic attributes and seasonal extended attributes are also pre-defined for each entity node. The relationship connection types include access permission relationships, data flow relationships, threat propagation relationships, and seasonal association sensitive relationships; the preset basic attributes include the sensitivity level of data assets and the access scope of the access subject; the seasonal extended attributes include the quarterly access peak of data assets and the quarterly abnormal operation frequency of the access subject; the seasonal association sensitive relationship is used to adjust the quarterly authorization strength coefficient of the access target entity according to the access time attribute of the access subject.
5. The data leakage judgment method as described in claim 4, characterized in that, The construction of the seasonal data breach monitoring map also includes: Obtain operation logs, permission change records, and vulnerability intelligence logs corresponding to historical leakage cases within a preset time period under different data leakage scenarios. Separate trend items and seasonal items through time series decomposition algorithm, extract seasonal characteristics of different entity nodes and seasonal correlation leakage risk characteristics between them and the remaining entity nodes, and obtain the seasonal risk coefficient of different entity nodes under the corresponding scenario. Based on entity nodes, the corresponding relationship connection types between different entity nodes, and the seasonal risk coefficient in the current scenario, an initial seasonal data leakage monitoring map is constructed using graph algorithms. Based on the initial seasonal data breach monitoring map, a seasonal data breach monitoring map is obtained by combining a breach risk logic graph obtained based on historical breach paths and preset trigger rules through a community algorithm.
6. The data leakage judgment method as described in claim 5, characterized in that, The process of obtaining the positive visualization leakage discrimination path sequence includes: Based on the data leakage risk type, abnormal score of the access subject's operation behavior, access path, and leakage risk logical path identification accuracy in the first leakage risk logical path, the product of the abnormal score of the access subject's operation behavior and the leakage risk logical path identification accuracy is mapped to the corrected observation state. The data leakage risk type, seasonal extension attribute, and associated time-sensitive weight of the entity node are set as hidden states. A hidden Markov model is trained to obtain the leakage behavior state transition probability matrix. Based on the state transition probability matrix of leakage behavior combined with the positive infectious disease model, and combined with the real-time access paths marked as susceptible, infected or recovered, the corrected risk transmission probability between entity nodes is obtained.
7. The data leakage judgment method as described in claim 6, characterized in that, The process of obtaining the positive visualization leakage discrimination path sequence also includes: Based on the preset basic attributes of each entity node in the current leakage risk logic graph in the seasonal data leakage monitoring map, the graph traversal algorithm is used to obtain the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node, starting from the preset initial risk node in the current leakage risk logic graph. At the same time, the number of entity nodes contained in the path and the risk propagation probability between every two entity nodes in each path and the original length of the corresponding associated connection are counted to obtain the leakage risk path length. Based on the positive predicted risk leakage probability of each entity node in each path and the risk propagation probability between entity nodes, a weighted average algorithm is used to obtain the positive comprehensive predicted risk probability. Based on the positive predicted risk leakage probability of each entity node in each path between the access subject node and the target entity node, and the corrected risk propagation probability between the corresponding entity nodes of each path, combined with the mapping rules constructed by the preset HSV color space and the leakage risk path length, a positive visual leakage discrimination path sequence is obtained.
8. The data leakage judgment method as described in claim 7, characterized in that, The process of obtaining the reverse prediction risk leakage probability and the reverse comprehensive prediction risk probability for each entity node for accessing the target entity includes: Based on the data sensitivity level and exposure risk of the target entity and the abnormal operation score of the current accessing subject, the base risk value of the target entity is calculated. The reverse comprehensive prediction risk probability of the target entity is obtained by multiplying the basic risk value of the target entity, the probability of occurrence at the current time output by the probability distribution function of the leakage type, and the accuracy of the identification of the leakage risk logical path. Starting from the target entity, traverse the entity nodes in the path in reverse, and filter the effective paths between the accessing main node and the target entity node by combining the preset time constraints. Based on the leakage type distribution probability function of each entity in the effective path, and combining the reverse comprehensive prediction risk probability of the target entity with the access risk probability between the two nodes in the corresponding effective path, the reverse prediction risk leakage probability of each entity node in the effective path between the target entity node and any access subject is obtained.
9. A data leakage detection system, used to implement the data leakage detection method according to any one of claims 1-8, characterized in that, include: Monitoring and classification module, forward discrimination module, and reverse discrimination module; The monitoring and classification module is used to respond to the preset seasonal data leakage monitoring map, collect the access status data and interaction information of each entity in the dynamic data leakage monitoring map in real time, and combine them with the preset leakage risk logic map to obtain the first leakage risk logic path. The positive discrimination module, based on the data leakage risk type, the access subject and the abnormal score of the access subject's operation behavior, the access path and the recognition accuracy of the leakage risk logical path, combined with the hidden Markov algorithm and the positive infectious disease model, obtains a positive visual leakage discrimination path sequence. The reverse discrimination module, based on the data leakage risk type, the access target entity and the data sensitivity level and exposure risk of the access target entity, the access path and the accuracy of the leakage risk logical path identification, combined with the reverse path reasoning algorithm and the comprehensive leakage risk assessment model configured with the leakage type distribution probability function, obtains the reverse visualized leakage discrimination path sequence.
10. The data leakage detection system as described in claim 9, characterized in that, The leakage detection system also includes: a simulation module, a risk determination module, an adjustment monitoring module, and an adjustment feedback module; The simulation module, based on the forward visualization leakage discrimination path sequence combined with the reverse visualization leakage discrimination path sequence and the dynamic data leakage monitoring map, uses a simulation algorithm combined with a leakage risk threshold to perform forward and reverse data leakage discrimination simulation, and obtains the data leakage risk discrimination result of the same access subject to the target entity under different leakage risk logical paths, as well as the corresponding simulated forward and reverse visualization leakage discrimination paths and the overlap of forward and reverse paths; The risk determination module determines the core risk nodes of data leakage based on the overlap of forward and reverse paths and the forward and reverse visualized leakage identification paths, combined with the time interval of the corresponding leakage risk type. The adjustment monitoring module generates corresponding defense strategies based on the core risk nodes of data leakage, the data leakage risk judgment results, and the preset anti-leakage strategy execution library. It then feeds the generated defense strategies back to the simulation algorithm for adjustment and monitors the change in leakage risk probability of the corresponding positive and negative visualized leakage judgment paths after adjustment. The adjustment feedback module determines a secondary data leakage defense strategy based on the change in leakage risk probability of the corresponding positive and negative visualized leakage discrimination paths after adjustment, until the leakage risk probability and risk level of the accessed target entity meet the data leakage security requirements.
Citation Information
Patent Citations
Establishment method of data leakage prevention system
CN119377995A
Vision-based industrial safety production full-period monitoring method and system
CN121258269A
Risk behavior identification method and apparatus, storage medium, and computer device
WO2025086900A1