Privacy leakage tracing method based on data consanguinity map

By generating a data kinship map in the blockchain network and predicting the probability of leakage, the problem of difficult privacy data leakage is solved, and efficient and accurate privacy data leakage investigation and risk assessment are achieved.

CN120387167APending Publication Date: 2025-07-29LINGSHU TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510439825.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, privacy data leakage is difficult to quickly locate the source, the investigation process lacks accuracy, and data security management lacks effective means to accurately assess the risk of leakage.

Method used

By obtaining any node in the blockchain network, blood relationship analysis is performed to generate a data blood relationship map, trace the starting point of privacy leakage, generate a visual trace map, and use the leakage probability prediction function to generate a check priority list for traceability.

Benefits of technology

It realizes efficient traceability and precise investigation of privacy data leakage, and improves data security management capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387167A_ABST
    Figure CN120387167A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy leakage traceability method based on a data consanguinity map, and relates to the technical field of privacy leakage traceability, and the method comprises the steps: obtaining any node in a block chain network, and enabling the node to correspond to any enterprise; performing blood relationship analysis on the plurality of privacy data of any enterprise to generate a data blood relationship map; obtaining leaked privacy data, recording the data as a tracing starting point, and performing tracing analysis to obtain a tracing record; generating a visual traceability chart; performing prediction analysis on a first tracing path in the plurality of tracing paths to obtain a first leakage probability; and generating a troubleshooting priority list, and performing leakage troubleshooting and traceability on the private data. The technical problems that in the prior art, it is difficult to quickly locate the source for privacy data leakage, the checking process is lack of accuracy, and data security management is lack of effective means for accurately evaluating the leakage risk are solved, and the technical effects of achieving efficient tracing and accurate checking of privacy data leakage and improving the data security management capacity are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of privacy leakage traceability, and particularly to a privacy leakage traceability method based on a data lineage graph. Background Art

[0002] In today's digital age, data has become a key asset for the development of enterprises and society, and the security of privacy data is of utmost importance. In the existing data security management system, many severe challenges are faced. On the one hand, traditional data storage and management methods lack effective tracking of the entire process of data flow. When privacy data is leaked, it is difficult to quickly and accurately trace the leakage source, resulting in the inability to take timely measures to contain the further leakage of data and the expansion of damage. On the other hand, existing investigation means often lack accuracy, mostly relying on experience and post-event review, with low efficiency and prone to missing key clues. In addition, it is difficult for enterprises to quantitatively analyze the possibility of data leakage and reasonably determine the investigation priority, making data security management in a passive defensive state.

[0003] The prior art has technical problems such as difficulty in quickly locating the source of privacy data leakage, lack of accuracy in the investigation process, and lack of effective means in data security management to accurately assess the leakage risk. Summary of the Invention

[0004] The present application provides a privacy leakage traceability method based on a data lineage graph, which is used to solve the technical problems in the prior art that it is difficult to quickly locate the source of privacy data leakage, the investigation process lacks accuracy, and data security management lacks effective means to accurately assess the leakage risk.

[0005] In view of the above problems, the present application provides a privacy leakage traceability method based on a data lineage graph, and the method includes:

[0006] Obtain any node in the blockchain network, and the any node corresponds to any enterprise, wherein the blockchain network is a data storage network of a trusted data space; perform lineage analysis on multiple privacy data of the any enterprise according to a predetermined lineage analysis mechanism to generate a data lineage graph; obtain the leaked privacy data in a privacy leakage event, denoted as the traceability starting point, and perform trace analysis on the data lineage graph based on the traceability starting point to obtain a trace record; generate a visual trace graph according to the trace record, wherein the visual trace graph includes multiple trace paths; introduce a leakage probability prediction function to perform prediction analysis on a first trace path among the multiple trace paths to obtain a first leakage probability; generate an investigation priority list with the first leakage probability as a constraint, and perform leakage investigation and traceability on the privacy data according to the investigation priority list.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0008] Obtain any node in the blockchain network, and the any node corresponds to any enterprise; perform blood relationship analysis on multiple privacy data of the any enterprise according to a predetermined blood relationship analysis mechanism to generate a data blood relationship graph; obtain the privacy data leaked in a privacy leakage event, denoted as the traceability starting point, and perform traceability analysis on the data blood relationship graph based on the traceability starting point to obtain a trace record; generate a visual traceability graph according to the trace record; introduce a leakage probability prediction function to perform prediction analysis on the first trace path in the multiple trace paths to obtain a first leakage probability; generate a troubleshooting priority list with the first leakage probability as a constraint, and perform leakage troubleshooting and traceability on the privacy data according to the troubleshooting priority list. The technical effect of realizing efficient traceability and accurate troubleshooting of privacy data leakage and improving data security management ability is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of a privacy leakage traceability method based on a data blood relationship graph provided by an embodiment of this application;

[0011] Figure 2 It is a flowchart of generating a data blood relationship graph in a privacy leakage traceability method based on a data blood relationship graph provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] This application provides a privacy leakage traceability method based on a data blood relationship graph, which is used to solve the technical problems in the prior art that it is difficult to quickly locate the source of privacy data leakage, the troubleshooting process lacks accuracy, and there is a lack of effective means for data security management to accurately evaluate the leakage risk.

[0013] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.

[0014] Embodiment, such as Figure 1As shown in the figure, the present application provides a method for tracing the source of privacy leakage based on a data lineage graph, and the method includes:

[0015] Step S100: Obtain any node in the blockchain network, and the any node corresponds to any enterprise, where the blockchain network is a data storage network of a trusted data space.

[0016] Specifically, under the architecture of the trusted data space, the blockchain network undertakes the important function of data storage. At this time, a node is randomly selected from this blockchain network, and each node corresponds to a specific enterprise. This selection process is based on the distributed characteristics of the blockchain network and is implemented through a corresponding node indexing algorithm or data access interface. The randomly selected any node contains various data information stored on the blockchain by the corresponding enterprise, and these information are the basic data sources for subsequent analysis of the blood relationship of privacy data and tracing the source of privacy leakage. By obtaining this node, it is possible to delve into the data field of a specific enterprise and provide a data entry and support for a series of subsequent analysis and tracing operations around the enterprise's privacy data, ensuring that the entire tracing process has a clear goal and data basis.

[0017] Step S200: Analyze the blood relationship of multiple privacy data of the any enterprise according to a predetermined blood relationship analysis mechanism to generate a data lineage graph.

[0018] Specifically, after obtaining the blockchain network node corresponding to any enterprise, according to the predetermined blood relationship analysis mechanism, conduct a blood relationship analysis on multiple privacy data of this enterprise. First, extract the first privacy data and find its first generation source, and then use this source as the starting point of propagation and determine the first propagation path including multiple propagation nodes according to the predetermined mechanism. Then set the first propagation node in the first propagation path as the secondary propagation starting point to further obtain the second propagation path. Finally, combine the first propagation path and the second propagation path to form the first blood relationship graph of the first privacy data. Many such first blood relationship graphs together constitute a complete data lineage graph, which clearly presents the transfer trajectory and mutual association of the enterprise's privacy data and provides a key basic framework for subsequent tracing analysis of privacy leakage.

[0019] Step S300: Obtain the privacy data leaked in the privacy leakage event, denoted as the tracing starting point, and perform a tracing analysis on the data lineage graph based on the tracing starting point to obtain a tracing record.

[0020] Specifically, when a privacy leak occurs, the first thing to do is to accurately obtain the leaked private data from the data involved in the leak, and this part of the data will be marked as the starting point for tracing back to the source. With the help of the already constructed data lineage map, starting from the just-determined tracing back to the starting point, a comprehensive and in-depth tracing back analysis of the data flow path is carried out according to the data lineage relationship recorded in the map. During the tracing process, the transmission and processing of the data in each link will be sorted out in detail, including information such as the flow direction of the data, the nodes handled, and the corresponding operations. Through such meticulous analysis, a detailed tracing record is finally generated. This record contains the entire data flow process starting from the tracing back to the starting point, providing a crucial basis for the subsequent identification of the specific cause of the privacy leak and the responsible party.

[0021] Step S400: generating a visual traceability graph according to the traceability record, wherein the visual traceability graph includes multiple traceability paths.

[0022] Specifically, data visualization technology is used to visualize the data information in the traceability records, generating a visual traceability diagram. This visual traceability diagram is presented in an intuitive and easy-to-understand graphical format, containing multiple traceability paths. These traceability paths are generated based on the flow of data in the traceability records. They clearly illustrate the propagation path of the leaked private data from the source through various links and nodes. Through these paths, relevant personnel can clearly see the flow of data at different stages, quickly sorting out the scope and key links of the possible data leak, and providing an intuitive and important reference for subsequent in-depth analysis of the cause of the private data leak, determining the responsible party, and formulating preventive measures.

[0023] Step S500: introducing a leakage probability prediction function to perform prediction analysis on a first tracing path among the multiple tracing paths to obtain a first leakage probability.

[0024] Specifically, after completing the previous steps to generate a visual traceability graph, we first construct a set of processing steps for the first traceability path. This set covers the first and second steps, which are consecutive processing steps. Next, we compare the timestamps of the first step with the second step to determine the dwell time of the first step. We then combine the operation type and dwell time of the first step to form a processing feature set for the first step. Finally, we apply this processing feature set to the leakage probability prediction function for computational analysis. Through the function's operation, we ultimately derive the first probability of data leakage along this first traceability path, providing a key quantitative basis for subsequent investigation priority determination.

[0025] Step S600: Generate a priority list for investigation based on the first leakage probability as a constraint, and perform leakage investigation and tracing on the private data according to the priority list for investigation.

[0026] Specifically, find the first entities corresponding to the first link. These entities may be personnel, systems, or other entities involved in data processing. Then, based on the first leakage probability, sort these first entities in descending order. The higher the probability, the higher the corresponding entity's ranking in the list, thus forming a troubleshooting priority list. Before sorting, obtain the historical operation records of the first entities, including the set of historical operation types and the set of historical residence durations. By determining whether the current first operation type belongs to the set of historical operation types and the belonging situation of the first residence duration in the residence duration judgment set after removing the maximum value in the set of historical residence durations, comprehensively obtain the first influence factor of the first entity, and then adjust the first leakage probability with this factor as the weight to make the sorting more accurate. After the sorting is completed to generate the troubleshooting priority list, read the predetermined leakage probability threshold. If the first leakage probability corresponding to a certain first entity does not reach this threshold, remove it from the list. Finally, based on the troubleshooting priority list, start from the entity with the highest possibility and conduct troubleshooting and analysis one by one to comprehensively trace the root cause of the privacy data leakage to determine the specific link and responsible entity where the leakage occurred.

[0027] In a possible implementation manner, as Figure 2 shown, step S200 further includes:

[0028] Step S210: Extract the first privacy data from the multiple privacy data and obtain the first generation source of the first privacy data.

[0029] Step S220: According to the predetermined blood relationship analysis mechanism, use the first generation source as the propagation starting point to obtain a first propagation path, where the first propagation path includes multiple propagation nodes.

[0030] Step S230: Use the first propagation node among the multiple propagation nodes as the secondary propagation starting point to obtain a second propagation path.

[0031] Step S240: Combine the first propagation path and the second propagation path to obtain the first blood relationship map of the first privacy data and form the data blood relationship map.

[0032] Specifically, first, when the system obtains multiple privacy data from an enterprise, it will select one of these data as the first privacy data. This selection can be based on the chronological order of the data, importance indicators, or random selection, etc., specifically depending on the predetermined blood relationship analysis mechanism. After determining the first privacy data, along the generation context of the data, through in-depth mining and analysis of the data records, operation logs, and relevant metadata in the blockchain network, the first generation source of the data is searched for. In a blockchain environment, the generation of data is often accompanied by a series of operation records. For example, operation information such as data creation, collection, and entry is completely recorded on the chain. Based on these records, combined with key elements such as timestamps and operation subject information, the formation process of the data is traced backwards until the initial generation point of the first privacy data is found, and this point is the first generation source. Obtaining the first generation source is the basis for constructing the data blood relationship, laying a key starting point for subsequent analysis of the data propagation path and blood relationship map, helping to more clearly sort out the data flow trajectory and attribution relationship, and thus providing strong support for privacy leakage tracing.

[0033] According to the predetermined blood relationship analysis mechanism, search for the propagation path of the first privacy data from the first generation source. The predetermined blood relationship analysis mechanism is a set of pre-set rules and algorithm systems, which are designed based on the characteristics of data operations in the blockchain network, the logic of data flow, and the requirements of privacy data management. This mechanism uses various data operation information recorded on the blockchain, such as timestamps, operation subjects, operation contents, etc. generated by data replication, sharing, modification, and other operations, to accurately identify and track the propagation direction of the data. Taking the first generation source as the propagation starting point, comprehensively and carefully scan and analyze the data records on the blockchain according to the predetermined blood relationship analysis mechanism. It will check which operations involve the first privacy data after the first generation source. Each operation involving this data corresponds to a potential propagation node, and these nodes are connected in sequence according to the chronological order and operation logic. For example, if at a certain moment, the first privacy data is copied from the generation source to another storage location, then the node corresponding to this copy operation will be included in the propagation path; if the data is then shared with other enterprises or departments, the node corresponding to the sharing operation will also be recorded. After such a series of analysis and sorting, a first propagation path containing multiple propagation nodes is finally formed. This path clearly shows the various links and locations that the first privacy data has experienced from the generation source through a series of operations and transfers. It provides important clues and basis for subsequent further analysis of the data usage situation and investigation of possible privacy leakage links, and is one of the key steps in constructing a complete data blood relationship map.

[0034] Select the first propagation node from multiple propagation nodes of the first propagation path. This selection is not arbitrary and is based on factors such as the importance of the propagation node, the sensitivity of the operation, or the chronological order. For example, if a propagation node involves cross-enterprise sharing of data or sensitive operations, it is very likely to be selected as the first propagation node. Once the first propagation node is determined, the predetermined lineage analysis mechanism is enabled again with it as the new secondary propagation starting point to deeply explore the data operation records related to the first propagation node and subsequent in the blockchain network. Just like in the construction process of the first propagation path, according to the detailed data operation meta-information on the blockchain, such as the timestamps, operation subjects, etc. generated by operations like data movement, copying, modification, access, etc., to clarify the subsequent flow of data starting from the first propagation node. Through careful sorting and logical association of these data operation records, a series of new propagation nodes derived from the first propagation node are identified and connected in the order of data flow, thus forming the second propagation path. This path further shows the subsequent propagation trajectory of the first private data based on the first propagation path, provides richer information for accurately positioning the possible links of privacy leakage, and helps to more efficiently carry out the work of tracing privacy leakage.

[0035] Sort out and merge all the propagation nodes and the data flow relationships between nodes in the first propagation path and the second propagation path, remove duplicate nodes and redundant relationships, and make the data flow logic clearer. Starting from the first generation source, fuse information such as the data flow direction and the connection relationships between nodes in the first propagation path and the second propagation path, and draw a map reflecting the whole process of the first private data from generation to subsequent flow, which is the first lineage map of the first private data. With similar processing of multiple private data, integrate the first lineage maps of numerous first private data, and finally form a complete data lineage map, which comprehensively shows the flow network of enterprise private data and provides an important basis for subsequent tracing of privacy leakage.

[0036] In a possible implementation manner, step S500 further includes:

[0037] Step S510: Form a set of processing links of the first trace path, where the set of processing links includes a first link and a second link, and the first link and the second link are consecutive processing links.

[0038] Step S520: Compare the first timestamp of the first link with the second timestamp of the second link to obtain the first residence duration of the first link.

[0039] Step S530: Based on the first operation type and the first residence duration of the first link, form a first set of processing characteristics of the first link.

[0040] Step S540: Analyze the first processing feature set according to the leakage probability prediction function to obtain the first leakage probability.

[0041] Specifically, first, by parsing the storage data structure of the visualization traceability graph, the detailed information of the first trace path is located. With the help of data mining algorithms, traverse the data processing records of the first trace path. According to the time sequence and operation association relationship, identify two adjacent and consecutive data processing stages, and define them as the first link and the second link respectively. Use the list data structure of a programming language (such as Python) to integrate the relevant information of these two links (such as link number, operation type, data identifier involved, etc.) to construct a processing link set. During the integration process, standardize the information of each link to ensure the consistency of the data format, which is convenient for subsequent steps to analyze and operate on the processing link set.

[0042] Extract the timestamp data corresponding to the first link and the second link respectively from the processing link set. The timestamp is used as an identifier to accurately record the data processing time. The first timestamp marks the moment when the first link starts to process the data, and the second timestamp records the time point when the second link starts to process the data. Use a time calculation tool to subtract the first timestamp from the second timestamp. Through this calculation of the time difference, the first residence duration of the first link can be accurately obtained. This duration reflects the time span during which the data stays and is processed in the first link, providing a key time dimension quantification index for subsequent evaluation of the security, stability, and potential risks of the data processing process in the first link, and playing an important role in subsequent analysis of the data leakage probability.

[0043] Clarify the first operation type implemented in the first link, which includes various operations such as data reading, writing, modification, transmission, deletion, etc. Different operation types reflect the processing methods and potential risks that the data experiences in this link. At the same time, combine the first residence duration of the first link calculated previously. This duration reflects the processing time of the data in this link and can indirectly reflect the complexity of the operation and possible abnormal situations. Subsequently, integrate the first operation type and the first residence duration to jointly form the first processing feature set of the first link with these two key elements. This feature set comprehensively and accurately depicts the data processing characteristics of the first link, providing key data support for subsequent analysis of the privacy leakage possibility of this link using the leakage probability prediction function.

[0044] After completing the construction of the first processing feature set of the first link, call the leakage probability prediction function to analyze it. This leakage probability prediction function is Among them, P(L|F,T) refers to the first leakage probability under the first operation type F and the first residence duration T. β0 refers to the intercept term, representing the baseline probability of the model. β1 refers to the first value of the first operation type F on the first leakage probability, and β2 refers to the second value of the first residence duration T on the first leakage probability. Through function operations and comprehensive consideration of these factors, a specific value, that is, the first leakage probability, is finally output. This probability quantifies the possibility of privacy leakage occurring in the first link of the first traceability path, providing important data support for subsequent privacy data leakage investigation and traceability based on the level of risk.

[0045] In a possible implementation manner, step S540 further includes:

[0046] Step S541: The expression of the leakage probability prediction function is:

[0047]

[0048] Among them, P(L|F,T) refers to the first leakage probability under the first operation type F and the first residence duration T. β0 refers to the intercept term, representing the baseline probability of the model. β1 refers to the first value of the first operation type F on the first leakage probability, and β2 refers to the second value of the first residence duration T on the first leakage probability.

[0049] Specifically, the leakage probability prediction function is Among them, β0 refers to the intercept term, which represents the baseline probability of the model in the logistic regression model. From the principle of the logistic regression model, it is used to model binary classification problems and predict the probability that the dependent variable belongs to a certain category by analyzing a series of independent variables. In this privacy leakage tracing scenario, the dependent variable is whether the privacy data is leaked (which can be regarded as two categories: leaked or not leaked), and the independent variables are the first operation type F and the first residence duration T. And β0 is the intercept term, which is the probability of privacy data leakage predicted by the model when all independent variables take the value of 0. It is a basic reference value of the model, representing a potential possibility of privacy data leakage without considering the influence of specific operation types and residence durations and other factors. This baseline probability provides a benchmark for comprehensively considering the influence of other factors on the leakage probability in the future. When calculating the first leakage probability, β0 plays a key starting reference role and is an indispensable part of the whole model. β1 reflects the first value of the first operation type F on the first leakage probability, that is, the difference in the influence degree of different operation types (such as data reading, modification, transmission, etc.) on the privacy leakage probability; β2 reflects the second value of the first residence duration T on the first leakage probability, indicating the influence degree of the length of data residence time in a certain link on the privacy leakage risk. By substituting the first operation type and the first residence duration of the first link into this function for calculation, a quantitative first leakage probability value can be obtained, providing a key quantitative basis for subsequent privacy data investigation and tracing according to the level of leakage risk, making the tracing work more scientific and efficient.

[0050] In a possible implementation manner, step S541 further includes:

[0051] Step S5411: Read a predetermined value evaluation function, and perform value quantification evaluation and analysis on the first operation type and the first leakage probability according to the predetermined value evaluation function to obtain the first value.

[0052] The expression of the predetermined value evaluation function is:

[0053]

[0054] Among them, IG(F) refers to the first value of the first operation type F, H(L) refers to the entropy of the first leakage probability, p i is the probability that the first leakage probability takes the i-th value, n is the total number of all values of the first leakage probability, H(L|F) is the conditional entropy of the first operation type F, p(x) is the probability that the first operation type F takes the value x, and H(L|F) = x) is the entropy of the first leakage probability under the condition of F = x.

[0055] Specifically, after obtaining the first leakage probability, this step first reads a predetermined value evaluation function and performs a value quantification evaluation and analysis on the first operation type and the first leakage probability. The function expression is: Among them, IG(F) represents the first value of the first operation type F, which reflects the importance of the operation type's impact on privacy data leakage. H(L) is the entropy of the first leakage probability, used to measure the uncertainty of the first leakage probability; p i is the probability that the first leakage probability takes the i-th value, and n is the total number of all values of the first leakage probability; H(L|F) is the conditional entropy of the first operation type F, p(x) is the probability that the first operation type F takes the value x, and H(L|F=x) is the entropy of the first leakage probability under the condition of F=x. Through the operation of these parameters, the quantitative evaluation of the relationship between the first operation type and the first leakage probability is realized, so as to obtain a more reference-worthy first value and provide more accurate data support for the subsequent privacy leakage traceability work.

[0056] In a possible implementation manner, step S600 further includes:

[0057] Step S610: Obtain the first entity corresponding to the first link.

[0058] Step S620: Sort in descending order based on the first leakage probability to obtain the sorting result of the first entity, denoted as the investigation priority list.

[0059] Specifically, first, collect relevant data from various record sources in the data processing flow. These record sources include system operation logs, database transaction records, and data access audit reports, etc. The system operation log details the time, operator identification, and operation content of each operation step. The database transaction record can reflect the data flow situation in different links. The data access audit report clarifies the data access behaviors of each entity. Use data mining and analysis techniques to process the collected data to write a special script program to parse the operation log and screen out the operation records related to the first link. At the same time, use the database query language to extract the data flow information of the first link from the database transaction record. Then, according to the preset rules and mapping relationships, match the screened information with the predefined entity information. This entity information can be stored in a special user information database or permission management system, which contains the detailed identification and attributes of entities such as personnel, departments, and system processes. Finally, through matching and verification, determine the first entity corresponding to the first link. In this way, the acquisition of the first entity can be completed efficiently and accurately, providing a key basis for the subsequent privacy leakage investigation and traceability work.

[0060] Arrange all the obtained first entities in descending order according to the values of the first leakage probabilities. The first leakage probability reflects the likelihood of privacy data leakage during the data processing process. Arrange the first entities in descending order, so that those first entities with a greater likelihood of privacy data leakage during the data processing will be ranked at the top. Organize and record this arrangement result, and the formed list is the troubleshooting priority list. Subsequently, based on this troubleshooting priority list, targeted leakage troubleshooting and tracing work can be carried out on the privacy data processing links involved in each first entity in descending order, giving priority to checking the links with higher leakage risks, and improving the efficiency and accuracy of the tracing work.

[0061] In a possible implementation manner, step S610 further includes:

[0062] Step S611: Obtain the historical operation records of the first entity, where the historical operation records include a set of historical operation types and a set of historical residence durations.

[0063] Step S612: Determine whether the first operation type belongs to the set of historical operation types to obtain a first judgment result.

[0064] Step S613: Remove the longest historical residence duration from the set of historical residence durations to obtain a set of residence duration judgments.

[0065] Step S614: Determine whether the first residence duration belongs to the set of residence duration judgments to obtain a second judgment result.

[0066] Step S615: Obtain the first influence factor of the first entity according to the first judgment result and the second judgment result.

[0067] Step S616: Adjust the first leakage probability with the first influence factor as the weight.

[0068] Specifically, obtain the historical operation records of the first entity from the system's operation record database or log file. These records cover all the operations of the entity in the past data processing process and are organized into a set of historical operation types and a set of historical residence durations. The set of historical operation types records various operation types that the entity has performed in the past, while the set of historical residence durations records the time spent on each operation.

[0069] Accurately extract the set of historical operation types from the data storage of historical operation records. This set details all the operation types of the first entity in the past. Then, compare the first operation type with each element in the set of historical operation types one by one. If the first operation type exactly matches a certain element in the set of historical operation types, then the first judgment result is "yes", indicating that this first operation type belongs to the historical operation scope of the first entity, meaning that the current operation is an operation habitually performed by the first entity and is an operation that conforms to its normal business process. On the contrary, if the first operation type is not found to have a matching item in the set of historical operation types, the first judgment result is "no", indicating that the current operation may be an abnormal operation of the first entity, with a potential risk of privacy data leakage. The first judgment result is of great significance for subsequent evaluation of the impact of the first entity on privacy data leakage in the current link and is a key basis for further analysis and decision-making.

[0070] In the obtained historical operation records, the set of historical residence durations records the time spent on each operation by the first entity during past operations. However, in actual operations, there will be some special situations that cause the residence duration of a certain operation to be abnormally long. This may be caused by external interference, system failures, or other accidental factors. Such abnormally long residence durations will mislead the subsequent judgment of whether the first residence duration in the current first link is normal. Therefore, in order to make the subsequent judgment more accurate and reliable, it is necessary to remove the longest historical residence duration from the set of historical residence durations. The specific approach is to traverse and compare all the residence duration data in the set of historical residence durations, find the one with the largest value, that is, the longest historical residence duration, and then remove it from the set of historical residence durations. After such processing, the remaining historical residence duration data forms the residence duration judgment set. This residence duration judgment set can more truly reflect the operation residence duration pattern of the first entity under normal circumstances and provide a more effective reference basis for subsequent judgment of whether the first residence duration conforms to the norm.

[0071] Accurately obtain the specific value of the first residence duration from the relevant records of the first link. At the same time, extract the residence duration judgment set from the storage location. This set is obtained after removing the longest historical residence duration and represents the residence duration set under the normal operation of the first subject. Then, compare the first residence duration with each element in the residence duration judgment set. If the first residence duration is the same as a certain element in the residence duration judgment set or within the duration range covered by this set, then the second judgment result is "yes", which means that the residence duration of the current first link is within the normal operation duration range of the first subject, the residence duration of this operation is relatively regular, and the risk of abnormal data leakage is relatively low. On the contrary, if the first residence duration does not match all elements in the residence duration judgment set and is not within the duration range covered by this set, the second judgment result is "no", indicating that the first residence duration exceeds the normal operation time range of the first subject, the residence duration of this operation is abnormal, and there may be a potential risk of privacy data leakage. The second judgment result is very crucial in the entire privacy data leakage tracing process. It works together with the first judgment result to provide an important basis for determining the first influencing factor of the first subject and helps to more accurately evaluate the data leakage risk.

[0072] Determine the first influencing factor of the first subject based on these two judgment results. The first judgment result reflects whether the first operation type belongs to the historical operation type set of the first subject, reflecting the regularity of the operation type; the second judgment result indicates whether the first residence duration is within the residence duration judgment set after removing the longest historical residence duration, reflecting the rationality of the operation duration. If the first judgment result is "yes" and the second judgment result is also "yes", it means that both the type and duration of the current operation conform to the historical operation habits of the first subject, and the first subject's current operation belongs to the normal business process, with less impact on privacy data leakage. At this time, a relatively low first influencing factor can be assigned, such as 0.2. If the first judgment result is "yes" but the second judgment result is "no", that is, the operation type is regular but the operation duration is abnormal, there may be certain potential risks, and the first influencing factor can be set to 0.5. If the first judgment result is "no" but the second judgment result is "yes", indicating that the operation type is abnormal but the duration is normal, there is also a risk of privacy data leakage, and the first influencing factor can be set to 0.6. If both the first judgment result and the second judgment result are "no", it means that both the operation type and the operation duration deviate from the historical habits of the first subject, and there is very likely a risk of privacy data leakage. At this time, a relatively high first influencing factor should be assigned, such as 0.8. By comprehensively considering the first judgment result and the second judgment result in this way, the first influencing factor of the first subject can be accurately obtained, which provides an important basis for subsequent adjustment of the first leakage probability and more accurate assessment of the privacy data leakage risk.

[0073] Review the previously calculated first leakage probability, which reflects the possibility of data leakage without considering the operation characteristics of the first entity. The first impact factor is obtained by comprehensively considering the first operation type and the matching degree between the first residence duration and the historical operation habits of the first entity, and is used to measure the impact degree of the first entity on the data leakage risk. Next, use the first impact factor as a weight to perform an operation with the first leakage probability. If the first impact factor is relatively high, for example, 0.8, it means that the current operation of the first entity is quite different from the historical habits and there is a relatively high risk. Then the first leakage probability will be correspondingly increased on the original basis, that is, through multiplication operation, so that the adjusted first leakage probability is closer to the real risk level. On the contrary, if the first impact factor is relatively low, such as 0.2, it indicates that the current operation of the first entity is relatively normal and has little impact on the leakage risk. The adjusted first leakage probability will be reduced on the original basis to more accurately reflect the actual situation. By adjusting the first leakage probability with the first impact factor as the weight, the assessment of the data leakage risk can be made more in line with the actual situation, providing a more reliable basis for the subsequent investigation priority and tracing work formulated according to the leakage probability.

[0074] In a possible implementation manner, step S620 further includes:

[0075] Step S621: Read a predetermined leakage probability threshold.

[0076] Step S622: If the first leakage probability does not meet the predetermined leakage probability threshold, remove the first entity from the investigation priority list.

[0077] Specifically, in the entire privacy data security management system, the predetermined leakage probability threshold is an important indicator set in advance, and its role is to provide a benchmark for subsequent judgment of the data leakage risk level. The setting of this threshold usually comprehensively considers various factors, such as the analysis of the enterprise's past data leakage cases, the data security standards of the industry where it is located, the data security risk level that the enterprise itself can bear, etc. In actual operation, this threshold will be stored in a specific location. It may be in a dedicated configuration file, clearly recorded in text form for easy viewing and modification by management personnel; it may also be stored in a specific table field of the database and managed together with other relevant security configuration information for quick invocation by the system during operation. Access the location storing this threshold according to the preset path, and accurately extract this predetermined leakage probability threshold through the corresponding reading program or instruction, so as to prepare for the next operation of comparing the first leakage probability with it and judging whether to remove the first entity.

[0078] Compare the previously obtained first leakage probability with the just-read predetermined leakage probability threshold. The predetermined leakage probability threshold is a risk boundary determined based on multiple factors such as enterprise data security policies, industry standards, and historical data, representing the degree of data leakage risk that the enterprise can accept. If the first leakage probability is lower than this threshold, it means that in the current operating scenario of the first entity, the possibility of privacy data leakage is relatively low and within the tolerable risk range of the enterprise. Considering the efficiency of the privacy data leakage investigation work, to avoid spending too much effort on low-risk links, the first entity will be automatically removed from the investigation priority list, and only those entities with higher data leakage risks will be retained in the investigation priority list. Subsequently, the investigation work can be more targeted at these key objects, and a detailed inspection and analysis will be carried out on them first, greatly improving the work efficiency and accuracy of privacy data leakage tracing.

[0079] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be beneficial.

[0080] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0081] This specification and the drawings are only exemplary descriptions of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.

Claims

1. A privacy leakage tracing method based on a data lineage graph, characterized in that, Including: Obtain any node in the blockchain network, and the any node corresponds to any enterprise, where the blockchain network is a data storage network of a trusted data space; Perform blood relationship analysis on multiple privacy data of the any enterprise according to a predetermined blood relationship analysis mechanism to generate a data blood relationship graph; Obtain the privacy data leaked in a privacy leakage event, denoted as the tracing starting point, and perform tracing analysis on the data blood relationship graph based on the tracing starting point to obtain a tracing record; Generate a visual tracing graph according to the tracing record, where the visual tracing graph includes multiple tracing paths; Introduce a leakage probability prediction function to perform prediction analysis on the first tracing path among the multiple tracing paths to obtain a first leakage probability; Generate a troubleshooting priority list with the first leakage probability as a constraint, and perform leakage troubleshooting and tracing on the privacy data according to the troubleshooting priority list.

2. The method for tracing the source of privacy leakage based on the data lineage graph according to claim 1, characterized in that, Performing blood relationship analysis on multiple privacy data of the any enterprise according to a predetermined blood relationship analysis mechanism to generate a data blood relationship graph includes: Extract the first privacy data from the multiple privacy data, and obtain the first generation source of the first privacy data; According to the predetermined blood relationship analysis mechanism, use the first generation source as the propagation starting point to obtain a first propagation path, where the first propagation path includes multiple propagation nodes; Use the first propagation node among the multiple propagation nodes as the secondary propagation starting point to obtain a second propagation path; Combine the first propagation path and the second propagation path to obtain the first blood relationship graph of the first privacy data, and form the data blood relationship graph.

3. The privacy leakage traceability method based on the data lineage graph according to claim 1, wherein Introducing a leakage probability prediction function to perform prediction analysis on the first tracing path among the multiple tracing paths to obtain a first leakage probability includes: Form a processing link set of the first tracing path, where the processing link set includes a first link and a second link, and the first link and the second link are consecutive processing links; Compare the first timestamp of the first link with the second timestamp of the second link to obtain the first residence duration of the first link; Based on the first operation type of the first link and the first residence duration, form a first processing feature set of the first link; Analyze the first processing feature set according to the leakage probability prediction function to obtain the first leakage probability.

4. The privacy leakage tracing method based on the data lineage graph according to claim 3, wherein The expression of the leakage probability prediction function is: Where P(L|F,T) refers to the first leakage probability under the first operation type F and the first residence duration T, β0 refers to the intercept term, representing the model baseline probability, β1 refers to the first value of the first operation type F on the first leakage probability, and β2 refers to the second value of the first residence duration T on the first leakage probability.

5. The method for tracing the source of privacy leakage based on the data lineage graph according to claim 4, wherein It also includes: Read a predetermined value evaluation function, and perform value quantification evaluation analysis on the first operation type and the first leakage probability according to the predetermined value evaluation function to obtain the first value; The expression of the predetermined value evaluation function is: Among them, IG(F) refers to the first value of the first operation type F, H(L) refers to the entropy of the first leakage probability, p i refers to the probability that the first leakage probability takes the i-th value, n refers to the total number of all values of the first leakage probability, H(L|F) refers to the conditional entropy of the first operation type F, p(x) refers to the probability that the first operation type F takes the value x, and H(L|F = x) refers to the entropy of the first leakage probability under the condition of F = x.

6. The method for tracing privacy leakage based on a data lineage graph according to claim 3, wherein Generate a troubleshooting priority list with the first leakage probability as a constraint, and perform leakage troubleshooting and tracing on the privacy data according to the troubleshooting priority list, including: Obtain the first entity corresponding to the first link; Arrange in descending order based on the first leakage probability to obtain the sorting result of the first entity, denoted as the troubleshooting priority list.

7. The method for tracing privacy leakage based on the data lineage graph according to claim 6, wherein Before arranging in descending order based on the first leakage probability to obtain the sorting result of the first entity, denoted as the troubleshooting priority list, it further includes: Obtain the historical operation records of the first entity, where the historical operation records include a set of historical operation types and a set of historical residence durations; Judge whether the first operation type belongs to the set of historical operation types to obtain a first judgment result; Remove the longest historical residence duration from the set of historical residence durations to obtain a set of residence duration judgments; Judge whether the first residence duration belongs to the set of residence duration judgments to obtain a second judgment result; Obtain the first influence factor of the first entity according to the first judgment result and the second judgment result; Adjust the first leakage probability with the first influence factor as the weight.

8. The method for tracing privacy leakage based on the data lineage graph according to claim 6, wherein After arranging in descending order based on the first leakage probability to obtain the sorting result of the first entity, denoted as the troubleshooting priority list, it further includes: Read a predetermined leakage probability threshold; If the first leakage probability does not meet the predetermined leakage probability threshold, remove the first entity from the troubleshooting priority list.

Citation Information

Cited By

  • Data full life cycle privacy protection encryption management method and system

    CN122197080A

  • Data privacy protection encryption management method and system throughout the entire lifecycle

    CN122197080B