Privacy data processing method and device
By semantic analysis of SQL statements in the source database, the blood relationship information between private data is obtained, and the graph data pattern is created in the graph database, the complex problem of privacy data management in the existing technology is solved, and refined and efficient privacy data management is achieved.
Patent Information
- Application Number
- CN202110932430.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-08-13
AI Technical Summary
Existing privacy data management solutions are difficult to achieve refined and efficient management of privacy data, especially when there are multi-level relationships between privacy data, the backtracking process is complicated.
By semantic analysis of SQL statements involving private data executed by the source database, the blood relationship information between the private data is obtained, and a graph data pattern representing these relationships is created in the target graph database, and auxiliary information is stored for more granular management.
It realizes refined management of private data, improves management efficiency and convenience, and can make more convenient and quick use of the blood relationship between private data for error correction, traceability and compliance judgment.
Smart Images

Figure CN113672977B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a method and device for processing privacy data. Background Art
[0002] When privacy data circulates between different business units, its processing and use may cause compliance issues. Therefore, privacy data needs to be managed so that it can be quickly traced and corrected.
[0003] At present, most traditional privacy data management solutions use regular expressions, syntax tree parsing or related keyword matching to obtain the upstream and downstream relationships between privacy data, and store the upstream and downstream relationships between privacy data through relational databases, and then use table queries to achieve the management and backtracking functions of related privacy data. However, this method is essentially a "point" oriented data management method with a coarse management granularity. Different privacy data are separated, and when there are multi-level relationships between privacy data, the backtracking process of the source of privacy data is highly complex.
[0004] Based on this, there is an urgent need for a privacy data processing solution that can achieve refined and efficient management of privacy data. Summary of the invention
[0005] The purpose of the embodiments of this specification is to provide a privacy data processing method and device to achieve refined and efficient management of privacy data.
[0006] In order to achieve the above purpose, the embodiments of this specification adopt the following technical solutions:
[0007] In a first aspect, a privacy data processing method is provided, comprising:
[0008] Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0009] Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0010] Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0011] Based on the atlas data pattern and the auxiliary information, the auxiliary information is stored in the target graph database.
[0012] In a second aspect, a privacy data processing device is provided, comprising:
[0013] A first acquisition unit is configured to acquire a structured query SQL statement involving private data and auxiliary information related to the private data, which is executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0014] A parsing unit, which performs semantic parsing on the acquired SQL statement to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0015] A creation unit, based on the blood relationship information, creates a graph data model representing the blood relationship between the private data in a target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0016] A storage unit stores the auxiliary information in the target graph database based on the graph data pattern and the auxiliary information.
[0017] According to a third aspect, an electronic device is provided, including:
[0018] Processor; and
[0019] a memory arranged to store computer executable instructions which, when executed, cause the processor to:
[0020] Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0021] Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0022] Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0023] Based on the atlas data pattern and the auxiliary information, the auxiliary information is stored in the target graph database.
[0024] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device performs the following operations:
[0025] Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0026] Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0027] Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0028] Based on the atlas data pattern and the auxiliary information, the auxiliary information is stored in the target graph database.
[0029] The scheme of the embodiments of this specification obtains the blood relationship information between the private data in the source database, including the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields where the private data is located and the data tables, by semantically parsing the SQL statements executed on the source database and involving the private data. The obtained blood relationship information can more finely reflect the blood relationship between the private data; based on the blood relationship information, a graph data model (Schema) representing the blood relationship between the private data is created in the target graph database, and the nodes in the graph data model represent fields or data tables, and the edges in the graph data model represent the association relationship between connected nodes. Further, based on the graph data model, the auxiliary information is stored in the target graph database, so that the blood relationship between the private data can be stored in the form of a knowledge graph, so as to realize the management of the private data from "point" to "surface", and then it is more convenient and quick to use the blood relationship between the private data to implement error correction, traceability, compliance judgment, etc. on the private data, thereby improving the efficiency and convenience of the management of the private data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0031] Figure 1 A schematic diagram of the overall solution flow of a privacy data processing method according to an embodiment of this specification;
[0032] Figure 2 A schematic diagram of the overall solution flow of a privacy data processing method according to another embodiment of the present specification;
[0033] Figure 3 This is a schematic diagram of the overall solution flow of a privacy data processing method according to another embodiment of the present specification;
[0034] Figure 4 A flowchart of a privacy data processing method provided for one embodiment of this specification;
[0035] Figure 5 A schematic diagram of a graph data mode provided for one embodiment of the present specification;
[0036] Figure 6 A flowchart of a privacy data processing method provided for another embodiment of this specification;
[0037] Figure 7 A schematic diagram of the structure of a privacy data processing device provided for one embodiment of this specification;
[0038] Figure 8 A schematic diagram of the structure of an electronic device provided for one embodiment of the present specification. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this document.
[0040] Description of some concepts:
[0041] Metadata: also known as intermediary data or relay data, is data about data. It is mainly used to describe the properties of data and to support functions such as indicating storage location, historical data, resource search, and file records. Metadata is a kind of electronic catalog. In order to achieve the purpose of cataloging, it is necessary to describe and collect the content or characteristics of the data, thereby achieving the purpose of assisting data retrieval.
[0042] Bloodline relationship: used to describe the upstream and downstream relationships between data.
[0043] Graph database: a non-relational database that uses graph theory to store relationship information between entities, such as the relationship between people in a social network. Compared with relational databases, the unique design of graph databases can make up for the defects of relational databases when storing "relational" data, such as complex, slow, and unexpected queries.
[0044] Private Data: Secret data refers to data that one does not want to be known by others or irrelevant persons. From the perspective of the owner of the privacy, private data can be divided into personal private data and common private data. Personal private data includes information that can be used to locate or identify individuals (such as phone numbers, addresses, credit card numbers, etc.) and sensitive information (such as personal health conditions, financial information, important company documents, etc.); common private data mainly includes family privacy, such as family annual income, etc. The leakage and abuse of private data can easily lead to various personal and public safety issues.
[0045] Knowledge Graph: It is called knowledge domain visualization or knowledge domain mapping map in the library and information industry. It is a series of various graphs that show the development process and structural relationship of knowledge. It uses visualization technology to describe knowledge resources and their carriers, and to mine, analyze, construct, draw and display knowledge and the related connections between them.
[0046] As mentioned above, most traditional privacy data management solutions use regular expressions, syntax tree parsing, or related keyword matching to obtain the upstream and downstream relationships between privacy data, and store the upstream and downstream relationships between privacy data through a relational database, as shown in Table 1 below, and then use table query to implement the management and backtracking functions of related privacy data. However, this method is essentially a "point" oriented data management method with a coarse management granularity, and different privacy data are separated. When there are multi-level relationships between privacy data, the backtracking process of the source of privacy data is highly complex.
[0047] Table 1
[0048]
[0049]
[0050] To this end, the embodiments of this specification aim to provide a privacy data processing solution, by semantically parsing SQL statements executed on the source database and involving privacy data, to obtain the blood relationship information between the privacy data in the source database, including the association relationship between the fields where the privacy data is located, the association relationship between the data tables where the privacy data is located, and the association relationship between the fields where the privacy data is located and the data tables. The obtained blood relationship information can more finely reflect the blood relationship between the privacy data; based on the blood relationship information, a graph data model (Schema) representing the blood relationship between the privacy data is created in the target graph database, and the nodes in the graph data model represent fields or data tables, and the edges in the graph data model represent the association relationship between connected nodes. Further, based on the graph data model, auxiliary information is stored in the target graph database, so that the blood relationship between the privacy data can be stored in the form of a knowledge graph, so as to realize the management of privacy data from "point" to "surface", and then it can be more convenient and quick to use the blood relationship between the privacy data to implement error correction, traceability, compliance judgment, etc. on the privacy data, thereby improving the efficiency and convenience of privacy data management.
[0051] It should be understood that the privacy data processing method provided in the embodiments of this specification can be executed by an electronic device or software installed in an electronic device, and can be specifically executed by a terminal device or a server device.
[0052] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0053] For ease of understanding, the following is a brief introduction to the overall solution process of the privacy data processing method provided in the embodiments of this specification. Figures 1 to 3 , is a schematic diagram of the overall solution flow of the privacy data processing method of the embodiment of this specification. Figure 1 As shown, the overall solution process of the privacy data processing method of the embodiment of this specification includes solutions at the data layer, the middle layer, the knowledge graph layer, etc.
[0054] In the data layer, the source database stores business data, metadata information of business data, operation log information, Structured Query Language (SQL) execution information, etc., which provide input for the middle layer and the knowledge graph layer at the upper layer. Among them, the source database can include, for example, but is not limited to at least one of the following databases: Open Data Processing Service (ODPS) database, Data Resource Platform (DataQ), MySQL, Qracle, etc. The operation log information is used to record additional information such as the owner and execution time of the SQL statement executed on the source database, the SQL execution information is used to record the SQL statement executed on the source database, etc., and the metadata information of the business data is used to describe the data table where the business data is located and the properties of each field in the data table, etc.
[0055] In the middle layer, by performing a privacy data scan on the business data stored in the source database, the privacy-related business data (hereinafter referred to as "privacy data") stored in the source database can be identified; by parsing the metadata database information stored in the source database, the attribute information of the privacy data (such as the data type of the privacy data, etc.) can be obtained; by parsing the SQL execution information of the data layer, the association relationships between the fields, between the data tables, and between the fields and the data tables described in the SQL statement can be obtained; by parsing the operation log information stored in the source database, additional information such as the owner of the SQL statement executed on the source database and the execution time can be obtained.
[0056] Furthermore, if Figure 1 and Figure 2 As shown, by performing blood relationship discovery on the association relationship between the fields where the privacy data is located, the association relationship between the data tables where the privacy data is located, and the association relationship between the fields where the privacy data is located and the data tables, the blood relationship information between the privacy data can be obtained.
[0057] Alternatively, if Figure 1 and Figure 2As shown in the figure, considering that source databases such as ODPS themselves store the stock association relationship information between table-level privacy data, the blood relationship information obtained through blood relationship discovery can also be merged with the stock association relationship information stored in the source database. In order to improve the accuracy and reliability of the obtained blood relationship information, the merged blood relationship information is further corrected.
[0058] In the knowledge graph layer, by performing graph calculation on the blood relationship information between private data, a graph data model (Schema) for characterizing the blood relationship between private data can be created in the target graph database, wherein the graph Schema includes multiple nodes and edges connecting different nodes, the nodes represent fields or data tables, and the edges represent the association relationship between connected nodes; further, based on the constructed graph Schema, the additional information parsed in the data layer, the attribute information of the private data, etc. are stored in the target graph database, thereby allowing the blood relationship between the private data to be stored in the form of a knowledge graph. Among them, the target graph database can include, for example, but is not limited to, at least one of the following graph databases: Alibaba Cloud GDB, Gea Base, TuGraph, Neo4J, etc. The graph calculation of blood relationship information can be implemented by calling at least one of the following graph computing application programming interfaces (APIs): Spark GraphX, PandaGraph, etc.
[0059] Optionally, a timer is deployed in the middle layer, such as Figure 1 and Figure 3 As shown, the timer can periodically start the scheduled task, that is, the operation log information, metadata information, SQL statement execution information, etc. stored in the source database are parsed through the pre-configured user defined functions (UDF), and the business data stored in the source database are scanned for privacy data, etc., to obtain the incremental operation information involving privacy data executed on the source data and the auxiliary information related to the privacy data (such as the attribute information of the privacy data, the owner of the SQL statement, the execution time, and other additional information), and further determine the incremental blood relationship information between the privacy data based on the obtained information; then, based on the difference information between the incremental blood relationship information and the stock association relationship information stored in the source database, the stock association relationship information in the source database and the target graph database are updated, so that the stock association relationship information in the source database and the blood relationship information in the target graph database are kept synchronized, and the blood relationship between the privacy data is obtained offline. Further, the above process can be implemented in a batch processing manner.
[0060] Specifically, the stock association relationship information is stored in the stock association relationship table of the source database. Accordingly, an incremental blood relationship table can be generated based on the incremental blood relationship information, and further based on the difference set between the incremental blood relationship table and the stock association relationship table, the difference information between the incremental blood relationship information and the stock association relationship information can be determined.
[0061] Optionally, the data layer is decoupled from the middle layer, and the middle layer can shield the differences between various storage environment interfaces in the data layer and provide a standard API interface to provide functions such as data reading.
[0062] Optionally, the knowledge graph layer is decoupled from the middle layer, and the knowledge graph layer can be privatized or publicized based on different usage environments. The knowledge graph layer can be responsible for shielding the differences in insertion, deletion, modification or reading of blood relationship information between its upper and lower layers.
[0063] Please refer to Figure 4 , is a flowchart of a privacy data processing method provided by an embodiment of this specification, and the method may include the following steps:
[0064] S402, obtaining SQL statements involving privacy data and auxiliary information related to the privacy data executed on the source database.
[0065] The auxiliary information related to the privacy data is used to describe the attributes of the fields and data tables where the privacy data is located. Optionally, the auxiliary information may include, but is not limited to, metadata information of the fields and data tables where the privacy data is located, the owner of the SQL statement, the execution time, and other additional information. Specifically, the auxiliary information may be obtained by performing a privacy data scan on the business data stored in the source database and parsing the metadata information, operation log information, etc. stored in the source database.
[0066] S404: semantically parse the acquired SQL statement to obtain the blood relationship information between the private data in the source database.
[0067] The blood relationship information between the private data is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields where the private data is located and the data tables.
[0068] Optionally, the association relationship between the fields where the private data is located may include but is not limited to at least one of the following relationships: copy relationship, truncation relationship (substr) and concatenation relationship (concat), etc. For example, the field identify_no in table Table1 is obtained by copying the field identify_no in table Table2, the field tenantId is concatenated by the field identify_no and the field mobile_no, and the field Id is obtained by truncating the field bankcard_no. The association relationship between the data tables where the private data is located may include a dependency relationship (depend), for example, the data table Table1 depends on the data table Table2. The association relationship between the field where the private data is located and the data table may include a subordinate relationship (belong), for example, the field identify_no belongs to the data table Table1.
[0069] SQL statements are a language used to operate on a database. Usually, an SQL statement is composed of multiple semantic units and grammatical orders between different semantic units. Specifically, the above-mentioned grammatical units refer to words that meet the semantic rules, such as keywords such as "SELECT", "FROM", operations, or identification information for identifying fields, data tables, etc., contained in SQL statements; the above-mentioned grammatical order refers to the grammatical rule order of semantic units agreed in SQL statements. By semantically parsing SQL statements according to the grammatical rules of SQL statements, we can obtain the associations between fields, between fields and data tables, and between data tables described in the SQL statements. It should be noted that in actual applications, the semantic analysis of SQL statements can be implemented using various existing tools, such as parsers.
[0070] For example, the SQL statement involving the private data ID number and mobile phone number is:
[0071] Create Table2 as
[0072] Select identify_no,mobile_no
[0073] From Table 1;
[0074] In the above SQL statement, the field identify_no is used to store the user's ID number, and the field mobile_no is used to store the user's mobile phone number. By semantically parsing the above SQL statement, the blood relationship between the private data ID number and the mobile phone number is obtained: (1) The field identify_no in the data table Table1 is copied from the field identify_no in the data table Table2; (2) The field mobile_no in the data table Table1 is copied from the field mobile_no in the data table Table2; (3) The data table Table1 depends on the data table Table2; (4) The field identify_no in the data table Table1 belongs to the data table Table1; (5) The field mobile_no in the data table 1 belongs to the data table Table1; (6) The field identify_no in the data table Table2 belongs to the data table Table2; (7) The field mobile_no in the data table 2 belongs to the data table Table2.
[0075] Among them, the above-mentioned relationships (1) and (2) are association relationships between the fields where the private data is located, the above-mentioned association relationship (3) is the association relationship between the data tables where the private data is located, and the above-mentioned association relationships (4) to (7) are the association relationships between the fields where the private data is located and the data tables.
[0076] In order to facilitate the storage and use of the blood relationship information obtained through analysis, in an optional manner, the blood relationship information between the private data can be converted into triple structure data (source_node, target_node, relation), where source_node represents the source node, target_node represents the target node, and relation represents the association relationship between the source node and the target node. For example, taking the blood relationship information obtained through analysis of the above SQL statement as an example, the above blood relationship information is converted into triple structure data as follows:
[0077] (Table2.identify_no,Table1.identify_no,copy)
[0078] (Table2.mobile_no,Table1.identify_no,copy)
[0079] (Table2,Table1,depend)
[0080] (Table1.identify_no,Table1,belong)
[0081] (Table1.mobile_no,Table1,belong)
[0082] (Table2.identify_no,Table2,belong)
[0083] (Table2.mobile_no,Table2,belong)
[0084] S406, based on the blood relationship information between the private data, create a graph data model representing the blood relationship between the private data in the target graph database.
[0085] In the graph database, the graph data model is used to describe the organization and structure of the graph database, and can reflect the data in the graph database and the relationship between them. The above graph data model includes multiple nodes and edges connecting different nodes. The nodes represent fields or data tables, and the edges represent the associations between connected nodes.
[0086] For example, Figure 5 An example of a graph data model is shown, such as Figure 5As shown in the figure, Node, Table and Column are nodes in the graph data model, where Node is the source node, and the id, nodeType and nodeName below it are auxiliary information used to describe the Node. Specifically, id is the unique identifier of the Node, nodeType is the type of the node, which indicates that the node is a data table or field, and nodeName is the name of the node; Table is the target node of the Node, its type is a data table, and Table depends on Node. The tenantId, projectName, tableName, tableOwner, and originSql below it are auxiliary information used to describe the node. Specifically, temamtId is the user identifier, projectName indicates the project to which the Table belongs, tableName is the name of the Table, tableOwner indicates the owner of the Table, and originSql indicates the SQL statement for generating and modifying the Table, so as to query the generation logic of the Table when user data exceptions occur; Column is a field, which is subordinate to the data table Table, and the tenantId, projectName, tabl eName, columnName, dataType, columnComment, sensitiveLevel and sensitiveType are auxiliary information used to describe the field Column. Specifically, tenantId is the user ID, projectName indicates the project to which the field Column belongs, tableName indicates the data table to which the field Column belongs, dataType indicates the data type of the field Column, columnComment indicates the description information of the field Column, sensitiveLevel indicates the sensitivity level of the field Column, and sensitiveType indicates the sensitive data type of the field Column; Edge is an edge, which is used to describe the relationship between nodes. The fromId, told, dependType and edgeId below it are auxiliary information used to describe the edge Edge. Specifically, fromId is the unique identifier of the source node, told is the unique identifier of the target node, dependType indicates that the relationship between the source node and the target node is a dependency relationship, that is, the target node depends on the source node, and edgeId is the unique identifier of the edge.
[0087] In order to ensure that the generated graph data pattern can accurately represent the blood relationship between the private data, in an optional method, S406 may include the following steps: converting the blood relationship information between the private data into triple structure data; generating nodes in the graph data pattern based on the source nodes and target nodes indicated by the triple structure data; generating edges connecting different nodes based on the association relationship between the source nodes and the target nodes indicated by the triple structure data.
[0088] For example, taking the triple structure data (Table2, Table1, depend) and (Table1.mobile_no, Table1, belong) mentioned above as examples, Table2 is the node Node in the above graph data pattern, Table1 is the node Table in the above graph data pattern, and Table1.mobile_no is the node Column in the above graph data pattern. Based on the association relationship depend between Table2 and Table1, an edge connecting the node Node and the node Table can be generated, and based on the association relationship belong between Table1.mobile and Table1, an edge connecting the node Table and the node Column can be generated.
[0089] Of course, it should be understood that the above-mentioned graph data model can also be created through various other existing methods, and the embodiments of this specification do not specifically limit this.
[0090] S408, based on the graph data model, storing the auxiliary information related to the privacy data into the target graph database.
[0091] Specifically, auxiliary information related to privacy data can be added to the graph data model, for example, auxiliary information used to describe fields can be added to the nodes corresponding to the fields, and auxiliary information used to describe data tables can be added to the nodes corresponding to the data tables, etc., so that the blood relationship information between privacy data can be stored in the target graph database in the form of a knowledge graph, which can not only clearly and accurately reflect the upstream and downstream relationships between data tables, between fields, and between data tables and fields, but also reflect the respective attributes of fields and data tables through auxiliary information, thereby improving the availability of blood relationship between privacy data, and enabling convenient and quick error correction, traceability, and visualization of privacy data based on the target graph database.
[0092] For example, in the application scenario of data query, if you need to query the distribution of all bank card numbers under a certain user's project, you can execute a graph query statement on the target graph database, such as the Gremlin statement: gV().has("sensitiveType", "bankcard_no").bothE(). Of course, the graph query statement can also be, for example, an OpenCypher statement, etc. The embodiment of this specification does not specifically limit the type of graph query statement.
[0093] For example, in the application scenario of correcting privacy data, if the field bancard_no in the data table Table1 is mistakenly identified as an ID card number, the blood relationship recorded in the target graph database can be used to determine that the bankcard_no in the data table Table2 originates from the data table Table1. Therefore, the privacy data type of the field bankcard_no in the data table Table1 can be automatically modified through blood relationship diffusion, without the need for users to correct each piece of data one by one, thereby improving the efficiency of privacy data correction and the accuracy and recall rate of privacy data identification results.
[0094] It should be noted that the privacy data processing method of the embodiments of this specification can also be used in business scenarios such as privacy data tracing, blood relationship visualization, and compliance determination of privacy data processing. The embodiments of this specification do not make specific limitations on this and will not be elaborated in detail here.
[0095] The privacy data processing method provided in the embodiments of this specification performs semantic analysis on SQL statements involving privacy data executed in the source database to obtain the blood relationship information between the privacy data in the source database, including the association relationship between the fields where the privacy data is located, the association relationship between the data tables where the privacy data is located, and the association relationship between the fields where the privacy data is located and the data tables. The obtained blood relationship information can more finely reflect the blood relationship between the privacy data; based on the blood relationship information, a graph data model (Schema) representing the blood relationship between the privacy data is created in the target graph database, and the nodes in the graph data model represent fields or data tables, and the edges in the graph data model represent the association relationship between connected nodes. Further, based on the graph data model, auxiliary information is stored in the target graph database, so that the blood relationship between the privacy data can be stored in the form of a knowledge graph, so as to realize the management of privacy data from "point" to "surface", and then the blood relationship between the privacy data can be used more conveniently and quickly to implement error correction, traceability, compliance judgment, etc. on the privacy data, thereby improving the efficiency and convenience of privacy data management.
[0096] Considering that the source database itself may store the existing association relationship information between the table-level private data, in order to ensure that the blood relationship information between the private data is accurate and reliable, in another embodiment of this specification, before the above S406, the privacy data processing method of this embodiment also includes: querying whether the source database stores the existing association relationship information between the private data. Accordingly, in the above S406, if the source database stores the existing association relationship information, the obtained blood relationship information and the existing association relationship information are fused, and further based on the fused blood relationship information, a graph data model is created in the target graph database.
[0097] It should be noted that the fusion of blood relationship information and stock association relationship information can be achieved based on various existing data fusion technologies, and the embodiments of this specification do not specifically limit this.
[0098] Furthermore, since there may be some conflicts between the blood relationship information obtained by parsing the SQL statement and the existing association relationship information, especially the fields with a copy relationship, in order to further improve the accuracy and reliability of the blood relationship information finally obtained, in another embodiment of the present specification, before creating the graph data model in the target graph database based on the fused blood relationship information, the fused blood relationship information can also be corrected. Specifically, the privacy data processing method of the embodiment of the present specification also includes: based on the fused blood relationship information, obtaining the metadata of the first privacy field and the second privacy field, respectively, wherein the first privacy field and the second privacy field are different fields where the privacy data is located and the association relationship is a copy relationship; then, the metadata of the first privacy field and the metadata of the second privacy field are compared to obtain the difference degree, and if the difference degree exceeds the difference degree threshold, the association relationship information between the first privacy field and the second privacy field in the fused blood relationship information is deleted. The difference degree threshold can be preset according to actual needs, and the embodiment of the present specification does not specifically limit the value of the difference degree threshold.
[0099] For example, the fused blood relationship information indicates that the first privacy field Table1.mobile_no is copied from the second privacy field Table2.nobile_no. If the metadata of the first privacy field Table1.mobile_no indicates that the data type of the field is text, and the metadata of the second privacy field Table2.mobile_no indicates that the data type of the field is string, it can be determined that the copy relationship between the first privacy field and the second privacy field is incorrect, and the copy relationship between the first privacy field Table1.mobile_no and the second privacy field Table2.mobile_no is deleted in the fused blood relationship information.
[0100] Considering that the blood relationship information between the private data in the source database changes dynamically, in order to further improve the accuracy and reliability of the blood relationship stored in the target graph database, in another embodiment of the present specification, after the auxiliary information related to the private data is stored in the target graph database, the changes of the blood relationship information in the source database can be monitored periodically in an offline manner, and then the local offline stock association relationship table and the target graph database are updated based on the monitored changes, so that the stock association relationship information stored in the source database is synchronized with the blood relationship information stored in the target graph database. Specifically, Figure 6 As shown, after S408, the privacy data processing method provided in the embodiment of this specification further includes:
[0101] S410: Obtain incremental operation information involving privacy data performed on a source database at a preset time interval.
[0102] Specifically, the incremental operation information may include but is not limited to newly added SQL statements involving privacy data executed on the source database, incremental operation log information, etc.
[0103] It should be noted that the length of the preset time interval can be set according to actual needs, and the embodiments of this specification do not specifically limit this.
[0104] S412: Determine incremental blood relationship information between the private data based on the incremental operation information and the auxiliary information related to the private data.
[0105] The incremental blood relationship information between the private data may include information that has changed compared to the blood relationship information between the existing private data, which is used to reflect the changes in the blood relationship between the private data.
[0106] Specifically, based on the attribute information of the fields and data tables described by the auxiliary information related to the privacy data, the changes in the fields and data tables where the privacy data is located can be determined, such as which fields and / or data tables have been added, which fields and / or data tables have been deleted, and which fields and / or data tables have changed their privacy data types, sensitivity levels, etc. By semantically parsing the SQL statements indicated by the incremental operation information, the changes in the associations between the changed fields, between the data tables, and between the fields and the data tables can be determined, thereby obtaining the incremental blood relationship information between the privacy data.
[0107] S414, obtaining the difference information between the incremental blood relationship information and the stock relationship information stored in the source database.
[0108] Taking into account that the stock relationship information is usually stored in the form of a relational data table in the stock association relationship table of the source database, in order to improve the accuracy and convenience of obtaining difference information, in an optional manner, S414 may include: generating an incremental blood relationship table based on the incremental blood relationship information. For example, the incremental blood relationship information indicates that the field bankcard_no where the newly added privacy data is located belongs to the data table Table2. Then, based on the incremental blood relationship information, a relational data table of the association relationship between the corresponding field bankcard_no and the data table Table2 is created to obtain the incremental blood relationship table; further, based on the difference set between the incremental blood relationship table and the stock association relationship table, the difference information between the incremental blood relationship information and the stock association relationship information is determined.
[0109] S416, based on the difference information, update the stock association relationship information and the target graph database.
[0110] Specifically, based on the difference information, at least one first operation statement such as adding, deleting, and modifying that needs to be executed in the target graph database can be determined, and by executing the first operation statement on the target graph database, the target graph database can be updated. Similarly, based on the difference information, at least one second operation statement such as adding, deleting, and modifying that needs to be executed in the source database can be determined, and by executing the second operation statement on the source database, the existing association relationship information can be updated.
[0111] It can be understood that through the above-mentioned privacy data processing method, the existing association relationship information stored in the source database and the blood relationship information stored in the target graph database are kept synchronized, and then in subsequent applications, the blood relationship between the privacy data can be obtained from the source database or the target graph database according to actual needs.
[0112] In addition, Figure 4 Corresponding to the privacy data processing method shown, the embodiment of this specification also provides a privacy data processing device. Figure 7 700 is a schematic diagram of a privacy data processing device 700 provided in an embodiment of this specification, including:
[0113] A first acquisition unit 710 acquires a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0114] The parsing unit 720 performs semantic parsing on the acquired SQL statement to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0115] A creation unit 730, based on the blood relationship information, creates a graph data model representing the blood relationship between the private data in a target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0116] The storage unit 740 stores the auxiliary information in the target graph database based on the graph data model.
[0117] Optionally, the device further comprises:
[0118] A query unit, before the creation unit 730 creates a graph data model representing the blood relationship between the private data in the target graph database based on the blood relationship information, querying whether the source database stores the existing association relationship information between the private data;
[0119] The creation unit 730, if the existing association relationship information is stored in the source database, fuses the blood relationship information with the existing association relationship information, and creates the graph data model in the target graph database based on the fused blood relationship information.
[0120] Optionally, the device further comprises:
[0121] A second acquisition unit, before the creation unit 730 creates the graph data model in the target graph database based on the fused blood relationship information, acquires metadata of each of the first privacy field and the second privacy field based on the fused blood relationship information, wherein the first privacy field and the second privacy field are different fields where the privacy data is located and the association relationship is a copy relationship;
[0122] A comparison unit, comparing the metadata of the first privacy field with the metadata of the second privacy field to obtain a difference degree;
[0123] A deleting unit is configured to delete the association relationship information between the first privacy field and the second privacy field in the fused blood relationship information if the difference exceeds a difference threshold.
[0124] Optionally, the device further comprises:
[0125] A third acquisition unit, after the storage unit 740 stores the auxiliary information in the target graph database based on the graph data mode, acquires incremental operation information involving privacy data performed on the source database at a preset time interval;
[0126] an incremental lineage determination unit, which determines incremental lineage relationship information between the private data based on the incremental operation information and auxiliary information related to the private data;
[0127] A fourth acquisition unit is configured to acquire difference information between the incremental blood relationship information and the stock association relationship information stored in the source database;
[0128] An updating unit updates the stock association relationship information and the target graph database based on the difference information.
[0129] Optionally, the stock association relationship information is stored in a stock association relationship table of the source database;
[0130] The fourth acquisition unit generates an incremental blood relationship table based on the incremental blood relationship information, and determines the difference information between the incremental blood relationship information and the stock association relationship information based on the difference set between the incremental blood relationship table and the stock association relationship table.
[0131] Optionally, the creation unit 730 converts the blood relationship information between the privacy data into triple structure data, generates nodes in the graph data pattern based on the source nodes and target nodes indicated by the triple structure data, and generates edges connecting different nodes based on the association relationship between the source nodes and the target nodes indicated by the triple structure data.
[0132] Optionally, the association relationship between the fields where the privacy data is located includes at least one of the following relationships: a copy relationship, a truncation relationship, and a splicing relationship;
[0133] The association relationship between the data tables where the privacy data is located includes a dependency relationship;
[0134] The association relationship between the field and the data table includes a subordinate relationship.
[0135] The privacy data device provided in the embodiments of the present specification obtains the blood relationship information between the privacy data in the source database, including the association relationship between the fields where the privacy data is located, the association relationship between the data tables where the privacy data is located, and the association relationship between the fields where the privacy data is located and the data tables, by semantically parsing the SQL statements executed on the source database and involving the privacy data. The obtained blood relationship information can more finely reflect the blood relationship between the privacy data; based on the blood relationship information, a graph data model (Schema) representing the blood relationship between the privacy data is created in the target graph database, and the nodes in the graph data model represent fields or data tables, and the edges in the graph data model represent the association relationship between connected nodes. Further, based on the graph data model, auxiliary information is stored in the target graph database, so that the blood relationship between the privacy data can be stored in the form of a knowledge graph, so as to realize the management of privacy data from "point" to "surface", and then it can more conveniently and quickly use the blood relationship between the privacy data to implement error correction, traceability, compliance judgment, etc. on the privacy data, thereby improving the efficiency and convenience of privacy data management.
[0136] Obviously, the privacy data processing device in the embodiment of this specification can be used as the above Figure 4 The execution subject of the privacy data processing method shown in FIG. Figure 4 The functions realized are as follows. Since the principles are the same, they will not be described here.
[0137] Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of this specification. Figure 8 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.
[0138] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0139] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.
[0140] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a privacy data processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0141] Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0142] Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0143] Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0144] Based on the atlas data pattern, the auxiliary information is stored in the target graph database.
[0145] The above is as in this manual Figure 4The method performed by the privacy data processing device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or instructions in software form. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of this specification can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of this specification can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0146] It should be understood that the electronic device of the embodiment of this specification can implement the privacy data processing device in Figure 4 Functions of the embodiments shown. Since the principles are the same, the embodiments of this specification will not be described in detail here.
[0147] Of course, in addition to software implementation, the electronic device of this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0148] The embodiment of the present specification also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by a portable electronic device including a plurality of application programs, enable the portable electronic device to execute Figure 4 The method of the embodiment shown is specifically used to perform the following operations:
[0149] Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located;
[0150] Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables;
[0151] Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes;
[0152] Based on the atlas data pattern, the auxiliary information is stored in the target graph database.
[0153] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0154] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
[0155] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer.
[0156] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0157] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0158] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
Claims
1. A method for processing private data, comprising: Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located; Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables; Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes; Based on the atlas data pattern, the auxiliary information is stored in the target graph database.
2. The method according to claim 1, before creating a graph data pattern representing the blood relationship between the private data in a target graph database based on the blood relationship information, the method further comprises: Querying whether the source database stores the existing association relationship information between the private data; The step of creating a graph data model representing the blood relationship between the private data in a target graph database based on the blood relationship information includes: If the existing association relationship information is stored in the source database, the blood relationship information and the existing association relationship information are merged; Based on the fused blood relationship information, the atlas data model is created in the target graph database.
3. The method according to claim 2, before creating the graph data model in the target graph database based on the fused blood relationship information, further comprises: Based on the fused blood relationship information, metadata of a first privacy field and a second privacy field are obtained, where the first privacy field and the second privacy field are different fields where the privacy data is located and the association relationship is a copy relationship; Comparing the metadata of the first privacy field with the metadata of the second privacy field to obtain a difference degree; If the difference exceeds a difference threshold, the association relationship information between the first privacy field and the second privacy field in the fused blood relationship information is deleted.
4. The method according to claim 1, after storing the auxiliary information in the target graph database based on the graph data mode, the method further comprises: At preset time intervals, obtaining incremental operation information involving privacy data performed on the source database; Determining incremental blood relationship information between the private data based on the incremental operation information and the auxiliary information related to the private data; Obtaining difference information between the incremental blood relationship information and the stock association relationship information stored in the source database; Based on the difference information, the stock association relationship information and the target graph database are updated.
5. The method according to claim 4, wherein the stock association relationship information is stored in the stock association relationship table of the source database; The obtaining of difference information between the incremental blood relationship information and the stock association relationship information stored in the source database includes: Based on the incremental blood relationship information, generating an incremental blood relationship table; Based on the difference set between the incremental blood relationship table and the existing association relationship table, the difference information between the incremental blood relationship information and the existing association relationship information is determined.
6. The method according to claim 1, wherein creating a graph data model representing the blood relationship between the private data in a target graph database based on the blood relationship information comprises: Converting the blood relationship information between the private data into triple structure data; Generate nodes in the graph data pattern based on the source nodes and the target nodes indicated by the triple structure data; Based on the association relationship between the source node and the target node indicated by the triple structure data, edges connecting different nodes are generated.
7. The method according to any one of claims 1 to 6, wherein the association relationship between the fields where the privacy data is located comprises at least one of the following relationships: a copy relationship, a truncation relationship, and a splicing relationship; The association relationship between the data tables where the privacy data is located includes a dependency relationship; The association relationship between the field and the data table includes a subordinate relationship.
8. A privacy data processing device, comprising: A first acquisition unit is configured to acquire a structured query SQL statement involving private data and auxiliary information related to the private data, which is executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located; A parsing unit, which performs semantic parsing on the acquired SQL statement to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables; A creation unit, based on the blood relationship information, creates a graph data model representing the blood relationship between the private data in a target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes; A storage unit stores the auxiliary information into the target graph database based on the graph data pattern.
9. An electronic device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located; Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables; Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes; Based on the atlas data pattern, the auxiliary information is stored in the target graph database.
10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to perform the following operations: Obtaining a structured query SQL statement involving private data and auxiliary information related to the private data executed on a source database, wherein the auxiliary information is used to describe the attributes of the field and data table where the private data is located; Performing semantic analysis on the acquired SQL statements to obtain the blood relationship information between the private data in the source database, wherein the blood relationship information is used to indicate the association relationship between the fields where the private data is located, the association relationship between the data tables where the private data is located, and the association relationship between the fields and the data tables; Based on the blood relationship information, a graph data model representing the blood relationship between the private data is created in the target graph database, wherein the graph data model includes a plurality of nodes and edges connecting different nodes, wherein the nodes represent fields or data tables, and the edges represent association relationships between connected nodes; Based on the atlas data pattern, the auxiliary information is stored in the target graph database.
Citation Information
Patent Citations
Intelligent collaborative cloud architecture based on data driving in universal network convergence scene
CN111371830A
Data consanguinity analysis method, device and equipment and computer readable storage medium
CN111694858A