Method for exploring data consanguinity through AI technology
Through AI technology and circular dependency exploration model, abnormal dependency nodes in the data blood relationship map are automatically identified and positioned, which solves the problem that existing technology cannot automatically detect and locate, and achieves more efficient data governance.
Patent Information
- Application Number
- CN202411826974.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art cannot automatically identify and locate abnormally dependent nodes in the data blood relationship map, and cannot explore in advance the data nodes in the data blood relationship map before the data quality problems are exposed.
Using AI technology, by obtaining the adjacency matrix and node attribute characteristics generated by the hierarchical government blood relationship map, input them into the cyclic dependence exploration model, and using the attention module, feature aggregation module and classification prediction module to explore and identify cyclic dependence nodes.
It realizes automatic detection and positioning of circular dependency problems before data quality problems are exposed, reducing data governance costs and improving the efficiency of data governance personnel to audit data quality problems.
Smart Images

Figure CN119989187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data lineage analysis, and in particular to a method for detecting data lineage using AI technology. Background Art
[0002] In the existing data lineage system in the government sector, SQL statements in the ETL (Extract, Transform, and Load) process of parsing data development are stored in the graph database and processed into an effective data lineage graph. Due to improper handling in the development and design stages, it is very easy to cause circular dependency nodes, null dependency nodes, and missing dependency nodes in the data lineage graph, which will cause data quality problems.
[0003] At present, when data quality problems occur, data quality management personnel use graph search syntax and human eye recognition to trace all related nodes in the data lineage map, determine whether there are abnormal dependency nodes, and locate and analyze abnormal dependency nodes to solve data quality problems. Abnormal dependency nodes include circular dependency nodes, null value dependency nodes, and missing dependency nodes. However, it is currently impossible to automatically identify abnormal dependency nodes, and it is impossible to detect the lineage of data nodes in the data lineage map and locate abnormal dependency nodes before data quality problems are exposed. Summary of the invention
[0004] The present invention aims to at least solve the technical problems in the prior art that it is impossible to automatically identify and locate abnormal dependent nodes, and it is impossible to detect the lineage of data nodes in the data lineage map and locate the abnormal dependent nodes before data quality problems are exposed, and provide a method for detecting data lineage using AI technology.
[0005] In order to achieve the above-mentioned purpose of the present invention, the present invention provides a method for exploring data lineage using AI technology, including: obtaining a hierarchical government lineage map composed of table nodes and field nodes, generating an adjacency matrix and attribute characteristics of each node based on the hierarchical government lineage map; inputting the adjacency matrix and the attribute characteristics of each node into a circular dependency detection model to obtain a circular dependency detection result; the circular dependency detection model includes: an attention module, which uses an attention weight matrix to process the attribute characteristics of each node to obtain the attention characteristics of each node, and calculates the attention coefficient between any two nodes based on the attention characteristics of the nodes and the adjacency matrix; a feature aggregation module, which obtains the aggregated characteristics of each node based on the attention weight matrix, the attribute characteristics of the nodes and the attention coefficients between the nodes; and a classification prediction module, which classifies the aggregated characteristics of each node to obtain the circular dependency detection result of each node.
[0006] The beneficial technical effects of the present invention are as follows: the adjacency matrix generated based on the hierarchical government blood relationship map and the attribute characteristics of each node are input into a pre-trained circular dependency detection model; the attention module of the circular dependency detection model uses the attention weight matrix and the adjacency matrix to obtain the attention coefficient between nodes, which can effectively capture the complex and heterogeneous dependencies between nodes, rather than relying solely on the topological structure, so that the circular dependency detection model can learn which adjacent nodes are more important for the classification of the current graph node, so that the attention coefficient fully reflects the dependency relationship between nodes; the feature aggregation module of the circular dependency detection model obtains the aggregated features of each node based on the attention weight matrix, the attribute characteristics of the node and the attention coefficient between the nodes, and each node will aggregate its neighbor nodes according to the calculated attention coefficient. The adjacency matrix ensures that attribute features are aggregated only between actually connected node pairs. The feature update of each node depends on the features of its neighboring nodes. The attention weight matrix and attention coefficient provided by the attention module capture the mutual influence between nodes, highlighting the cyclic dependency characteristics between nodes, so as to facilitate subsequent accurate cyclic dependency classification. Finally, the classification prediction module is used to classify the nodes based on their aggregated features to obtain accurate detection results of whether the nodes are cyclic dependency nodes or non-cyclic dependency nodes. The present invention can automatically detect and locate cyclic dependency problems before data quality problems are exposed, reducing data governance costs, and is more conducive to data governance personnel auditing data quality problems developed, effectively preventing and controlling data quality problems caused by data cyclic dependency. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 It is a flow chart of a method for detecting data lineage using AI technology in a preferred embodiment of the present invention;
[0008] Figure 2 It is a structural schematic diagram of a circular dependency detection model in a preferred embodiment of the present invention;
[0009] Figure 3 It is a partial schematic diagram of a hierarchical government bloodline map in an example of the present invention;
[0010] Figure 4 It is a schematic diagram of the entire process in an application scenario of the present invention. DETAILED DESCRIPTION
[0011] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0012] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0013] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0014] The execution subject of the method for detecting data lineage using AI technology provided by the present invention includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, a method for detecting data lineage using AI technology can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0015] The present invention discloses a method for detecting data lineage using AI technology. In a preferred implementation, its flow chart is as follows: Figure 1 As shown, including:
[0016] Step S1, obtain a hierarchical government lineage graph consisting of table nodes and field nodes, and generate an adjacency matrix and attribute features of each node based on the hierarchical government lineage graph.
[0017] In this embodiment, the hierarchical government lineage map can be obtained by extracting the network topology of the table node level and the field node level from the constructed government data lineage map. The government data lineage map is stored in a tree, and its node set includes library nodes, table nodes, file nodes, field nodes, file element nodes, etc. The library node may include multiple file nodes, multiple table nodes, multiple file element nodes and multiple field nodes, the file node may include multiple field nodes, multiple table nodes, multiple file element nodes, etc., and the table node may include multiple field nodes. The library node can represent various government databases, such as population databases, legal person databases, natural resources databases, economic databases, employment and social security databases, medical and health databases, etc. The file node represents various text, csv, word, pdf, excel and other files in the process of government affairs, finance, and data development, as well as the programmer's code files. The file element node is mainly an information set such as keywords, key paragraphs, information summaries, etc. that represent files or data tables. The table node mainly represents various government tables, such as population statistics tables, employment status tables, education data tables, etc.
[0018] In this embodiment, the node also includes multiple attribute information. In the government data lineage map, the node attributes include:
[0019] 1) Node category (node_category): database (database), table (table), field (field), file name (file_name), file element (file_keyword);
[0020] 2) Node sequence type (node_sequence): source (is_source), intermediate (is_intermediate), target (is_target);
[0021] 3) Node update frequency (node_update_frequency): day, hour, minute, real-time;
[0022] 4) Creator: the creator of the database, table, or field;
[0023] 5) Level: According to the data warehouse hierarchical design, it is divided into business layer (original), post source layer (ODS), detail layer (DWD), light convergence layer (DWS), data service layer (ADS), and others (other);
[0024] 6) Edge in-degree value (in_degree): Dependent on edge in-degree (in_degree);
[0025] 7) Out-degree value (out_degree): Dependent edge out-degree (out_degree);
[0026] 8) Edge in-degree and out-degree difference (minus_degree): dependent edge in-degree (in_degree) - dependent edge out-degree (out_degree);
[0027] 9) Special definition attributes of each type of node:
[0028] The special definition properties of the library node, i.e. the database, are:
[0029] Number of tables (table_counts);
[0030] Number of fields (filed_counts);
[0031] Number of data rows (data_rows);
[0032] Data size (data_sizes);
[0033] The table node, i.e. the table's special definition attributes, are:
[0034] Number of fields (filed_counts);
[0035] Number of data rows (data_rows);
[0036] Data size (data_sizes);
[0037] Table comment (table_comment);
[0038] Table name (table_name);
[0039] Create table statement (table_ddl);
[0040] The special definition attributes of field nodes are:
[0041] Field name (filed_name);
[0042] Field comment (field_comment);
[0043] Field type (column_type);
[0044] Is it a key value (is_key);
[0045] Is it the default value (is_default): yes / no;
[0046] Default content value (default_value): default content value;
[0047] The special definition attributes of file nodes, i.e. files, are:
[0048] File name (file_name);
[0049] File comment (file_comment);
[0050] File type (file_type): PPTX, DOCX, EXCEL, TXT, EXCEL, PNG, etc.
[0051] Abstract content (abstract_data);
[0052] Default content value (default_value): default content value;
[0053] The special definition attributes of the file feature node are:
[0054] Element segment name (keyword_name);
[0055] Feature comments (keyword_comment);
[0056] Element content (keyword_data).
[0057] In this embodiment, the government data lineage graph also includes directed edges connecting nodes, which are used to represent the dependency relationship between nodes. For example, if the direction of the edge is from table A to table B, then table B is considered to be dependent on table A. In one example, the directed edges connect nodes in the following manner: E1->E2->E3, E2->E4, indicating that the pre-dependency of node E2 is E1, the pre-dependency of E3 is E2, and the pre-dependency of E4 is E2. In addition, the in-degree of E1 is 0 and the out-degree is 1; the in-degree of E2 is 1 and the out-degree is 2. The in-degree of E3 is 1 and the out-degree of E3 is 0. The in-degree of E4 is 1 and the out-degree is 0.
[0058] In one example, a partial diagram of the hierarchical government lineage map is shown below: Figure 3 As shown in Figure 1, it includes multiple table nodes and field nodes, as well as multiple directed edges. Figure 3 In, let table A be the source table, containing fields (C1, C2, C3, C4); table B is the intermediate table, containing fields (L1, L2, L3); table C is the target table, containing fields (M1, M2, M3); table D is the intermediate table, containing field (N1); memory table V, containing field (O1) - the memory table V is temporarily stored in the memory, and will not exist after use, and is a transient table; table E is the intermediate table, containing field (P3). Figure 3In the example, L2 of intermediate table B is generated by C3 of source table A, M1 of target table C is generated by L2 of intermediate table B, and C3 of source table A depends on M1 of target table, resulting in a circular dependency, which can be expressed as:
[0059] A->B,B->C,C->A.
[0060] C3->L2, L2->M1, M1->C3.
[0061] Therefore, a circular dependency relationship can be defined as a closed loop formed by directed edges between multiple nodes (multiple field nodes or multiple table nodes) in the hierarchical government lineage graph. In this case, multiple nodes are considered to be in a circular dependency relationship. These nodes are called circular dependency nodes, such as Figure 3 In the figure, nodes A, B, C, C3, L2, and M1 are all circularly dependent nodes, where nodes A, B, and C are a group of circularly dependent nodes, and nodes C3, L2, and M1 are another group of circularly dependent nodes.
[0062] In this embodiment, the adjacency matrix represents the connection relationship between nodes in the hierarchical government lineage map. The adjacency matrix can reflect the dependency relationship of the nodes, which is convenient for detecting circular dependency nodes. Figure 3 Construct the adjacency matrix A as shown in Table 1. The adjacency matrix A is a square matrix that represents the connection relationship between the nodes in the graph. In the cyclic dependency detection model, the feature update of each node depends on the features of the neighboring nodes obtained by the adjacency matrix, and ensures that only the actually connected node pairs will perform feature aggregation in the feature aggregation module.
[0063] Table 1
[0064] A B C ... A 0 1 0 B 0 0 1 C 1 0 0 ..
[0065] A->B:A[0,1]=1
[0066] B->C:A[1,2]=1
[0067] C->A:A[2,0]=1
[0068] The circular dependency A->B->C->A is clearly shown.
[0069] In this embodiment, the attribute characteristics of each node in the hierarchical government lineage map can be obtained by arranging the attribute values of all attributes of the node to form a column vector or a row vector. For example, in the above example, Figure 3 The attribute characteristics of the nodes A, B, and C in the table are shown in Table 2 below:
[0070] Table 2
[0071]
[0072]
[0073] Among them, the levels are: 1 (original), 2 (ODS), 3 (DWD), 4 (DWS), 5 (ADS), 6 (other); in-degree and out-degree difference = in-degree - out-degree; in-degree ratio = in-degree / number of fields; out-degree ratio = out-degree / number of out-degrees.
[0074] For example, for the attribute matrix h of node A A The size is 13×1, which is represented as follows:
[0075]
[0076] Step S2, input the adjacency matrix and the attribute features of each node into the circular dependency detection model to obtain the circular dependency detection result. The circular dependency detection model can be implemented based on the Graph Attention Networks (GAT) structure, such as Figure 2 As shown, the circular dependency detection model includes:
[0077] The attention module uses the attention weight matrix to process the attribute features of each node to obtain the attention features of each node, and calculates the attention coefficient between any two nodes based on the node's attention features and the adjacency matrix. The attention module helps capture the complex and heterogeneous dependencies between nodes, rather than relying solely on the topological structure. Through the attention mechanism, the model can learn which neighbor nodes are more important for the current graph node classification.
[0078] The feature aggregation module obtains the aggregated features of each node based on the attention weight matrix, the attribute features of the nodes, and the attention coefficients between the nodes. The feature aggregation module aggregates homogeneous features together. For example, the adjacent nodes of node A may have the characteristics of circular dependency. Then, during model training, when the features of node A are similar or identical to those of more adjacent nodes, node A will also be classified as a node with circular dependency.
[0079] The classification prediction module classifies the aggregated features of each node to obtain the circular dependency detection results of each node.
[0080] In this embodiment, the feature update of each node depends on the features of its neighboring nodes, and this dependency can capture the cyclic dependency pattern in the graph. The cyclic dependency detection model can identify this cyclic dependency through the attention mechanism because it can capture the mutual influence between nodes. For example, if nodes A, B, and C have similar out-degree, in-degree, and other features, the cyclic dependency detection model can learn this complex dependency pattern and take this cyclic dependency into account when classifying.
[0081] In this embodiment, after obtaining the circular dependency detection result, the data attribute of the node in the hierarchical government bloodline map is added: whether there is a circular dependency (is_loop): yes (1) / no (0).
[0082] In this embodiment, preferably, the classification prediction module performs:
[0083] Step C1, multiply the classification weight matrix by the aggregated features of each node to obtain the prediction score of each node belonging to each node category, and the node category includes circular dependency and non-circular dependency. The classification weight matrix is obtained by training and optimizing the circular dependency detection model. Specifically:
[0084] z i =Bh” i
[0085] Among them, B represents the classification weight matrix, whose dimension is F c ×F',F c represents the number of node categories, which can be 2 here. Therefore, z i is the prediction score vector of node i, including the prediction scores of two node categories, namely, the prediction score that node i belongs to the circular dependency node category and the prediction score that node i belongs to the non-circular dependency node category.
[0086] Step C2, determine the predicted probability of each node belonging to each node category according to the predicted score of each node belonging to each node category. The predicted probability of node i belonging to node category c is:
[0087]
[0088] Among them, z ic Represents the predicted score vector z of node i i The prediction score of node category c in c∈C exp(z ic ) represents the predicted score vector z of node i i The cumulative sum of the prediction scores of all node categories c in . C represents the set of node categories.
[0089] In this embodiment, during the attention module processing, the attention weight matrix is assumed to be W, which is obtained through the training optimization of the cyclic dependency exploration model. The calculation formula for the attention weight matrix to process the attribute features of each node is:
[0090] h' i =W*h i ;
[0091] Where i is the node index, which is a positive integer; h iRepresents the attribute characteristics of node i, with dimension F×1; h′ i represents the attention feature of node i, and its dimension is F′×1. The attention weight matrix W is used to map the attribute features of the node from the original feature space to a unique feature space, and its dimension is F′×F. In the above example, h A The dimension is 13×1, then h′ A The dimension is 3×1, and the dimension of the attention weight matrix is 3×13.
[0092] In a preferred embodiment, in order to make the attention coefficient accurately express the dependency relationship between two nodes, in the step of calculating the attention coefficient between any two nodes based on the attention features of the nodes and the adjacency matrix, the attention coefficient calculation process between node i and node j includes:
[0093] Step A1, concatenate the attention feature h′ of node i i and the attention feature h′ of node j h Get the first splicing feature (h′ i ||h′ j ), the first splicing feature (h′ i ||h′ j ) and the auxiliary matrix a T Multiply to obtain the first conversion feature (a T (h′ i ||h′ j ), the first conversion feature is processed by the first function to obtain the first function processed value; the auxiliary matrix a T It is obtained through the optimization of the circular dependency detection model training, and its dimension is 2F'×F. || means that the matrix is linked head to tail, such as h′ i and h′ j The dimensions of are all 3×1, then (h′ i ||h′ j ) has a dimension of 6×1.
[0094] Step A2: Determine the neighbor node set N(i) of node i based on the adjacency matrix, and convert the attention feature h′ of node i into i The attention feature h′ of each neighbor node (such as neighbor node k) in the neighbor node set is k Concatenate to obtain multiple second concatenation features, multiply the multiple second concatenation features with the auxiliary matrix respectively to obtain multiple second conversion features, perform second function processing on the multiple second concatenation features respectively to obtain multiple second function processing values, and accumulate the multiple second function processing values to obtain a second function processing accumulated value; the neighbor node set N(i) includes a set of nodes that directly have connecting edges with node i, including nodes that depend on node i and nodes that node i depends on.
[0095] Step A3, obtaining the attention coefficient between node i and node j by dividing the first function processing value by the second function processing accumulated value; wherein i and j are both node indexes, and i and j are both positive integers. Preferably, the first function processing includes a process of first processing with a first activation function and then processing with an exponential function exp(·), and the first activation function is preferably but not limited to a leakyReLU(·) function. k represents a neighbor node index.
[0096] In this embodiment, in order to better express the above attention coefficient calculation process, the attention coefficient a between node i and node j is ij The calculation formula is:
[0097]
[0098] In a preferred embodiment, in order to make the aggregated features highlight the cyclic dependency characteristics between nodes, the aggregation module obtains the aggregated features of node i based on the attention weight matrix, the attribute characteristics of the nodes and the attention coefficients between the nodes as follows:
[0099] Step B1, determining the neighbor node set N(i) of node i based on the adjacency matrix.
[0100] Step B2, multiply the attention coefficient between node i and each neighbor node, the attribute characteristics of each neighbor node, and the attention weight matrix to obtain the fusion feature value of node i and each neighbor node. The fusion feature value of node i and neighbor node k in the neighbor node set N(i) is: ik W k .h k The attribute characteristics represented by a ik Represents the attention coefficient between node i and its neighbor node k.
[0101] Step B3, accumulating the fusion feature values of node i and all neighboring nodes to obtain the accumulated fusion feature value, which can be expressed as: ∑ k∈N(i) a ik W k .
[0102] Step B4: Perform a second function processing on the accumulated fusion feature value to obtain the aggregate feature of node i. The second function processing is an activation function processing, and the activation function is preferably but not limited to a ReLU function.
[0103] The aggregate feature h″ of node i i =σ(∑ k∈N(i) a ik W k ). σ(·) represents the second function processing.
[0104] In a preferred embodiment, the training process of the cyclic dependency detection model includes:
[0105] Step D1, obtain a training-level government affairs lineage graph, and set true labels for nodes of the training-level government affairs lineage graph, where the true labels include circular dependencies and non-circular dependencies.
[0106] The hierarchical government lineage graph used for training is called the training hierarchical government lineage graph. The training hierarchical government lineage graphs of multiple time periods can be obtained for training. The true labels are mainly used to mark the node categories of the nodes in the training hierarchical government lineage graph, including cyclic dependencies and non-cyclic dependencies. For the convenience of calculation, the label values of cyclic dependencies and non-cyclic dependencies can be set to different values.
[0107] In the above example, if Figure 3 As shown in the figure, Table A, Table B, and Table C are loop-dependent nodes, and their true label values (is_loop) are all set to 1, that is, is_loop = 1. Table D and Table E are non-loop-dependent nodes, that is, the true label value is_loop = 0. {A, B, C}∈C loop ,{D,E}∈C n-loop . C loop Represents a set of circular dependency nodes, C n-loop Represents a collection of acyclic dependency nodes.
[0108] Step D2, based on the training hierarchical government relations graph, generates a training adjacency matrix and training attribute features of each node. Refer to the above-mentioned adjacency matrix and node attribute feature acquisition process, which will not be repeated here.
[0109] Step D3: construct a circular dependency detection model and randomly initialize the attention weight matrix W and auxiliary matrix a T and classification weight matrix B, preferably but not limited to using Xavier uniform initialization (also known as Glorot initialization) to initialize the attention weight matrix W and the auxiliary matrix a T and the classification weight matrix B.
[0110] Step D4, iteratively train the cyclic dependency detection model using the training adjacency matrix and the training attribute features of each node until the training stop condition is reached. In the iterative training, the loss function is calculated, and the attention weight matrix, the auxiliary matrix, and the classification weight matrix are updated according to the gradient of the loss function (such as the gradient descent method). The loss function includes the standard loss term Loss s :
[0111] Loss s =-∑ i∈V ∑ c∈C y ic log(p ic )
[0112] Among them, V represents the node set of the training-level government lineage graph, i is the node index in the training-level government lineage graph, C represents the node category set, c represents the node category index, and y ic represents the true label value of node i in node category c, p ic Represents the predicted probability of node i in node category c output by the circular dependency detection model.
[0113] In this embodiment, the loss function is used to identify and process cyclic dependencies and perform accurate node classification. s It is used to measure the difference between the predicted probability distribution of the circular dependency detection model and the true label, and improve the accuracy of node classification.
[0114] In this embodiment, the training stop condition is preferably, but not limited to, that the number of iterative training times reaches a preset maximum number of training times, or the reduction value of the loss function value is less than or equal to a preset reduction threshold. During the training process, start with a small learning rate (such as 0.001) and adjust according to the training effect.
[0115] In a preferred embodiment, the loss function also includes a cyclic dependency loss f c , sparsity loss f s and directional loss f d The weighted sum of , the loss function is expressed as:
[0116] Loss total =Loss s +λLoss ex
[0117] Among them, Loss ex Represents the cycle dependency auxiliary loss term, specifically the cycle dependency loss f c , sparsity loss f s and directional loss f d λ represents the input value parameter, which is used to adjust the Loss ex importance.
[0118] In this embodiment, Loss ex =W c f c +W s f s +W d f d
[0119] Among them, W c It is expressed as the penalty weight of circular dependency, which ensures the accuracy of circular dependency detection; W s W represents the sparsity loss weight and optimizes the attention distribution; dRepresents the direction consistency loss weight to ensure the correct training direction. Initial setting {W c :1.0,W s :0.2,W d :0.5}, fine-tuning can be done during actual training.
[0120] In this embodiment, the cycle dependency loss f c It is the cumulative sum of the minimum attention coefficients within the group of all group circular dependency nodes detected in the circular dependency detection model training (such as this iteration training or this batch training). The calculation formula is:
[0121] f c =∑ m∈M min (i,j∈Vm) a ij
[0122] Among them, M groups of circular dependency nodes are detected in the circular dependency detection model training, m is the group index, 1≤m≤M; Vm represents the mth group of circular dependency nodes; min (i,j∈Vm) a ij It means to find the minimum attention coefficient between nodes in the mth group of circular dependency nodes.
[0123] Cyclic dependency loss f c In order to identify circular dependencies in graph data, by penalizing high attention coefficients between node pairs, a more reasonable dependency relationship between nodes can be learned, which helps to improve the accuracy of circular dependencies.
[0124] In this embodiment, the sparsity loss f s is the average value of the absolute value of the attention coefficient between the nodes at both ends of all edges in the training level government lineage graph, and the calculation formula is:
[0125]
[0126] Among them, (i, j) represents the edge connecting node i and node j in the edge set E of the training hierarchical government lineage graph, N e Represents the total number of edges included in the edge set E. Sparsity loss f s The loop dependency detection model is made to learn a sparse attention coefficient distribution and penalize those non-zero attention coefficients. It forces the loop dependency detection model to only focus on the most important neighbor nodes, thereby improving the interpretability and efficiency of the loop dependency detection model.
[0127] In this embodiment, the directivity loss f d It is the cumulative sum of the attention coefficients between nodes in the node set of the training hierarchical government lineage graph. The calculation formula is:
[0128] f d =∑i,j∈V a ij
[0129] Among them, V represents the node set of the training level government lineage graph, f d All attention coefficients from high level to low level are accumulated. Directional loss f d It helps the circular dependency detection model to learn the attention coefficient distribution that conforms to the direction of the graph structure in the directed graph (training hierarchical government bloodline graph), and better capture the hierarchical and directional information in the graph.
[0130] exist Figure 3 In the example shown, the loop dependency node detection effect of the loop dependency detection model obtained through training is experimentally verified, and the attention coefficient matrix is obtained as shown in Table 3:
[0131] Table 3
[0132] From / To A B C D E ... A 0.00 0.55 0.90 0.55 0.55 ... B 0.90 0.00 0.55 0.45 0.45 ... C 0.55 0.90 0.00 0.90 0.90 ... D 0.01 0.01 0.01 0.00 0.01 ... E 0.01 0.01 0.01 0.01 0.01 ... ... ... ... ... ... ... 0.00
[0133] From the analysis of Table 3, we can see that the closed-loop path formed by high attention coefficients may be a set of circular dependency nodes. As mentioned above:
[0134] A->B:0.9
[0135] B->C:0.9
[0136] C->A:0.9
[0137] The high attention coefficient (≥0.8) of this path clearly indicates the existence of circular dependency, which shows that the circular dependency detection model provided by the present invention can accurately identify circular dependency nodes.
[0138] In a preferred embodiment, in step S1, obtaining a hierarchical government lineage graph consisting of table nodes and field nodes includes:
[0139] Step S11, constructing a government data lineage graph, wherein the node set of the government data lineage graph at least includes field nodes, table nodes, file nodes, and database nodes;
[0140] Step S12, extracting a hierarchical government affairs lineage graph consisting of table nodes and field nodes from the government affairs data lineage graph.
[0141] In this embodiment, compared with the traditional timed analysis lineage solution, in order to present the lineage information of government data in a more real-time manner, preferably, in step S11, constructing a lineage map of government data includes:
[0142] Step E1, real-time collection of development data generated by government data developers during the development process. Preferably, the development data includes multiple data scripts and various semi-structured and unstructured data. The data scripts may include SQL scripts and SHELL scripts.
[0143] Step E2, using the message middleware to receive and parse the development data, obtain a node set, node attributes and edge set; the edge in the edge set is a directed edge, which is used to represent the dependency relationship between two connected nodes. The message middleware is preferably, but not limited to, kafka and rabbitMQ, which is used to decouple the entire collection and data parsing process, parse the development data into libraries, tables, fields, files, file keywords, and their inputs and outputs for graph structure storage.
[0144] Step E3, input the node set, node attributes and edge set into the graph database to obtain the government data lineage graph.
[0145] In this embodiment, the process of constructing the government data lineage map can refer to Figure 4 This embodiment utilizes traditional data lineage combined with message middleware to achieve decoupled operations and real-time analysis of receiving data scripts, parsing data, and writing into the graph database.
[0146] In a preferred embodiment, in order to automatically identify null value dependencies and missing dependencies, after obtaining the node set, node attributes and edge set, the method further includes:
[0147] For the field node in the node set, determine whether the field name of the field node is NULL. If the field name of the field node is NULL, the node that depends on the field node is considered to be a null value dependent node.
[0148] and / or,
[0149] For a table node in a node set, the table name of the table node is queried in the metadata database. If the query fails, the table node is considered to be a transient table, and the nodes that depend on the table node and the nodes that depend on the field nodes related to the table node are considered as missing dependent nodes. The field nodes related to the table node represent the fields in the transient table corresponding to the table node.
[0150] In this embodiment, referring to the above Figure 3In the example, the field M2 of the target table C is generated by L3 of the intermediate table B, N1 of the intermediate table D, and NULL value. NULL means empty and does not contain any information. Therefore, the table node C and the field node M2 are considered to be null-dependent nodes, expressed as: B->C, D->C, Null->C; (L3, N1, Null)->M2. Therefore, null-dependency is defined as the node dependency on the NULL field in the hierarchical government lineage graph. It is considered that there is a null-dependency, and the table nodes and field nodes that depend on the NULL field are called null-dependent nodes, such as Figure 3 The middle node M2 and the table node C are null value dependent nodes.
[0151] In this embodiment, referring to the above Figure 3 For example, field M3 of target table C is generated by field O1 of memory table V. When the task is completed, the memory data will disappear, and the persistent information M3 cannot be found in data table C. It is considered that there is a missing dependency, which is expressed as: V->C, E->C; O1->M3, P3->M3. In fact, this V table does not exist in the database. Therefore, missing dependency can be defined as: if there are nodes in the hierarchical government lineage graph that depend on the transient table and the fields in the transient table, it is considered that there is a missing dependency relationship, and the nodes that depend on the transient table and the fields in the transient table are regarded as missing dependency nodes, such as Figure 3 The table node C and the field node M3 are both missing dependent nodes. A transient table refers to a table that cannot persist and does not exist in the database.
[0152] There are still different entity names in the data lineage map, but their meanings are almost the same in actual work, which makes it impossible for data quality managers to unify and standardize entities when auditing data, and thus cannot solve data quality problems. Since it is impossible to determine whether these entities have the same meaning, it is impossible to effectively track and identify them at the source to achieve accurate data fusion. Therefore, in a preferred implementation, after obtaining the government data lineage map, it also includes:
[0153] Step F1, perform entity recognition on multiple data scripts, determine the domain to which the recognized entity belongs, and perform domain marking on the corresponding nodes (i.e., the nodes corresponding to the entity names) in the government data lineage map according to the domain to which they belong.
[0154] In this embodiment, it is preferred but not limited to using an existing NLP named entity recognition method, such as an entity recognition algorithm based on a BiLSTM-CRF network. In one example, the following fields are included:
[0155] Basic information: name, mobile phone number, address, ID card, etc.
[0156] Financial field: borrower, loan amount, loan time, loan type, depositor, deposit amount, etc.
[0157] Telecommunications field: communication number, communication type, communication duration, etc.
[0158] Set domain attributes in the government data lineage graph and label the node domains:
[0159] Domain: basic information, finance, telecommunications, urban construction, etc. The markup can be:
[0160] 'BASE':'name','sex','berth','basic_info','address','phone','phone_number' and other words;
[0161] 'FINANCE':'loan_name','loan_sex','loan_amount','loan_basic','loan_time','loan_type','deposit_name','deposit_time', etc.;
[0162] 'TELECOM':'telecom_name','telecom_sex','telecom_long_time','telecom_type','telecom_number', etc.
[0163] By identifying entities, determining the domain of the identified entities, and marking the determined domains at the nodes corresponding to the entities in the graph storage structure, it helps to identify whether the node dependencies are normal.
[0164] Step F2, calculate the similarity of entities in the same field. When the similarity of multiple entities in the same field reaches a preset similarity threshold, the multiple entities are recorded as multiple similar entities, the standard entity names of the multiple similar entities are determined, and the standard entity names are marked at the corresponding nodes of multiple similar entities in the government data lineage map.
[0165] In this embodiment, the preset similarity threshold can be set based on experience, and the similarity preferably adopts, but is not limited to, an existing text similarity algorithm, which will not be described in detail here.
[0166] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "an implementation", "a preferred implementation" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0167] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for detecting data lineage using AI technology, characterized in that: include: Obtain a hierarchical government relationship graph consisting of table nodes and field nodes, and generate an adjacency matrix and attribute features of each node based on the hierarchical government relationship graph; Input the adjacency matrix and the attribute features of each node into the circular dependency detection model to obtain the circular dependency detection results; The circular dependency detection model includes: The attention module uses the attention weight matrix to process the attribute features of each node to obtain the attention features of each node, and calculates the attention coefficient between any two nodes based on the node attention features and the adjacency matrix; Feature aggregation module, which obtains the aggregated features of each node based on the attention weight matrix, the attribute features of the node and the attention coefficient between nodes; The classification prediction module classifies the aggregated features of each node to obtain the circular dependency detection results of each node.
2. The method for detecting data lineage using AI technology as claimed in claim 1, characterized in that: In the step where the attention module calculates the attention coefficient between any two nodes based on the attention features of the nodes and the adjacency matrix, the attention coefficient calculation process between node i and node j includes: Concatenate the attention feature of node i and the attention feature of node j to obtain a first concatenated feature, multiply the first concatenated feature by the auxiliary matrix to obtain a first transformed feature, and perform a first function processing on the first transformed feature to obtain a first function processed value; Determine a set of neighbor nodes of node i based on the adjacency matrix, concatenate the attention feature of node i with the attention feature of each neighbor node in the neighbor node set to obtain a plurality of second concatenated features, multiply the plurality of second concatenated features with the auxiliary matrix to obtain a plurality of second converted features, perform second function processing on the plurality of second concatenated features to obtain a plurality of second function processed values, and accumulate the plurality of second function processed values to obtain a second function processed accumulated value; The attention coefficient between node i and node j is obtained by dividing the value processed by the first function by the accumulated value processed by the second function; Among them, i and j are node indexes, and both i and j are positive integers.
3. The method for detecting data lineage using AI technology as claimed in claim 1, characterized in that: The feature aggregation module obtains the aggregated feature of node i based on the attention weight matrix, the attribute features of the node and the attention coefficient between nodes: Determine the set of neighbor nodes of node i based on the adjacency matrix; Multiply the attention coefficient between node i and each neighbor node, the attribute characteristics of each neighbor node, and the attention weight matrix to obtain the fusion feature value of node i and each neighbor node; Accumulate the fused feature values of node i and all neighboring nodes to obtain the accumulated fused feature value; The accumulated fusion feature values are processed by the second function to obtain the aggregated feature of node i.
4. The method for detecting data lineage using AI technology as claimed in claim 1, characterized in that: Classification prediction module execution: Multiply the classification weight matrix by the aggregated features of each node to obtain the prediction score of each node belonging to each node category, including circular dependencies and non-circular dependencies; The predicted probability of each node belonging to each node category is determined based on the predicted score of each node belonging to each node category.
5. A method for detecting data lineage using AI technology as described in any one of claims 1 to 4, characterized in that: The training process of the circular dependency detection model includes: Obtain a training-level government affairs lineage graph, and set true labels for nodes of the training-level government affairs lineage graph, where the true labels include circular dependencies and non-circular dependencies; Generate a training adjacency matrix and training attribute features of each node based on the training hierarchical government bloodline graph; Build a circular dependency detection model; The cyclic dependency detection model is iteratively trained using the training adjacency matrix and the training attribute features of each node until the training stop condition is reached. In the iterative training, the loss function is calculated, and the attention weight matrix, auxiliary matrix and classification weight matrix are updated according to the gradient of the loss function. The loss function includes the standard loss term Loss s : Loss s =-∑ i∈V ∑ c∈C y ic log(p ic ) Among them, V represents the node set of the training-level government lineage graph, i is the node index in the training-level government lineage graph, C represents the node category set, c represents the node category index, and y ic represents the true label value of node i in node category c, p ic Represents the predicted probability of node i in node category c output by the circular dependency detection model.
6. The method for detecting data lineage using AI technology as claimed in claim 5, characterized in that: The loss function also includes the cyclic dependency loss f c , sparsity loss f s and directional loss f d The weighted sum of Cyclic dependency loss f c It is the cumulative sum of the minimum attention coefficients within the group of all group circular dependency nodes detected during the circular dependency detection model training; Sparsity loss f s is the average of the absolute values of the attention coefficients between the nodes at both ends of all edges in the training-level government relations graph; Directional loss f d It is the cumulative sum of attention coefficients between nodes in the node set of the training hierarchical government lineage graph.
7. A method for detecting data lineage using AI technology as described in claim 1, 2, 3, 4, or 6, characterized in that: The step of obtaining a hierarchical government affairs lineage graph consisting of table nodes and field nodes includes: Constructing a government data lineage graph, the node set of which at least includes field nodes, table nodes, file nodes, and database nodes; A hierarchical government lineage graph consisting of table nodes and field nodes is extracted from the government data lineage graph.
8. The method for detecting data lineage using AI technology as claimed in claim 7, characterized in that: The construction of the government data lineage map includes: Real-time collection of development data generated by government data developers during the development process; Utilize the message middleware to receive and parse the development data, and obtain the node set, node attributes, and edge set; the edges in the edge set are directed edges, which are used to represent the dependency relationship between two connected nodes; Input the node set, node attributes and edge set into the graph database to obtain the government data lineage graph.
9. The method for detecting data lineage using AI technology as claimed in claim 8, characterized in that: After obtaining the node set, node attributes and edge set, it also includes: For the field node in the node set, determine whether the field name of the field node is NULL. If the field name of the field node is NULL, the node that depends on the field node is considered to be a null value dependent node. and / or, For a table node in a node set, the table name of the table node is queried in the metadata database. If the query fails, the table node is considered to be a transient table, and the nodes that depend on the table node and the nodes that depend on the related field nodes of the table node are regarded as missing dependent nodes.
10. The method for detecting data lineage using AI technology as claimed in claim 8 or 9, characterized in that: After obtaining the government data lineage map, it also includes: Perform entity recognition on multiple data scripts, determine the domain to which the recognized entities belong, and mark the corresponding nodes in the government data lineage map according to the domain to which they belong; Calculate the similarity of entities in the same field. When the similarity of multiple entities in the same field reaches a preset similarity threshold, record the multiple entities as multiple similar entities, determine the standard entity names of the multiple similar entities, and mark the standard entity names at the corresponding nodes of multiple similar entities in the government data lineage map.