Data asset graph sampling method
By performing multi-dimensional calculation marking and multiple sub-random wandering on the data asset graph, the problem of missing important assets and destroying business structure in the sampling results in the existing technology is solved, and the effective simplification of the data asset graph and the improvement of data management efficiency is achieved.
Patent Information
- Application Number
- CN202510513612.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-19
AI Technical Summary
The existing data asset graph sampling method fails to effectively consider the business characteristics of the data asset graph, resulting in the sampling results being prone to loss of important assets and destroying the business structure. At the same time, it is very random and lacks a biased selection mechanism for important nodes.
By performing multi-dimensional calculation marking of data table nodes, screening seed nodes, and using multiple seed random walk methods for sampling, a more random walk strategy is designed to ensure that the sampling results retain key business characteristics and structural integrity.
It effectively simplifies the data asset graph, reduces visual confusion, improves data management and analysis efficiency, while ensuring data security and privacy, and retains key information and business structures.
Smart Images

Figure CN120509466A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visualization of data governance, and in particular relates to a method for sampling a data asset graph. Background Art
[0002] In the field of data processing and analysis, graphs, as powerful data structures, are often used to represent complex data asset relationships. Data asset graphs can represent the relationships between data assets and the business structure, making them a key tool for data governance. Visualizing data asset graphs using node connection diagrams can help users intuitively analyze asset relationships and business structures. However, current data asset graphs are large in scale and complex in business structure, and the resulting visualizations often exhibit visual clutter, such as point-edge clustering, overlap, and overlays.
[0003] As data volumes continue to expand, effectively simplifying data asset graphs to improve computational and visualization efficiency becomes increasingly important. Existing techniques utilize graph sampling as the primary means of simplifying data asset graphs at the data level. Graph sampling methods can be categorized into random and feature-driven graph sampling algorithms, each employing different strategies to select nodes and edges within the graph to achieve simplification. Random graph sampling methods offer the advantages of high sampling efficiency and low algorithmic complexity, but they can lead to significant sample randomness and lack a biased selection mechanism for important nodes, resulting in significant differences in the resulting data asset graphs from each sampling. Feature-driven graph sampling algorithms, on the other hand, excel at preserving the original graph's features. For example, algorithms such as IS, RD, and SV effectively preserve important nodes and topological structures in the graph through various approaches. While graph sampling methods can simplify data asset graphs at the data level, reducing visual clutter, existing methods fail to consider the business characteristics of data asset graphs, leading to the sampling results being prone to missing important assets and disrupting business structures. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for sampling a data asset graph to reduce the scale of the data asset graph, thereby reducing visual clutter, improving the visualization effect of the graph, and enhancing the efficiency of data management and analysis.
[0005] The technical solution adopted by the present invention is a method for sampling a data asset graph, comprising the following steps:
[0006] Step S1: Acquire asset data and perform data cleaning and desensitization processing;
[0007] Step S2, performing calculation marking on the data table nodes;
[0008] Step S3, screening seed nodes;
[0009] Step S4: Design the walking process of seed node sampling.
[0010] Furthermore, the specific steps of S2 are as follows:
[0011] S21, calculate the number of other data table nodes that a data table node is connected to through the job node or directly connected to, find other data tables that the data table node is connected to through the job node or directly connected to, count the number of data tables connected to the data table node, and record it as an attribute of the data table node. The specific formula is as follows:
[0012] C lineage (T i )=count(T connected (T i ))
[0013] Among them, C lineage (T i ) represents the data table node T i The number of nodes connected to other data tables, T connected (T i ) represents the data table node T i Connected data table node, count(·) indicates statistics;
[0014] S22, calculate the number of field nodes connected to a data table node, count the number of fields in the cluster structure where it is located, and record it as the attribute of the data table node. The specific formula is as follows:
[0015] C capacity (T i )=count(F connected (T i ))
[0016] Among them, C capacity (T i ) represents the data table node T i The number of connected field nodes, F connected (T i ) represents the data table node T i Connected field nodes, count(·) indicates statistics;
[0017] S23 calculates the number of key fields in a data table node. Field nodes in the same cluster structure are marked with the same group attribute to indicate that they belong to the same data table. Primary keys and foreign keys are directly identified through node attributes in the data set. Key field attribute identifiers are added to primary and foreign key field nodes, and non-key field attribute identifiers are added to other field nodes. The number of key fields is counted and recorded as the attribute of the data table node. The specific formula is as follows:
[0018] C key (T i)=count(K connected (T i ))
[0019] Among them, C key (T i ) represents the data table node T i The number of key fields, K connected (T i ) represents the data table node T i Connected key field nodes, count(·) indicates statistics;
[0020] S24, calculate the data flow structure. For a data table node, check whether it is a data table end node in the data flow structure. If it is, add a data flow structure attribute identifier and traverse the job nodes associated with it. For the job nodes connected to the same data table node, mark them as the same group of attributes. Count the number of data flow structures and data flow substructures in the data flow structure and record them as the attributes of the data table node. If it is not, check the next data table node. The specific formula is as follows:
[0021] C flow (T i )=count(FL connected (T i ))
[0022] C subflow (T i )=count(J connected (T i ))
[0023] Among them, C flow (T i ) represents the data table node T i The number of associated data flow structures, FL connected (T i ) represents the data table node T i The associated data flow structure, C subflow (T i ) represents the data table node T i The number of associated data flow substructures, J connected (T i ) is the node T of the data table i Connected data operation nodes, count(·) indicates the total number of statistics;
[0024] S25, calculate the data interpretation structure, for a data table node, detect whether it is a data table end node in the data interpretation structure, if it is detected as yes, add a data interpretation structure attribute identifier, and mark its associated field node, corresponding business attribute node and logical entity node with the same group attribute, if it is detected as no, continue to detect the next data table node.
[0025] Furthermore, the specific process of screening in S3 is as follows:
[0026] S31, calculate the importance of the data table node, the specific formula is as follows:
[0027] I=μ1·L+μ2·C+μ3·K+μ4·F+μ5·D
[0028] Among them, I represents the importance of the data table node, L represents the data table node T i Normalized value of the number of connected data table nodes, C represents the data table node T i Normalized value of the number of connection fields, K represents the data table node T i Normalized value of the number of key fields, F represents the data table node T i Normalized value of the number of data flow substructures, D represents the data table node T i Whether to associate the data interpretation structure, D is 0 for no association, D is 1 for association;
[0029] S32, calculate the number of seed nodes, the number of seed nodes should meet Among them, seed represents the number of seed nodes, |·| represents the absolute value, and Φ represents the sample space capacity;
[0030] S33, screening seed nodes, sorting the initial candidate seed nodes in descending order according to the node importance I in the data table, and screening the first seed nodes as the final seed nodes.
[0031] Furthermore, in S4, a multi-seed random walk method is adopted. Each seed node sends a sampling thread to perform a random walk. When a sampling thread reaches the upper limit of its sample space, it stops sampling. The sample graph formed by sampling multiple seed nodes individually is constructed into an induced subgraph. The specific walking process of seed node sampling is as follows:
[0032] S41, random walk centered on the data table, sampling other types of nodes that are structurally attached to the data table nodes;
[0033] S42, data table node importance guided biased random walk;
[0034] S43, random sampling of business structures.
[0035] Furthermore, in S43, the specific process of sampling the business structure is as follows:
[0036] S43a, for the data table node currently visited, determine whether there is an associated business structure based on the attribute tag of the data table node;
[0037] S43b, for the cluster structure, key field nodes are added to the sample space, and the remaining field nodes are randomly selected according to the sampling rate to be added to the sample space;
[0038] S43c, for the data flow structure, randomly select job nodes according to the sampling rate to add to the sample space;
[0039] S43d, for the data interpretation structure, add the logical entity end node at the other end to the sample space, add the key field node and its corresponding business attribute node to the sample space, and randomly select the remaining pairs of field nodes and business attribute nodes according to the sampling rate, and retain at least one substructure.
[0040] The beneficial effects of the present invention are:
[0041] 1. Through multi-dimensional calculation and comprehensive evaluation, the present invention can accurately identify and retain the key business characteristics in the data asset graph, ensuring that the sampled induced subgraph truly reflects the business logic and structure of the original data.
[0042] 2. The BDAG algorithm of the present invention adopts a rule-guided biased sampling strategy, which effectively avoids the problem of random selection algorithms losing important nodes due to lack of directionality. It can maintain the connectivity of the business structure during the sampling process. At the same time, it ensures that the sampling covers the entire graph through a multi-seed random walk mechanism, overcoming the defect of traversal algorithms such as FF and BF that they lose entire areas due to local sampling.
[0043] 3. In terms of data security, the present invention effectively prevents the leakage of sensitive information through desensitization processing, ensuring data security and privacy.
[0044] 4. The method of the present invention can make the sampling results cover the entire map and retain key information, preventing the business structure from being lost or destroyed due to incomplete sampling, while maintaining the business characteristics and statistical characteristics of the data asset map, improving the performance of data processing and analysis, and providing enterprises and organizations with more advanced data governance means. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 It is a flow chart of the present invention.
[0047] Figure 2 This is a diagram of the data assets, main node types, and business structure of the present invention.
[0048] Figure 3 This is a schematic diagram of a biased random walk guided by the importance of a data table according to the present invention.
[0049] Figure 4 It is a schematic diagram of the business structure sampling of the present invention.
[0050] Figure 5 a~f are the original image of dataset 8 and the sample images obtained by using the RNE algorithm, RD algorithm, RW algorithm, BF algorithm and the BDAG algorithm of the present invention.
[0051] Figure 6 a~f are the original image of dataset 5 and the sample images obtained by using the RN algorithm, TIES algorithm, FF algorithm, DPL algorithm and the BDAG algorithm of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] Example
[0054] The embodiment of the present invention provides a method for sampling a data asset graph, the flow chart of which is as follows: Figure 1 As shown, the steps include:
[0055] Step S1: Obtain asset data and perform data cleansing and desensitization. Asset data covers typical application scenarios requiring data governance, such as infrastructure management and customer operations services. This raw data contains numerous node types and involves actual business confidentiality, so data cleansing and desensitization are necessary.
[0056] Step S2: Calculate and mark the data table nodes so that the seed nodes can better prevent the business characteristics of the data asset graph from being destroyed during the screening and sampling process. The specific steps for calculating and marking the data table nodes are as follows:
[0057] S21 calculates the number of other data table nodes that a data table node is connected to through a job node or directly connected to (this reflects the importance and influence of the data table in the entire business process), finds other data tables that the data table node is connected to through a job node or directly connected to, counts the number of data tables connected to the data table node, and records it as an attribute of the data table node. The specific formula is as follows:
[0058] C lineage (T i )=count(T connected (T i ))
[0059] Among them, C lineage (T i ) represents the data table node T i The number of nodes connected to other data tables, T connected (T i ) represents the data table node T i Connected data table node, count(·) indicates statistics.
[0060] S22 calculates the number of field nodes connected to a data table node, counts the number of fields in the cluster structure where it is located, and records it as the attribute of the data table node. Data table nodes with more field nodes usually contain more information or support more business operations and may be more important in the business process. The specific formula is as follows:
[0061] C capacity (T i )=count(F connected (T i ))
[0062] Among them, C capacity (T i ) represents the data table node T i The number of connected field nodes, F connected (T i ) represents the data table node T i Connected field nodes, count(·) indicates statistics.
[0063] S23, calculate the number of key fields of a data table node. Data table nodes and connected field nodes usually form a cluster structure. For a cluster structure, field nodes on the same cluster structure are marked with the same group attribute to indicate that they belong to the same data table. Primary keys and foreign keys are directly identified through node attributes in the data set. Key field attribute identifiers are added to primary and foreign key field nodes, and non-key field attribute identifiers are added to other field nodes. The number of key fields is counted and recorded as the attribute of the data table node. The specific formula is as follows:
[0064] Ckey (T i )=count(K connected (T i ))
[0065] Among them, C key (T i ) represents the data table node T i The number of key fields, K connected (T i ) represents the data table node T i Connected key field nodes, count(·) indicates statistics.
[0066] S24, calculate the data flow structure. The data in the data table is extracted, converted, and loaded to form a new data table. This flow process is usually represented in the data asset diagram as a data flow structure connecting two large cluster structures. A data flow structure usually contains multiple data flow substructures with the same data table end node. A data flow substructure contains a job node and two data table nodes. There are usually multiple data flow substructures between the two data table nodes, such as Figure 2 As shown, there are four data flow substructures between data table nodes T1 and T3.
[0067] For a data table node, check whether it is a data table end node in the data flow structure. If so, add a data flow structure attribute identifier and traverse the job nodes associated with it. For the job nodes connected to the same data table node, mark them with the same group of attributes. Count the number of data flow structures and data flow substructures in the data flow structure and record them as the attributes of the data table node. If not, check the next data table node. The specific formula is as follows:
[0068] C flow (T i )=count(FL connected (T i ))
[0069] C subflow (T i )=count(J connected (T i ))
[0070] Among them, C flow (T i ) represents the data table node T i The number of associated data flow structures, FL connected (T i ) represents the data table node T i The associated data flow structure, C subflow (T i ) indicates the data table node Ti The number of associated data flow substructures, J connected (T i ) is the node T of the data table i Connected data job nodes, count(·) indicates the total number of statistics.
[0071] S25, calculate the data interpretation structure. Logical entities are used to interpret the meaning of data tables, and business attributes are used to interpret the meaning of fields. In the data asset diagram, it is usually represented as a data interpretation structure. A data interpretation structure usually contains multiple data interpretation substructures with the same data table end node and entity logical end node. A data interpretation substructure contains a data table node, a field node, a business attribute node, and a logical entity node, such as Figure 2 As shown, there are 8 data interpretation substructures between the data table node T8 and its logical entity node.
[0072] For each data table node, check whether it is a data table end node in the data interpretation structure. If so, add a data interpretation structure attribute identifier and mark the associated field node, corresponding business attribute node, and logical entity node with the same group attribute. If not, continue checking the next data table node. Considering that the number of data interpretation substructures is equivalent to the number of fields in a data table node, there is no need to count the number of data interpretation substructures associated with the data table node.
[0073] Step S3: Screening seed nodes. Screening seed nodes is an important step in walking graph sampling. The algorithm of the present invention screens out an appropriate number of data table nodes from the original graph as seed nodes. This is mainly due to the fact that the data asset graph is an association graph with the data table as the core. The calculation of all business rules is centered on the data table nodes, and other types of nodes are attached to the data table nodes. Screening seed nodes first requires calculating the importance of the data table nodes, then calculating the number of seed nodes, and finally screening out the seed nodes. The specific screening process is as follows:
[0074] S31, calculate the importance of the data table node, based on the five evaluation indicators obtained in step S2: the number of data table nodes connected to other data table nodes, the number of data table node connection fields, the number of key fields of the data table node, the number of data flow substructures of the data table node, and whether the data interpretation structure is associated, and derive the weight μ of each group of evaluation indicators based on the attention given by the user i (i=1,2,3,4,5), used to calculate the importance of the data table node. The specific formula is as follows:
[0075] I=μ1·L+μ2·C+μ3·K+μ4·F+μ5·D
[0076] Among them, I represents the importance of the data table node, L represents the data table node Ti Normalized value of the number of connected data table nodes, C represents the data table node T i Normalized value of the number of connection fields, K represents the data table node T i Normalized value of the number of key fields, F represents the data table node T i Normalized value of the number of data flow substructures, D represents the data table node T i Whether to associate the data interpretation structure. D is 0 for no association, and D is 1 for association.
[0077] S32, calculate the number of seed nodes. The number of seed nodes is related to the sampling rate. The higher the sampling rate, the larger the sample space, and the more seed nodes that can be accommodated. However, seed nodes will occupy the sample space. Too many seed nodes will cause the sampling process to end prematurely, making it difficult to effectively access the entire graph. Therefore, the number of seed nodes cannot be more than the sampling quota that can be evenly allocated to each seed. For example, if the sample space is 49, the number of seed nodes should be less than 7, so that the average number of samples allocated to each seed is greater than 7. Therefore, the number of seed nodes should satisfy Among them, seed represents the number of seed nodes, |·| represents the absolute value, and Φ represents the sample space capacity.
[0078] S33, screening seed nodes, according to the data table node importance I, the initial candidate seed nodes are sorted in descending order, and the first seed nodes are screened out as the final seed nodes. The seed nodes determined in this step have high importance in terms of business characteristics, which is conducive to maintaining the business characteristics of the original graph during sampling.
[0079] Step S4, design the walk process of seed node sampling. Based on the seed nodes determined in S3, a multi-seed random walk method is adopted. Each seed node sends a sampling thread to perform a random walk. When a sampling thread reaches the upper limit of its sample space, it stops sampling. The sample graph formed by sampling multiple seed nodes separately is constructed into an induced subgraph, such as Figure 4 As shown in the right part, the induced subgraph is the final sample graph. The specific walking process of seed node sampling is as follows:
[0080] S41,random walk with data table as the center, data table nodes are associated with business rules, which can cover all business structures, such as Figure 3 As shown, the walk is performed on the data table nodes T6 and T8. The data table nodes are associated with the structures DF1, C1, DF2, DF3, DF4, and DE1, and sampling of other types of nodes that are structurally attached to the data table nodes can be achieved.
[0081] S42, the importance of the data table nodes guides the biased random walk. The importance of the data table nodes is calculated using the formula in step S31. The importance of the data table nodes is used to guide the walk process, so that the data table nodes with high importance ranking have a higher access probability, and more nodes in the important business structure are retained to promote the maintenance of the business characteristics distribution of the original graph. The walk process is as follows Figure 3 As shown in the figure, the size of the data table node indicates its importance. The greater the importance, the larger the radius. Figure 3 It can be seen that the importance of data table node T8 is greater than that of data table node T3. The current walk is to data table node T6. The next walk will be guided to walk to T8 based on the importance ranking.
[0082] S43, random sampling of business structures, performing a biased random walk with the data table node as the center, achieving traversal and sampling of the data table nodes, but not sampling the structures attached to the data table nodes. Therefore, the present invention samples the business structure based on the walk with the data table node as the center. The specific process is as follows:
[0083] S43a, for the data table node currently wandered to, determine whether there is an associated business structure based on the attribute tag of the data table node in S2, and use different random sampling strategies for different business structures.
[0084] S43b, for the cluster structure, the key field nodes are added to the sample space, and the remaining field nodes are randomly selected according to the sampling rate to add to the sample space, such as Figure 3 The middle cluster structure C1 contains 8 field nodes, one of which is a key field node. If the sampling rate is 40%, one keyword node is retained and two common field nodes are randomly retained.
[0085] S43c, for the data flow structure, randomly select the job nodes according to the sampling rate to add to the sample space (at least one substructure is retained), and the business structure sampling is as follows Figure 4 As shown, Figure 4 The data flow structure DF3 in contains 4 data flow substructures. If the sampling rate is 40%, 2 job nodes are randomly retained.
[0086] S43d, for the data interpretation structure, add the logical entity end node at the other end to the sample space, the key field node and its corresponding business attribute node to the sample space, and randomly select the remaining pairs of field nodes and business attribute nodes according to the sampling rate. At least one substructure must be retained, such as Figure 4 As shown, Figure 4 The data explanation structure DE1 contains 8 data explanation substructures. If the sampling rate is 40%, 3 pairs of field nodes and business attribute nodes are randomly retained.
[0087] After detailed processing and calculation of the above steps S1 to S4, the present invention successfully constructed an induced subgraph that not only retains the core business characteristics of the original data asset graph but also meets data governance requirements, and obtained a highly optimized data asset graph sampling result.
[0088] This invention achieves effective sampling of complex data asset graphs by calculating the business influence of data table nodes across multiple dimensions, comprehensively assessing node importance, and designing a biased random walk strategy. This method not only preserves the key business characteristics of the original graph but also ensures data security and privacy. Ultimately, this invention achieves its goal of optimizing data governance processes and improving data utilization efficiency, providing powerful data support and governance tools for application scenarios such as infrastructure management and customer operations services.
[0089] Experimental verification
[0090] In terms of important table retention, the BDAG method of the present invention avoids the problem of random selection algorithms (such as RN and RNE) losing important nodes due to lack of directionality through rule-guided biased sampling. At the same time, BDAG adopts a multi-seed random walk mechanism to ensure that the sampling covers the entire graph, overcoming the defect of traversal algorithms such as FF and BF losing the entire area due to local sampling. In terms of key field retention, BDAG first selects the nodes connected to the key fields, avoiding the problem of algorithms such as RNE losing key fields due to randomness (such as Figure 5 As shown, Figure 5 The original graph (a) of Example 8 and the sample graphs obtained by using the RNE algorithm (b), RD algorithm (c), RW algorithm (d), BF algorithm (e) and the BDAG algorithm (f) of the present invention are included in the figure. The node positions of the sample graphs remain the same as the original graph. Figure 1 The original graph has 366 nodes, 405 edges, and a sampling rate of 40%) or traversal algorithms such as BF and RD miss key fields due to local sampling. In addition, BDAG's multi-seed strategy makes it superior to single-seed random walk algorithms (such as RW) in covering the key fields of the entire graph. In terms of business structure integrity, BDAG maintains the connectivity of the business structure during the sampling process through data flow compression rules and data interpretation compression rules, avoiding the destruction of the business structure by algorithms such as RN due to the lack of induced subgraph steps (such as Figure 6 As shown, Figure 6 The original graph (a) of Example 5 and the sample graphs obtained by using the RN algorithm (b), TIES algorithm (c), FF algorithm (d), DPL algorithm (e) and the BDAG algorithm (f) of the present invention are included in the graph. The node positions of the sample graphs remain the same as the original graph. Figure 1 The original graph has 1122 nodes, 1301 edges, and a sampling rate of 40%.
[0091] At the same time, BDAG's multi-seed random walk mechanism ensures that the entire graph's business structure is not missed, while traversal algorithms such as FF may cause the business structure of an entire area to be lost due to incomplete sampling. In addition, BDAG performs best in statistical property preservation indicators (BCD, RCCD, JI), thanks to its multi-seed random walk and induced subgraph strategies, which can better maintain network connectivity, neighbor similarity, and betweenness centrality. Although it performs moderately in degree distribution indicators (DD, NSDD), this is a reasonable trade-off made to prioritize the protection of key business nodes and structures. In terms of time performance, although BDAG is not as efficient as random selection algorithms, its time consumption is significantly lower than algorithms such as DPL and SGP that rely on complex calculations (such as community detection).
[0092] Overall, BDAG, through its rule-guided biased sampling, full-graph coverage with multi-seed random walks, and business-structure-first design, comprehensively outperforms existing methods in core metrics such as important table retention, key field preservation, and business structure integrity, while also maintaining statistical properties, making it a superior solution for sampling complex data graphs. Its superiority is reflected not only in experimental results but also in its flexible optimizability, providing an efficient and reliable sampling method for real-world business scenarios.
[0093] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0094] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A method for sampling a data asset graph, characterized in that the steps include: Step S1: Acquire asset data and perform data cleaning and desensitization processing; Step S2, performing calculation marking on the data table nodes; Step S3, screening seed nodes; Step S4: Design the walking process of seed node sampling.
2. A data asset graph sampling method according to claim 1, characterized in that: The specific steps of S2 are as follows: S21, calculate the number of other data table nodes that a data table node is connected to through the job node or directly connected to, find other data tables that the data table node is connected to through the job node or directly connected to, count the number of data tables connected to the data table node, and record it as an attribute of the data table node. The specific formula is as follows: C lineage (T i )=count(T connected (T i )) Among them, C lineage (T i ) represents the data table node T i The number of nodes connected to other data tables, T connected (T i ) represents the data table node T i Connected data table node, count(·) indicates statistics; S22, calculate the number of field nodes connected to a data table node, count the number of fields in the cluster structure where it is located, and record it as the attribute of the data table node. The specific formula is as follows: C capacity (T i )=count(F connected (T i )) Among them, C capacity (T i ) represents the data table node T i The number of connected field nodes, F connected (T i ) represents the data table node T i Connected field nodes, count(·) indicates statistics; S23 calculates the number of key fields in a data table node. Field nodes in the same cluster structure are marked with the same group attribute to indicate that they belong to the same data table. Primary keys and foreign keys are directly identified through node attributes in the data set. Key field attribute identifiers are added to primary and foreign key field nodes, and non-key field attribute identifiers are added to other field nodes. The number of key fields is counted and recorded as the attribute of the data table node. The specific formula is as follows: C key (T i )=count(K connected (T i )) Among them, C key (T i ) represents the data table node T i The number of key fields, K connected (T i ) indicates the data table node T i Connected key field nodes, count(·) indicates statistics; S24, calculate the data flow structure. For a data table node, check whether it is a data table end node in the data flow structure. If it is, add a data flow structure attribute identifier and traverse the job nodes associated with it. For the job nodes connected to the same data table node, mark them as the same group of attributes. Count the number of data flow structures and data flow substructures in the data flow structure and record them as the attributes of the data table node. If it is not, check the next data table node. The specific formula is as follows: C flow (T i )=count(FL connected (T i )) C subflow (T i )=count(J connected (T i )) Among them, C flow (T i ) indicates the data table node T i The number of associated data flow structures, FL connected (T i ) indicates the data table node T i The associated data flow structure, C subflow (T i ) indicates the data table node T i The number of associated data flow substructures, J connected (T i ) is the node T of the data table i Connected data operation nodes, count(·) indicates the total number of statistics; S25, calculate the data interpretation structure, for a data table node, detect whether it is a data table end node in the data interpretation structure, if it is detected as yes, add a data interpretation structure attribute identifier, and mark its associated field node, corresponding business attribute node and logical entity node with the same group attribute, if it is detected as no, continue to detect the next data table node.
3. The method for data asset graph sampling according to claim 1, characterized in that: The specific process of screening in S3 is as follows: S31, calculate the importance of the data table node, the specific formula is as follows: I=μ1·L+μ2·C+μ3·K+μ4·F+μ5·D Among them, I represents the importance of the data table node, L represents the data table node T i Normalized value of the number of connected data table nodes, C represents the data table node T i Normalized value of the number of connection fields, K represents the data table node T i Normalized value of the number of key fields, F represents the data table node T i Normalized value of the number of data flow substructures, D represents the data table node T i Whether to associate the data interpretation structure, D is 0 for no association, D is 1 for association; μ i , where i=1,2,3,4,5 are weights; S32, calculate the number of seed nodes, the number of seed nodes should meet Among them, seed represents the number of seed nodes, |·| represents the absolute value, and Φ represents the sample space capacity; S33, screening seed nodes, sorting the initial candidate seed nodes in descending order according to the node importance I in the data table, and screening the first seed nodes as the final seed nodes.
4. The method for data asset graph sampling according to claim 1, characterized in that: In S4, a multi-seed random walk method is adopted. Each seed node sends a sampling thread to perform a random walk. When a sampling thread reaches the upper limit of its sample space, it stops sampling. The sample graph formed by sampling multiple seed nodes individually is constructed into an induced subgraph. The specific walking process of seed node sampling is as follows: S41, random walk centered on the data table, sampling other types of nodes that are structurally attached to the data table nodes; S42, data table node importance guided biased random walk; S43, random sampling of business structures.
5. A data asset graph sampling method according to claim 4, characterized in that: In S43, the specific process of sampling the business structure is as follows: S43a, for the data table node currently visited, determine whether there is an associated business structure based on the attribute tag of the data table node; S43b, for the cluster structure, key field nodes are added to the sample space, and the remaining field nodes are randomly selected according to the sampling rate to be added to the sample space; S43c, for the data flow structure, randomly select job nodes according to the sampling rate to add to the sample space; S43d, for the data interpretation structure, add the logical entity end node at the other end to the sample space, add the key field node and its corresponding business attribute node to the sample space, and randomly select the remaining pairs of field nodes and business attribute nodes according to the sampling rate, and retain at least one substructure.