A method and device for determining a root cause quality index and a computer device
By acquiring and analyzing the poor-quality metadata of enterprise data assets and their dependencies, generating data analysis graphs, and selecting root cause indicators of poor quality, the problems of chaotic data standards and inconsistent quality are solved, achieving the effect of rapid location and cost reduction in data management.
Patent Information
- Application Number
- CN202210945799.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-08-08
AI Technical Summary
Enterprise data assets are distributed across multiple systems, resulting in inconsistent data standards and quality, making it difficult to quickly locate problematic data sources and their impact, thus affecting data management efficiency and risk control.
By acquiring the dependencies of poor-quality metadata, relevant dependent metadata and metadata indicators are determined. A data analysis graph is generated using a relational graph algorithm. Root cause poor-quality indicators are selected in response to user commands, and data relationships are displayed through visualization tools.
Quickly locate data problems, improve processing efficiency, reduce processing costs, mitigate data risks, and enhance the accuracy and efficiency of data management.
Smart Images

Figure CN116991951B_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, and computer device for determining a root cause quality index. [Background Technology]
[0002] Data assets help improve enterprise decision-making capabilities and enhance competitiveness. By effectively managing data assets, enterprises can provide better products and services, reduce data asset management costs, and mitigate data risks. As data assets grow in scale and complexity, their management becomes increasingly difficult, while enterprises' reliance on them continues to increase, leading to a growing demand for data management. Because enterprise data assets are typically distributed across multiple systems, they suffer from problems such as inconsistent data standards, varying data quality, and severe data silos between systems. When data asset issues arise, enterprises often cannot analyze the impact and scope of these issues on subsequent processes, hindering effective data asset management.
[0003] In the data processing process, from the source of the data to the final data generation, each step can potentially lead to data quality issues. For example, if the source data is of low quality, and subsequent processing steps fail to perform data quality checks and adjustments, the generated data will also be of low quality. Alternatively, inappropriate processing at any stage can also result in low-quality data generated in subsequent stages.
[0004] Effective data asset management is key to fully unlocking the value of data. Currently, data resources are distributed across multiple enterprise systems, lacking a unified data view. Data managers cannot quickly and accurately locate the data they need, nor can they obtain a macro-level understanding of the quantity and distribution of their data assets. When auditing data quality, quickly identifying problematic data sources and predicting their impact remains a significant challenge for current data operators. [Summary of the Invention]
[0005] In view of this, embodiments of the present invention provide a method, apparatus and computer equipment for determining root cause quality poor index, in order to solve the problems of low data quality and high data risk in the data processing process of the prior art.
[0006] In a first aspect, embodiments of the present invention provide a method for determining a root cause quality index, the method comprising:
[0007] Dependencies for obtaining poor-quality metadata;
[0008] Based on the poor quality metadata and the dependency relationship, the dependency metadata related to the poor quality metadata is determined;
[0009] Multiple metadata metrics corresponding to the poor quality metadata and the dependent metadata were identified;
[0010] A data analysis graph is generated based on the multiple metadata indicators using a relational graph algorithm.
[0011] In response to the user's selection command, the root cause quality index is selected from the data analysis graph.
[0012] In one possible implementation, after selecting the root cause quality index from the data analysis graph in response to a user-input selection command, the method further includes:
[0013] Early warning information is generated based on the root cause quality poor index.
[0014] In one possible implementation, the data analysis graph includes an impact analysis graph, and the root cause quality index includes the metadata index with the most connected branches in the impact analysis graph.
[0015] In one possible implementation, the data analysis graph includes a lineage analysis graph;
[0016] The step of selecting the root cause quality index from the data analysis chart in response to the user's input selection command includes:
[0017] Select the root node from the bloodline analysis diagram;
[0018] Determine whether the root node is a poor quality indicator;
[0019] If the root node is determined to be a poor quality index, then the root node is identified as a root cause poor quality index.
[0020] In one possible implementation, the data analysis graph includes data relationships between at least one metadata indicator;
[0021] After generating the data analysis graph based on the multiple metadata indicators using the relationship graph algorithm, the process further includes:
[0022] A visual data analysis graph is generated based on the data analysis graph using the backpropagation BP neural network algorithm.
[0023] In response to a user-input query command, the data relationship between the at least one metadata indicator is queried through the visual data analysis graph.
[0024] One possible implementation also includes:
[0025] An end-to-end metadata information chain is generated based on the quality difference metadata, the dependency metadata, and the dependency relationship.
[0026] In one possible implementation, the quality difference metadata includes first technical metadata or first business metadata, and the dependency metadata includes second technical metadata or second business metadata;
[0027] The step of generating the end-to-end metadata information chain based on the quality difference metadata, the dependency metadata, and the dependency relationship includes:
[0028] Generate a technology metadata chain based on the first technology metadata and / or the second technology metadata;
[0029] Generate a business metadata chain based on the first business metadata and / or the second business metadata;
[0030] The technical metadata chain and the business metadata chain are merged according to the aforementioned dependencies;
[0031] If the fusion is successful, an end-to-end metadata information chain will be generated;
[0032] If the fusion fails, the technical metadata chain and / or business metadata chain that failed to merge will be repaired through a deep learning model to generate a repaired technical metadata chain and / or business metadata chain.
[0033] Generate an end-to-end metadata information chain based on the repaired technical metadata chain and / or business metadata chain.
[0034] Secondly, embodiments of the present invention provide an apparatus for determining a root cause quality index, the apparatus comprising:
[0035] The acquisition module is used to obtain the dependencies of poor-quality metadata;
[0036] The first determining module is used to determine the dependency metadata related to the quality poor metadata based on the quality poor metadata and the dependency relationship;
[0037] The second determining module is used to determine multiple metadata metrics corresponding to the poor quality metadata and the dependent metadata;
[0038] The first generation module is used to generate a data analysis graph based on the multiple metadata indicators using a relational graph algorithm;
[0039] The selection module is used to select the root cause quality index from the data analysis chart in response to the selection command input by the user.
[0040] Thirdly, embodiments of the present invention provide a computer device, including: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the computer device, cause the computer device to perform a method for determining a root cause quality index in the first aspect or any possible implementation of the first aspect.
[0041] Fourthly, embodiments of the present invention provide a computer-readable storage medium comprising a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to execute the method for determining the root cause quality index in the first aspect or any possible implementation thereof.
[0042] The technical solution provided in this invention involves: acquiring the dependency relationships of poor-quality metadata; determining the dependent metadata related to the poor-quality metadata based on the poor-quality metadata and the dependency relationships; determining multiple metadata indicators corresponding to the poor-quality metadata and dependent metadata; generating a data analysis graph based on the multiple metadata indicators using a relational graph algorithm; and selecting the root cause poor-quality indicator from the data analysis graph in response to a user-input selection command. By using the data analysis graph to determine the root cause poor-quality indicator, the computer device can quickly locate data problems, improve the efficiency of data problem processing, and reduce the cost of data problem processing. [Attached Image Description]
[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 A flowchart illustrating a method for determining a root cause quality index provided in an embodiment of the present invention;
[0045] Figure 2 A flowchart illustrating a method for generating an end-to-end metadata information chain according to an embodiment of the present invention;
[0046] Figure 3 A flowchart illustrating a method for determining predictive triples provided in an embodiment of the present invention;
[0047] Figure 4 A schematic diagram of a triple provided in an embodiment of the present invention;
[0048] Figure 5 A flowchart illustrating a method for determining the shortest path provided in an embodiment of the present invention;
[0049] Figure 6 A schematic diagram of a central indicator and an endpoint indicator provided in an embodiment of the present invention;
[0050] Figure 7 A flowchart illustrating a method for determining the importance of metadata metrics provided in an embodiment of the present invention;
[0051] Figure 8 A schematic diagram of a metadata indicator connection branch provided in an embodiment of the present invention;
[0052] Figure 9 A schematic diagram of an influence analysis diagram provided in an embodiment of the present invention;
[0053] Figure 10 A schematic diagram of a device for determining the root cause quality index provided in an embodiment of the present invention;
[0054] Figure 11 This is a schematic diagram of the structure of a selection module provided in an embodiment of the present invention;
[0055] Figure 12 This is a schematic diagram of the structure of a fourth generation module provided in an embodiment of the present invention;
[0056] Figure 13 This is a schematic diagram of a computer device provided in an embodiment of the present invention.
Detailed Implementation Methods
[0057] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0058] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0059] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0060] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0061] Figure 1 A flowchart illustrating a method for determining a root cause quality index provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method includes:
[0062] Step 101: Collect at least one piece of metadata.
[0063] In this embodiment of the invention, each step is executed by a computer device. Specifically, in this embodiment of the invention, each step is executed by the computer device through a metadata governance tool. The metadata governance tool can provide an intuitive visual interface, allowing data managers and users to search and browse metadata according to different categories and usage scenarios, thereby achieving metadata information sharing.
[0064] In this step, the AKKA scheduler controls the computer device to collect at least one metadata record from the Hadoop Distributed File System (HDFS), a database, or an Extract-Transform-Load (ETL) scheduling system, and controls the computer device to store at least one metadata record into a MySQL database according to storage rules. Storage rules include MySQL database storage rules. The at least one metadata record includes a dataset, data model, streaming data, or operational data. The computer device provides a user interface (UI) and an application programming interface (API) for user interaction.
[0065] Step 102: Select the poor-quality metadata from at least one metadata.
[0066] In this step, the computer device sequentially performs data quality checks on each piece of metadata, generates a quality score for the metadata, and determines whether the quality score is less than a set threshold. If the quality score is less than the set threshold, the metadata is identified as poor quality metadata; if the quality score is greater than or equal to the set threshold, the metadata is identified as normal metadata.
[0067] Step 103: Obtain the dependencies of the poor quality metadata.
[0068] In this step, the dependencies of the poor-quality metadata are obtained by analyzing its metadata information. The metadata information includes the data processing procedures, data tables, and field dependencies.
[0069] Step 104: Determine the dependency metadata related to the poor quality metadata based on the poor quality metadata and the dependency relationship.
[0070] In this step, dependent metadata includes metadata that has a dependency relationship with poor-quality metadata during the data processing of poor-quality metadata. Specifically, dependent metadata includes metadata that has a dependency relationship with poor-quality metadata upstream of poor-quality metadata during the data processing of poor-quality metadata, as well as metadata that has a dependency relationship with poor-quality metadata downstream of poor-quality metadata.
[0071] As an alternative, after step 104, the computer device can also perform step S1.
[0072] Step S1: Generate an end-to-end metadata information chain based on the quality difference metadata, dependency metadata, and dependency relationships.
[0073] Figure 2 A flowchart illustrating a method for generating an end-to-end metadata information chain according to an embodiment of the present invention is shown below. Figure 2 As shown, step S1 specifically includes:
[0074] Step S11: Generate a technology metadata chain based on the first technology metadata and / or the second technology metadata.
[0075] In this step, if the poor quality metadata includes the first technical metadata and the dependent metadata includes the second business metadata, then a technical metadata chain is generated based on the first technical metadata.
[0076] Step S12: Generate a business metadata chain based on the first business metadata and / or the second business metadata.
[0077] In this step, if the poor quality metadata includes the first technical metadata and the dependent metadata includes the second business metadata, then a business metadata chain is generated based on the second business metadata.
[0078] Step S13: Merge the technical metadata chain and the business metadata chain according to the dependency relationship. If the fusion is successful, proceed to step S14; if the fusion fails, proceed to step S15.
[0079] Step S14: Generate an end-to-end metadata information chain.
[0080] In this step, after the computer device generates an end-to-end metadata information chain, it can display the metadata information chain through metadata governance tools. As an optional solution, the computer device can train a deep learning model based on the user's viewing habits and preferences for metadata, generating a trained deep learning model, and then displaying the metadata that the user is interested in through this trained deep learning model.
[0081] In this embodiment of the invention, the deep learning model may include a recurrent neural network (RNN) model or a convolutional neural network (CNN) model.
[0082] Step S15: Repair the failed technical metadata chain and / or business metadata chain using a deep learning model to generate a repaired technical metadata chain and / or business metadata chain.
[0083] Step S16: Generate an end-to-end metadata information chain based on the repaired technical metadata chain and / or business metadata chain.
[0084] In the technical solution of the end-to-end metadata information chain generation method provided in this invention, a business metadata chain is generated based on first business metadata and / or second business metadata; the technical metadata chain and the business metadata chain are merged according to dependencies; if the fusion is successful, an end-to-end metadata information chain is generated; if the fusion fails, the failed technical metadata chain and / or business metadata chain are repaired using a deep learning model to generate a repaired technical metadata chain and / or business metadata chain, and the end-to-end metadata information chain is generated based on the repaired technical metadata chain and / or business metadata chain. Using a deep learning model to repair the failed technical metadata chain and / or business metadata chain improves repair efficiency and reduces repair costs.
[0085] As an alternative, after step S1, the computer device can also perform step S2.
[0086] Step S2: Determine the predicted triples based on the end-to-end metadata information chain using the Translating Embedding (TransE) algorithm.
[0087] Figure 3 A flowchart of a method for determining predicted triples provided in an embodiment of the present invention is shown below. Figure 3 As shown, step S2 specifically includes:
[0088] Step S21: Determine if the entity part of the triple is missing. If the entity part of the triple is found to be missing, proceed to step S22; if the entity part of the triple is not found to be missing, the process ends.
[0089] In this step, the end-to-end metadata information chain includes at least one triple. A complete triple includes two entities and the dependency relationship between the two entities. The two entities include two metadata that have a dependency relationship. The two metadata that have a dependency relationship constitute the head entity and the tail entity of the triple, respectively. The dependency relationship between the two entities includes the dependency relationship between the two metadata that have a dependency relationship.
[0090] Figure 4 A schematic diagram of a triple provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the triplet consists of the triplet (h, r, t). Here, h, r, and t are all low-dimensional vectors, h is the head entity, t is the tail entity, and r is the dependency between the head entity and the tail entity. In this case, the triplet satisfies the relation: t≈h+r, that is, head entity ≈ tail entity + dependency between head entity and tail entity.
[0091] Step S22: Sort the metadata metrics in the end-to-end metadata information chain according to the triple entity attributes.
[0092] In this step, the triple entity attribute includes either a head entity or a tail entity. The computer device sorts the metadata in the end-to-end metadata information chain in the order of head entity first, tail entity last.
[0093] Step S23: Use the sorted metadata indicators as the missing entities in the triples to generate at least one predicted triple.
[0094] Step S24: Generate triplet scores based on the predicted triplets.
[0095] In this step, at least one triplet score is generated based on at least one predicted triplet, with each predicted triplet corresponding to a triplet score. The higher the triplet score, the closer the dependencies between the head entity, tail entity, and head entity and tail entity in the predicted triplet are to the triplet's relational expression.
[0096] Step S25: Determine whether the triplet score is greater than or equal to the set threshold. If the triplet score is greater than or equal to the set threshold, proceed to step S26; if the triplet score is less than the set threshold, proceed to step S27.
[0097] Step S26: Determine that the predicted triplet is the correct triplet.
[0098] Step S27: Determine that the predicted triplet is not a correct triplet.
[0099] In the technical solution of the method for determining predicted triples provided in this invention, a triple score is generated based on the predicted triples. The method then determines whether the triple score is greater than or equal to a set threshold. If the triple score is greater than or equal to the set threshold, the predicted triple is determined to be a correct triple; if the triple score is less than the set threshold, the predicted triple is determined to be an incorrect triple. Determining the correctness of triples using a binary classification method improves the efficiency of the determination.
[0100] Step 105: Identify multiple metadata metrics related to poor-quality metadata and dependent metadata.
[0101] In this step, multiple metadata metrics include those corresponding to poor-quality metadata and those corresponding to dependent metadata. Metadata metrics are the data information of the metadata. As an optional approach, the data information can be displayed through data tables. For example, each metadata item corresponds to a data table, and each data table includes at least one metadata metric corresponding to that metadata item.
[0102] Step 106: Generate a data analysis graph based on multiple metadata indicators using the relational graph algorithm.
[0103] In this step, the data analysis graph includes a lineage analysis graph and an influence analysis graph. Each data analysis graph shows the data relationships between at least one metadata indicator, including lineage or influence relationships. The generation of the data analysis graph incorporates the temporal order and progressive relationships between multiple metadata indicators, thus ensuring that the data analysis graph fully reflects the data relationships between these indicators.
[0104] As an alternative, after step 106, the computer device can also perform step S3.
[0105] Step S3: Query the data relationship between at least one metadata indicator based on the data analysis chart.
[0106] For example, computer devices can use the back propagation (BP) neural network algorithm to query the data relationships between at least one metadata indicator based on the data analysis graph.
[0107] In this embodiment of the invention, step S3 specifically includes:
[0108] Step S31: Generate a visual data analysis graph based on the data analysis graph using the backpropagation neural network algorithm.
[0109] Step S32: In response to the query command input by the user, query the data relationship between at least one metadata indicator through a visual data analysis graph.
[0110] In this step, the computer device provides query services to the user based on the BP neural network algorithm. The BP neural network is a multi-layer neural network, comprising three or more layers, each consisting of several neurons. Specifically, in response to the user's input query command, the computer device retrieves the data relationships between metadata indicators through a visualized data analysis graph. The query command may include an instruction to query the data relationship of a specific metadata indicator in the data analysis graph, or an instruction to query the root cause quality poorness indicator in the data analysis graph; however, this is not limited in this embodiment of the invention.
[0111] In this embodiment of the invention, a backpropagation neural network algorithm is used to generate a visualized data analysis graph based on the data analysis graph. Responding to a user-inputted query command, the visualized data analysis graph allows for the querying of data relationships between at least one metadata indicator. The computer device, through the BP neural network algorithm, visualizes and queries the data relationships between different metadata indicators, making the flow of data during use clearer and more intelligent, thus improving query efficiency.
[0112] As an alternative, after step 106, the computer device can also perform step S4.
[0113] S4. Determine the shortest path between the selected central indicator and each endpoint indicator based on the data analysis chart.
[0114] For example, computer devices can use Dijkstra's algorithm to determine the shortest paths between selected central indicators and various endpoint indicators based on data analysis graphs.
[0115] Figure 5 A flowchart illustrating a method for determining the shortest path provided in an embodiment of the present invention is shown below. Figure 5 As shown, step S4 specifically includes:
[0116] Step S41: In response to the user's input selection instruction, select the central indicator from multiple metadata indicators.
[0117] In this step, the computer device responds to the user's selection command and selects a central indicator from multiple metadata indicators in the data analysis graph. The central indicator includes metadata indicators for which the shortest path needs to be determined with respect to other metadata indicators in the data analysis graph. The computer device then assigns a number to the central indicator, generating a central indicator number.
[0118] Figure 6 This is a schematic diagram of a central indicator and a endpoint indicator provided in an embodiment of the present invention, as shown below. Figure 6As shown, the central indicator number includes indicator 1, and the endpoint indicator numbers include indicator 2, indicator 3, indicator 4, indicator 5, and indicator 6. The shortest path between indicator 1 and indicator 2 is 1, between indicator 1 and indicator 3 is 12, between indicator 2 and indicator 3 is 9, between indicator 2 and indicator 4 is 3, between indicator 3 and indicator 4 is 4, between indicator 4 and indicator 5 is 13, between indicator 4 and indicator 6 is 15, and between indicator 5 and indicator 6 is 4.
[0119] Step S42: Generate a central indicator number based on the central indicator, and construct the first array based on the central indicator number.
[0120] In this step, the first array includes the numbers of the metadata metrics for which the shortest path has been determined. Initially, the first array only includes the central metric number. For example, if the central metric number includes metric 1, then the first array is constructed based on metric 1.
[0121] Step S43: Determine the endpoint indicator based on the central indicator.
[0122] In this step, the endpoint metrics include metadata metrics in the data analysis chart, excluding the central metrics.
[0123] Step S44: Generate endpoint indicator number based on endpoint indicator, and construct a second array based on endpoint indicator number.
[0124] In this step, the computer device assigns numbers to the endpoint indicators, generating endpoint indicator numbers. The second array includes the numbers of metadata indicators for which the shortest path has not yet been determined. Initially, the second array contains all endpoint indicator numbers. For example, if the endpoint indicator numbers include indicator 2, indicator 3, indicator 4, indicator 5, and indicator 6, then the second array is constructed based on indicator 2, indicator 3, indicator 4, indicator 5, and indicator 6.
[0125] Step S45: Select the nearest indicator number from the endpoint indicator number.
[0126] Step S46: Update the first and second arrays according to the proximity indicator numbers.
[0127] In this step, the computer device removes the near-field indicator number from the second array and adds the near-field indicator number to the first array.
[0128] Since the first array and the second array are updated simultaneously based on the proximity index in this embodiment of the invention, the second array is an empty array when the first array contains all the endpoint index numbers.
[0129] Step S47: Determine whether the shortest paths between the central indicator and the endpoint indicator have all been determined. If it is determined that the shortest paths between the central indicator and the endpoint indicator have all been determined, proceed to step S48; if it is determined that the shortest paths between the central indicator and the endpoint indicator have not all been determined, proceed to step S45.
[0130] As an alternative, the computer device determines whether the first array contains all the endpoint indicator numbers. If it determines that the first array contains all the endpoint indicator numbers, then step S48 is executed; if it determines that the first array does not contain all the endpoint indicator numbers, then step S45 is executed.
[0131] As an alternative, the computer device determines whether the second array is an empty array. If the second array is determined to be an empty array, step S48 is executed; if the second array is determined not to be an empty array, step S45 is executed.
[0132] S48. Determine the shortest path between the central indicator and each endpoint indicator based on the shortest path between the central indicator and the nearest indicators.
[0133] In this embodiment of the invention, in response to a user-input selection command, a central indicator is selected from multiple metadata indicators; a first array is constructed based on the central indicator; a destination indicator is determined based on the central indicator; a second array is constructed based on the destination indicator; a proximity indicator is selected from the destination indicators; and the first and second arrays are updated based on the proximity indicator. By updating the first and second arrays, the computer device determines the shortest path between the central indicator and the destination indicator, avoiding the problem of repeatedly determining the shortest path and improving the efficiency of shortest path determination.
[0134] As an alternative, after step 106, the computer device can also perform step S5.
[0135] Step S5: Determine the importance score of metadata indicators based on the data analysis chart.
[0136] For example, computer devices can use the PageRank algorithm to determine the importance score of metadata metrics based on data analysis graphs.
[0137] Figure 7 A flowchart illustrating a method for determining the importance of metadata metrics provided in an embodiment of the present invention is shown below. Figure 7 As shown, step S5 specifically includes:
[0138] Step S51: Assign the same initial score to each metadata metric in the data analysis graph.
[0139] Step S52: Using the PageRank algorithm, generate multiple quantity scores based on the number of connection branches for each metadata metric.
[0140] In this step, the more connection branches a metadata metric has, the higher its quantity score, and the more important that metadata metric is.
[0141] Figure 8 This is a schematic diagram of a metadata indicator connection branch provided in an embodiment of the present invention, as shown below. Figure 8 As shown, the metadata metrics include node0, node1, node2, node3, and node4. Node0 has 1 connection branch, node1 has 4 connection branches, node2 has 1 connection branch, node3 has 2 connection branches, and node4 has 3 connection branches. Since node1 has the most connection branches, it has the highest score among the five metadata metrics, making it the most important.
[0142] Step S53: Using the PageRank algorithm, generate multiple quality scores based on the connection branch quality of each metadata metric.
[0143] In this step, the higher the quality of the connection branch of the metadata metric, the higher the quality score of the metadata metric, and the more important the metadata metric is.
[0144] Step S54: Iteratively update the initial score of each metadata indicator based on the quantity score and quality score, and use the updated metadata indicator score as the importance score of the metadata indicator.
[0145] In this step, the computer device updates the initial score of each metadata indicator based on the quantity score and quality score using an iterative recursive algorithm until the score stabilizes. The score after the last update is then used as the final importance score for the metadata indicator.
[0146] In the technical solution of the method for determining the importance of metadata indicators provided in this invention, the initial score of each metadata indicator is iteratively updated based on quantity and quality scores, and the updated metadata indicator score is used as the importance score of the metadata indicator. Determining the importance of metadata indicators based on multiple factors makes the importance scores of metadata indicators more accurate.
[0147] Step 107: In response to the user's input selection command, select the root cause quality index from the data analysis chart.
[0148] In this step, after analyzing the data analysis chart, the user determines the root cause quality index. The computer device responds to the user's selection command and selects the root cause quality index from the data analysis chart. There can be one or more root cause quality indices.
[0149] As an optional approach, the data analysis graph includes a lineage graph, with the low-quality metadata as the endpoint. Therefore, the lineage graph can illustrate the source of the data and the data processing preceding the low-quality metadata. The lineage graph includes the data relationships between the low-quality metadata and its upstream dependent metadata during the data processing of the low-quality metadata. A root node is selected from the lineage graph, and it is determined whether the root node is a low-quality indicator. If the root node is determined to be a low-quality indicator, it is identified as a root cause low-quality indicator; otherwise, it is identified as not a root cause low-quality indicator.
[0150] As an alternative, the data analysis graph includes an impact analysis graph. The poor-quality metadata serves as the starting point for the impact analysis graph, thus illustrating the data flow and the data processing after the poor-quality metadata. The impact analysis graph shows the data relationships between the poor-quality metadata and its downstream dependent metadata during the data processing process. The root cause poor-quality indicator includes the metadata indicator with the most connecting branches in the impact analysis graph. The number of connecting branches includes both direct and indirect connections. Direct connections include those directly connected to the root cause poor-quality indicator, with each direct connection linking to both the root cause poor-quality indicator and a directly affecting indicator. Indirect connections include those indirectly connected to the root cause poor-quality indicator, with each indirect connection linking to both a directly affecting indicator and an indirectly affecting indicator.
[0151] Figure 9 This is a schematic diagram of an influence analysis diagram provided in an embodiment of the present invention, such as... Figure 9 As shown, the impact analysis diagram includes one root cause quality poor indicator, seven impact indicators, and eight connection branches. Specifically, indicator 1 is the root cause quality poor indicator. The seven impact indicators include four directly impacted indicators and three indirectly impacted indicators. Among them, indicators 2, 3, 4, and 5 are directly impacted indicators; indicators 6, 7, and 8 are indirectly impacted indicators. The eight connection branches include four directly connected branches and four indirectly connected branches. Since the impact indicators are metadata indicators downstream of the root cause quality poor indicator, they are more likely to have poor data quality due to the influence of the root cause quality poor indicator.
[0152] In this embodiment of the invention, the computer device can also perform difference analysis on metadata indicators through data analysis graphs to obtain the differences between metadata indicators. For example, the differences include differences between names or differences between attributes. Through difference analysis, business personnel can analyze multiple metadata indicators with small differences from multiple directions such as business definition and data generation to determine the differences between these multiple metadata indicators; technical personnel can identify information based on metadata indicators with small differences.
[0153] Step 108: Generate early warning information based on the root cause quality poor index.
[0154] In this step, if the data analysis graph includes a lineage analysis graph, then a warning message is generated based on the root cause quality poor index, and the warning message includes the root cause quality poor index; if the data analysis graph includes an impact analysis graph, then a warning message is generated based on the root cause quality poor index and the impact index, and the warning message includes the root cause quality poor index and the impact index.
[0155] The technical solution of the method for determining root cause quality issues provided in this invention involves: acquiring the dependency relationships of quality issue metadata; determining the dependent metadata related to the quality issue metadata based on the quality issue metadata and the dependency relationships; determining multiple metadata indicators corresponding to the quality issue metadata and dependent metadata; generating a data analysis graph based on the multiple metadata indicators using a relational graph algorithm; and selecting the root cause quality issue indicator from the data analysis graph in response to a selection command input by the user. By determining the root cause quality issue indicator, computer devices can quickly locate data problems, improve the efficiency of data problem processing, and reduce the cost of data problem processing.
[0156] In this embodiment of the invention, early warning information is generated based on the root cause quality poor index to provide early warning of data risks. This helps data managers to quickly locate metadata indicators that may have problems and to handle them in a timely manner, thereby avoiding potential data problems, preventing losses caused by data problems, and improving the efficiency of data problem handling.
[0157] As an optional solution, metadata governance tools can also be used to improve the description of data assets, processing and organizing data into unambiguous data assets to ensure data understandability. Data asset descriptions include at least one of the following: basic attributes, business definitions of metrics, technical definitions of metrics, related reports, dependency models, dependency metrics, version change history, field attributes, and data distribution. Table 1 below shows the relevant report descriptions in the data asset description.
[0158] Table 1
[0159] Serial Number Report Name Report path Report Description Report coding 1 Community Equipment Resource Qualification Rate Report Resource data accuracy Description 1 TA098766 2 Community Address Resource Qualification Rate Report Resource data accuracy Description 2 TA986544
[0160] As shown in Table 1 above, the relevant report descriptions in the data asset description include serial number, report name, report path, report description, and report code. As shown in Table 1, in the relevant report descriptions, the report name corresponding to serial number 1 is "Community Equipment Resource Qualification Rate Report", the report path is "Resource Data Accuracy Rate", the report description is "Description 1", and the report code is TA098766.
[0161] As an optional solution, metadata governance tools can also be used to generate data asset catalogs. Computer devices use metadata governance tools to categorize data according to data type, data distribution, and data source, generating a data asset catalog. This catalog provides suitable asset catalogs to data asset developers, enabling them to publish data assets, understand the overall status of their data assets, and allow users to quickly locate the data assets they need. A data asset catalog includes at least one of the following: serial number, asset number, asset type, asset name, category, subcategory, source system, layer, registrant, and asset launch time. Table 2 below illustrates the data asset catalog.
[0162] Table 2
[0163] Serial Number Asset Number Asset types Asset Name Category Registrant Asset Launch Time 1 20190907 Data Model Model 1 resource Wang Yiyi 2020-02-01 2 20190908 index Indicator 1 resource Li Yiyi 2020-03-01 3 20190909 index Indicator 2 business Zhao Yiyi 2020-05-01
[0164] As shown in Table 2 above, the data asset catalog includes serial number, asset number, asset type, asset name, category, registrant, and asset launch date. As shown in Table 2, in the data asset catalog, serial number 1 corresponds to asset number 20190907, asset type is data model, asset name is Model 1, category is resource, registrant is Wang Yiyi, and asset launch date is 2020-02-01.
[0165] As an optional solution, metadata governance tools can also be used to generate asset maps. Data asset maps can display data from multiple dimensions such as asset type, asset classification, and asset stratification, providing multi-level and multi-perspective data assets including total data volume, data growth, distribution of data assets, and data relationships between systems. A data asset map includes at least one of the following: total asset volume, data model, indicator asset volume, data source interface, data sharing service, classification management data of data models and indicators, multiple data source systems, business asset changes, and daily asset access volume. Computer devices can display data asset maps in any data representation format, and this embodiment of the invention does not limit this. For example, computer devices can display data asset maps using tables, bar charts, or line charts. This embodiment of the invention describes the display of data asset maps using tables as an example. As shown in Table 3 below, Table 3 shows the classification management data of data models and indicators in the data asset map.
[0166] Table 3
[0167]
[0168]
[0169] As shown in Table 3 above, the data models and indicators in the data asset map are categorized and managed as follows: serial number, data type, data name, and data quantity. For example, in the data models and indicators categorized and managed in the data asset map, serial number 2 corresponds to the data type "indicator," the data name "performance business indicator," and the data quantity is 80876.
[0170] As an optional solution, metadata governance tools can also be used for data asset valuation. Computer equipment uses metadata governance tools to establish a data asset valuation system, automatically collecting and analyzing data asset usage, access frequency, and call patterns and trends. This allows for the assessment of the current value of the data assets and the provision of corresponding data asset processing recommendations based on the valuation results. Data asset valuation data includes at least one of the following: serial number, Chinese model name, English model name, subject domain, subject subdomain, number of calls in the past two years, number of calls in the past year, number of records from the previous year, storage statistics from the previous year, analysis of reasons for decommissioning, and decommissioning operations. Table 4 below shows the data asset valuation data, enabling computer equipment to complete data asset valuation using this data.
[0171] Table 4
[0172]
[0173] As shown in Table 4 above, the data asset valuation data includes the serial number, Chinese name of the model, English name of the model, subject domain, subject subdomain, number of records before 2019, storage statistics before 2019, and analysis of reasons for shutdown. As shown in Table 4 above, in the data asset valuation data, the Chinese name of the model corresponding to serial number 2 is "Grid Personnel Information Table," the English name of the model is "BROADBAND," the subject domain is "Resources," the subject subdomain is "Resources," the number of records before 2019 is 0, the storage statistics before 2019 are 15, and the analysis of reasons for shutdown indicates no records were recorded in 2019.
[0174] As an alternative, metadata governance tools can also be used for data retrieval. Computer devices can use metadata governance tools to provide users with fast retrieval services and dataset sharing services, improving the depth of information search and sharing within the system and revitalizing data assets.
[0175] Figure 10 This is a schematic diagram of a device for determining root cause quality index provided in an embodiment of the present invention, as shown below. Figure 10As shown, the device includes: an acquisition module 11, a first determination module 12, a second determination module 13, a first generation module 14, and a selection module 15. The acquisition module 11 is connected to the first determination module 12, the first determination module 12 is connected to the second determination module 13, the second determination module 13 is connected to the first generation module 14, and the first generation module 14 is connected to the selection module 15. The acquisition module 11 is used to acquire the dependency relationships of the poor quality metadata; the first determination module 12 is used to determine the dependent metadata related to the poor quality metadata based on the poor quality metadata and the dependency relationships; the second determination module 13 is used to determine multiple metadata indicators corresponding to the poor quality metadata and the dependent metadata; the first generation module 14 is used to generate a data analysis graph based on the multiple metadata indicators using a relational graph algorithm; and the selection module 15 is used to select the root cause poor quality indicator from the data analysis graph in response to a selection command input by the user.
[0176] In this embodiment of the invention, the device further includes a second generation module 16. The second generation module 16 is connected to the selection module 15. The second generation module 16 is used to generate early warning information based on the root cause quality poor index.
[0177] Figure 11 A schematic diagram of the structure of a selection module provided in an embodiment of the present invention, such as... Figure 11 As shown, the selection module 15 includes a selection unit 151, a judgment unit 152, and a determination unit 153. The selection unit 151 is connected to the judgment unit 152, and the judgment unit 152 is connected to the determination unit 153. The selection unit 151 is used to select the root node from the lineage analysis diagram; the judgment unit 152 is used to determine whether the root node is a poor quality indicator; the determination unit 153 is used to determine the root node as a root cause poor quality indicator if the judgment module 152 determines that the root node is a poor quality indicator.
[0178] In this embodiment of the invention, the device further includes a third generation module 17 and a query module 18, with the third generation module 17 connected to the first generation module 14 and the query module 18. The third generation module 17 is used to generate a visualized data analysis graph based on the data analysis graph using a backpropagation BP neural network algorithm; the query module 18 is used to query the data relationship between at least one metadata indicator through the visualized data analysis graph in response to a query command input by the user.
[0179] In this embodiment of the invention, the device further includes a fourth generation module 19, which is connected to the acquisition module 11 and the first determination module 12. The fourth generation module 19 is used to generate an end-to-end metadata information chain based on the quality difference metadata, dependency metadata, and dependency relationships.
[0180] Figure 12 This is a schematic diagram of the structure of a fourth generation module provided in an embodiment of the present invention, as shown below. Figure 12As shown, the fourth generation module 19 includes a first generation unit 191, a second generation unit 192, a fusion unit 193, a third generation unit 194, a fourth generation unit 195, and a fifth generation unit 196. The first generation unit 191 is connected to the fusion unit 193, the second generation unit 192 is connected to the fusion unit 193, the fusion unit 193 is connected to the third generation unit 194 and the fourth generation unit 195, and the fourth generation unit 195 is connected to the fifth generation unit 196. The first generation unit 191 is used to generate a technical metadata chain based on the first technical metadata and / or the second technical metadata; the second generation unit 192 is used to generate a business metadata chain based on the first business metadata and / or the second business metadata; the fusion unit 193 is used to fuse the technical metadata chain and the business metadata chain according to the dependency relationship; the third generation unit 194 is used to generate an end-to-end metadata information chain if the fusion unit 193 successfully fuses; the fourth generation unit 195 is used to repair the reasons for the failed fusion of the technical metadata chain and / or business metadata chain through a deep learning model if the fusion unit 193 fails, and generate a repaired technical metadata chain and / or business metadata chain; the fifth generation unit 196 is used to generate an end-to-end metadata information chain based on the repaired technical metadata chain and / or business metadata chain.
[0181] The technical solution of the root cause quality defect determination device provided in this invention involves: acquiring the dependency relationship of quality defect metadata; determining the dependent metadata related to the quality defect metadata based on the quality defect metadata and the dependency relationship; determining multiple metadata indicators corresponding to the quality defect metadata and dependent metadata; generating a data analysis graph based on the multiple metadata indicators using a relational graph algorithm; and selecting the root cause quality defect indicator from the data analysis graph in response to a selection command input by the user. The computer device determines the root cause quality defect indicator through the data analysis graph, which is beneficial for quickly locating data problems, improving the efficiency of data problem processing, and reducing the processing cost of data problems.
[0182] This invention provides a computer-readable storage medium including a stored program, wherein the program, when running, controls the device where the computer-readable storage medium is located to execute the aforementioned method for determining root cause quality indicators.
[0183] Figure 13 A schematic diagram of a computer device provided in an embodiment of the present invention, such as... Figure 13 As shown, the computer device 3 in this embodiment includes a processor 31, a memory 32, and a computer program 33 stored in the memory 32 and executable on the processor 31. When the computer program 33 is executed by the processor 31, it implements the method for determining the root cause quality index in the embodiment. To avoid repetition, it will not be described in detail here.
[0184] Computer device 3 includes, but is not limited to, processor 31 and memory 32. Those skilled in the art will understand that... Figure 13 This is merely an example of computer device 3 and does not constitute a limitation on computer device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, a network device may also include input / output devices, network access devices, buses, etc.
[0185] The processor 31 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0186] The memory 32 can be an internal storage unit of the computer device 3, such as a hard disk or RAM of the computer device 3. The memory 32 can also be an external storage device of the computer device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or FlashCard equipped on the computer device 3. Furthermore, the memory 32 can include both internal and external storage units of the computer device 3. The memory 32 is used to store computer programs and other programs and data required by network devices. The memory 32 can also be used to temporarily store data that has been output or will be output.
[0187] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0188] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0189] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.
[0190] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0191] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0192] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for determining a root cause quality index, characterized in that, The method includes: Dependencies for obtaining poor-quality metadata; Based on the poor quality metadata and the dependency relationship, the dependency metadata related to the poor quality metadata is determined; Multiple metadata metrics corresponding to the poor quality metadata and the dependent metadata were identified; A data analysis graph is generated based on the multiple metadata indicators using a relational graph algorithm. In response to the user's selection command, the root cause quality index is selected from the data analysis chart; The method further includes: Generate an end-to-end metadata information chain based on the quality difference metadata, the dependency metadata, and the dependency relationship; The quality-poor metadata includes first technical metadata or first business metadata, and the dependency metadata includes second technical metadata or second business metadata. The step of generating the end-to-end metadata information chain based on the quality difference metadata, the dependency metadata, and the dependency relationship includes: Generate a technology metadata chain based on the first technology metadata and / or the second technology metadata; Generate a business metadata chain based on the first business metadata and / or the second business metadata; The technical metadata chain and the business metadata chain are merged according to the aforementioned dependencies; If the fusion is successful, an end-to-end metadata information chain will be generated; If the fusion fails, the technical metadata chain and / or business metadata chain that failed to merge will be repaired through a deep learning model to generate a repaired technical metadata chain and / or business metadata chain. Generate an end-to-end metadata information chain based on the repaired technical metadata chain and / or business metadata chain.
2. The method according to claim 1, characterized in that, After selecting the root cause quality index from the data analysis chart in response to the user's selection instruction, the method further includes: Early warning information is generated based on the root cause quality poor index.
3. The method according to claim 1, characterized in that, The data analysis graph includes an impact analysis graph, and the root cause quality difference index includes the metadata index with the most connected branches in the impact analysis graph.
4. The method according to claim 1, characterized in that, The data analysis graph includes a lineage analysis graph; The step of selecting the root cause quality index from the data analysis chart in response to the user's input selection command includes: Select the root node from the bloodline analysis diagram; Determine whether the root node is a poor quality indicator; If the root node is determined to be a poor quality index, then the root node is identified as a root cause poor quality index.
5. The method according to claim 1, characterized in that, The data analysis graph includes data relationships between at least one metadata indicator; After generating the data analysis graph based on the multiple metadata indicators using the relationship graph algorithm, the process further includes: A visual data analysis graph is generated based on the data analysis graph using the backpropagation BP neural network algorithm. In response to a user-input query command, the data relationship between the at least one metadata indicator is queried through the visual data analysis graph.
6. A device for determining a root cause quality index, characterized in that, The device includes: The acquisition module is used to obtain the dependencies of poor-quality metadata; The first determining module is used to determine the dependency metadata related to the quality poor metadata based on the quality poor metadata and the dependency relationship; The second determining module is used to determine multiple metadata metrics corresponding to the poor quality metadata and the dependent metadata; The first generation module is used to generate a data analysis graph based on the multiple metadata indicators using a relational graph algorithm; The selection module is used to select the root cause quality index from the data analysis chart in response to the selection command input by the user. The device also includes a fourth generation module; The fourth generation module is used to generate an end-to-end metadata information chain based on the quality difference metadata, the dependency metadata, and the dependency relationship; The quality-poor metadata includes first technical metadata or first business metadata, and the dependency metadata includes second technical metadata or second business metadata. The fourth generation module includes a first generation unit, a second generation unit, a fusion unit, a third generation unit, a fourth generation unit, and a fifth generation unit; The first generation unit is used to generate a technical metadata chain based on the first technical metadata and / or the second technical metadata; the second generation unit is used to generate a business metadata chain based on the first business metadata and / or the second business metadata; the fusion unit is used to fuse the technical metadata chain and the business metadata chain according to the dependency relationship; the third generation unit is used to generate an end-to-end metadata information chain if the fusion unit successfully fuses the data; the fourth generation unit is used to repair the reasons for the failed fusion of the technical metadata chain and / or business metadata chain through a deep learning model if the fusion unit fails to fuse the data, and generate a repaired technical metadata chain and / or business metadata chain; the fifth generation unit is used to generate an end-to-end metadata information chain based on the repaired technical metadata chain and / or business metadata chain.
7. A computer device, characterized in that, include: One or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, the one or more computer programs including instructions that, when executed by the computer device, cause the computer device to perform the method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method and device and computer storage medium
CN112491636A
Abnormal root cause analysis method and device and storage medium
CN112882796A