Data governance method and device

By obtaining the genetic attribute information of the new data standard and establishing non-genetic attribute information, and generating and adding it to the existing standards, the problems of inefficiency and conflict with the existing standards are solved, and the efficiency and accuracy of data governance are improved.

CN114722110BActive Publication Date: 2025-06-27INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210434570.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-06-27
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The existing data governance methods are less efficient in creating new data standards when new data is introduced, and new standards are prone to conflict with existing standards, resulting in low data governance efficiency and high error rate, affecting banking business.

Method used

Genetic attribute information is obtained based on the new identification and existing standards of the new data standard, and non-genetic attribute information is established based on the newly introduced data, and new data standards are generated, and new data standards are added to the existing standards to determine the governance scope of data governance, so as to carry out data governance.

Benefits of technology

It improves the generation efficiency of new data standards and compatibility with existing standards, reduces the error rate during data governance, improves the efficiency of data governance, and ensures the normal operation of the data system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114722110B_ABST
    Figure CN114722110B_ABST
Patent Text Reader

Abstract

The present invention provides a data governance method and apparatus, particularly relating to the field of big data technology. The method includes: obtaining the genetic attribute information of the newly created data standard according to the newly created identifier and the existing standard of the newly created data standard; establishing the non-genetic attribute information of the newly created data standard according to the newly introduced data; generating a newly created data standard according to the genetic attribute information and the non-genetic attribute information, and adding the newly created data standard to the existing standard to obtain an updated total standard, so as to determine the governance scope of data governance according to the updated total standard, and thereby perform data governance according to the governance scope. The present invention can improve the efficiency of data governance and reduce the probability of errors occurring during data governance, thus being beneficial to the normal operation of relevant data systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, particularly to the technical field of big data, and more particularly to a data governance method and apparatus. Background Art

[0002] In the process of data governance for banking operations, it often involves introducing new data into relevant data systems. When the new data cannot be classified into the existing data standards, new data standards need to be created for the new data so that the new data can also be governed subsequently. However, the existing data governance methods have low efficiency in creating new data standards for the introduced new data, and the newly created standards are prone to conflict with the existing data standards in the relevant data systems, resulting in low efficiency and prone to errors when conducting data governance according to the newly created standards and the existing standards, which is not conducive to the normal operation of the relevant data systems and has an adverse impact on banking operations. Summary of the Invention

[0003] An object of the present invention is to provide a data governance method to solve the problems that the existing data governance has low efficiency and a high probability of errors during data governance, which is not conducive to the normal operation of the relevant data systems and further has an adverse impact on banking operations. Another object of the present invention is to provide a data governance apparatus. Still another object of the present invention is to provide a computer device. Yet another object of the present invention is to provide a readable medium.

[0004] To achieve the above objects, one aspect of the present invention discloses a data governance method, the method comprising:

[0005] Obtaining genetic attribute information of the new data standard according to a new creation identifier of the new data standard and the existing standards;

[0006] Establishing non-genetic attribute information of the new data standard according to the newly introduced data;

[0007] Generating a new data standard according to the genetic attribute information and the non-genetic attribute information, adding the new data standard to the existing standards to obtain an updated total standard, determining a governance scope of data governance according to the updated total standard, and thus conducting data governance according to the governance scope.

[0008] Optionally, the obtaining genetic attribute information of the new data standard according to a new creation identifier of the new data standard and the existing standards includes:

[0009] Obtaining an approximate identifier of an approximate data standard corresponding to the new creation identifier according to the new creation identifier and the existing standards;

[0010] Obtain the superior identifier of the superior data standard corresponding to the newly created identifier according to the approximate identifier;

[0011] Obtain the genetic attribute information of the newly created data standard according to the superior identifier.

[0012] Optionally, the obtaining the approximate identifier of the approximate data standard corresponding to the newly created identifier according to the newly created identifier and the existing standard includes:

[0013] Perform semantic analysis and index analysis on the newly created identifier, and obtain the approximate identifier of the approximate data standard corresponding to the newly created identifier from the existing standard.

[0014] Optionally, the obtaining the superior identifier of the superior data standard corresponding to the newly created identifier according to the approximate identifier includes:

[0015] Obtain the relevant nodes of the approximate data standard in the data standard pedigree according to the approximate identifier;

[0016] Query the root node of the relevant nodes in the data standard pedigree, obtain the root node identifier of the root node data standard corresponding to the root node, and use the root node identifier as the superior identifier.

[0017] Optionally, the obtaining the genetic attribute information of the newly created data standard according to the superior identifier includes:

[0018] Obtain the superior data standard according to the superior identifier;

[0019] Obtain the genetic attribute information according to the superior data standard.

[0020] Optionally, the establishing the non-genetic attribute information of the newly created data standard according to the newly introduced data includes:

[0021] Obtain the new attribute corresponding to the newly introduced data according to the newly introduced data;

[0022] Establish the non-genetic attribute information of the newly created data standard according to the new attribute.

[0023] Optionally, the generating the newly created data standard according to the genetic attribute information and the non-genetic attribute information includes:

[0024] Obtain the non-genetic attribute and the first sub-standard of the non-genetic attribute information according to the non-genetic attribute information;

[0025] Obtain the genetic attribute and the second sub-standard of the genetic attribute information according to the genetic attribute information;

[0026] Generate a new data standard according to the genetic attribute, the second sub-criterion, the non-genetic attribute, and the first sub-criterion.

[0027] Optionally, adding the new data standard to the existing standard to obtain an updated total standard includes:

[0028] Obtain the parent node of the relevant node in the data standard pedigree according to the relevant node;

[0029] Add the new data standard as another child node of the parent node to the data standard pedigree to obtain an updated total standard;

[0030] Wherein, the set of data standards of all nodes in the data standard pedigree before adding the other child node is the existing standard.

[0031] Optionally, it further includes:

[0032] After the staff performs data governance according to the governance scope, obtain the total number of data standard attributes included in all data standards, the total number of data table attributes of all data tables, the number of first data records whose data record attributes conform to the currently existing data standards, the number of second data records without missing attribute values in the first data records, and the number of third data records whose all attribute values conform to the existing data standards in the second data records;

[0033] Obtain a data governance quality indicator according to the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records, and feedback the data governance quality indicator to the staff so that the staff can improve the data governance quality according to the data governance quality indicator.

[0034] Optionally, obtaining the data governance quality indicator according to the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records includes:

[0035] Obtain a coverage rate according to the total number of data standard attributes and the total number of data table attributes;

[0036] Obtain a completeness rate according to the number of first data records and the number of second data records;

[0037] Obtain an accuracy rate according to the number of first data records and the number of third data records;

[0038] Obtain a data governance quality indicator according to the coverage rate, the completeness rate, and the accuracy rate.

[0039] To achieve the above object, another aspect of the present invention discloses a data governance device, including:

[0040] A genetic attribute information determination module, configured to obtain the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standard;

[0041] A non-genetic attribute information determination module, configured to establish the non-genetic attribute information of the new data standard according to the newly introduced data;

[0042] A new module, configured to generate a new data standard according to the genetic attribute information and the non-genetic attribute information, add the new data standard to the existing standard to obtain an updated total standard, determine the governance scope of data governance according to the updated total standard, and thus perform data governance according to the governance scope.

[0043] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.

[0044] The present invention also discloses a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, the above method is implemented.

[0045] The data governance method and device provided by the present invention can obtain the genetic attribute information of the newly created data standard according to the newly created identifier and the existing standards of the newly created data standard. It can obtain the general attribute information applicable to the newly created data standard, that is, the genetic attribute information, from the existing data standards according to the characteristics of the newly created data standard reflected by the newly created identifier, so that there is no need to perform additional design and establishment of the genetic attribute information when establishing the newly created data standard, thereby improving the efficiency of generating the newly created data standard and the compatibility between the newly created data standard and the existing standards, further improving the efficiency of data governance and reducing the probability of errors occurring during data governance. By establishing the non-genetic attribute information of the newly created data standard according to the newly introduced data, when the newly introduced data cannot be classified into the existing data standards, non-genetic attribute information can be established for the characteristics of the new data that cannot be classified into the existing data standards, thereby improving the coverage rate of the newly created data standard generated in the subsequent steps, further reducing the omission of data during subsequent data governance based on the newly created data standard and the existing data standards, and further reducing the probability of errors occurring during data governance. By generating the newly created data standard according to the genetic attribute information and the non-genetic attribute information, adding the newly created data standard to the existing standards to obtain the updated total standard, and determining the governance scope of data governance according to the updated total standard, so as to perform data governance according to the governance scope, it can make the newly created data standard cover the newly introduced data and the existing data as much as possible, reduce the probability of errors occurring during subsequent data governance based on the newly created data standard and the existing data standards, and can realize more convenient data governance according to the newly created data standard and the existing data standards, further improving the efficiency of data governance. In summary, the present invention can improve the efficiency of data governance and reduce the probability of errors occurring during data governance, which is beneficial to the normal operation of related data systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 The flowchart of a data governance method according to an embodiment of the present invention is shown;

[0048] Figure 2 The schematic diagram of an optional step for obtaining genetic attribute information according to an embodiment of the present invention is shown;

[0049] Figure 3 The schematic diagram of an optional step for establishing non-genetic attribute information according to an embodiment of the present invention is shown;

[0050] Figure 4 shows a schematic diagram of an optional step for generating a new data standard according to an embodiment of the present invention;

[0051] Figure 5 shows a schematic diagram of an optional step for obtaining data governance quality indicators according to an embodiment of the present invention;

[0052] Figure 6 shows a schematic diagram of modules of a data governance device according to an embodiment of the present invention;

[0053] Figure 7 shows a schematic diagram of the structure of a computer device suitable for implementing an embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Regarding the "first", "second",... used herein, they do not particularly refer to the order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0056] Regarding the "including", "comprising", "having", "containing", etc. used herein, they are all open-ended terms, that is, they mean including but not limited to.

[0057] Regarding the "and / or" used herein, it includes any one or all combinations of the described things.

[0058] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solution of the present invention all comply with the relevant regulations of national laws and regulations.

[0059] An embodiment of the present invention discloses a transfer processing method based on a blockchain, as Figure 1 shown, the method specifically includes the following steps:

[0060] S101: Obtain the genetic attribute information of the new data standard according to the new creation identifier and the existing standard of the new data standard.

[0061] S102: Establish the non-genetic attribute information of the new data standard according to the newly introduced data.

[0062] S103: Generate a new data standard based on the genetic attribute information and non-genetic attribute information, add the new data standard to the existing standard to obtain an updated overall standard, determine the governance scope of data governance according to the updated overall standard, and thus perform data governance according to the governance scope.

[0063] Exemplarily, determining the governance scope of data governance according to the updated overall standard may be, but is not limited to, obtaining the governance attributes of all data that need to be governed according to the updated overall standard, and then searching in relevant data systems to obtain all data tables to be governed that include all or part of the governance attributes. The data tables to be governed are the governance scope. Determining the governance scope can be achieved manually or through relevant software, programs, etc. It should be noted that the specific implementation method of determining the governance scope of data governance according to the updated overall standard can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0064] Exemplarily, performing data governance according to the governance scope can be achieved by relevant staff through manual means or by software, programs, etc. that can be used for data governance. It should be noted that the specific implementation method of performing data governance according to the governance scope can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0065] The data governance method and device provided by the present invention can obtain the genetic attribute information of the new data standard by according to the new identifier of the newly created data standard and the existing standards. It can obtain the general attribute information applicable to the new data standard, that is, the genetic attribute information, from the existing data standards according to the characteristics of the new data standard reflected by the new identifier, so that there is no need to perform additional design and establishment of the genetic attribute information when establishing the new data standard, thereby improving the efficiency of generating the new data standard and the compatibility between the new data standard and the existing standards, and further improving the efficiency of data governance and reducing the probability of errors occurring during data governance. By establishing the non-genetic attribute information of the new data standard according to the newly introduced data, when the newly introduced data cannot be classified into the existing data standards, the non-genetic attribute information can be established according to the characteristics of the new data that cannot be classified into the existing data standards, thereby improving the coverage rate of the new data standard generated in the subsequent steps, and further reducing the omission of data during subsequent data governance according to the new data standard and the existing data standards, and further reducing the probability of errors occurring during data governance. By generating the new data standard according to the genetic attribute information and the non-genetic attribute information, adding the new data standard to the existing standards to obtain the updated total standard, and determining the governance scope of data governance according to the updated total standard, so as to perform data governance according to the governance scope, it can make the new data standard cover the newly introduced data and the existing data as much as possible, reduce the probability of errors occurring during subsequent data governance according to the new data standard and the existing data standards, and can realize more convenient data governance according to the new data standard and the existing data standards, further improving the efficiency of data governance. In summary, the present invention can improve the efficiency of data governance and reduce the probability of errors occurring during data governance, which is beneficial to the normal operation of relevant data systems.

[0066] In an alternative embodiment, as Figure 2 shown, the step of obtaining the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standards includes the following steps:

[0067] S201: According to the new identifier and the existing standards, obtain the approximate identifier of the approximate data standard corresponding to the new identifier.

[0068] S202: According to the approximate identifier, obtain the superior identifier of the superior data standard corresponding to the new identifier.

[0069] S203: According to the superior identifier, obtain the genetic attribute information of the new data standard.

[0070] Exemplarily, the newly created identifier may be, but is not limited to, text format, string format, integer number format, etc. For example, the newly created identifier may be, but is not limited to, "ID card", "ID card", or "13520", etc. The specific content of the newly created identifier is not limited to the above examples and may be determined by the staff according to the new data introduced into the relevant data system, or automatically generated by relevant programs, software, etc. according to the new data. It should be noted that the format and content of the newly created identifier can be determined by those skilled in the art according to the actual situation. The above description is only for illustration and does not constitute a limitation thereto.

[0071] Exemplarily, the approximate identifier may be, but is not limited to, text format, string format, integer number format, etc.

[0072] Exemplarily, the superior identifier may be, but is not limited to, text format, string format, integer number format, etc.

[0073] Through steps S201 to S203, it is possible to obtain an approximate data standard that is relatively close to the newly created data standard from the existing data standards according to the characteristics of the newly created data standard reflected by the newly created identifier, and obtain genetic attribute information according to the superior identifier of the superior data standard of the approximate data standard, so that the genetic attribute information better meets the requirements of the newly created data standard and indirectly improves the compatibility of the newly created data standard generated in the subsequent steps with the existing data standards and relevant data systems, thereby reducing the probability of errors occurring during subsequent data governance.

[0074] In an optional implementation manner, obtaining an approximate identifier of an approximate data standard corresponding to the newly created identifier according to the newly created identifier and the existing standard includes:

[0075] Performing semantic analysis and index analysis on the newly created identifier, and obtaining an approximate identifier of an approximate data standard corresponding to the newly created identifier from the existing standard.

[0076] Exemplarily, for a newly created identifier "bank VIP customer", after semantic analysis and index analysis, an approximate identifier "bank ordinary customer" can be obtained. It should be noted that the specific implementation manner of performing semantic analysis and index analysis on the newly created identifier and obtaining an approximate identifier of an approximate data standard corresponding to the newly created identifier from the existing standard can be determined by those skilled in the art according to the actual situation. The above description is only for illustration and does not constitute a limitation thereto.

[0077] Through the above steps, the proximity between the approximate identifier and the newly created identifier can be improved, thereby enhancing the relevance between the determined approximate data standard and the newly created data standard intended to be established. Further, the genetic attribute information obtained in subsequent steps can better meet the requirements of the newly created data standard and indirectly improve the compatibility of the newly created data standard generated in subsequent steps with the existing data standard and relevant data systems, thereby reducing the probability of errors occurring during subsequent data governance.

[0078] In an optional implementation manner, obtaining the superior identifier of the superior data standard corresponding to the newly created identifier according to the approximate identifier includes:

[0079] Obtaining the relevant nodes of the approximate data standard in the data standard pedigree according to the approximate identifier;

[0080] Querying the root node of the relevant nodes in the data standard pedigree to obtain the root node identifier of the root node data standard corresponding to the root node, and using the root node identifier as the superior identifier.

[0081] Exemplarily, in a banking business scenario, if there is the following data standard pedigree (i.e., the pedigree of the existing data standard):

[0082]

[0083] Exemplarily, the data standard pedigree can be, but is not limited to, a tree structure or a graph structure, etc.

[0084] Among them, if the approximate identifier is "bank ordinary customer", then from the above pedigree, the relevant nodes (i.e., the nodes where "bank ordinary customer" is located in the pedigree) can be easily obtained. The node includes the data standard of the node. It should be noted that the specific implementation manner of obtaining the relevant nodes of the approximate data standard in the data standard pedigree according to the approximate identifier can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0085] Exemplarily, querying the root node of the relevant nodes in the data standard pedigree is a conventional technical means in the art and will not be elaborated here. For example, if the approximate identifier is "bank ordinary customer", then from the above pedigree, the root node, that is, the node where "business handling information" is located, can be easily obtained, and the root node identifier is "business handling information".

[0086] Through the above steps, the proximity between the superior identifier and the newly created identifier can be improved, thereby enhancing the relevance between the determined superior data standard and the newly created data standard intended to be established. Further, the genetic attribute information obtained in the subsequent steps can better meet the requirements of the newly created data standard and indirectly improve the compatibility of the newly created data standard generated in the subsequent steps with the existing data standards and related data systems, thereby reducing the probability of errors occurring during subsequent data governance. Moreover, by accessing the data standard pedigree, the time and difficulty required to determine the superior identifier can be reduced, the efficiency of determining the superior identifier can be improved, and thus the efficiency of the overall data governance process can be indirectly enhanced.

[0087] In an alternative embodiment, obtaining the genetic attribute information of the newly created data standard according to the superior identifier includes:

[0088] Obtaining a superior data standard according to the superior identifier;

[0089] Obtaining the genetic attribute information according to the superior data standard.

[0090] Exemplarily, if the superior identifier is "business processing information", the superior standard with the superior identifier of "business processing information" can be obtained by, but not limited to, querying relevant data systems or the data standard pedigree. Among them, the superior standard can be, but not limited to, Table 1:

[0091] Table 1

[0092]

[0093] It should be noted that the specific implementation method of obtaining the superior data standard according to the superior identifier and the specific content of the superior data standard can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0094] Exemplarily, obtaining the genetic attribute information according to the superior data standard can be, but not limited to, obtaining each data attribute included in the superior data standard and the corresponding specific sub - standard for each data attribute according to the superior data standard, and integrating each data attribute and its corresponding specific sub - standard to obtain multiple genetic attribute information. For example, the genetic attribute information obtained according to the standard in Table 1 can be, but not limited to:

[0095] Genetic attribute: Name of the handler. Corresponding sub - standard: It should be a Chinese character string, and the string length is 2 to 10, and

[0096] Genetic attribute: Mobile phone number of the handler. Corresponding sub - standard: The string type should be numbers + symbols, and the string length is 14.

[0097] It should be noted that the specific implementation method of obtaining the genetic attribute information according to the superior data standard can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0098] Through the above steps, it is possible to parse the superior standard, extract the genetic attribute information required for generating the new identification, so as to further make the genetic attribute information more in line with the requirements of the new data standard and indirectly improve the compatibility of the new data standard generated in the subsequent steps with the existing data standard and relevant data systems, thereby reducing the probability of errors occurring during subsequent data governance. Moreover, by directly extracting the superior data standard, it is possible to reduce the time and difficulty of determining the genetic attribute information, improve the efficiency of determining the genetic attribute information, and thus indirectly improve the efficiency of the overall data governance process.

[0099] In an alternative embodiment, as Figure 3 shown, the establishment of the non-genetic attribute information of the new data standard according to the newly introduced data includes the following steps:

[0100] S301: Obtain a new attribute corresponding to the newly introduced data according to the newly introduced data.

[0101] S302: Establish the non-genetic attribute information of the new data standard according to the new attribute.

[0102] Exemplarily, if there is the following newly introduced data:

[0103] "Zhang San", "Li Si", "+8612345678900", "+8600987654321", "male", "female".

[0104] At this time, if it is possible to map "Zhang San", "Li Si", "+8612345678900" and "+8600987654321" in the newly introduced data to the range of genetic attributes, while the data "male" and "female" still cannot be mapped to the range of genetic attributes, then it is necessary to determine the new attribute "gender" according to the data "male" and "female".

[0105] It should be noted that the specific implementation method of obtaining a new attribute corresponding to the newly introduced data according to the newly introduced data can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0106] Exemplarily, the establishment of the non-genetic attribute information of the new data standard according to the new attribute can be achieved by manual means, or by means of software, programs, etc. in combination with the actual business processing requirements. For example, for a new attribute "gender", the following non-genetic attribute information can be obtained:

[0107] Non-genetic attribute: Gender. Corresponding sub-criterion: It should be a Chinese character string with a length of 1.

[0108] It should be noted that the specific implementation of establishing the non-genetic attribute information of the new data standard according to the new attribute can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation.

[0109] Through step S301 and step S302, based on the newly introduced data, the new attributes that need to be newly introduced are determined, and then based on the new attributes, the non-genetic attribute information is established, which can make the established non-genetic attribute information more in line with the characteristics of the newly introduced data, so that the new data standard generated in the subsequent steps is more consistent with the new data, thereby improving the accuracy of subsequent data governance.

[0110] In an optional implementation manner, as Figure 4 shown, generating the new data standard according to the genetic attribute information and the non-genetic attribute information includes the following steps:

[0111] S401: Obtain the non-genetic attribute and the first sub-criterion of the non-genetic attribute information according to the non-genetic attribute information.

[0112] S402: Obtain the genetic attribute and the second sub-criterion of the genetic attribute information according to the genetic attribute information.

[0113] S403: Generate a new data standard according to the genetic attribute, the second sub-criterion, the non-genetic attribute, and the first sub-criterion.

[0114] Exemplarily, if there is the following non-genetic attribute information:

[0115] Non-genetic attribute: Gender. Corresponding sub-criterion: It should be a Chinese character string with a length of 1.

[0116] Then the obtained non-genetic attribute is "Gender", and the first sub-criterion is "It should be a Chinese character string with a length of 1".

[0117] Since the non-genetic attribute information includes the non-genetic attribute and the first sub-criterion, the non-genetic attribute and the first sub-criterion of the non-genetic attribute information can be directly obtained according to the non-genetic attribute information.

[0118] It should be noted that the specific implementation of obtaining the non-genetic attribute and the first sub-criterion of the non-genetic attribute information according to the non-genetic attribute information can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation.

[0119] Exemplarily, if there is the following genetic attribute information:

[0120] Genetic attribute: Name of the handler. Corresponding sub - standard: It should be a Chinese character string, and the string length is from 2 to 10, and

[0121] Genetic attribute: Mobile phone number of the handler. Corresponding sub - standard: The string type should be numbers + symbols, and the string length is 14.

[0122] Then the genetic attributes obtained are "Name of the handler" and "Mobile phone number of the handler", and the obtained second sub - standards include "It should be a Chinese character string, and the string length is from 2 to 10" and "The string type should be numbers + symbols, and the string length is 14".

[0123] The genetic attribute information includes genetic attributes and second sub - standards, so the genetic attributes and the second sub - standards of the genetic attribute information can be directly obtained according to the genetic attribute information.

[0124] Exemplarily, generating a new data standard according to the genetic attribute, the second sub - standard, the non - genetic attribute and the first sub - standard may include, but is not limited to, integrating the genetic attribute, the second sub - standard, the non - genetic attribute and the first sub - standard to generate a new data standard.

[0125] For example, the generated new data standard is shown in Table 2 below:

[0126] Table 2

[0127]

[0128] It should be noted that the specific implementation method of generating a new data standard according to the genetic attribute, the second sub - standard, the non - genetic attribute and the first sub - standard can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation.

[0129] Through steps S401 to S403, the content of the new data standard can be made clearer, thereby improving the governance effect and governance speed when performing data governance according to the new data standard subsequently, and further improving the efficiency of data governance. Moreover, while improving the compatibility of the new data standard with the existing data standard and related data systems, the relevance between the new data standard and the newly introduced data can be improved, thereby reducing the probability of errors when performing data governance according to the new data standard and the existing data standard subsequently.

[0130] In an alternative embodiment, adding the new data standard to the existing standard to obtain an updated total standard includes:

[0131] Based on the relevant node, obtain the parent node of the relevant node in the data standard pedigree;

[0132] Add the newly created data standard as another child node of the parent node to the data standard pedigree to obtain the updated overall standard;

[0133] Among them, the set of data standards of all nodes in the data standard pedigree before adding the other child node is the existing standard.

[0134] Exemplarily, in a banking business scenario, if there is the following data standard pedigree (i.e., the pedigree of existing data standards):

[0135]

[0136] If the relevant node is the node where "ordinary bank customers" are located in the pedigree, then from the pedigree, the parent node can be obtained as the upper-level node of the relevant node, that is, the node where "user-side processing information" is located. It should be noted that the specific implementation manner of obtaining the parent node of the relevant node in the data standard pedigree according to the relevant node can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0137] Exemplarily, if the identifier of the newly created data standard is "bank VIP customers", the updated overall standard can be represented as the following pedigree:

[0138]

[0139] It should be noted that the specific implementation manner of adding the newly created data standard as another child node of the parent node to the data standard pedigree to obtain the updated overall standard can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0140] Through the above steps, the newly created data standard can be added to the appropriate position in the pedigree according to the relevant logical relationship of the data standard pedigree, which is beneficial for the staff to maintain the data standard pedigree, and is also beneficial for the search and retrieval of the newly created data standard, thereby reducing the difficulty of data governance based on the updated overall standard and improving the efficiency of data governance.

[0141] In an alternative embodiment, it further includes:

[0142] After the staff perform data governance according to the governance scope, obtain the total number of data standard attributes included in all data standards, the total number of data table attributes of all data tables, the number of first data records whose data record attributes conform to the current existing data standards, the number of second data records without missing attribute values in the first data records, and the number of third data records whose all attribute values in the second data records conform to the existing data standards;

[0143] According to the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records, obtain a data governance quality indicator, and feedback the data governance quality indicator to the staff, so that the staff can improve the data governance quality according to the data governance quality indicator.

[0144] Exemplarily, the total number of data standard attributes is the non-repeated data standard attributes. For example, if the attribute "card number" appears in one data standard and also appears in another data standard, the total number of data standard attributes will be incremented by 1 instead of 2 when counting.

[0145] Exemplarily, the total number of data table attributes is the non-repeated data table attributes. For example, if the attribute "card type" appears in one data table and also appears in another data standard, the total number of data table attributes will be incremented by 1 instead of 2 when counting.

[0146] Exemplarily, the number of first data records whose data record attributes conform to the current existing data standards is specifically the number of data records whose attributes are all reflected in the current existing data standards (including newly created data standards). For example, for a data record with three attributes: name, gender, and age, if all three attributes are recorded in the current existing data standards and have corresponding sub-standards, then this data record is a first data record whose data record attributes conform to the current existing data standards; if one or more of these three attributes are not recorded in the current existing data standards or do not have corresponding sub-standards, then this data record is not a first data record whose data record attributes conform to the current existing data standards.

[0147] Exemplarily, the number of second data records with no missing attribute values in the first data records specifically refers to the number of data records in the first data records that have corresponding attribute values for each data attribute. For example, for a data record with three attributes: name, gender, and age, if the attribute value corresponding to the name attribute of this data record is "Wang Wu", the attribute value corresponding to the gender attribute is "male", and the attribute value corresponding to the age attribute is "20", then this data record is a second data record with no missing attribute values. If one or more attributes in this data record do not have corresponding attribute values, then this data record is not a second data record with no missing attribute values.

[0148] Exemplarily, improving the data governance quality according to the data governance quality indicator may include, but is not limited to, determining whether the data governance quality indicator is less than or equal to an indicator threshold. If so, re-perform data governance according to all data standards and conduct inspections on each piece of data after governance to improve the data governance quality. Among them, the indicator threshold can be determined according to the actual situation, such as being determined to be, but not limited to, 90%, 80%, etc.

[0149] It should be noted that after the staff perform data governance according to the governance scope, obtaining the total number of data standard attributes included in all data standards, the total number of data table attributes of all data tables, the number of first data records whose data record attributes conform to the current existing data standards, the number of second data records with no missing attribute values in the first data records, and the number of third data records whose all attribute values in the second data records conform to the existing data standards; obtaining the data governance quality indicator based on the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records, and feeding back the data governance quality indicator to the staff, so that the staff can improve the data governance quality according to the data governance quality indicator. The specific implementation manner can be determined by those skilled in the art according to the actual situation. The above description is only an example and does not constitute a limitation thereto.

[0150] Through the above steps, the quality of data governance can be improved, which is further more conducive to the normal operation of the relevant data system.

[0151] In an alternative embodiment, as Figure 5 shown, obtaining the data governance quality indicator based on the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records includes the following steps:

[0152] S501: Obtain the coverage rate based on the total number of data standard attributes and the total number of data table attributes.

[0153] S502: Obtain the completeness rate based on the first data record count and the second data record count.

[0154] S503: Obtain the accuracy rate based on the first data record count and the third data record count.

[0155] S504: Obtain the data governance quality indicator based on the coverage rate, the completeness rate, and the accuracy rate.

[0156] Exemplarily, the obtaining of the coverage rate based on the total number of data standard attributes and the total number of data table attributes may, but is not limited to, dividing the total number of data standard attributes by the total number of data table attributes to obtain the coverage rate.

[0157] Exemplarily, the obtaining of the completeness rate based on the first data record count and the second data record count may, but is not limited to, dividing the second data record count by the first data record count to obtain the completeness rate.

[0158] Exemplarily, the obtaining of the accuracy rate based on the first data record count and the third data record count may, but is not limited to, dividing the third data record count by the first data record count to obtain the accuracy rate.

[0159] Exemplarily, the obtaining of the data governance quality indicator based on the coverage rate, the completeness rate, and the accuracy rate may, but is not limited to, taking the average value of the coverage rate, the completeness rate, and the accuracy rate as the data governance quality indicator, or integrating the coverage rate, the completeness rate, and the accuracy rate to obtain the data governance quality indicator, so that the data governance quality indicator includes the coverage rate, the completeness rate, and the accuracy rate.

[0160] Through steps S501 to S504, based on the data governance situation, quantitative data governance sub - indicators are obtained, and then these sub - indicators are processed to obtain the data governance quality indicator, which can make the obtained data governance quality indicator more in line with the actual situation of data governance, improve the accuracy of the data governance quality indicator, and thus be more conducive to improving the data governance quality according to the data governance quality indicator.

[0161] Based on the same principle, an embodiment of the present invention discloses a data governance device 600, as Figure 6 shown. The data governance device 600 includes:

[0162] A genetic attribute information determination module 601, configured to obtain the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standard.

[0163] A non - genetic attribute information determination module 602, configured to establish the non - genetic attribute information of the new data standard according to the newly introduced data.

[0164] A new module 603 is used to generate a new data standard according to the genetic attribute information and non-genetic attribute information, add the new data standard to the existing standard to obtain an updated total standard, determine the governance scope of data governance according to the updated total standard, and thus perform data governance according to the governance scope.

[0165] In an optional implementation manner, the genetic attribute information determination module 601 is used to:

[0166] According to the new identifier and the existing standard, obtain an approximate identifier of an approximate data standard corresponding to the new identifier;

[0167] According to the approximate identifier, obtain a superior identifier of a superior data standard corresponding to the new identifier;

[0168] According to the superior identifier, obtain the genetic attribute information of the new data standard.

[0169] In an optional implementation manner, the genetic attribute information determination module 601 is used to:

[0170] Perform semantic analysis and index analysis on the new identifier, and obtain an approximate identifier of an approximate data standard corresponding to the new identifier from the existing standard.

[0171] In an optional implementation manner, the genetic attribute information determination module 601 is used to:

[0172] According to the approximate identifier, obtain relevant nodes of the approximate data standard in the data standard pedigree;

[0173] Query the root node of the relevant node in the data standard pedigree, obtain the root node identifier of the root node data standard corresponding to the root node, and use the root node identifier as the superior identifier.

[0174] In an optional implementation manner, the genetic attribute information determination module 601 is used to:

[0175] According to the superior identifier, obtain the superior data standard;

[0176] According to the superior data standard, obtain the genetic attribute information.

[0177] In an optional implementation manner, the non-genetic attribute information determination module 602 is used to:

[0178] According to the newly introduced data, obtain new attributes corresponding to the newly introduced data;

[0179] According to the new attributes, establish the non-genetic attribute information of the new data standard.

[0180] In an alternative embodiment, the new creation module 603 is configured to:

[0181] Obtain a first sub - standard of the non - genetic attribute and the non - genetic attribute information according to the non - genetic attribute information;

[0182] Obtain a second sub - standard of the genetic attribute and the genetic attribute information according to the genetic attribute information;

[0183] Generate a new data standard according to the genetic attribute, the second sub - standard, the non - genetic attribute, and the first sub - standard.

[0184] In an alternative embodiment, the new creation module 603 is configured to:

[0185] Obtain the parent node of the relevant node in the data standard pedigree according to the relevant node;

[0186] Add the new data standard as another child node of the parent node to the data standard pedigree to obtain an updated overall standard;

[0187] Wherein, the set of data standards of all nodes in the data standard pedigree before adding the other child node is the existing standard.

[0188] In an alternative embodiment, it further includes a data governance improvement module, configured to:

[0189] After the staff performs data governance according to the governance scope, obtain the total number of data standard attributes included in all data standards, the total number of data table attributes of all data tables, the number of first data records whose data record attributes conform to the currently existing data standards, the number of second data records without missing attribute values in the first data records, and the number of third data records in which all attribute values in the second data records conform to the existing data standards;

[0190] Obtain a data governance quality index according to the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records, and feedback the data governance quality index to the staff, so that the staff can improve the data governance quality according to the data governance quality index.

[0191] In an alternative embodiment, it further includes a data governance improvement module, configured to:

[0192] Obtain a coverage rate according to the total number of data standard attributes and the total number of data table attributes;

[0193] Obtain a completeness rate according to the number of first data records and the number of second data records;

[0194] Based on the first data record number and the third data record number, an accuracy rate is obtained.

[0195] Based on the coverage rate, integrity rate, and accuracy rate, a data governance quality indicator is obtained.

[0196] Since the principle of the data governance device 600 for solving problems is similar to the above method, the implementation of this data governance device 600 can refer to the implementation of the above method, which will not be elaborated here.

[0197] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0198] In a typical example, the computer device specifically includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-described method is implemented.

[0199] Next, refer to Figure 7 , which shows a schematic structural diagram of a computer device 700 suitable for implementing the embodiments of the present application.

[0200] As Figure 7 shown, the computer device 700 includes a central processing unit (CPU) 701, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0201] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom can be installed in the storage section 708 as needed.

[0202] Specifically, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711.

[0203] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0204] For convenience of description, the above-described apparatus is described by functionally dividing it into various units. Of course, when implementing the present application, the functions of the various units can be implemented in one or more software and / or hardware.

[0205] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0206] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0207] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0208] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.

[0209] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0210] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0211] Each embodiment in this specification is described in a progressive manner. For parts that are the same or similar among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0212] The above description is only for the embodiments of this application and is not intended to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A data governance method, characterized in that, Including: Obtain the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standard; Establish the non-genetic attribute information of the new data standard according to the newly introduced data; Generate a new data standard according to the genetic attribute information and the non-genetic attribute information, add the new data standard to the existing standard to obtain an updated total standard, and determine the governance scope of data governance according to the updated total standard, so as to perform data governance according to the governance scope; Among them, obtaining the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standard includes: Perform semantic analysis and index analysis on the new identifier, and obtain an approximate identifier of an approximate data standard corresponding to the new identifier from the existing standard; Obtain relevant nodes of the approximate data standard in the data standard pedigree according to the approximate identifier; Query the root node of the relevant node in the data standard pedigree, obtain the root node identifier of the root node data standard corresponding to the root node, and use the root node identifier as the superior identifier; Obtain the genetic attribute information of the new data standard according to the superior identifier.

2. The method according to claim 1, wherein Obtaining the genetic attribute information of the new data standard according to the superior identifier includes: Obtain the superior data standard according to the superior identifier; Obtain the genetic attribute information according to the superior data standard.

3. The method according to claim 1, wherein Establishing the non-genetic attribute information of the new data standard according to the newly introduced data includes: Obtain a new attribute corresponding to the newly introduced data according to the newly introduced data; Establish the non-genetic attribute information of the new data standard according to the new attribute.

4. The method according to claim 1, wherein Generating a new data standard according to the genetic attribute information and the non-genetic attribute information includes: Obtain a first sub-standard of the non-genetic attribute and the non-genetic attribute information according to the non-genetic attribute information; Obtain a second sub-standard of the genetic attribute and the genetic attribute information according to the genetic attribute information; Generate a new data standard according to the genetic attribute, the second sub-standard, the non-genetic attribute, and the first sub-standard.

5. The method according to claim 1, characterized in that, Adding the new data standard to the existing standard to obtain an updated total standard includes: Obtain the parent node of the relevant node in the data standard pedigree according to the relevant node; Add the new data standard as another sub-node of the parent node to the data standard pedigree to obtain an updated total standard; Among them, the set of data standards of all nodes in the data standard pedigree before adding the other sub-node is the existing standard.

6. The method according to claim 1, wherein Further including: After performing data governance according to the governance scope, obtain the total number of data standard attributes included in all data standards, the total number of data table attributes of all data tables, the number of first data records whose data record attributes conform to the current existing data standards, the number of second data records without vacancies in the attribute values in the first data records, and the number of third data records whose all attribute values conform to the existing data standards in the second data records; Based on the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records, a data governance quality indicator is obtained, and the data governance quality indicator is fed back to the staff so that the staff can improve the data governance quality according to the data governance quality indicator.

7. The method according to claim 6, characterized in that, The obtaining of the data governance quality indicator based on the total number of data standard attributes, the total number of data table attributes, the number of first data records, the number of second data records, and the number of third data records includes: Obtaining a coverage rate based on the total number of data standard attributes and the total number of data table attributes; Obtaining a completeness rate based on the number of first data records and the number of second data records; Obtaining an accuracy rate based on the number of first data records and the number of third data records; Obtaining a data governance quality indicator based on the coverage rate, the completeness rate, and the accuracy rate.

8. A data governance device, characterized in that, Including: A genetic attribute information determination module for obtaining the genetic attribute information of the new data standard according to the new identifier of the new data standard and the existing standard; A non-genetic attribute information determination module for establishing the non-genetic attribute information of the new data standard according to the newly introduced data; A new module for generating a new data standard according to the genetic attribute information and the non-genetic attribute information, adding the new data standard to the existing standard to obtain an updated total standard, determining the governance scope of data governance according to the updated total standard, and thus performing data governance according to the governance scope; The genetic attribute information determination module is used to perform semantic analysis and index analysis on the new identifier, obtain an approximate identifier of an approximate data standard corresponding to the new identifier from the existing standard; according to the approximate identifier, obtain relevant nodes of the approximate data standard in the data standard pedigree; query the root node of the relevant nodes in the data standard pedigree to obtain the root node identifier of the root node data standard corresponding to the root node, and use the root node identifier as the superior identifier; Obtaining the genetic attribute information of the new data standard according to the superior identifier.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1-7 is implemented.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Data governance method based on configuration management system and related equipment

    CN113590230A

  • Data standard management method and device and electronic equipment

    CN113849480A