Data management method and device, equipment and medium

By building data source containers and determining the main data sources and unique related fields, the data silos between different business systems are solved and efficient data governance is achieved.

CN120371840APending Publication Date: 2025-07-25HANGZHOU GONGHUI DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510856144.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There are data island problems between different business systems, which leads to high difficulty and low efficiency in data governance.

Method used

Build a data source container to connect to all data sources, record the data source of the attribute fields, determine the main data source and unique associated fields, and perform data governance through the data feature standard table.

Benefits of technology

It improves the convenience and accuracy of data acquisition, reduces the difficulty and complexity of data governance, and improves the efficiency of data governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371840A_ABST
    Figure CN120371840A_ABST
Patent Text Reader

Abstract

The invention discloses a data management method and device, equipment and a medium. The method comprises the following steps: acquiring all data tables in each data source, and loading each data table into a pre-constructed data source container as a to-be-processed table; verifying whether each to-be-processed table exists in a pre-constructed standard data element container, fusing each to-be-processed table into the standard data element container according to an existing result, and recording all data source information corresponding to all attribute fields in each to-be-processed table; determining a main data source and a unique associated field of the standard data according to the information of each data source; determining a data element standard table according to the main data source and the unique associated field; and performing data management according to the data element standard table. According to the technical scheme provided by the embodiment of the invention, the difficulty and complexity of data governance are reduced, and the efficiency of data governance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data governance, and in particular, to a data governance method, apparatus, device, and medium. Background Art

[0002] With the continuous development of science and technology, more and more industries have begun to improve their productivity and social competitiveness through computer technology and Internet technology. In this process, a large amount of data is bound to appear that needs to be processed. As the basic information for production and life, this data must be properly stored and processed. Therefore, the necessity of data governance is recognized by the majority of relevant technical personnel.

[0003] Currently, for any production entity, there are basically multiple business systems. It is easy to have the problem of data islands between different business systems, that is, the data is not directly interconnected, and the data elements and data lineage relationships are complex, resulting in high data governance difficulty and poor data governance efficiency. Summary of the Invention

[0004] This application provides a data governance method, apparatus, device, and medium to reduce the difficulty and complexity of data governance and improve the efficiency of data governance.

[0005] According to one aspect of this application, a data governance method is provided, including:

[0006] Obtain all data tables in each data source, and load each data table into a pre-constructed data source container as a table to be processed;

[0007] Verify whether each table to be processed exists in a pre-constructed standard data element container. According to the existence result, fuse each table to be processed into the standard data element container, and record all data source information corresponding to all attribute fields in each table to be processed;

[0008] Determine the main data source and the unique association field of the standard data according to each data source information;

[0009] Determine a data element standard table according to the main data source and the unique association field;

[0010] Perform data governance according to the data element standard table.

[0011] According to another aspect of this application, a data governance apparatus is provided, including:

[0012] Obtain all data tables in each data source, and load each data table into a pre-constructed data source container as a table to be processed;

[0013] Verify whether each table to be processed exists in the pre-constructed standard data element container. According to the existence results, fuse each table to be processed into the standard data element container, and record all data source information corresponding to all attribute fields in each table to be processed.

[0014] Determine the main data source and the unique association field of the standard data according to each data source information.

[0015] Determine the data element standard table according to the main data source and the unique association field.

[0016] Perform data governance according to the data element standard table.

[0017] According to another aspect of the present application, there is provided an electronic device, which includes:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data governance method according to any embodiment of the present application.

[0021] According to another aspect of the present application, there is provided a computer-readable storage medium, which stores computer instructions for causing a processor to implement the data governance method according to any embodiment of the present application when executed.

[0022] According to another aspect of the present application, there is provided a computer program product, which includes a computer program that implements the data governance method according to any embodiment of the present application when executed by a processor.

[0023] In the technical solution of the embodiment of the present application, constructing a data source container can be connected to all data sources, thereby improving the convenience of obtaining data; after fusing each data into the standard data element container, record the data sources corresponding to all attribute fields to ensure the comprehensiveness of the attribute fields; determining the main data source and the unique association field can improve the accuracy of determining the standard data; performing data governance according to the data element standard table can further reduce the difficulty and complexity of data governance and improve the efficiency of data governance.

[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for description in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0026] Figure 1 is a flowchart of a data governance method provided in Embodiment 1 of the present application;

[0027] Figure 2 is a schematic diagram of a data governance process applicable to Embodiment 2 of the present application;

[0028] Figure 3 is a schematic structural diagram of a data governance device provided in Embodiment 3 of the present application;

[0029] Figure 4 is a schematic structural diagram of an electronic device for implementing the data governance method of the embodiments of the present application. Detailed Embodiments

[0030] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0032] Embodiment 1

[0033] Figure 1The present invention provides a flowchart of a data governance method for the first embodiment of the present application. This embodiment is applicable to scenarios such as data governance and data source restoration operations on data in multiple different business systems. This method can be executed by a data governance device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device. As Figure 1 shown, the method includes:

[0034] S110. Obtain all data tables in each data source, and load each data table into a pre-constructed data source container as a table to be processed.

[0035] Among them, the data source can be the source from which data is obtained during the data governance process, that is, where to obtain data. Exemplarily, it can be multiple business systems owned by the production entity, and these business systems perform different work contents and respectively store data corresponding to the work contents. During the data governance process, various types of data are obtained from these business systems, so these business systems can be used as data sources. Of course, it is possible to directly obtain data from the databases in the business systems or through the external interfaces of these business systems. The present application embodiment does not limit the acquisition method. A data table can be a unit for storing data and usually exists in the form of a table. Moreover, when obtaining all data tables of the data source, all data in the table is actually obtained.

[0036] The data source container can be a unit for connecting to all data sources and managing the tables and data obtained from each data source, and can be implemented based on a server. One data source container can correspond to (connect to) multiple data sources (business systems).

[0037] When the data source container obtains all data tables in each data source, it also obtains all data, which can include but is not limited to table names, all fields in the table, field types, description information of the fields, and all data corresponding to the fields. The tables to be processed are the states of the tables of each data source in the data source container, and these tables to be processed serve as the basis for subsequent data governance.

[0038] Load the data tables and data obtained from each data source into the data source container, and form a path index, such as "data source → table → field", to record the tables to which each field belongs and the data sources of each table.

[0039] S120. Verify whether each table to be processed exists in a pre-constructed standard data element container. According to the existence result, fuse each table to be processed into the standard data element container, and record all data source information corresponding to all attribute fields in each table to be processed.

[0040] It should be noted that the standard data can actually be understood as data that has all the dimensional information of a certain object. This standard data can exist in the form of a table. Exemplarily, assuming that the data object corresponding to a certain standard data is a user (such as a natural person or a unit, institution, etc., which will not be elaborated here), taking a natural person as an example here, the dimensions describing this natural person can include but are not limited to name, age, native place, ID number, contact information, home address, etc., which contains all the information of this user and can be used as the standard data of the user.

[0041] Correspondingly, the standard data element container can be a container for storing and managing standard data, which can be implemented through a server, for example. Then, first determine whether the table to be processed already exists in this standard data element container, and fuse each table to be processed into the standard data element container according to the result of whether it already exists. For example, if the table to be processed already exists in the standard data element container, it is possible to compare whether the data between the same tables to be processed is the same. If there are different data, the newly emerged data can be added to the standard data element container and become part of the standard data; correspondingly, if the table to be processed does not exist in the standard data element container, then directly insert the table to be processed into the standard data element container.

[0042] At the same time, when fusing new data into the standard data element container, record the data source information corresponding to all attribute fields in all tables to be processed. The attribute field can be the field name of the description dimension of the data object. Continuing the previous example, for example, "name, age, address" of a natural person, etc. It can be understood that recording the data sources of all fields provides a basis for subsequent data governance. It can be understood that there may be many data sources for the data fused in the standard data element container, and the same data may have multiple data sources, and all these existing data sources are recorded.

[0043] S130. Determine the main data source and the unique associated field of the standard data according to each data source information.

[0044] Among them, the main data source can be the most reliable data source in the data in the standard data element container, that is to say, the main data source is the data source with the highest authenticity and reliability. All the data sources in the standard data are recorded in the previous steps, and the one with the highest authenticity and reliability is selected from these data sources as the main data source.

[0045] The unique associated field can be the field information that has a unique association with the standard data. That is, among multiple fields in the standard data, if there is a certain field that has a unique corresponding relationship with the standard data, then this field can be used as the unique associated field.

[0046] Exemplarily, for the continuity potential, assuming that the object of the standard data is a natural person (user), then the information of all dimensions describing this user can be used as the standard data. For a natural person, there is one and only one ID number, so the "ID number" can be used as the unique associated field corresponding to the standard data of this user.

[0047] S140. Determine the data element standard table according to the main data source and the unique associated field.

[0048] Among them, the data element standard table can be a table for storing all standard data. That is, for the data object, a standard of data elements is established, and all description dimension information of the data object can be queried from this data element standard table.

[0049] Since the main data source has the highest accuracy and reliability, it can be understood that when multiple data sources have different description information for the same dimension of the data object, it can be directly determined according to the main data source. Exemplarily, continuing the previous example, when the contact phone numbers in the personal information recorded for a certain natural person user in different business systems are completely different, at this time, directly store the contact phone number corresponding to the main data source determined in the previous step as the standard data in the data element standard table. Similarly, since the unique associated field corresponds uniquely to the data object, it can be used to verify the existence of error information to ensure that the information in the data element standard table is all accurate content.

[0050] S150. Perform data governance according to the data element standard table.

[0051] Based on the data element standard table determined in the previous steps, data governance can be performed. Since the data in the data element standard table integrates all data sources (such as business systems), the data in the data element standard table is the most comprehensive and accurate. Therefore, data synchronization, data correction, data backtracking, etc. can be performed on each data source (such as business systems) according to the data content in the data element standard table. The embodiments of the present application do not make limitations in this regard.

[0052] In the technical solution of the embodiments of the present application, constructing a data source container can be connected to all data sources, thereby improving the convenience of obtaining data; after integrating each data into the standard data element container, record the data sources corresponding to all attribute fields to ensure the comprehensiveness of the attribute fields; determining the main data source and the unique associated field can improve the accuracy of determining the standard data; performing data governance according to the data element standard table can further reduce the difficulty and complexity of data governance and improve the efficiency of data governance.

[0053] In an alternative embodiment, the step of loading each data table into a pre-constructed data source container as a table to be processed in S110 may include:

[0054] S111. Split the name of a field in any data table according to a preset rule to obtain at least one field name splitting result.

[0055] Among them, the preset rule may be a splitting rule for the field name, that is, how the field name needs to be split into multiple parts. The field name splitting result may be the part obtained by splitting the name of the field in the data table, and generally there are multiple ones.

[0056] It should be noted that there are generally naming rules for the fields in the data table, and the corresponding preset rule may be a splitting rule corresponding to the naming rule. For example, if the characters or strings before and after in the field are connected by an underscore "_", then the preset rule may split the field name by identifying the underscore to obtain multiple characters or strings as the field name splitting result.

[0057] S112. In response to any field name splitting result referring to the name of another data table, determine the reference relationship between the data tables.

[0058] Among them, if the split field name splitting result refers to the name of another data table, it means that this data table also refers to other data tables, so the reference relationship between the data tables can be determined. Exemplarily, the table names of all data tables stored in the data source container can be marked with a root model (assuming the table is the root), and then the fields of the table are judged to determine whether there is a situation of referring to the id (Identity document) of other tables. If the field name includes the id of other tables, judge whether all fields have a self-reference situation (that is, whether they refer to the id of the table where the field is located). If it is not a self-reference situation, it is determined that the id of other tables is referred to. Then, the field name is split according to a certain recognition rule to determine the name of the table referring to the id field, and an affiliated model object is established on the referred table, and the table name and table information referring to this table are filled into the affiliated model object. Once it becomes an affiliated model object, it loses the situation of assuming the root model. Thus, a relationship model tree with the reference relationship between tables can be obtained. That is to say, this relationship model tree can have the reference relationship between all tables.

[0059] Of course, the splitting rule can be set by relevant technical personnel according to the actual situation or manual experience, and the method of judging the reference relationship is not limited to determining by identifying the id. The embodiments of the present application are only for illustration and are not limited.

[0060] S113. Load each data table recording the reference relationship into the data source container as a table to be processed.

[0061] Continuing from the previous example, the data table with the relational model tree is saved in the data source container as the table to be processed, serving as the basis for subsequent data governance.

[0062] In the above implementation, by splitting the field names in the table and identifying the reference relationships between the tables, clarifying the reference relationships between different tables can effectively help accurately find the required data during the data governance process, contributing to improving the accuracy and efficiency of data governance.

[0063] In an alternative implementation, verifying whether each table to be processed exists in the pre-constructed standard data element container in S120 may include: in response to the name of the table to be processed being in Chinese pinyin or English abbreviation, converting the name of the table to be processed into an English name according to the preset annotation information of the table to be processed; in response to the name of the table to be processed being an English name, verifying whether the table to be processed exists in the standard data element container according to the similarity between the English name and the name of the table to be processed.

[0064] It should be noted that when querying whether a table to be processed exists in the standard data element container, it is identified by the English name of the table to be processed. Then there is a situation where the name of the table is not an English name, but a Chinese name or an English abbreviation. In this case, the name of the table needs to be converted into an English name before querying. In this implementation, the name of the table is translated into an English name according to the pre-annotated annotation information in each table. The preset annotation information is used to make remarks on the table, which can directly remark the English name of the table or describe the table through remarks to assist in translating the table name. The method of translating the table name can choose to adopt a translation large model in the relevant field, which is not limited in this embodiment of the present application.

[0065] After converting the names of all tables into English names, the similarity between the names of the tables to be processed in the standard data element container and the English names of the tables to be stored in the standard data element container is identified. Any similarity calculation method in the related technology can be used, such as using cosine similarity calculation, to determine whether the table name to be stored in the standard data element container already exists. For example, if the similarity is between 80% - 100%, it is considered that the table corresponding to the table name already exists in the standard data element container.

[0066] In the above implementation, by verifying whether a certain table to be processed already exists in the standard data element container, it can provide a basis for whether the table needs to be fused or inserted subsequently, contributing to timely supplementing or expanding the data content in the standard data element container and providing strong support for subsequent data governance.

[0067] In an alternative embodiment, the step of fusing each table to be processed into the standard data element container according to the existence result in S120 may include:

[0068] S121. In response to the table to be processed existing in the standard data element container, convert the name of the field to be processed in the table to English to obtain the English field name.

[0069] Among them, if the table to be processed is confirmed to already exist in the standard data element container, the names of all fields (i.e., the fields to be processed) in the table to be processed are converted to English. The conversion method to English may be the same as or different from the method for converting the table name in the above embodiment. This application embodiment does not make a limitation in this regard. That is, the field annotation according to the preset annotation can be used to translate the field name. Translate the names of all fields into English to obtain the English field name.

[0070] S122. According to the similarity between the English field name and the field name in the existing table, determine whether each field to be processed exists in the existing table.

[0071] It is the same as the method for judging whether the table name already exists described in the foregoing embodiment. Determine the similarity between the English field name obtained by translating the field name in the foregoing step and the field name of the table already existing in the standard data element container, and judge whether the field to be processed already exists in the table according to this similarity, that is, whether it already exists in the standard data element container.

[0072] Of course, the similarity can be determined by any one of the calculation methods in the related art. For example, by calculating the cosine similarity, this application embodiment does not make a limitation in this regard. Exemplarily, if the similarity is between 80% and 100%, it can be considered that the field to be processed already exists in the table, that is, it already exists in the standard data element container; otherwise, it is considered that the field to be processed does not exist in the table, that is, it does not exist in the standard data element container.

[0073] S123. In response to the field to be processed not existing in the existing table, insert and store the field to be processed and the corresponding data into the existing table.

[0074] It can be understood that if the field to be processed does not exist in the existing table, it means that the field to be processed and its related data are missing in the standard data. Just insert and store the field to be processed and its related data into this table. Correspondingly, if the field to be processed exists in the existing table, then the data corresponding to the field to be processed can be compared, and the non-existing data can be inserted and stored to ensure that the data content corresponding to the field to be processed also exists in the standard data element container.

[0075] S124. In response to the pending table not existing in the standard data element container, insert and store the pending table in the standard data element container.

[0076] It can be understood that if the pending table does not exist in the standard data element container, the pending table can be directly inserted and stored in the standard data element container, so as to ensure that the data content covered in the standard data element container is the most comprehensive.

[0077] In the above embodiment, by verifying whether the fields and data exist in the tables in the standard data element container, and then determining whether to insert and store them in the standard data element container, the comprehensiveness of the data content in the standard data element container is ensured, providing a favorable basis for subsequent data governance in accordance with the standard data.

[0078] In another alternative embodiment, the determining the main data source and the unique association field of the standard data in S130 according to the information of each data source may include:

[0079] S131. Determine the candidate sources in each data source according to the field width and data volume in the data source information.

[0080] Among them, the field width may be the number of field types in the data source. The more the number of field types, the wider the field width. Exemplarily, continuing the previous example, a natural person user has three dimensions of field types (such as name, age, and native place) in a certain business system A, and five dimensions of field types (such as name, age, native place, address, and contact information) in another business system B. Then it can be considered that the field width of the description of the user by business system B is wider than that of business system A. In fact, the field width can be understood as the number of description dimensions of the data object.

[0081] The data volume of the data source may be the number or volume of the existing data in the data source. The more data in the business system, the more it proves the importance of the business system.

[0082] Determine the candidate sources from all data sources according to the field width and data volume. Among them, the candidate sources are more important and have higher data reliability compared to other data sources.

[0083] Exemplarily, the reliability index of each data source can be calculated according to the weighted sum of the field width and the data volume. For example, first normalize and / or standardize the values of the field width and the data volume, and then calculate the weighted sum according to the preset weights. A part of the data sources with higher obtained reliability indices can be used as candidate sources. Of course, the weights can be set by those skilled in the art according to a large number of experiments or actual situations. Preferably, the weight of the field width is greater than the weight of the data volume. It should be noted that a large data volume does not mean a more comprehensive description of the data object, while the field width means the number of dimensions of the description of the data object. The wider the field width, the more dimensions of the description, the more comprehensive the description of the object, and then the more reliable the data in this data source.

[0084] S132. Determine the main data source according to the field storage rate of each candidate source.

[0085] Among them, the field storage rate can be the value of the data volume stored corresponding to a field. If the field storage rate is 0, it means that no data is stored under this field. It can be understood that even if the field width is wide, but no data is stored under the field, it is meaningless and the description of the data object is not accurate enough. Therefore, calculate the field storage rates of each candidate source determined in the foregoing steps respectively, and use the candidate source with the highest field storage rate as the main data source. Of course, the calculation method of the field storage rate can adopt any one in the related technologies, and the embodiments of the present application do not make limitations in this regard.

[0086] Then, it can be understood that among numerous data sources, several candidate sources with higher data reliability are screened out by the field width and the data volume. Using the candidate source with the highest field storage rate as the main data source indicates that the data in this data source is the most reliable, which helps to provide an accurate basis for subsequent data governance.

[0087] S133. In response to all data sources of any field in the standard data existing and being non-duplicate, use any field as the unique association field; wherein, the unique association field has a unique corresponding relationship with the standard data.

[0088] It should be noted that if all data sources corresponding to a field in the standard data exist and are non-duplicate, then it can be considered that there is a unique corresponding relationship between this field and the standard data.

[0089] It is rational. For a natural person user, when querying the "ID number" field related to this user in all business systems, it is found that the ID number appears in multiple business systems, but each business system is different. Then, the "ID number" can be used as the unique association field. In another example, for a natural person user, when querying the "contact phone number" field related to this user in all business systems, it is found that there are duplicates in the business systems where the contact phone number of this user exists. The reason is that different users can set the same contact phone number (for example, different family members in the same family all set the landline phone at home). Therefore, the "contact phone number" cannot be used as the unique association field.

[0090] In yet another alternative implementation, the data governance includes data synchronization and data sourcing back.

[0091] In S150, the data governance according to the data element standard table may include: determining the modification and update frequency of each data source according to the corresponding data sources in the data element standard table; synchronizing the data element standard table with each data source according to the modification and update frequency; in response to data changes in the data element standard table, pushing the changed data to each data source according to the association relationship between the standard data element container and each data source to complete data sourcing back.

[0092] Among them, the modification and update frequency can be, for each data source, the modification frequency of the data stored by itself. Exemplarily, during the operation of a certain business system, new data may be stored or old data may be modified and updated due to certain transaction processes. These frequencies of addition and modification can be used as the modification and update frequency.

[0093] The association relationship between the standard data element container and each data source can be the interaction relationship established between the standard data element container and each data source at the beginning of construction. Based on these interaction relationships, the standard data element container can apply to each data source to obtain their data.

[0094] Then, according to the modification and update frequency, apply to each data source for their data to ensure that the data in the standard data element container is the latest.

[0095] At the same time, once the data in the standard data element container (that is, the data in the data element standard table) changes, it means that the standard data has changed. Other data sources (business systems) that have not been updated should be updated in time to ensure data accuracy. Therefore, when the standard data changes, synchronize the changed data to all data sources to achieve data sourcing back.

[0096] In the above embodiments, based on the interaction relationship between the standard data element container and each data source, the updated data in the data source is synchronized to the standard data element container, and then the standard data element container is respectively sourced back to other data sources, so that the data of all data sources can be updated in a timely manner, thereby ensuring the real-time and accuracy of the data.

[0097] In yet another alternative embodiment, the data governance includes data quality verification;

[0098] In S150, the data governance according to the data element standard table may include: inputting each field to be verified in the standard data element container into a pre-set big data model for question-and-answer reasoning to obtain the field data verification rules for the field to be verified; in response to the completion of data synchronization in the data element standard table, determining whether the data of the field to be verified is normal according to the field data verification rules; in response to the data of the field to be verified not conforming to the field data verification rules, determining the data of the field to be verified as abnormal data, recording and outputting an alarm log.

[0099] Among them, the big data model may be a large model for question-and-answer reasoning. The field data verification rules may be rules for verifying a field to verify whether the data in the field is normal. The fields to be verified are actually all the fields in the standard data element container, and the data corresponding to these fields all need to be verified. First, the name of the field to be verified is input to the big data model for question-and-answer reasoning to obtain how to verify the data corresponding to this field. Exemplarily, inputting "ID number" into the big data model, the big data model will output the judgment basis according to the characteristics of the ID number, such as "the ID number has 18 digits and starts with a number".

[0100] When the data element standard table receives the updated data content from the data source, first complete the data synchronization (such as the data synchronization process described in the above embodiments), and then verify the data under each field according to the field data verification rules output by the large model. Continuing the previous example, if it is found that the data corresponding to the "ID number" field is not 18 digits and / or does not start with a number, it is considered that the data in this field does not conform to the field data verification rules, and it is processed as abnormal data. The processing method may be to record the content of the abnormal data and output the corresponding alarm log.

[0101] In the above embodiments, a practical method for data quality verification is provided, which can timely check according to the field data verification rules after data synchronization, thereby ensuring the accuracy of the data and helping to provide accurate data when sourcing back data subsequently.

[0102] Embodiment II

[0103] Figure 2The flowchart of a data governance method provided in the second embodiment of this application. The embodiment of this application is a preferred example provided on the basis of the foregoing embodiments and implementation manners. As Figure 2 shown, it specifically includes:

[0104] By constructing a data source container, establish data source connection management for all business system sources. At the same time, according to each connection of the data source container, collect the metadata of the data source. The collected data includes table name, fields, types, descriptions, and data volume. After all the metadata is collected, build the association relationship from the data source to the metadata in the data source container (data source → table → field). After the data source container is constructed, the following processing is performed:

[0105] S210. Construct the relationship of business data element models.

[0106] Determine the construction of the core business object relationship of each data source according to the metadata of each data source in the data source container. First, mark the root model for the table name for all metadata in the data source container. Then, judge the fields of the table to determine whether there is a situation of referring to the id of other tables. If the field name contains id, it is necessary to judge according to the association of all fields to determine whether the field is self-referencing. The characteristics of self-referencing are: the appearance of fields such as parent_id (parent name), path (path), ids (identifier), etc. If it is not a self-referencing situation, split the field name according to the id field, and determine the table name of the table that references the id field according to the split name. Establish an affiliated model object in this table, and put the current table name and table information into it, and remove this table from the data source container;

[0107] S220. Construct a standard data element container.

[0108] Synthesize the data element models of the data sources in the data source container, obtain the tables of each data, and judge whether they exist in the standard data element container. If they do not exist, insert them. The main judgment logic for judging whether they exist in the standard data element container is as follows: First, infer according to the table name to determine the naming method of the current table name, whether it belongs to Chinese pinyin, letter abbreviation, or English naming. If it is Chinese pinyin or letter abbreviation, it is necessary to perform English translation according to the annotation of the table, and uniformly search in the standard data element container through English naming. If the table name cannot be located, infer according to the similarity of the fields to determine whether they are the same. If the same table name can be found in the standard data element container, it is necessary to perform field fusion. The fusion judgment method is similar to that of the table name. Determine the naming method of the field name, uniformly fuse it in English, and supplement it to the standard data element container. At the same time, in the standard data element container, record the source of each field. If there are the same tables and fields in multiple data sources, record multiple sources.

[0109] S230. Determine the main data sources.

[0110] Through the processing in S220, the standards of data elements of the entire business system are formed in the standard data element container. However, there may be multiple sources for the data in the tables in the standard data element container. In this case, it is necessary to determine the source of each element field. By polling the standard data element container and judging whether there are multiple sources for the source of each table, if so, obtain all the sources, and perform a priority sorting according to the data volume of the tables of each source and the field width of the tables. The tables with wider widths and larger data volumes are the candidate sources. After determining the candidate sources, it is necessary to judge again according to the storage rate of the field data of the corresponding tables. If there is no data (storage rate is 0) or the storage rate is relatively low in the fields of the data table, it will be screened out, and the candidate source with a relatively high storage rate is selected as the main data source. When determining the unique associated field of the data element, for the standard data element container model, if all sources of each model field exist and are not repeated, this field can be marked as the unique associated field.

[0111] S240. Execution of data governance tasks.

[0112] After the establishment of the data element standard, create a data element model table in the data governance library, and at the same time construct the data governance tasks of the data element model. In the data governance tasks, it mainly includes data synchronization, data quality tasks, and data source subscription. Complete the automatic construction of data tasks:

[0113] Data synchronization: For data synchronization, it is necessary to detect the time according to the model data source table, determine the time trend of data increase and modification, such as the data modification frequency per day, week, or month, and establish an incremental data synchronization task according to the modification frequency.

[0114] Data quality tasks: For the establishment of data quality tasks, mainly obtain the data model tables of the standard data element container, and perform large model question-answering reasoning according to the field names of the tables to obtain the field data verification rules and convert them into regular expressions (for example: the rules of email or ID card number, etc. Regular expressions are used for identification and execution at the code level). After the data synchronization task is completed, the data quality task is automatically triggered for quality verification. If it is found that the data does not meet the quality requirements, an alarm log of the data is output.

[0115] Data source subscription task: According to the data model relationship of the standard data element container, construct a data source subscription task listener. When the data changes, automatically associate the data source according to the data fields of the standard data element container and push the latest data.

[0116] Embodiment III

[0117] Figure 3 This is a schematic structural diagram of a data governance device provided in Embodiment 3 of the present application. As Figure 3 shown, the device 300 includes:

[0118] A data source table acquisition module 310, configured to acquire all data tables in each data source, and load each data table into a pre-constructed data source container as a table to be processed;

[0119] A data source record module 320, configured to verify whether each table to be processed exists in a pre-constructed standard data element container, and according to the existence result, fuse each table to be processed into the standard data element container, and record all data source information corresponding to all attribute fields in each table to be processed;

[0120] A main source determination module 330, configured to determine the main data source and the unique association field of the standard data according to each data source information;

[0121] A data standard determination module 340, configured to determine a data element standard table according to the main data source and the unique association field;

[0122] A data governance module 350, configured to perform data governance according to the data element standard table.

[0123] In the technical solution of the embodiment of the present application, constructing a data source container can be connected to all data sources, thereby improving the convenience of obtaining data; after fusing each data into the standard data element container, recording the data sources corresponding to all attribute fields to ensure the comprehensiveness of the attribute fields; determining the main data source and the unique association field can improve the accuracy of determining the standard data; performing data governance according to the data element standard table can further reduce the difficulty and complexity of data governance, and at the same time improve the efficiency of data governance.

[0124] In an alternative embodiment, the data source table acquisition module 310 may include:

[0125] A field splitting unit, configured to split the name of the field in any data table according to a preset rule to obtain at least one field name splitting result;

[0126] A reference confirmation unit, configured to determine the reference relationship between data tables in response to a field name splitting result referring to the name of another data table;

[0127] A table loading unit, configured to load each data table recording the reference relationship into the data source container as a table to be processed.

[0128] In an alternative embodiment, the data source record module 320 may include:

[0129] A name conversion unit, configured to, in response to the name of the table to be processed being in Chinese pinyin or English abbreviation, convert the name of the table to be processed into an English name according to the preset annotation information of the table to be processed;

[0130] An existence verification unit, configured to, in response to the name of the table to be processed being an English name, verify whether the table to be processed exists in the standard data element container according to the similarity between the English name and the name of the table to be processed.

[0131] In an alternative embodiment, the data source recording module 320 may include:

[0132] An English field conversion unit, configured to, in response to the table to be processed existing in the standard data element container, convert the name of the field to be processed in the table to be processed into English to obtain an English field name;

[0133] A field verification unit, configured to determine whether each field to be processed exists in the existing table according to the similarity between the English field name and the field names in the existing table;

[0134] A field fusion unit, configured to, in response to the field to be processed not existing in the existing table, insert and store the field to be processed and the corresponding data into the existing table;

[0135] A table fusion unit, configured to, in response to the table to be processed not existing in the standard data element container, insert and store the table to be processed in the standard data element container.

[0136] In an alternative embodiment, the main source determination module 330 may include:

[0137] A candidate source determination unit, configured to determine candidate sources in each data source according to the field width and data volume in the data source information;

[0138] A main data source determination unit, configured to determine the main data source according to the field storage rate of each candidate source;

[0139] A unique association field determination unit, configured to, in response to all data sources of any field in the standard data existing and being non-duplicate, use any field as the unique association field; wherein, the unique association field has a unique corresponding relationship with the standard data.

[0140] In an alternative embodiment, the data governance includes data synchronization and data sourcing;

[0141] The data governance module 350 may include:

[0142] An update frequency determination unit, configured to determine the modification and update frequencies of each data source according to each data source corresponding to the data element standard table;

[0143] A data synchronization unit for synchronizing the data element standard table with each data source according to the modified update frequency;

[0144] A data source return unit for pushing the changed data to each data source in response to data changes in the data element standard table according to the association relationship between the standard data element container and each data source, so as to complete the data source return.

[0145] In an optional implementation manner, the data governance includes data quality verification;

[0146] The data governance module 350 may include:

[0147] A verification rule determination unit for inputting each field to be verified in the standard data element container into a pre-set big data model for question-and-answer reasoning to obtain the field data verification rule for the field to be verified;

[0148] A data quality inspection unit for determining whether the data of the field to be verified is normal according to the field data verification rule in response to the completion of data synchronization of the data element standard table;

[0149] An abnormal data alarm unit for determining that the data of the field to be verified is abnormal data in response to the data of the field to be verified not conforming to the field data verification rule, and recording and outputting an alarm log.

[0150] The data governance device provided by the embodiments of the present application can execute the data governance method provided by any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing each data governance method.

[0151] Embodiment 4

[0152] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0153] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0154] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0155] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data governance method.

[0156] In some embodiments, the data governance method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data governance method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data governance method in any other appropriate manner (e.g., by means of firmware).

[0157] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0158] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0159] In the context of this application, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0162] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0163] The embodiment of the present application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the data governance method provided in any embodiment of the present application. This program product and the data governance methods disclosed in the embodiments of the present application belong to the same inventive concept, so details are not described herein.

[0164] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved, and no limitation is made herein.

[0165] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A data governance method, characterized in that, Including: Obtain all data tables in each data source, and load each of the data tables into a pre-constructed data source container as a table to be processed; Verify whether each of the tables to be processed exists in a pre-constructed standard data element container. According to the existence result, fuse each of the tables to be processed into the standard data element container, and record all data source information corresponding to all attribute fields in each of the tables to be processed; Determine the main data source and the unique association field of the standard data according to each of the data source information; Determine a data element standard table according to the main data source and the unique association field; Perform data governance according to the data element standard table.

2. The method according to claim 1, wherein The step of loading each of the data tables into a pre-constructed data source container as a table to be processed includes: Split the name of the field in any one of the data tables according to a preset rule to obtain at least one field name splitting result; In response to any one of the field name splitting results referring to the name of another data table, determine the reference relationship between the data tables; Load each of the data tables recording the reference relationship into the data source container as a table to be processed.

3. The method according to claim 1, wherein The step of verifying whether each of the tables to be processed exists in a pre-constructed standard data element container includes: In response to the name of the table to be processed being in Chinese pinyin or English abbreviation, convert the name of the table to be processed into an English name according to the preset annotation information of the table to be processed; In response to the name of the table to be processed being an English name, verify whether the table to be processed exists in the standard data element container according to the similarity between the English name and the name of the table to be processed.

4. The method according to claim 1, wherein The step of fusing each of the tables to be processed into the standard data element container according to the existence result includes: In response to the table to be processed existing in the standard data element container, convert the name of the field to be processed in the table to be processed into English to obtain an English field name; Judge whether each of the fields to be processed exists in the existing table according to the similarity between the English field name and the field name in the existing table; In response to the field to be processed not existing in the existing table, insert and store the field to be processed and the corresponding data into the existing table; In response to the table to be processed not existing in the standard data element container, insert and store the table to be processed into the standard data element container.

5. The method according to claim 1, characterized in that, The step of determining the main data source and the unique association field of the standard data according to each of the data source information includes: Determine the candidate sources in each data source according to the field width and data volume in the data source information; Determine the main data source according to the field storage rate of each of the candidate sources; In response to all data sources of any field in the standard data existing and being non-repetitive, use the any field as the unique association field; wherein, the unique association field has a unique corresponding relationship with the standard data.

6. The method according to claim 1, wherein The data governance includes data synchronization and data return; The step of performing data governance according to the data element standard table includes: Determine the modification and update frequencies of each data source corresponding to the data element standard table; Synchronize the data element standard table with each data source according to the modification and update frequencies; In response to data changes in the data element standard table, push the changed data to each data source according to the association relationship between the standard data element container and each data source to complete data return to the source.

7. The method according to claim 6, characterized in that The data governance includes data quality verification; Performing data governance according to the data element standard table includes: Input each to-be-verified field in the standard data element container into a pre-set big data model for question-and-answer reasoning to obtain the field data verification rules for the to-be-verified field; In response to the data element standard table completing the data synchronization, determine whether the data of the to-be-verified field is normal according to the field data verification rules; In response to the data of the to-be-verified field not conforming to the field data verification rules, determine that the data of the to-be-verified field is abnormal data, record and output an alarm log.

8. A data governance device, characterized in that, Including: A data source table acquisition module, configured to acquire all data tables in each data source and load each data table into a pre-constructed data source container as a to-be-processed table; A data source record module, configured to verify whether each to-be-processed table exists in a pre-constructed standard data element container, and according to the existence result, fuse each to-be-processed table into the standard data element container and record all data source information corresponding to all attribute fields in each to-be-processed table; A main source determination module, configured to determine the main data source and the unique association field of the standard data according to each data source information; A data standard determination module, configured to determine a data element standard table according to the main data source and the unique association field; A data governance module, configured to perform data governance according to the data element standard table.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data governance method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the data governance method according to any one of claims 1-7 when executed.

Citation Information

Patent Citations

  • Data similarity calculation method, device, readable medium and electronic equipment

    CN113837307A

  • Data asset standardization management method and device

    CN115982166A

  • Data association method and device, electronic equipment and medium

    CN116483823A

  • Intelligent data management method of full-data AI management and control system

    CN119149922A