Data management method and device, equipment and storage medium

The use of a visual interface to generate and execute data governance tasks through a DAG improves efficiency by enabling task reuse and simplifying design, addressing inefficiencies in current data governance methods.

CN120316097APending Publication Date: 2025-07-15QINGDAO HISENSE TRANS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410031811.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, data governance demanders need to design a set of software for each data governance requirement, resulting in inefficiency.

Method used

By generating a directed acyclic graph (DAG), data governance nodes and their connection relationships are drawn on the visual human-computer interactive interface, target data governance tasks are generated, and configured node operations are performed in turn to realize data governance.

Benefits of technology

It improves the reusability and development efficiency of data governance tasks, and reduces the difficulty of data governance logic design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316097A_ABST
    Figure CN120316097A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data governance method and device, equipment and a storage medium, and the method comprises the steps: generating a data governance directed acyclic graph (DAG) in response to an editing instruction triggered on a visual human-computer interaction interface; wherein the data governance DAG comprises a plurality of data governance nodes, and the connection relationship between the data governance nodes represents the execution sequence of the data governance nodes in the target data governance task; generating the target data governance task according to the data governance DAG; in response to a starting instruction for a target data governance task, after it is determined that configuration of the to-be-configured data governance nodes in the target data governance task is completed, operation corresponding to the data governance nodes of the target data governance task is executed in sequence, and data governance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, and particularly to a data governance method, device, equipment and storage medium. Background Art

[0002] With the development of big data technology, more and more data is collected in different industries and dimensions. Valuable information is often hidden in this data, and the hidden information can be mined from big data through data governance, so as to provide effective analysis suggestions for business research, public management and other needs. At present, when different data governance demanders conduct data governance, the software used to implement data governance still stays in the traditional software development process of "requirement analysis - design - implementation - testing - delivery". Data governance personnel need to design a set of data governance software for each data governance requirement respectively, and the efficiency of data governance is relatively low. Summary of the Invention

[0003] Embodiments of the present invention provide a data governance method, device, electronic equipment and computer-readable storage medium, so as to provide a data governance solution with higher development efficiency.

[0004] Embodiments of the present invention provide a data governance method, including:

[0005] Responding to an editing instruction triggered on a visual human-computer interaction interface, generating a data governance directed acyclic graph (DAG); wherein, the data governance DAG includes a plurality of data governance nodes, and the connection relationship between the data governance nodes represents the execution order of the data governance nodes in the target data governance task;

[0006] Generating the target data governance task according to the data governance DAG;

[0007] Responding to a start instruction for the target data governance task, after determining that each to-be-configured data governance node in the target data governance task is respectively configured, sequentially executing the operations corresponding to the data governance nodes of the target data governance task to implement data governance.

[0008] Furthermore, the target data governance task at least includes a data collection node and a first processing node.

[0009] In the process of sequentially executing the operations corresponding to the data governance nodes of the target data governance task to implement data governance:

[0010] The data collection node is used to configure a target physical table containing data to be governed and the original database source of the target physical table;

[0011] The first data processing node is used to obtain the target physical table from the original database source, map the data to be governed in the target physical table to the first virtual table according to the mapping relationship between the first virtual table configured for the first data processing node and the target physical table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations to the first governance database in the form of a processed physical table.

[0012] Further optionally, the data governance node further includes at least one of the following:

[0013] The data profiling node after the data collection node is used to obtain the target physical table from the original database source, map the data to be governed in the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data profiling node and the target physical table, and perform at least one format check on the data to be governed in the second virtual table;

[0014] The data fusion node after the first data processing node is used to respectively determine the corresponding processed physical tables according to the multiple first data processing nodes connected to the data fusion node, obtain the multiple processed physical tables from the first governance database, map the data to be governed in the multiple processed physical tables to the third virtual table according to the mapping relationship between the third virtual table configured for the data fusion node and the multiple processed physical tables, and output a fused physical table to the second governance database according to the third virtual table;

[0015] The data verification node after the data fusion node is used to determine the corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed in the fused physical table to the fourth virtual table, and perform at least one format check on the data to be governed in the fourth virtual table;

[0016] The second data processing node after the data verification node is used to determine the corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed in the fused physical table to the fifth virtual table, and perform at least one data processing operation on the data to be governed in the fifth virtual table and output the result.

[0017] Optionally, obtaining a physical table from a database includes:

[0018] Obtaining a physical table from a database through JDBC;

[0019] Wherein, if the physical table is the target physical table, the database is the original database;

[0020] If the physical table is the processed physical table, then the database is the first governance database;

[0021] If the physical table is the integrated physical table, then the database is the second governance database.

[0022] Optionally, the configuration instructions of different data governance nodes are generated by different users;

[0023] The user who triggers the editing instruction is different from the user who configures the mapping relationship for the data governance node.

[0024] Based on the same inventive concept, an embodiment of the present invention further provides a data governance device, including:

[0025] A visual editing module, configured to generate a data governance DAG in response to an editing instruction triggered on a visual human-computer interaction interface; wherein, the data governance DAG includes a plurality of data governance nodes, and the connection relationship between the data governance nodes represents the execution order of each data governance node in the target data governance task;

[0026] A task generation module, configured to generate the target data governance task according to the data governance DAG;

[0027] An execution module, configured to, in response to a start instruction for the target data governance task, after determining that each to-be-configured data governance node in the target data governance task is respectively configured, sequentially execute the operations corresponding to each data governance node of the target data governance task to implement data governance.

[0028] Furthermore, the target data governance task at least includes a data collection node and a first processing node.

[0029] In the process of sequentially executing the operations corresponding to each data governance node of the target data governance task to implement data governance:

[0030] The data collection node is configured to configure a target physical table containing data to be governed and the original database source of the target physical table;

[0031] The first data processing node is configured to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the first virtual table according to the mapping relationship between the first virtual table configured for the first data processing node and the target physical table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations in the form of a processed physical table to the first governance database.

[0032] Further optionally, the data governance node further includes at least one of the following:

[0033] A data exploration node after the data collection node, configured to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data exploration node and the target physical table, and perform at least one format check on the data to be governed of the second virtual table;

[0034] A data fusion node after the first data processing node, configured to respectively determine corresponding processed physical tables according to multiple first data processing nodes connected to the data fusion node, obtain the multiple processed physical tables from the first governance database, map the data to be governed of the multiple processed physical tables to the third virtual table according to the mapping relationship between the third virtual table configured for the data fusion node and the multiple processed physical tables, and output a fused physical table to the second governance database according to the third virtual table;

[0035] A data verification node after the data fusion node, configured to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to the fourth virtual table, and perform at least one format check on the data to be governed of the fourth virtual table;

[0036] A second data processing node after the data verification node, configured to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to the fifth virtual table, and perform at least one data processing operation on the data to be governed of the fifth virtual table and output the result.

[0037] Based on the same inventive concept, an embodiment of the present invention further provides a device, including a processor and a memory for storing instructions executable by the processor;

[0038] Wherein, the processor is configured to execute the instructions to implement the data governance method.

[0039] Based on the same inventive concept, an embodiment of the present invention further provides a storage medium storing program code, which, when run on a computer, causes the computer to execute the data governance method.

[0040] The beneficial effects of the present invention are as follows:

[0041] The data governance method, device, equipment and storage medium provided by the embodiments of the present invention generate a data governance task composed of multiple data governance nodes combined in a certain order, and at least some of the data governance nodes can be configured and adjusted according to the configuration instructions of the user, so that the data governance requirements for some data to be governed that are different but have the same data governance logic can be realized using the same data governance task, improving the reusability of the data governance task and thus improving the development efficiency of data governance. In addition, in the data governance method provided by the embodiments of the present invention, the data governance task is generated by a DAG drawn by the user on a visual human-computer interaction interface, so that users who do not need to design data governance logic do not need to master program programming development knowledge, reducing the design difficulty of data governance logic. Description of the Drawings

[0042] Figure 1 It is a flowchart of the design stage of the data governance method provided by the embodiments of the present invention;

[0043] Figure 2 It is one of the schematic diagrams of the visual human-computer interaction interface in the embodiments of the present invention;

[0044] Figure 3 It is another schematic diagram of the visual human-computer interaction interface in the embodiments of the present invention;

[0045] Figure 4 It is still another schematic diagram of the visual human-computer interaction interface in the embodiments of the present invention;

[0046] Figure 5 It is a flowchart of the execution stage of the data governance method provided by the embodiments of the present invention;

[0047] Figure 6 It is one of the data governance DAGs in the embodiments of the present invention;

[0048] Figure 7 It is a schematic diagram of the mapping relationship configuration interface between the first virtual table and the target physical table in the embodiments of the present invention;

[0049] Figure 8 It is another data governance DAG in the embodiments of the present invention;

[0050] Figure 9 It is a schematic diagram of the user configuration permissions of different data governance nodes in the embodiments of the present invention;

[0051] Figure 10 It is a schematic diagram of the structure of the data governance device provided by the embodiments of the present invention;

[0052] Figure 11 It is a schematic diagram of the structure of the electronic device provided by the embodiments of the present invention;

[0053] Figure 12 Schematic diagram of the computer-readable storage medium provided by the embodiment of the present invention. Detailed implementation manners

[0054] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described below with reference to the accompanying drawings and embodiments. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments described herein; on the contrary, these embodiments are provided to make the present invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings represent the same or similar structures, so their repeated descriptions will be omitted. The words expressing positions and directions in the present invention are all described by taking the accompanying drawings as examples, but can be changed according to needs, and all changes are included in the protection scope of the present invention. The accompanying drawings of the present invention are only used to illustrate the relative position relationship and do not represent the actual proportion.

[0055] It should be noted that specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementation manners disclosed below. The subsequent description of the specification is the preferred implementation manner for implementing the present application, but the description is for the purpose of illustrating the general principles of the present application and is not intended to limit the scope of the present application. The protection scope of the present application shall be determined by the scope defined by the appended claims.

[0056] The data governance method, apparatus, device, and storage medium provided by the embodiment of the present invention will be specifically described below with reference to the accompanying drawings. It should be stated that in the technical solutions involved in the embodiments of the present invention, the collection, dissemination, use, etc. of data all comply with the requirements of relevant national laws and regulations.

[0057] The embodiment of the present invention provides a data governance method, including a design process and an execution process. Among them, as Figure 1 shown, the design process includes:

[0058] S110. Display a visual human-computer interaction interface (User Interface, UI) for generating a data governance task.

[0059] S120. Respond to an edit instruction triggered on the visual human-computer interaction interface to generate a data governance directed acyclic graph (Directed Acyclic Graph, DAG). Among them, the data governance DAG includes a plurality of data governance nodes, and the connection relationship between the data governance nodes represents the execution order of the data governance nodes in the target data governance task.

[0060] DAG is a common data structure in graph theory. DAG consists of a set of nodes and directed edges, where nodes represent data in the graph and directed edges represent the dependencies between the data corresponding to the nodes. Unlike traditional directed graphs, DAG does not allow loops. DAG can describe nonlinear structures, that is, the calculations between different nodes in DAG can be performed in parallel, so DAG can be applied to the field of parallel computing to improve computing efficiency.

[0061] In the specific implementation process, Figures 2 to 4 As shown, the visual human-computer interaction interface may include a drawing area A1 and a data management node selection area A2. The data management node selection area A2 is pre-set with a plurality of node icon elements E for user selection, and each node icon element E corresponds to a data management node. Figure 2 As shown, the user can select the data management node corresponding to the dragged target node icon element E on the visual human-computer interaction interface by dragging the target node icon element E in the data management node selection area A2 to the drag command in the drawing area A1. Figure 3 As shown, the user can set the execution order between the data governance nodes corresponding to each target icon element by connecting the connection lines of each target node icon element in the drawing area A1. Figure 4 As shown, the user can set the governance logic of the corresponding data governance node by selecting instructions in the setting options of the target node icon element E.

[0062] S130. Generate the target data governance task according to the data governance DAG.

[0063] Since each DAG node is represented by a corresponding target node icon element E, and each target node icon element E corresponds to a data governance node, and the directed edges between DAG nodes represent the execution order relationship of the corresponding data governance nodes, and in the data governance DAG diagram in the embodiment of the present invention, each node is provided with attributes such as identification-id, node name-name, node type-type, storage governance logic-data, and the connection relationship is provided with the previous node-from and next node-to attributes of the connection. Therefore, the execution order between each data governance node can be determined through the data governance DAG, and the codes corresponding to each data governance node can be combined according to the execution order to generate the code of the corresponding target data governance task.

[0064] Optionally, after step S130, the method further includes (not shown in the figure):

[0065] Submit the target data governance task to the reviewing user; in response to the passing instruction triggered by the reviewing user, save the target data governance task; or in response to the non-passing instruction triggered by the reviewing user, generate non-passing review information. Thereby, the management of the data governance task can be realized, and the stability and compliance of the data governance task can be ensured.

[0066] After completing the design process of the data governance method, the execution process of the data governance method can be carried out below to use the code of the target data governance task to implement the target data governance task. As Figure 5 shown, the execution process includes:

[0067] S210. In response to the start instruction for the target data governance task, enable the configuration permissions of each to-be-configured data governance node of the target data governance task.

[0068] In the specific implementation process, before executing the step S210, it should be determined that the target data governance task is the target data governance task saved in response to the passing instruction triggered by the reviewing user.

[0069] In the specific implementation process, the to-be-configured data governance node is at least part of the data governance nodes in the target data governance task.

[0070] S220. For any to-be-configured data governance node of the target governance task, configure the data governance node according to the configuration instruction for the data governance node.

[0071] S230. After determining that each to-be-configured data governance node in the target data governance task is respectively configured, sequentially execute the operations corresponding to each data governance node of the target data governance task to implement data governance.

[0072] In this way, the data governance method provided by the embodiment of the present invention generates a data governance task composed of a plurality of data governance nodes combined in a certain order, and at least part of the data governance nodes can be configured and adjusted according to the user's configuration instruction. Thereby, the data governance requirements for some to-be-governed data with different but the same data governance logic can be realized using the same data governance task, improving the reusability of the data governance task, and thus improving the development efficiency of data governance. In addition, in the data governance method provided by the embodiment of the present invention, the data governance task is generated by the user drawing a DAG on the visual human-computer interaction interface, so that users who do not need to design the data governance logic do not need to master program programming development knowledge, reducing the design difficulty of the data governance logic.

[0073] Furthermore, as Figure 6As shown, the target data governance task at least includes a data collection node J1 and a first processing node J3.

[0074] In step S230, when sequentially performing operations corresponding to each data governance node of the target data governance task to implement the data governance process:

[0075] (1) The data collection node J1 is used for: during the data governance process, configuring a target physical table containing the data to be governed and the original database source of the target physical table.

[0076] In the embodiment of the present invention, a physical table refers to a table that actually stores data in a database, that is, a table stored on a disk. The physical table contains actual data records, rather than a logical view or query result.

[0077] In the specific implementation process, the original database is used to store the data to be governed that has not been processed in the form of a physical table. These data to be governed can be collected from the corresponding third-party data platform in a manner compliant with relevant laws and regulations. For example, if the data to be governed includes commercial survey data collected by a third-party business consulting agency in a manner compliant with relevant laws and regulations, then the commercial survey data in the form of a physical table can be obtained from the data platform of the third-party business consulting agency in a manner compliant with relevant laws and regulations; another example is that if the data to be governed includes public management data collected by a public management department in a manner compliant with relevant laws and regulations, then the public management data in the form of a physical table can be obtained from the data platform of the public management department in a manner compliant with relevant laws and regulations. In the embodiment of the present invention, the original database needs to keep the data in its original state without any modification and serves as a backup of the data. Further, in order to improve the subsequent acquisition efficiency of the data to be governed in the original database, partitioned physical tables can be created for the original database according to certain rules to avoid full-scale scanning operations on the physical table in the future.

[0078] (2) The first data processing node J3 is used for: obtaining the target physical table from the original database source, and according to the mapping relationship between the first virtual table configured for the first data processing node and the target physical table (such as Figure 7As shown, map the data to be governed in the target physical table to the first virtual table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations in the form of a processed physical table to the first governance database. The specific objective of the first data processing node J3 is to govern the data to be governed from the target physical table, so as to achieve data governance operations such as changing the field names of the data to be governed in the original database to standardized names, filtering out data that does not meet expectations such as dirty data, non-compliant data, and garbage data, ensuring the uniqueness of the ID numbers in the data to be governed, and ensuring that the date of birth of the same person is less than their date of death.

[0079] Among them, the first governance database is used to store the data processed by the first data processing node, and the data stored in the first governance database in the form of a processed physical table needs to be separated from the data structure of the original database, so that subsequent data governance nodes can further associate and aggregate the data in the first governance database.

[0080] In the embodiments of the present invention, a virtual table refers to a normative constraint that ensures the consistency and accuracy of the internal and external use and exchange of data. Generally, data standards are analyzed and disassembled from the business, technical, and management dimensions: business standard specifications generally include the definition of the business, the name of the standard, the classification of the standard, etc.; technical standard specifications view data standards from a technical perspective and include the type, length, format, encoding rules, etc. of the data; management standard specifications view data standards from a management perspective. For example, who is the manager of the data standard, how to add, how to delete, and the access standard conditions, etc. are all data specification requirements.

[0081] Specifically, the first data processing node J3 may be composed of an input component C31, at least one first data processing component C32, and an output component C33.

[0082] Among them, when the input component C31 responds to the configuration instruction corresponding to the first data processing node J3, it is encapsulated and generated according to the target physical table configured by the data collection node, the information of the original database source, and the mapping relationship between the first virtual table and the target physical table configured for the first data processing node J3, and is used to obtain the target physical table configured by the data collection node and the original database source, and obtain the target physical table from the original database source through Java Database Connectivity (JDBC), and according to the mapping relationship between the first virtual table and the target physical table, map the data to be governed of the target physical table to the first virtual table. Among them, the input component C31 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data data, etc. The component data data attribute further includes the mapping relationship between the first virtual table and the target physical table stored in the input component C31 in the form of a JSON object. Each JSON object is composed of attributes such as the target physical table field name - standardColumnName, the target physical table field type - standardColumnType, the first virtual table field name - jrColumnName, and the first virtual table field type - jrColumnType. The mapping relationship between the first virtual table and the target physical table can be obtained by parsing the JSON object.

[0083] The first data processing component C32 is a component that has been pre - encapsulated before generating the data governance DAG and is set in the first data processing node according to the user's editing instruction in the step S120. For any one of the first data processing components C32, the first data processing component C32 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data data, etc. The component data data attribute further includes the first virtual table field name - columnName, the first virtual table field type - columnType, the first processing rule - notnull, the first processing rule - regex attribute, etc. For any one of the first data processing components C32, the data exploration component C32 is used to perform corresponding data processing operations on the data to be governed in the first virtual table.

[0084] Specifically, the first data processing component C32 may specifically include at least one of the following:

[0085] ① A normalization component, which is used to limit the value range, length, maximum value, and minimum value of each field in the first virtual table.

[0086] ②Sequential sorting component, used to sort the field values in the first virtual table.

[0087] ③Data volume statistics component, used to perform early warning statistics on the data volume of the field values in the first virtual table.

[0088] ④Cascading verification component, used to perform joint verification on multiple field values in the first virtual table. For example, perform joint verification on the field values of province, city, and district.

[0089] ⑤Null value replacement component, used to handle null values in the fields of the first virtual table. For example, replace null values with default values.

[0090] ⑥Date format conversion component, used to convert the date format of the date fields in the first virtual table. For example, convert numerical dates to character dates.

[0091] ⑦ID card standardization component, used to standardize the ID cards in the first virtual table. For example, convert 16-digit ID card numbers to 18-digit ID card numbers.

[0092] ⑧Telephone number standardization component, used to standardize the telephone numbers in the first virtual table. For example, add area codes to all telephone numbers.

[0093] ⑨Regular expression extraction component, used to process the field values in the first virtual table using regular expressions.

[0094] ⑩String conversion component, used to process the strings in the first virtual table.

[0095] Special character cleaning component, used to process special characters such as empty lines.

[0096] Custom mapping component, used to perform mapping processing on the dictionary values in the first virtual table. For example, perform binary mapping on the gender field according to male→0, female→1.

[0097] Column addition component, used to add new columns to the first virtual table.

[0098] Data deduplication component, used to remove duplicate data in the first virtual table.

[0099] Groovy script component, used to support Groovy scripts to process the field values in the first virtual table.

[0100] Encryption component, used to encrypt the data to be governed in the first virtual table.

[0101] A decryption component for decrypting the data to be governed in the encrypted state in the first virtual table. A data desensitization component for desensitizing the data to be governed in the first virtual table.

[0102] A filtering component for filtering the field values in the first virtual table.

[0103] The above-mentioned first data processing component C32 performs logical processing based on the field information of the first virtual table. For example, the filtering component is a JSON object composed of attributes such as id, name, data, and type. Among them, the data object contains columns and arr. The columns consist of a JSON array of the field information of the first virtual table. Arr is a JSON array composed of filtering objects. The filtering object is mainly composed of three attributes: name - the name of the virtual table field, filterValue - the filtering value, and filterCondition - the filtering condition, which is used to determine that the corresponding data is problematic data when the field value is equal to the filtering value.

[0104] When the output component C33 responds to the configuration instruction corresponding to the first data processing node J3, it is generated by encapsulation according to the mapping relationship between the first virtual table and the processing physical table configured by the first data processing node J3, and is used to output the data to be governed after all data processing operations are completed by recursively calling the first data processing component C32 of the first data processing node to the first governance database in the form of a processing physical table. Among them, the output component C33 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, and component data - data. Among them, the component data - data attribute further includes the mapping relationship between the first virtual table and the processing physical table stored in the output component C33 in the form of a JSON object. The JSON object is composed of attributes such as the name of the processing physical table field - standardColumnName, the type of the processing physical table field - standardColumnType, the name of the first virtual table field - jrColumnName, and the type of the first virtual table field - jrColumnType. The mapping relationship between the first virtual table and the processing physical table can be obtained by parsing the JSON object.

[0105] Optionally, the data governance node further includes at least one of the following ( Figure 8 indicating a data governance task including all the following data governance nodes):

[0106] (3) The data exploration node J2 after the data collection node J1 is used to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data exploration node and the target physical table, and perform at least one format check on the data to be governed in the second virtual table.

[0107] Specifically, the data exploration node J2 may be composed of an input component C21, at least one data exploration component C22, and an output component C23.

[0108] Among them, when the input component C21 responds to the configuration instruction corresponding to the data exploration node J2, it is encapsulated according to the target physical table configured by the data collection node J1, the information of the original database source, and the mapping relationship between the second virtual table configured for the data exploration node and the target physical table, and is used to obtain the target physical table configured by the data collection node and the original database source, and obtain the target physical table from the original database source through JDBC, and map the data to be governed of the target physical table to the second virtual table according to the mapping relationship between the second virtual table and the target physical table. Among them, the input component C21 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, and component data data. The component data data attribute further includes the mapping relationship between the second virtual table and the target physical table stored in the form of a JSON object in the input component C21. Each JSON object is composed of attributes such as the target physical table field name - standardColumnName, the target physical table field type - standardColumnType, the second virtual table field name - jrColumnName, and the second virtual table field type - jrColumnType. The mapping relationship between the second virtual table and the target physical table can be obtained by parsing the JSON object.

[0109] The data exploration component C22 is a pre - encapsulated component before generating the data governance DAG, and is set in the data exploration node according to the user's editing instruction in the step S120. For any of the data exploration components C22, the data exploration component C22 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data data, etc. The component data data attribute further includes attributes such as the second virtual table field name - columnName, the second virtual table field type - columnType, the verification rule - notnull, the verification rule - regex attribute, etc. For any of the data exploration components C22, the data exploration component C22 is used to perform corresponding format verification on the data to be governed in the second virtual table.

[0110] Specifically, the data exploration component C22 may specifically include at least one of the following:

[0111] ① Field length verification component, used to analyze whether the field length of the data to be governed in the second virtual table meets the preset requirements.

[0112] ② Null value verification component, used to analyze whether there are null values in the data to be governed in the second virtual table.

[0113] ③ Uniqueness verification component, used to analyze whether the data with uniqueness requirements (such as primary keys, candidate keys, etc.) in the data to be governed in the second virtual table is globally unique.

[0114] ④ Enumeration value verification component, used to analyze whether the enumeration values in the data to be governed in the second virtual table belong to the corresponding enumeration range.

[0115] The output component C23 is used to output the data to be governed after all format verification operations are completed by recursively exploring the data exploration component C22 of the data exploration node, and output the corresponding format verification result. The format verification result can be output to a preset location for the user to view.

[0116] (4) The data fusion node J4 after the data first - processing node J3 is used to respectively determine the corresponding processing physical tables according to the multiple data first - processing nodes connected to the data fusion node, obtain the multiple processing physical tables from the first governance database, map the data to be governed in the multiple processing physical tables to the third virtual table according to the mapping relationship between the third virtual table configured for the data fusion node and the multiple processing physical tables, and output the fused physical table to the second governance database according to the third virtual table.

[0117] Among them, the second governance database is used to store the data processed by the data fusion node J4, and the data stored in the second governance database in the form of a fused physical table needs to be separated from the data structure of the original database. The second governance database can be a population information governance database, an industrial and commercial information governance database, a natural resources information governance database, a macroeconomic information governance database, etc. In this way, the data stored in the second governance database can provide powerful decision-making assistance for relevant demand parties.

[0118] Specifically, the data fusion node J4 can be composed of an input component C41, a Join component C42, and an output component C43.

[0119] Among them, when the input component C41 executes in the step S220, according to the processing physical tables and the information of the first governance database respectively configured by each of the connected data first processing nodes J3, it performs encapsulation to generate, for obtaining the processing physical tables and the first governance database configured by the data first processing node J3, and obtains the corresponding processing physical tables from each of the first governance databases through JDBC.

[0120] The Join component C42 is generated by encapsulation in response to the configuration instruction corresponding to the data fusion node J4, according to the mapping relationship between the third virtual table configured by the data fusion node J4 and the multiple processing physical tables, for mapping the data to be governed of each of the processing physical tables to the third virtual table according to the mapping relationship between the third virtual table and the multiple processing physical tables. Among them, the Join component C42 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data - data, etc., where the component data - data attribute further includes the mapping relationship between the third virtual table and the multiple processing physical tables stored in the input component C42 in the form of a JSON object, and the mapping relationship between the third virtual table and the multiple processing physical tables can be obtained by parsing the JSON object.

[0121] The output component C43 is generated by encapsulation in response to the configuration instruction corresponding to the data fusion node J4 according to the mapping relationship between the third virtual table configured for the data fusion node J4 and the output fusion physical table. Among them, the output component C43 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data data, etc. The component data data attribute further includes the mapping relationship between the third virtual table and the fusion physical table stored in the output component C43 in the form of a JSON object. The JSON object consists of attributes such as the name of the fusion physical table field - standardColumnName, the type of the fusion physical table field - standardColumnType, the name of the third virtual table field - jrColumnName, and the type of the third virtual table field - jrColumnType. The mapping relationship between the third virtual table and the fusion physical table can be obtained by parsing the JSON object.

[0122] (5) The data verification node J5 after the data fusion node J4 is used to determine the corresponding fusion physical table according to the data fusion node, obtain the fusion physical table from the second governance database, map the data to be governed in the fusion physical table to the fourth virtual table, and perform at least one format verification on the data to be governed in the fourth virtual table.

[0123] Specifically, the data verification node J5 may be composed of an input component C51, at least one data verification component C52, and an output component C53.

[0124] Among them, when the input component C51 is executed in step S220, it is generated by encapsulation according to the fusion physical table configured by the connected data fusion node J1 and the information of the second governance database, and is used to obtain the fusion physical table configured by the data fusion node J4 and the source of the second governance database, and obtain the fusion physical table from the source of the second governance database through JDBC, and map the data to be governed in the fusion physical table to the fourth virtual table. Among them, the input component C51 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, component data data, etc.

[0125] The data verification component C52 is a pre-encapsulated component before generating the data governance DAG, and is set in the data exploration node according to the user's editing instruction in the step S120. For any of the data verification components C52, the data verification component C52 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, and component data data. The component data data attribute further includes attributes such as the name of the fourth virtual table field - columnName, the type of the fourth virtual table field - columnType, verification rule - notnull, and verification rule - regex. For any of the data verification components C52, the data verification component C52 is used to perform corresponding format verification on the data to be governed in the fourth virtual table.

[0126] Specifically, the data verification component C52 may specifically include at least one of the following:

[0127] ① Data integrity verification component, used to analyze whether there is at least one of the four situations of entity missing, attribute missing, record missing, and field value missing in the data to be governed in the fourth virtual table.

[0128] ② Data uniqueness verification component, used to analyze whether the data with uniqueness requirements (such as primary key, candidate key, etc.) in the data to be governed in the fourth virtual table is globally unique

[0129] ③ Data consistency verification component, used to analyze whether the data source, data storage location, and data caliber of the data to be governed in the fourth virtual table are consistent.

[0130] ④ Data accuracy verification component, used to verify whether the measurement error and measurement unit of the data to be governed in the fourth virtual table meet the preset requirements.

[0131] ⑤ Data legality verification component, used to verify whether the format, type, domain value, and business rules of the data to be governed in the fourth virtual table are legal.

[0132] ⑥ Data timeliness verification component, used to verify whether the generation time of the data to be governed in the fourth virtual table meets the preset time requirements.

[0133] The output component C53 is used to output the data to be governed after recursively completing all format verification operations by the data verification component C52 of the data exploration node, and output the corresponding format verification result. The format verification result can be output to a preset location for the user to view.

[0134] The second data processing node J6 after the data verification node J5 is used to determine the corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed in the fused physical table to the fifth virtual table, and perform at least one data processing operation on the data to be governed in the fifth virtual table and output the result.

[0135] Specifically, the second data processing node J6 may be composed of an input component C61, at least one second data processing component C62, and an output component C63.

[0136] Among them, when the input component C61 executes in step S220, it is encapsulated according to the fused physical table configured by the previous data fusion node J4 and the information of the second governance database, and is used to obtain the fused physical table configured by the data fusion node J4 and the source of the second governance database, and obtain the fused physical table from the source of the second governance database through JDBC, and map the data to be governed in the fused physical table to the fifth virtual table. Among them, the input component C61 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, and component data data.

[0137] The second data processing component C62 is a component that has been pre-encapsulated before generating the data governance DAG and is set in the second data processing node J6 according to the user's editing instruction in step S120. For any second data processing component C62, the second data processing component C62 is a JSON object, including attributes such as component identifier - id, component type - type, component name - name, and component data data. Among them, the component data data attribute further includes the fifth virtual table field name - columnName, the fifth virtual table field type - columnType, and the data processing script - standardSql. For any of the second data processing components C62, the second data processing component C62 is used to perform corresponding data processing on the data to be governed in the fifth virtual table. For example, calculating the average value, variance value, analyzing the fitting relationship, etc.

[0138] The output component C63 is used to output the corresponding format verification result for the data to be governed after all data processing operations are completed by recursively calling the second data processing component C62 of the second data processing node. The format verification result can be output to a preset location for the user to view.

[0139] As described above, the data collection node J1, data exploration node J2, first data processing node J3, and data fusion node J4 in the data governance node need to be configured in step S220 (for example, for the data collection node, configure the target physical table and the original database source, and for the data governance nodes other than the data collection node J1, configure the mapping relationship between the virtual table and the physical table). Optionally, in order to achieve permission isolation, as Figure 9 shown, the configuration instructions for different data governance nodes are generated by different users. For example, the 7 data governance nodes including the data collection node J1, data exploration node J2, first data processing node J3, and data fusion node J4 are respectively configured by Figure 9 the 7 different users U1 to U7 shown in, and the user corresponding to a certain data governance node has no right to configure other data governance nodes. In addition, the user who triggers the edit instruction is different from the user who configures the mapping relationship for the data governance node (for example, the user U8 who triggers the edit instruction is not Figure 9 any one of the users U1 to U7 shown in), and the user corresponding to a certain data governance node has no right to modify the data DAG, that is, the user corresponding to a certain data governance node has no right to modify the data governance task.

[0140] Based on the same inventive concept, an embodiment of the present invention further provides a data governance device, as Figure 10 shown, including:

[0141] A visual editing module M1, configured to generate a data governance DAG in response to an edit instruction triggered on a visual human-computer interaction interface; wherein, the data governance DAG includes a plurality of data governance nodes, and the connection relationship between the data governance nodes represents the execution order of the data governance nodes in the target data governance task;

[0142] A task generation module M2, configured to generate the target data governance task according to the data governance DAG;

[0143] An execution module M3, configured to, in response to a start instruction for the target data governance task, after determining that each to-be-configured data governance node in the target data governance task is configured, sequentially execute the operations corresponding to the data governance nodes of the target data governance task to implement data governance.

[0144] Furthermore, the target data governance task at least includes a data collection node and a first processing node.

[0145] In the process of sequentially executing the operations corresponding to the data governance nodes of the target data governance task to implement data governance:

[0146] The data collection node is used to configure a target physical table containing data to be governed and the original database source of the target physical table;

[0147] The first data processing node is used to obtain the target physical table from the original database source, map the data to be governed in the target physical table to the first virtual table according to the mapping relationship between the first virtual table configured for the first data processing node and the target physical table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations in the form of a processed physical table to the first governance database.

[0148] Further optionally, the data governance node further includes at least one of the following:

[0149] A data profiling node after the data collection node, which is used to obtain the target physical table from the original database source, map the data to be governed in the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data profiling node and the target physical table, and perform at least one format check on the data to be governed in the second virtual table;

[0150] A data fusion node after the first data processing node, which is used to respectively determine corresponding processed physical tables according to multiple first data processing nodes connected to the data fusion node, obtain the multiple processed physical tables from the first governance database, map the data to be governed in the multiple processed physical tables to the third virtual table according to the mapping relationship between the third virtual table configured for the data fusion node and the multiple processed physical tables, and output a fused physical table to the second governance database according to the third virtual table;

[0151] A data verification node after the data fusion node, which is used to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed in the fused physical table to the fourth virtual table, and perform at least one format check on the data to be governed in the fourth virtual table;

[0152] A second data processing node after the data verification node, which is used to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed in the fused physical table to the fifth virtual table, perform at least one data processing operation on the data to be governed in the fifth virtual table and output the result.

[0153] Optionally, obtaining a physical table from a database includes:

[0154] Obtaining a physical table from a database through JDBC;

[0155] Among them, if the physical table is the target physical table, the database is the original database;

[0156] If the physical table is the processed physical table, the database is the first governance database;

[0157] If the physical table is the integrated physical table, the database is the second governance database.

[0158] Optionally, the configuration instructions of different data governance nodes are generated by different users;

[0159] The user who triggers the editing instruction is different from the user who configures the mapping relationship for the data governance node.

[0160] In several embodiments provided in the present application, it should be understood that the above-described embodiments of the data governance device are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in electrical, mechanical or other forms. The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0161] Since the principle of the data governance device for solving problems is basically the same as that of the data governance method, the implementation of the data governance device can refer to the implementation of the data governance method, which will not be elaborated here.

[0162] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, as Figure 11 shown, including: a processor 110 and a memory 120 for storing executable instructions of the processor 110;

[0163] Among them, the processor 110 is configured to execute the instructions to implement the data governance method.

[0164] In a specific implementation process, the device may vary greatly due to configuration or performance differences, and may include one or more processors 110, a memory 120, and a computer-readable storage medium 130. One or more application programs 131 or data 132 are included in the memory 120 and / or the computer-readable storage medium 130. One or more operating systems 133 may also be included in the memory 120 and / or the computer-readable storage medium 130, such as Windows, Mac OS, Linux, IOS, Android, Unix, FreeBSD, etc. Among them, the memory 120 and the computer-readable storage medium 130 may be transient storage or persistent storage. The application program 131 may include one or more of the above-mentioned modules ( Figure 11 not shown in the figure), and each module may include a series of instruction operations. Further, the processor 110 may be set to communicate with the computer-readable storage medium 130 and execute a series of instruction operations in the computer-readable storage medium 130 on the device. The device may also include one or more power supplies ( Figure 11 not shown in the figure); one or more network interfaces 140, where the network interface 140 includes a wired network interface 141 and / or a wireless network interface 142; and one or more input / output interfaces 143.

[0165] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which, when the computer program code runs on a computer, enables the computer to implement the data governance method.

[0166] In a specific implementation process, as Figure 12 shown, the computer-readable storage medium may adopt a portable compact disc read-only memory (CD-ROM). However, the embodiment of the present invention is not limited thereto. In the embodiment of the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a computer program.

[0167] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable computer program code is carried. Such a propagated data signal may take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the above. The computer program included on the readable storage medium may be transmitted using any appropriate medium, including - but not limited to - wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0168] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product, which includes: computer program code that, when running on a computer, enables the computer to implement the data governance method described above.

[0169] In a specific implementation process, the computer program code of the computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0170] In a specific implementation process, a computer program for executing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language, assembly language, or similar programming languages.

[0171] The data governance method, device, equipment, and storage medium provided by the embodiments of the present invention generate a data governance task composed of multiple data governance nodes combined in a certain order, and at least some of the data governance nodes can be configured and adjusted according to the configuration instructions of the user. Therefore, the data governance requirements for some to-be-governed data that are different but have the same data governance logic can be realized using the same data governance task, improving the reusability of the data governance task and thus the development efficiency of data governance. In addition, in the data governance method provided by the embodiments of the present invention, the data governance task is generated by a DAG drawn by the user on a visual human-computer interaction interface, so that users who do not need to design data governance logic do not need to master program programming development knowledge, reducing the design difficulty of data governance logic.

[0172] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0173] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks. Figure 1 one process or a plurality of processes and / or Figure 1 one block or a plurality of blocks.

[0176] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to cover these changes and modifications.

Claims

1. A data governance method, characterized in that, Including: Generating a data governance directed acyclic graph (DAG) in response to an editing instruction triggered on a visual human-computer interaction interface; wherein, the data governance DAG includes multiple data governance nodes, and the connection relationship between the data governance nodes represents the execution order of each data governance node in the target data governance task; Generating the target data governance task according to the data governance DAG; In response to a start instruction for the target data governance task, after determining that each to-be-configured data governance node in the target data governance task is configured, sequentially performing the operations corresponding to each data governance node of the target data governance task to implement data governance.

2. The method according to claim 1, characterized in that, The target data governance task at least includes a data collection node and a first processing node. During the process of sequentially performing the operations corresponding to each data governance node of the target data governance task to implement data governance: The data collection node is used to configure a target physical table containing data to be governed and the original database source of the target physical table; The first data processing node is used to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the first virtual table according to the mapping relationship between the first virtual table configured for the first data processing node and the target physical table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations in the form of a processed physical table to a first governance database.

3. The method according to claim 2, characterized in that, The data governance node further includes at least one of the following: A data profiling node after the data collection node, which is used to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data profiling node and the target physical table, and perform at least one format check on the data to be governed in the second virtual table; A data fusion node after the first data processing node, which is used to respectively determine the corresponding processed physical tables according to the multiple first data processing nodes connected to the data fusion node, obtain the multiple processed physical tables from the first governance database, map the data to be governed of the multiple processed physical tables to the third virtual table according to the mapping relationship between the third virtual table configured for the data fusion node and the multiple processed physical tables, and output a fused physical table to a second governance database according to the third virtual table; A data verification node after the data fusion node, which is used to determine the corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to the fourth virtual table, and perform at least one format check on the data to be governed in the fourth virtual table; The second data processing node after the data verification node is used to determine the corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to the fifth virtual table, and perform at least one data processing operation on the data to be governed of the fifth virtual table and output the result.

4. The method according to claim 2 or 3, characterized in that, Obtaining the physical table from the database includes: Obtaining the physical table from the database through JDBC; Wherein, if the physical table is the target physical table, the database is the original database; If the physical table is the processed physical table, the database is the first governance database; If the physical table is the fused physical table, the database is the second governance database.

5. The method according to claim 2 or 3, characterized in that, The configuration instructions of different data governance nodes are generated by different users; The user who triggers the edit instruction is different from the user who configures the mapping relationship for the data governance node.

6. A data governance device, characterized in that, Including: A visual editing module for generating a data governance DAG in response to an edit instruction triggered on a visual human-computer interaction interface; wherein, the data governance DAG includes multiple data governance nodes, and the connection relationship between the data governance nodes represents the execution order of the data governance nodes in the target data governance task; A task generation module for generating the target data governance task according to the data governance DAG; An execution module for, in response to a start instruction for the target data governance task, after determining that each to-be-configured data governance node in the target data governance task is configured, sequentially executing the operations corresponding to each data governance node of the target data governance task to implement data governance.

7. The device according to claim 6, characterized in that, The target data governance task at least includes a data collection node and a first processing node. In the process of sequentially executing the operations corresponding to each data governance node of the target data governance task to implement data governance: The data collection node is used to configure the target physical table containing the data to be governed and the original database source of the target physical table; The data first processing node is used to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the first virtual table according to the mapping relationship between the first virtual table configured for the data first processing node and the target physical table, perform at least one data processing operation on the data to be governed in the first virtual table, and output the data to be governed that has completed all data processing operations in the form of a processed physical table to the first governance database.

8. The device according to claim 7, characterized in that The data governance node further includes at least one of the following: The data exploration node after the data collection node is used to obtain the target physical table from the original database source, map the data to be governed of the target physical table to the second virtual table according to the mapping relationship between the second virtual table configured for the data exploration node and the target physical table, and perform at least one format check on the data to be governed in the second virtual table; A data fusion node after the first data processing node is configured to respectively determine corresponding processed physical tables according to a plurality of first data processing nodes connected to the data fusion node, obtain the plurality of processed physical tables from the first governance database, map the data to be governed of the plurality of processed physical tables to a third virtual table according to a mapping relationship between a third virtual table configured for the data fusion node and the plurality of processed physical tables, and output a fused physical table to a second governance database according to the third virtual table; A data verification node after the data fusion node is configured to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to a fourth virtual table, and perform at least one format verification on the data to be governed of the fourth virtual table; A second data processing node after the data verification node is configured to determine a corresponding fused physical table according to the data fusion node, obtain the fused physical table from the second governance database, map the data to be governed of the fused physical table to a fifth virtual table, and perform at least one data processing operation on the data to be governed of the fifth virtual table and output a result.

9. A device, characterized in that, It includes a processor and a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the data governance method according to any one of claims 1-5.

10. A storage medium, characterized in that, The storage medium stores program code, and when the program code runs on a computer, the computer is caused to execute the data governance method according to any one of claims 1-5.