A Multi-Granularity Traceability Method for Data Integration
By constructing a multi-grained traceability model DI_PROV and traceability toolbox ProvToolBox, the problem that the existing data integration traceability method cannot meet the multi-source heterogeneous data integration task is solved, and the efficient interpretability, credibility and repeatability of the data integration process are achieved.
Patent Information
- Application Number
- CN202211545898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-05
AI Technical Summary
The existing data integration traceability method cannot effectively meet the needs of multi-source heterogeneous data integration tasks, and cannot play back the data integration process from multiple granularity, resulting in insufficient interpretability, credibility and repeatability of data integration.
Build a multi-grained traceability model DI_PROV for data integration, including node division and relationship division of traceability model, combining the data integration traceability toolbox ProvToolBox, data integration workflow model DI_Workflow and traceability meta information model ProvMeta_in, build a multi-grained traceability map and design a traceability query method to support the playback of data integration from the active level and the entity level.
It improves the interpretability, credibility and reusability of data integration, and realizes efficient traceability in the graph database Neo4j through multi-grained traceability query method, and supports multi-grained traceability query.
Smart Images

Figure BDA0003979791600000071 
Figure BDA0003979791600000101 
Figure BDA0003979791600000111
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data integration and data traceability, and particularly relates to a multi-granularity traceability method for data integration. Background Art
[0002] In the era of data empowerment, data integration, as a core technology in data science applications, has played an important role in enterprise decision-making. However, with the enrichment of data resources, the characteristics of multi-source data such as complexity, heterogeneity, and low quality are more prominent, which brings new challenges to data science applications based on data integration technology.
[0003] Data integration aims to fuse data resources in multi-source heterogeneous data sets. In terms of the data integration process, the data integration process is complex and usually covers multiple activities such as data discovery, schema alignment, entity resolution, and data repair, and it is impossible to guarantee the integration quality of multi-source data. In terms of data integration technology, the existing data integration technologies based on rules, similarity matching, etc. can no longer meet the needs of enterprises well, and the current data integration technology based on deep learning has poor solvability and it is difficult to guarantee the credibility of the integration result.
[0004] Traceability is a form of structured metadata used to record the source of information and helps to judge whether the information is credible. Data traceability technology focuses on analyzing and simulating the evolution process of data, and its main purpose is to perform real-time tracking and updating on the development process of the original data, and it can be used in business fields such as data quality assessment, data verification, and data recovery. Traceability can provide the following benefits: 1) It explains to users the contributions of each source and processing step to the final result; 2) It tracks the data processing process, enriches the data with information on how the data is obtained, and improves the credibility of the result; 3) Ideally, it can reconstruct the data integration process according to the traceability information embedded in the final result.
[0005] Due to the relevance of multi-source heterogeneous data and the complexity of the integration process, users have higher requirements for the interpretability, credibility, and reproducibility of the data integration process and the integration results. Interpretability emphasizes explaining the results or processes to make them transparent to certain audiences. The credibility of data refers to the integrity, consistency, and accuracy of the data, which is used to describe that all stored data is objectively true and depends on the data collection, processing, and analysis processes, etc. Data provenance traces the process of data integration, records relevant data operations, and enriches the data with information on how the data was obtained, which can improve the credibility of the integration results. The reproducibility of the results includes starting from the same materials and methods and checking whether the previous results can be confirmed. Since provenance can generally record anything that happened during the processing, it is naturally used for reproducibility. For example, the provenance of a recall can include the intermediate steps included in the process, what changes were made, how the parameters were set, and how they were changed, etc. Data quality includes various dimensions and metrics for evaluating data quality, including integrity, accuracy, timeliness, or credibility, etc. In this case, provenance is usually used to monitor and debug applications to evaluate and improve certain quality dimensions.
[0006] The existing traceability technologies are mainly divided into database query traceability and data science workflow traceability. Database query traceability replays the data evolution process at the data record level (fine-grained), and uses core algebraic data items to reason about the evolution to represent the derivation process of query results. Data science workflow traceability is used to explain the derivation process of a certain data set (coarse-grained) in the workflow. The mainstream workflow-oriented traceability methods take entities, activities, and agents as the core, and are used to capture information about the steps of the workflow and the use of data. The existing database traceability models and workflow traceability models can only replay the data evolution process from a single granularity, and cannot meet the traceability requirements of data integration tasks, because data traceability for data integration tasks is much more complex than workflow traceability. It is necessary to track the processing processes of multi-granularity objects such as data sets, data records, and data attributes, and be able to query traceability data in a simple and efficient manner. Early data traceability for data integration mainly tracked the traceability information during the database integration process. Trio (Jennifer Widom. Trio: A System for Integrated Management of Data, Accuracy, and Lineage[C] / / CIDR, 2005:262-276.) managed the correctness and lineage of data as components of the data together with the data. Perm (B. Glavic and G. Alonso. Perm: Processing Provenance and Data on the Same Data Model through Query Rewriting[C] / / ICDE, Shanghai, China, 29 March - 2 April 2009. Washington, DC, USA: IEEE Computer Society, 2009:174-185.) used query rewriting in relational databases to capture traceability information. Entity Resolution (ER) is a key task in data integration. Erprov (Oppold S., Herschel M. Provenance for Entity Resolution.[C] / / Provenance and Annotation of Data and Processes. IPAW 2018. Lecture Notes in Computer Science, Springer, Cham. 2018:226-230.) describes a traceability model for how to process data during entity resolution. First, the ER task is abstracted into algebraic operation operations, and then the traceability is defined on this abstract representation.Blast (G. Simonini, L. Gagliardelli, S. Zhu, et al. Enhancing Loosely Schema-aware Entity Resolution with User Interaction [C] / / 2018 International Conference on High Performance Computing & Simulation (HPCS). Orleans, France: IEEE, 2018: 860-864.) The extended entity resolution algorithm captures provenance information and visualizes the process of entity resolution results through provenance information. PROVDB (Hui M, Deshpande A. ProvDB: Provenance-enabled Lifecycle Management of Collaborative Data Analysis Workflows. [J]. ACM, 2018.) is a unified provenance and metadata management system that supports the lifecycle management of complex collaborative data science workflows. Provision (Sven Hertling, Heiko Paulheim. Provision and Usage of Provenance Data in the WebIsALOD Knowledge Graph [C] / / ISWC 2018.) supports provenance tracking for ETL and matching calculations, with database-style optimizability and the potential for on-demand computing. Although there are various provenance solutions for multiple applications, existing data provenance solutions only target individual tasks of data integration, do not cover the entire process, and are not fully applicable to data integration tasks. Summary of the Invention
[0007] In view of the deficiencies of existing data provenance methods for data integration, the present invention proposes a multi-granularity provenance method for data integration, including:
[0008] Step 1: Construct a multi-granularity traceability model DI_PROV for data integration tasks to achieve the playback of the data integration process at multiple granularities. At a coarse granularity, the sub-tasks of data integration are used as basic units to playback the data integration process at the table level. At a fine granularity, data entities and attributes are used as the core to playback the evolution process of data during the data integration process. The relationships in the traceability model DI_PROV include: Used, where the starting point is an activity and the ending point is an object; WasGeneratedBy, where the starting point is an object and the ending point is an activity; WasIncludedBy, where the starting point is an activity and the ending point is an activity, indicating that the parent activity includes the child activity; WasContributedTo, where the starting point is an activity and the ending point is an activity, indicating that the activity contributes to the activity; WasDerivedBy, where the starting point is an object and the ending point is an object, indicating that the object is generated by the object; WasAttributedTo, where the starting point is an object and the ending point is an agent, indicating that the agent is responsible for the object; ActedOnBehalfOf, where the starting point is an agent and the ending point is an agent, indicating the relationship between agents assuming different responsibilities; WasAssociatedWith, where the starting point is an activity and the ending point is an agent, indicating that the agent assumes responsibility in the activity; WasRecordOf, where the starting point is an entity and the ending point is an entity, indicating that the entity is a subset of the entity; HasConflictAttribute, where the starting point is an entity and the ending point is an entity, indicating that the entity has conflicting attributes; HasTruth, where the starting point is an entity and the ending point is an entity, indicating the true value of the entity attribute. The specific construction process includes traceability model node division and traceability model relationship division;
[0009] For the traceability model node division, the nodes in the traceability model are divided into three node types: object, activity, and agent. Objects correspond to data files, models, data entities, and attributes used and generated in data integration. Activities correspond to the sub-tasks of data integration, including the processes of schema matching (SM), entity resolution (ER), data integration (DI), and conflict resolution (CR). Agents correspond to the users responsible for completing data integration tasks, and agents are refined into general agents and sub-agents;
[0010] For the traceability model relationship division, based on inheriting the relationships of objects, activities, and agents in the PROV model, the following relationships are added:
[0011] 1) Data integration includes several sub-tasks. Inclusion relationships and contribution relationships are introduced to reflect the structural characteristics of the workflow;
[0012] 2) Agents are divided into sub-agents and general agents, and there is a representative relationship between general agents and sub-agents;
[0013] 3) A data file contains several data entities, and a sub-record relationship is added between the data file and the data entities;
[0014] 4) When data entities from different data files are integrated, attribute conflicts may occur. The data entities have conflicting attributes and attribute true values, and a has-true-value relationship and a has-conflicting-attribute relationship are added;
[0015] Step 2: Construct a data integration traceability process model, including a data integration traceability toolbox ProvToolBox composed of data integration sub-units DI_Blockbox, a data integration workflow model DI_Workflow, and a traceability meta-information model ProvMeta_in; According to the data integration task, the user constructs a data integration workflow DI_Workflow based on the data integration sub-units DI_Blockbox in the traceability toolbox ProvToolBox, and the data integration workflow execution generates data integration traceability meta-information ProvMeta_in;
[0016] The data integration traceability toolbox ProvToolBox is a collection of data integration sub-units DI_Blockbox. The data integration sub-unit DI_Blockbox is represented as DI_Blockbox(Activity, API, Paramter_type, Output_type), where Activity = {Activity_name: Activity_type}, Activity_name represents the name of the data integration sub-unit, and Activity_type represents the task type corresponding to the data integration sub-unit Activity; API = {api1, api2,... api j} represents the set of functions in this sub-unit, and api j represents the j-th function in the sub-unit; Paramter_type = {pt1, pt2,..., pt j} is the set of specific data types corresponding to the input parameters of the API. The input parameters of the API registered by the user are the data types corresponding to the actual parameters in the data integration workflow model DI_Workflow, and pt j represents the type of the j-th input parameter; Output_type = {ot1, ot2,..., ot m} is the set of data types corresponding to the output of the API, and ot m represents the type of the m-th output data;
[0017] The described data integration workflow model DI_Workflow is represented as DI_Workflow(Workflow_activity, Api_Para), where Workflow_activity = {a1, a2, …, a i} is a set of data integration sub-units of the data integration workflow, and a i corresponds to the data integration sub-unit DI_Blockbox in the data integration traceability toolbox ProvToolBox i ; Api_Para = {api_para1, api_para2, …, api_para j} is a set of functions in the data integration sub-unit DI_Blockbox. api_para j represents that the j-th function inputs a specific parameter value Para and obtains the corresponding output value Output. Here, api represents the processing function in the data integration sub-unit DI_Blockbox, where Para = {p1, p2, …, p j} is the user parameter of the data integration workflow, and Output = {o1, o2, …, o m} is the data generated by the functions in the data integration sub-unit DI_Blockbox;
[0018] The traceability meta-information model ProvMeta_in is represented as ProvMeta_in(Wfprov, Aprov, Sprov, Tprov), which is used to represent the traceability meta-information generated during the data integration process. Here, Wfprov is Workflow_activity in the data integration workflow model DI_Workflow, Aprov = {aprov1, aprov2, …, aprov j} corresponds to the set of functions API in the data integration workflow, Sprov = {sprov1, sprov2, …, sprov j} is the set of source objects for API operations, corresponding to Para in the data integration workflow model DI_Workflow. Each source object is represented as sprov j {sprov j : sprovt j}, and the type value sprovt j of sprov j is calculated by a function based on Paramter_type in the data integration sub-unit DI_Blockbox, and Tprov = {tprov1, tprov2, …, tprov m} is the set of target objects, corresponding to the Output in the data integration workflow. Each target object is represented as tprov m {tprov m :tprovt m}, where the type value tprovt m is calculated by a function based on the Output_type of the data integration subunit DI_Blockbox. The source object sprov in the provenance meta-information j , the target object tprov m and the function aprov j are identified using Identify(id, name, type). The id is used to uniquely identify an object or operation, the name represents the object name or method name, and the type is used to identify the type value of the object;
[0019] Step 3: Construct a multi-granularity provenance graph based on the multi-granularity provenance model DI_PROV and the provenance meta-information ProvMeta_in generated by executing the data integration workflow model DI_Workflow; specifically including:
[0020] Step 3.1: Initialize the multi-granularity provenance graph G(V, E);
[0021] Step 3.2: Construct the provenance relationship between objects and activities based on the provenance model DI-PROV and the provenance meta-information ProvMeta_in, and store it in the provenance relationship list Prov_doc. The nodes in the multi-granularity provenance graph are the set of source objects sprov j , target objects tprov m and functions aprov j in the provenance meta-information ProvMeta_in. The provenance relationship between nodes in the multi-granularity provenance graph is determined by the function aprov j in the provenance meta-information ProvMeta_in and the type of function parameters;
[0022] Step 3.3: Loop through each data record stored in the provenance relationship list Prov_doc, and create coarse-grained and fine-grained provenance graph nodes and the relationships between nodes according to the type of data objects;
[0023] Step 3.4: Create coarse-grained and fine-grained provenance graphs in the graph database Neo4j based on the relationships between nodes, associate the provenance graphs of different granularities and return;
[0024] The multi-granularity traceability graph G(V,E) is a directed graph, where V = {Node} is the node set of the multi-granularity traceability graph, corresponding to the objects, activities, and agents in the traceability model DI_PROV; E = {Edge} is the edge set of the multi-granularity traceability graph, corresponding to the edges in the traceability model DI_PROV, used to represent the relationships between activities, objects, and agents. The node Node is represented as Node(node_id, node_name, node_type), where node_id is used to represent the uniqueness of the node, node_name is the name of the node, and node_type identifies the type of the node; the edge Edge is represented as Edge(edge_id, edge_value, node1_id, node2_id), where edge_id represents the unique ID of the edge, edge_value is the value of the edge, node1_id represents the starting node of the edge, and node2_id represents the ending node of the edge; Coarse-grained and fine-grained traceability graph nodes and the relationships between nodes are created according to the types of data objects; Coarse-grained and fine-grained traceability graphs are created respectively based on the relationships between nodes, and a multi-granularity traceability graph is generated by associating traceability graphs of different granularities.
[0025] Step 4: Store the multi-granularity traceability graph in the Neo4j database, and use the Cypher query language to implement graph traversal queries and aggregation queries, including coarse-grained traceability queries and fine-grained traceability queries;
[0026] The coarse-grained traceability queries include:
[0027] Step A1: Initialize the traceability record;
[0028] Step A2: When the query type is traceability summary, first obtain the source data sets used in the entire data integration process, secondly obtain the main activities of the data integration, and finally traverse the multi-granularity traceability graph based on the source data sets and the main activities to obtain the traceability record and return it;
[0029] Step A3: When the query type is traceability segmentation, according to the user's query conditions, that is, a certain subtask of the data integration that the user wants to query, and then traverse the multi-granularity traceability graph to return the traceability relationships of all activities and objects of the subtask;
[0030] The fine-grained traceability queries include:
[0031] Step B1: Initialize the result set, convert the query result R into an entity e, and check whether conflict resolution has occurred during the data integration process for the entity e;
[0032] Step B2: If the entity e has had attribute conflicts, replay the conflict resolution process;
[0033] Step B3: If there has been no attribute conflict for entity e, first obtain the result data set where entity e is located, then perform a breadth-first traversal of the multi-granularity traceability graph starting from the result data set to obtain traceability records. The termination condition for the breadth-first traversal is the source data set. Finally, query the traceability relationship between the result entity and the entities in the source data set, obtain the final traceability records, and return them.
[0034] The beneficial effects of the present invention are:
[0035] The present invention proposes a multi-granularity traceability method for data integration. Based on the multi-granularity traceability model, a multi-granularity traceability graph is constructed in the graph database Neo4j, and a variety of traceability query methods are designed, which support replaying the data integration process at the activity level and entity level, and can improve the interpretability, credibility, and reusability of data integration. Brief Description of the Drawings
[0036] Figure 1 It is a schematic diagram of a multi-granularity traceability method for data integration in the present invention;
[0037] Figure 2 It is a conceptual diagram of the DI_PROV traceability model in the present invention;
[0038] Figure 3 It is an example of a multi-granularity traceability graph in the present invention;
[0039] Figure 4 It is an example of a summary query of coarse-grained traceability in the present invention;
[0040] Figure 5 It is an example of a segmented query of coarse-grained traceability in the present invention;
[0041] Figure 6 It is an example of a fine-grained traceability query in the present invention. Detailed Embodiments
[0042] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0043] For the traceability requirements of data integration tasks, data traceability users first propose a multi-granularity traceability model that supports replaying the data integration process from multiple granularities. Secondly, a data integration traceability process model is constructed, which mainly includes the data integration traceability toolbox ProvToolBox, the data integration workflow model DI_Workflow, and the traceability meta-information model ProvMeta_in. Data integration users select data integration subunits in the data traceability toolbox ProvToolBox to construct the data integration workflow DI_Workflow to integrate the dataset restaurant datasets Fodors and Zagats, and obtain the final data integration result. Sample data is shown in Tables 1, 2, and 3. Data traceability users execute the mapping relationship between the traceability toolbox and the data integration workflow constructed by data integration users to obtain traceability meta-information and construct it into a multi-granularity traceability graph. Data quality management users do not trust the final result given by data integration users and choose multi-granularity traceability query to query the traceability data in the data integration process to obtain the evolution process of data in the data integration process.
[0044] Table 1 Sample Data of Fodors
[0045] Identity ID Name Addr City Phone Type restraurant10 Maggie Riverside street-32 New York 328799414 snackery restraurant11 Sunset_restaurant Sunset Avenue New York 237681123 french_cuisine
[0046] Table 2 Sample Data of Zagats
[0047] Identity ID Name Addr City Phone Type Restraurant20 Maggie Riverside street-32 New York 328799414 snackery Restraurant21 Sunset_restaurant Sunset Avenue New York 237681123 german_cuisine
[0048] Table 3 Sample Data of Integration Result
[0049] Identity ID Name Addr City Phone Type Restraurant00 Maggie Riverside street-32 New York 328799414 snackery Restraurant01 Sunset_restaurant Sunset Avenue New York 237681123 german_cuisine
[0050] A multi-granularity traceability method for data integration, as Figure 1 shown, specifically includes:
[0051] Step 1: Construct a multi-granularity traceability model DI_PROV for data integration tasks. Establishing a data model is a key technology for data traceability. According to the characteristics of data integration tasks and the requirements of multi-granularity traceability, a traceability model DI_PROV that supports data integration is proposed based on the PROV traceability model, as Figure 2 shown. The DI_PROV traceability model supports replaying the data integration process from multiple granularities. At the coarse granularity, the data integration process is replayed from the table level with the subtasks of data integration (such as schema alignment, entity resolution, etc.) as the basic units. At the fine granularity, the evolution process of data in the data integration process is replayed with data entities and attributes as the core. Table 4 explains the nodes and relationships in the DI_PROV traceability model.
[0052] Table 4 Meanings of Relationships in the DI_PROV Traceability Model
[0053]
[0054] Node division of the traceability model: There are three types of nodes in the traceability model: objects, activities, and agents. Objects correspond to data files, models, data entities, and attributes used and generated in data integration; activities correspond to subtasks of data integration, mainly including processes such as schema alignment, entity resolution, and data fusion; agents correspond to users responsible for completing data integration tasks. Since the data integration process is complex, multiple groups of users are usually required to complete it together, and different tasks are assigned according to the skills and specialties of the users. Agents are refined into general agents and sub-agents;
[0055] Relationship division of the traceability model: On the basis of inheriting the relationships of objects, activities, and agents in the PROV model, the following relationships are added:
[0056] 1) Data integration includes several subtasks. The inclusion relationship (WasIncludedBy) and contribution relationship (WasContributedTo) are introduced to reflect the structural characteristics of the workflow.
[0057] 2) Agents are divided into sub-agents and general agents, and there is a representative relationship (ActedOnBehalfOf) between the general agent and the sub-agent.
[0058] 3) A data file contains several data entities, and the sub-record relationship (WasRecordOf) is added between the data file and the data entity.
[0059] 4) When data entities from different data files are fused, attribute conflicts may occur. There are conflict attributes and attribute truth values for the data entities. The has-truth relationship (HasTruth) and has-conflict-attribute relationship (HasConflictAttribute) are added.
[0060] Data traceability users analyze the characteristics of data integration tasks, divide the nodes and edges of the multi-granularity traceability model to construct the multi-granularity traceability model DI_PROV. The entities in the traditional workflow model PROV are extended to objects, including data files, models, data entities, and attributes used and generated in data integration; the activities are divided into subtasks of data integration, including schema alignment, entity resolution, data fusion, and conflict resolution; the agents are divided into general agents and sub-agents.
[0061] On the basis of inheriting the relationships of the traditional workflow model PROV, the relationships between nodes are divided. The inclusion relationship and contribution relationship are added between activities, the representative relationship is added between the general agent and the sub-agent, the sub-record relationship is added between the data file and the data entity, and the has-truth relationship and has-conflict-attribute relationship are added between the data entity and the attribute.
[0062] Step 2: Construct a data integration traceability process model, including a data integration traceability toolbox ProvToolBox composed of data integration subunits DI_Blockbox, a data integration workflow model DI_Workflow, and a traceability meta-information model ProvMeta_in; according to the data integration task, the user constructs a data integration workflow DI_Workflow based on the data integration subunits DI_Blockbox in the traceability toolbox ProvToolBox, and the data integration workflow execution generates data integration traceability meta-information ProvMeta_in;
[0063] Data integration traceability toolbox ProvToolBox: Typical data integration includes subtasks such as schema matching (SM), entity resolution (ER), data fusion (DF), and conflict resolution (CR). Each subtask is a data integration subunit DI_Blockbox. The set of DI_Blockbox constitutes the data integration traceability toolbox ProvToolBox, that is, ProvToolBox = {DI_Blockbox1, DI_Blockbox2, …, DI_Blockbox i};
[0064] The data integration subunit DI_Blockbox is defined as DI_Blockbox(Activity, API, Paramter_type, Output_type), where Activity = {Activity_name: Activity_type}, Activity_name represents the name of the data integration subunit, and Activity_type represents the task type corresponding to the Activity of the data integration subunit, that is, the subtask in the data integration process (usually including schema matching, entity resolution, entity fusion, etc.). API = {api1, api2, … api j} represents the processing functions in this subunit, Paramter_type = {pt1, pt2, …, pt j} is the set of specific data types corresponding to the input parameters (data sets, entity attributes, etc.) of API. The input parameters registered by the user for the API in the data integration subunit are the data types corresponding to the actual parameters in the data integration workflow model DI_Workflow. Output_type = {ot1, ot2, …, ot m} is the set of data types corresponding to the output of API;
[0065] Data integration workflow model DI_Workflow: Data integration usually covers data integration subtasks such as schema alignment, entity resolution, and conflict resolution. Users can choose to use any number of data integration subunits DI_Blockbox in the data integration provenance toolbox ProvToolBox i to build a data integration workflow;
[0066] Data integration workflow model DI_Workflow. The data integration workflow model is defined as DI_Workflow(Workflow_activity, Api_Para), where Workflow_activity = {a1, a2, …, a n} is the set of data integration subunits of the data integration workflow, and a i corresponds to the data integration subunit DI_Blockbox in the data integration provenance toolbox ProvToolBox i . Api_Para = {api_para1, api_para2, …, api_para j} is the set of functions in the data integration subunit DI_Blockbox. Each api_para i can be expressed as api_para i :{API:Para} → Output. API represents the processing function in the data integration subunit DI_Blockbox, where Para = {p1, p2, …, p j} are the user parameters of the data integration workflow, including data sets, entity attributes, etc., and Output = {o1, o2, …, o m} is the data generated by the functions in the data integration subunit DI_Blockbox.
[0067] Provenance meta-information model ProvMeta_in: Users use the registered data integration subunit DI_Blockbox in the data integration provenance toolbox ProvToolBox to build a data integration workflow DI_Workflow, and save the built data integration workflow and related configurations (configuration information such as the used subunit DI_Blockbox, data dependency set, user parameters, and output) to generate provenance meta-information about each executed step.
[0068] Provenance meta-information ProvMeta_in. The provenance meta-information is defined as ProvMeta_in(Wfprov, Aprov Sprov, Tprov) to represent the provenance meta-information generated by the data integration process, where Wfprov is the Workflow_activity in the data integration workflow, Aprov = {aprov1, aprov2, …, aprov j} corresponds to the set of functions API in the data integration workflow. Sprov = {sprov1, sprov2, …, sprov j} is the set of source objects operated on by the API, corresponding to Para in the data integration workflow DI_Workflow. Each source object is represented as sprov j {sprov j : sprovt j}, and the type value sprovt j of sprov j is calculated by a function sprovt j = ξ(sprov j ) according to the Paramter_type in the data integration subunit DI_Blockbox. Tprov = {tprov1, tprov2, …, tprov m} is the set of target objects, corresponding to Output in the data integration workflow. Each target object is represented as tprov m {tprov m : tprovt m}, and the type value tprovt m of tprov m is calculated by a function according to the Output_type of the data integration subunit DI_Blockbox. The source object sprov j , target object tprov m and function aprov j in the provenance meta-information are identified using Identify(id, name, type). The id is used to uniquely identify an object or operation, the name represents the object name or method name, and the type is used to identify the type value of the object.
[0069] Collecting provenance meta-information of the data integration process: The data integration user registers the data integration subunit DI_Blockbox online, clearly describing the atomic tasks, functions, inputs, outputs, and specific implementation methods of a specific subunit. In this example, the data integration subunit registered by the data integration user is DI_Blockbox{flexmatch, py_entitymatching, Voting, Data_fusion}. The data provenance user uses the data integration subunit to build the data provenance toolbox ProvToolBox, as shown in Table 5. It contains multiple data integration subunits DI_Blockbox{flexmatch, py_entitymatching, Voting, Data_fusion}. For example, the first sub-module in the data integration subunit py_entitymatching is represented as ({py_entitymatching, ER}, block_tables, {source_dataset, source_dataset, blocking_traindata, blocking_model}, candiate_entityset), which means that when the entity resolution task is completed through the data integration subunit py_entitymatching and the block_tables operation is used, the types of its first input parameter and second input parameter are the source dataset source_dataset, and the third and fourth parameters are the training set blocking_traindata and the blocking model blocking_model used. The data type corresponding to the output of the block_tables function is the candidate entity set candiate_entityset.
[0070] Table 5 Example of the provenance toolbox
[0071]
[0072] For example, a sub-module in the data integration subunit py_entitymatching is represented as ({py_entitymatching, ER}, entity_matching, candidate_dataset, matching_dataset), which means that when using the data integration subunit py_entitymatching to complete entity resolution and the entity_matching operation is used, the type of its input parameter is the candidate dataset candidate_dataset, and the data type corresponding to the output of the function is the matching entity set matching_dataset.
[0073] Data integration users select flexmatch for schema alignment of different data sources, which deals with the problem of matching multiple schemas to a single mediation schema. Then, py_entitymatching is used to group records that describe the same entity in different data sources to obtain an entity matching dataset. Finally, Data_fusion is adopted to fuse the entities.
[0074] To facilitate the description of data instances such as datasets and entity attributes involved in the data integration workflow process, variables are used to replace them in the following description. The variables and their corresponding data instances are shown in Table 6. Suppose there is a data integration workflow DI_Workflow(Workflow_activity, Api_Para) as shown in Table 7. Among them, Workflow_activity = {flexmatch, py_entitymatching, Data_fusion, Voting}, indicating that the data integration workflow DI_Workflow of this instance mainly includes the following data integration subunits: schema alignment flexmatch, entity resolution py_entitymatching, data fusion Data_fusion, and conflict resolution Voting. Api_Para = {schema_matching, {block_tables, entity_matching}, entity_fusion, get_realattribute} is a set of functions in the data integration subunits used in the data integration process. Each function in Api_Para has a parameter list.
[0075] Table 6 Data Variable Instance Table
[0076] Variable Definition <![CDATA[S1]]> Source dataset Sfordors.csv without schema alignment <![CDATA[S2]]> Source dataset Szagats.csv without schema alignment <![CDATA[T1]]> Source dataset fordors.csv with completed schema alignment <![CDATA[T2]]> Source dataset zagats.csv with completed schema alignment <![CDATA[S3]]> Dataset traindata.csv used by the chunking model <![CDATA[M1]]> Chunking model er <![CDATA[C1]]> Candidate dataset candiatedata.csv for entity resolution <![CDATA[T3]]> Matching result dataset matchingdata.csv for entity resolution <![CDATA[T4]]> Attribute conflict dataset conflictattribute.csv <![CDATA[T5]]> Data fusion result dataset resultdata.csv <![CDATA[E1]]> Sunset Restaurant (restraurant11, Sunset_restaurant, Sunset Avenue, New York, 237681123, french_cuisine) <![CDATA[E2]]> Sunset Restaurant (restraurant21, Sunset_restaurant, Sunset Avenue, New York, 237681123, german_cuisine) <![CDATA[E4]]> Sunset Restaurant (restraurant1, Sunset_restaurant, Sunset Avenue, New York, 237681123, french_cuisine) <![CDATA[A1]]> The type of the Sunset restaurant in the fordors.csv dataset is french_cuisine <![CDATA[A2]]> The type of the Sunset restaurant in the zagats.csv dataset is german_cuisine <![CDATA[A3]]> The type of the Sunset restaurant in the fordors.csv dataset is french_cuisine
[0077] Table 7 Data Integration Workflow Example
[0078] DI_Blockbox Api_Para Para Output flexmatch schema_matching <![CDATA[{S1,S2}]]> <![CDATA[{T1,T2}]]> py_entitymatching block_tables <![CDATA[{T1,T2,S3,M1}]]> <![CDATA[{C1}]]> py_entitymatching entity_matching <![CDATA[{C1}]]> <![CDATA[{T3}]]> Data_fusion entity_fusion <![CDATA[{T3}]]> <![CDATA[{T4,T5}]]> Voting get_realattribute <![CDATA[{A1,A2,E4}]]> <![CDATA[{A3}]]>
[0079] Table 8 Provenance Meta-Information Example
[0080]
[0081] Integrate the Fordors restaurant dataset S1 and the Zagats restaurant dataset S2. First, select the data integration subunit flexmatch for schema alignment. Call the function shcema_matching in schema alignment to calculate the similarity between strings, obtain the correspondence between the original schema and the target schema, and modify the schema of the source dataset. The resulting datasets are T1 and T2. Second, select the data integration subunit py_entitymatching for entity resolution. Entity resolution mainly includes the steps of block_tables and entity_matching. block_tables uses the block model M1 to block the source datasets T1 and T2. This model is trained using the training dataset S3. The similarity calculation algorithm is used to calculate the similarity of entities within the blocks to generate the candidate dataset C1. The entity_matching function groups the records describing the same restaurant in different data sources to obtain the matching entity set T3. Finally, select the data integration subunit Data_fusion for entity fusion. Call the entity_fusion function to fuse the data records describing the same restaurant in the matching dataset. If attribute conflicts occur during the entity fusion process, an attribute conflict dataset T4 will be generated to record the conflicting attributes of the entities. Select the data integration subunit Voting to resolve the attribute conflicts, and finally obtain the consistent and clean integrated data T5. T5 contains the data of several restaurants. E4 is a French cuisine restaurant named Sunset_restaurant, located on Sunset Avenue in New York. E4 is obtained from the restaurant E1 and restaurant E2 in the Zagats restaurant dataset S1 through the above data integration operations. The cuisine type of restaurant E1 recorded in the Zagats restaurant dataset S1 is french_cuisine, while the dataset Fordors restaurant dataset S2 shows that the cuisine type of E2 is german_cuisine. When fusing E1 and E2 to obtain E4, the attribute type conflicts, and the conflict values are A1 = french_cuisine and A2 = german_cuisine. By adopting the conflict resolution strategy, the true attribute value A3 = french_cuisine is obtained.
[0082] The data tracing user obtains the tracing meta-information by performing a mapping relationship between the tracing toolbox and the data integration workflow constructed by the data integration user. The tracing meta-information generated during the data integration process is represented by ProvMeta_in(Wfprov, Aprov Sprov, Tprov). The tracing meta-information generated in this instance is shown in Table 8. The tracing meta-information generated in the schema alignment phase ({flexmatch}, {schema_matching}, {(S1:unsmdataset), (S2:unsmdataset)}, {(T1:source_dataset), (T2:source_dataset)}), after the schema matching operation, the target schema is obtained. After aligning the source datasets S1 and S2 with the target schema, the resulting datasets are {(T1:source_dataset), (T2:source_dataset)}. The tracing meta-information generated in the chunking phase of entity resolution ({py_entitymatching}, {block_tables}, {(T1:source_dataset), (T2:source_dataset), (S3:blocking_traindata), (M1:blocking_model)}, {(C1:candiate_dataset)}), where {(T1:source_dataset), (T2:source_dataset), (S3:blocking_traindata), (M1:blocking_model)} is a set representing the source objects, and {(C1:candiate_dataset)} is the target object, and block_tables represents the chunking function. The tracing meta-information generated in the matching phase of entity resolution ({py_entitymatching}, {entity_matching}, {(C1:candiate_dataset)}, {(T3:matching_dataset)}). The entity matching is performed on the candidate dataset C1 generated in the chunking phase using entity_matching to obtain the matching dataset T3. The tracing meta-information generated in the data fusion phase is ({Data_fusion}, {entity_fusion}, {(T3:matching_dataset)}, {(T4:conflict_attributeset), (T5:target_dataset)}), where the conflict resolution strategy is adopted to resolve data conflicts to generate the conflict attribute set T4, and the data records describing the same restaurant in the matching dataset are fused to finally obtain the consistent and clean integrated data T5.The provenance meta-information generated in the conflict resolution phase is ({Voting}, {get_realattribute}, {(A1:conflict_attribute), (A2:conflict_attribute), (E4:conflict_entity)}, {(A3:real_attribute)}). E4 is obtained from restaurant E1 and restaurant E2 in Zagats restaurant dataset S1 through the above data integration operation. The cuisine of restaurant E4 recorded in Zagats restaurant dataset S1 is french_cuisine, while the Fordors restaurant dataset S2 shows that the cuisine of E4 is german_cuisine. There is a conflict in the type of E4, and the conflict values are A1 = french_cuisine and A2 = german_cuisine. By adopting the conflict resolution strategy, the true attribute value A3 = french_cuisine is obtained. The source object sprov in the provenance meta-information. j and the target object tprov m and the function aprov j Use Identify(id, name, type) for identification. id is used to uniquely identify an object or function, name represents the object name or method name, and type is used to identify the type value of the object. For example, (T1, fodors.csv, source_dataset) represents a source dataset fodors.csv in the data integration activity, and the corresponding entity identifier is T1; such as (activity_01, block_tables, activity) indicates that block_tables is an activity in the data integration process, and the activity identifier is activity_01.
[0083] Step 3: Construct a multi-granularity provenance graph based on the multi-granularity provenance model DI_PROV and the provenance meta-information ProvMeta_in generated by the user to construct the data integration workflow; construct the provenance relationship between objects and activities based on the DI-PROV provenance model. The nodes in the multi-granularity provenance graph are the data operation source objects sprov represented by Identify in the provenance meta-information ProvMeta_in j and the target object tprov m and the function aprov j . The provenance relationship between the nodes in the multi-granularity provenance graph is determined by the function aprov in the provenance meta-information ProvMeta_in j and the types of function inputs and outputs. Specifically, it includes:
[0084] Step 3.1: Initialize the multi-granularity traceability graph G(V, E);
[0085] Step 3.2: Construct the traceability relationship between objects and activities based on the traceability model DI-PROV and the traceability meta-information ProvMeta_in, and store it in the traceability relationship list Prov_doc. The nodes in the multi-granularity traceability graph are the source objects sprov in the traceability meta-information ProvMeta_in j , target objects tprov m and functions aprov j . The traceability relationship between nodes in the multi-granularity traceability graph is determined by the function aprov j in the traceability meta-information ProvMeta_in and the types of function parameters;
[0086] Step 3.3: Loop through each data record stored in the traceability relationship list Prov_doc, and create coarse-grained and fine-grained traceability graph nodes and the relationships between nodes according to the type of data objects;
[0087] Step 3.4: Create coarse-grained and fine-grained traceability graphs in the graph database Neo4j based on the nodes and the relationships between nodes, associate the traceability graphs of different granularities, and return them.
[0088] The multi-granularity traceability graph is a directed graph G(V, E), where V = {Node i} is the node set of the multi-granularity traceability graph, corresponding to objects, activities, and agents in the DI_PROV traceability model; E = {Edge j} is the edge set of the multi-granularity traceability graph, corresponding to the edges in the DI_PROV traceability model, used to represent the relationships between activities, objects, and agents. Define Node(node_id, node_name, node_type) to represent the nodes in the multi-granularity traceability graph, where node_id is used to represent the uniqueness of the node, node_name is the name of the node, and node_type identifies the type of the node. Define Edge(edge_id, edge_value, node1_id, node2_id) to represent the edges in the multi-granularity traceability graph, where edge_id represents the unique ID of the edge, edge_value is the value of the edge, node1_id represents the starting node of the edge, and node2_id represents the ending node of the edge. Create coarse-grained and fine-grained traceability graph nodes and the relationships between nodes according to the type of data objects; create coarse-grained and fine-grained traceability graphs respectively based on the nodes and the relationships between nodes, and generate and return the multi-granularity traceability graph by associating the traceability graphs of different granularities.
[0089] Taking the integration of the Fordors restaurant dataset S1 and the Zagats restaurant dataset S2 as an example, the construction of the multi-granularity provenance graph is illustrated. First, initialize the multi-granularity provenance graph G(V, E). Then, according to the multi-granularity provenance model DI_PROV, call the find_provobject function to obtain the objects in the provenance meta-information, call the find_provactivity to obtain the activities in the provenance meta-information, add the objects and activities to the node set of the provenance graph, call find_elementrelation based on the objects and activities to obtain the provenance relationship between the objects and activities, and add it to the edge set of the provenance graph. Repeat this process to obtain all the objects and activities involved in the data integration process, and construct the provenance relationship between the objects and activities. Based on the node set and edge set of the provenance graph, construct the coarse-grained and fine-grained provenance graphs in the graph database Neo4j, associate the provenance graphs of different granularities and return. The constructed multi-granularity provenance graph is as Figure 3 shown. The source objects included in the data integration process are {S1, S2, E1, E2, M1, S4, A1, A2}, the target objects are {T1, T2, C1, T3, T4, T5, A3, E4}, and the activities are {schema_matching, block_tables, entity_fusion, get_realattribute, entity_matching}.
[0090] For example, the provenance meta-information generated in the blocking stage of entity resolution ({py_entitymatching}, {block_tables}, {(T1:source_dataset), (T2:source_dataset), (S3:blocking_traindata), (M1:blocking_model)}, {(C1:candiate_dataset)}), identify the objects in DI_PROV: the source datasets T1 and T2, S3 and M1 represent the training set and blocking model used by block_tables, C1 is the candidate entity set output by the block_tables function, and the activity: the blocking function block_tables. The specific relationships between the activity and the objects are Used and WasGeneratedBy. The nodes and edges of the provenance graph in the blocking stage of entity resolution are shown in Table 9.
[0091] Table 9 Examples of nodes and edges of the provenance graph in the blocking stage of entity resolution
[0092] <![CDATA[Node Node1]]> <![CDATA[Node Node2]]> Edge <![CDATA[Node(T1,fordors.csv,source_dataset)]]> Node(activity_01,block_tables,activity) <![CDATA[Edge(Used_01,Used,T1,activity_01)]]> <![CDATA[Node(T2,zagats.csv,source_dataset)]]> Node(activity_01,block_tables,activity) <![CDATA[Edge(Used_02,Used,T2,activity_01)]]> <![CDATA[Node(S3,traindata.csv,blocking_traindata)]]> Node(activity_01,block_tables,activity) <![CDATA[Edge(Used_03,Used,S3,activity_01)]]> <![CDATA[Node(M1,er,blocking_model)]]> Node(activity_01,block_tables,activity) <![CDATA[Edge(Used_04,Used,M1,activity_01)]]> <![CDATA[Node(C1,candiatedata.csv,candidate_dataset)]]> Node(activity_01,block_tables,activity) <![CDATA[Edge(WasGeneratedBy_0,WasGeneratedBy,C1,activity_01)]]>
[0093] Step 4: Store the multi-granularity traceability graph in the Neo4j database, and use the Cypher query language to implement graph traversal queries and aggregation queries, including coarse-grained traceability queries and fine-grained traceability queries.
[0094] Coarse-grained traceability queries are divided into traceability summary queries and traceability segment queries. For traceability summary queries, first obtain the source datasets, i.e., the Fordors restaurant dataset S1 and the Zagats restaurant dataset S2, then obtain the main activities of data integration, including schema alignment, entity resolution, and conflict resolution, as well as the relationships between these activities, and finally obtain the unified and clean result dataset T5. The coarse-grained summary traceability query results for data quality management users are as Figure 4 shown. For traceability segment queries, segment the multi-granularity traceability graph according to the data integration sub-units of the user-input data integration workflow, and return the traceability records of the execution process of a single data integration sub-unit. The coarse-grained segment traceability query results for data quality management users are as Figure 5 shown.
[0095] Coarse-grained traceability query: Summarize and segment the multi-granularity traceability graph according to the conditions input by the user to return activity-level traceability information, and summarize or segment the traceability graph G(V,E) according to the query type query_type and query conditions (condition) input by the user. A1) First, initialize the traceability record; A2) When the query type is traceability summary, first obtain the source datasets used in the entire data integration process, then obtain the main activities of data integration. Here, the activities are the first-level activities of data integration, such as entity resolution, and do not include sub-activities of entity resolution, such as training matching, etc. Finally, traverse the multi-granularity traceability graph based on the source datasets and main activities to obtain the traceability record and return it; A3) When the query type is traceability segment, find the sub-task of data integration that meets the query conditions, i.e., the sub-task of entity resolution queried by the user, and then traverse the multi-granularity traceability graph to return the traceability relationships of all activities and objects of the entity resolution sub-task.
[0096] Fine-grained traceability query focuses on data entities and attributes, and replays the evolution of data entities during the data integration process. B1) Initialize the result set, convert the query result R into an entity e, and check whether conflict resolution has occurred for entity e during the data integration process; B2) If attribute conflicts have occurred for entity e, replay the conflict resolution process; B3) If no attribute conflicts have occurred for entity e, first obtain the result dataset where entity e is located, because most data processing during data integration is carried out in the form of datasets, then perform a breadth-first traversal of the multi-granularity traceability graph starting from the result dataset to obtain the traceability record. The termination condition for the breadth-first traversal is the source dataset. Finally, query the traceability relationship between the result entity and the entity in the source dataset to obtain the final traceability record and return it.
[0097] The fine-grained traceability query converts the query result into an entity e, and checks whether conflict resolution has occurred during the data integration process for e. If there have been attribute conflicts for e, replay the conflict resolution process. If there have been no attribute conflicts for entity e, first obtain the result data set where entity e is located, and then breadth-first traverse the multi-grained traceability graph G(V, E) starting from the result data set to obtain the traceability records, and return the final traceability records. The fine-grained traceability query result for the data management user's query of restaurant E3 is as Figure 6 shown. From the query result, the data management user knows that restaurant E1 in source data set S1 and restaurant E2 in S2 represent the same restaurant. Use entity fusion entity_fusion to fuse E1 and E2 to obtain restaurant E4. There have been attribute conflicts during the entity fusion process, and E4 has conflicting attributes A1 and A2. The type of restaurant E1 in source data set S1 is A1 = french_cuisine, while the type of restaurant E2 in S2 is A2 = german_cuisine. The conflict resolution strategy get_realattribute is used to obtain the attribute true value A3 = french_cuisine.
Claims
1. A multi-granularity traceability method for data integration, characterized in that Including: Step 1: Construct a multi-granularity traceability model DI_PROV for data integration tasks to realize the playback of the data integration process at multiple granularities. At the coarse granularity, the sub-tasks of data integration are used as the basic units to playback the data integration process at the table level. At the fine granularity, data entities and attributes are taken as the core to playback the evolution process of data during the data integration process; Step 2: Construct a data integration traceability process model, including a data integration traceability toolbox ProvToolBox composed of data integration subunits DI_Blockbox, a data integration workflow model DI_Workflow, and a traceability meta-information model ProvMeta_in. According to the data integration task, the user constructs a data integration workflow DI_Workflow based on the data integration subunit DI_Blockbox in the traceability toolbox ProvToolBox, and the execution of the data integration workflow generates data integration traceability meta-information ProvMeta_in; Step 3: Construct a multi-granularity traceability graph based on the multi-granularity traceability model DI_PROV and the traceability meta-information ProvMeta_in generated by the execution of the data integration workflow model DI_Workflow; Step 4: Store the multi-granularity traceability graph in the Neo4j database and implement graph traversal queries and aggregation queries using the Cypher query language, including coarse-grained traceability queries and fine-grained traceability queries.
2. The multi-granularity traceability method for data integration according to claim 1, wherein The relationships in the multi-granularity traceability model DI_PROV in Step 1 include: Used where the starting point is an activity and the ending point is an object, WasGeneratedBy where the starting point is an object and the ending point is an activity, WasIncludedBy where the starting point is an activity and the ending point is an activity (parent activity includes sub-activity), WasContributedTo where the starting point is an activity and the ending point is an activity (activity contributes to activity), WasDerivedBy where the starting point is an object and the ending point is an object (object is generated by object), WasAttributedTo where the starting point is an object and the ending point is an agent (agent is responsible for object), ActedOnBehalfOf where the starting point is an agent and the ending point is an agent (assuming different agent relationships), WasAssociatedWith where the starting point is an activity and the ending point is an agent (agent is responsible in the activity), WasRecordOf where the starting point is an entity and the ending point is an entity (entity is a subset of entity), HasConflictAttribute where the starting point is an entity and the ending point is an entity (entity has conflicting attributes), HasTruth where the starting point is an entity and the ending point is an entity (true value of entity attribute); The specific construction process includes traceability model node division and traceability model relationship division.
3. A multi-granularity traceability method for data integration according to claim 2, characterized in that, The above-mentioned traceability model node division divides the nodes in the traceability model into three node types: object, activity, and agent; the object corresponds to the data files, models, data entities, and attributes used and generated in data integration; the activity corresponds to the subtasks of data integration, including the processes of schema matching (SM), entity resolution (ER), data integration (DI), and conflict resolution (CR); the agent corresponds to the user responsible for completing the data integration task, and the agent is refined into the general agent and the agent.
4. A multi-granularity traceability method for data integration according to claim 2, characterized in that The above-mentioned traceability model relationship division adds the following relationships on the basis of inheriting the relationships of objects, activities, and agents in the PROV model: 1) Data integration includes several subtasks, and the inclusion relationship and contribution relationship are introduced to reflect the structural characteristics of the workflow; 2) The agent is divided into the sub-agent and the general agent, and there is a representative relationship between the general agent and the sub-agent; 3) A data file contains several data entities, and a sub-record relationship is added between the data file and the data entity; 4) When data entities from different data files are fused, attribute conflicts may occur. There are conflict attributes and attribute truth values for the data entities, and the has-truth-value relationship and the has-conflict-attribute relationship are added.
5. A multi-granularity traceability method for data integration according to claim 1, characterized in that, In step 2, the data integration traceability toolbox ProvToolBox is a collection of data integration subunits DI_Blockbox. The data integration subunit DI_Blockbox is represented as DI_Blockbox(Activity, API, Paramter_type, Output_type), where Activity = {Activity_name: Activity_type}, Activity_name represents the name of the data integration subunit, and Activity_type represents the task type corresponding to the data integration subunit Activity; API = {api1, api2, … api j} represents the set of functions in this subunit, api j represents the j-th function in the subunit; Paramter_type = {pt1, pt2, …, pt j} is the set of specific data types corresponding to the input parameters of the API. The input parameters of the API registered by the user in the data integration subunit are the data types corresponding to the actual parameters in the data integration workflow model DI_Workflow, and pt j represents the type of the j-th input parameter; Output_type = {ot1, ot2, …, ot m} is the set of data types corresponding to the output of the API, and ot m represents the type of the m-th output data.
6. A multi-granularity traceability method for data integration according to claim 1, characterized in that, In step 2, the data integration workflow model DI_Workflow is represented as DI_Workflow(Workflow_activity, Api_Para), where Workflow_activity = {a1, a2, …, a i} is the set of data integration subunits of the data integration workflow, and a i corresponds to the data integration subunit DI_Blockbox in the data integration traceability toolbox ProvToolBox i ; Api_Para = {api_para1, api_para2, …, api_para j} is the set of functions in the data integration subunit DI_Blockbox, and api_para j represents that the j-th function inputs a specific parameter value Para to obtain the corresponding output value Output; Among them, api represents the processing function in the data integration subunit DI_Blockbox, where Para = {p1, p2, …, p j} are the user parameters of the data integration workflow, and Output = {o1, o2, …, o m} is the data generated by the function in the data integration subunit DI_Blockbox.
7. A multi-granularity traceability method for data integration according to claim 1, characterized in that, In step 2, the provenance meta - information model ProvMeta_in is represented as ProvMeta_in(Wfprov, Aprov, Sprov, Tprov), which is used to represent the provenance meta - information generated during the data integration process. Here, Wfprov is the Workflow_activity in the data integration workflow model DI_Workflow, Aprov = {aprov1, aprov2, …, aprov j} corresponds to the set of functions API in the data integration workflow, Sprov = {sprov1, sprov2, …, sprov j} is the set of source objects operated by the API, corresponding to Para in the data integration workflow model DI_Workflow. Each source object is represented as sprov j {sprov j : sprovt j}, and the type value sprovt j of sprov j is calculated by a function based on Paramter_type in the data integration subunit DI_Blockbox. Tprov = {tprov1, tprov2, …, tprov m} is the set of target objects, corresponding to Output in the data integration workflow. Each target object is represented as tprov m {tprov m : tprovt m}, where the type value tprovt m is calculated by a function based on Output_type of the data integration subunit DI_Blockbox. The source object sprov j , target object tprov m and function aprov j in the provenance meta - information are identified using Identify(id, name, type). The id is used to uniquely identify an object or operation, the name represents the object name or method name, and the type is used to identify the type value of the object.
8. A multi-granularity traceability method for data integration according to claim 1, characterized in that The specific content of step 3 includes: Step 3.1: Initialize the multi-granularity traceability graph G(V,E); Step 3.2: Construct the provenance relationship between objects and activities based on the provenance model DI-PROV and the provenance meta-information ProvMeta_in, and store it in the provenance relationship list Prov_doc. The nodes in the multi-granularity provenance graph are the source objects sprov in the provenance meta-information ProvMeta_in j , target objects tprov m and functions aprov j . The provenance relationship between nodes in the multi-granularity provenance graph is determined by the function aprov j in the provenance meta-information ProvMeta_in and the types of function parameters; Step 3.3: Loop through and access each data record stored in the traceability relationship list Prov_doc, and create coarse-grained and fine-grained traceability graph nodes and the relationships between the nodes according to the type of data object; Step 3.4: Create coarse-grained and fine-grained traceability graphs in the graph database Neo4j based on the nodes and the relationships between the nodes, associate the traceability graphs of different granularities, and return them.
9. A multi-granularity traceability method for data integration according to claim 1, characterized in that In step 3, the multi-granularity traceability graph G(V, E) is a directed graph, where V = {Node} is the node set of the multi-granularity traceability graph, corresponding to the objects, activities, and agents in the traceability model DI_PROV; E = {Edge} is the edge set of the multi-granularity traceability graph, corresponding to the edges in the traceability model DI_PROV, used to represent the relationships between activities, objects, and agents. The node Node is represented as Node(node_id, node_name, node_type), where node_id is used to represent the uniqueness of the node, node_name is the name of the node, and node_type identifies the type of the node; the edge Edge is represented as Edge(edge_id, edge_value, node1_id, node2_id), where edge_id represents the unique ID of the edge, edge_value is the value of the edge, node1_id represents the starting node of the edge, and node2_id represents the ending node of the edge; coarse-grained and fine-grained traceability graph nodes and the relationships between nodes are created according to the type of data objects; based on the relationships between nodes, coarse-grained and fine-grained traceability graphs are created respectively, and a multi-granularity traceability graph is generated by associating traceability graphs of different granularities.
10. A multi-granularity traceability method for data integration according to claim 1, characterized in that In step 4, the coarse-grained traceability query includes: Step A1: Initialize the traceability record; Step A2: When the query type is traceability summary, first obtain the source data set used in the entire data integration process, secondly obtain the main activities of the data integration, and finally traverse the multi-granularity traceability graph based on the source data set and the main activities to obtain the traceability record and return it; Step A3: When the query type is traceability segmentation, according to the user's query condition, that is, a certain subtask of the data integration that the user wants to query, and then traverse the multi-granularity traceability graph to return the traceability relationships of all activities and objects of the subtask; The fine-grained traceability query includes: Step B1: Initialize the result set, convert the query result R into an entity e, and check whether conflict resolution has occurred during the data integration process for the entity e; Step B2: If the entity e has had an attribute conflict, replay the conflict resolution process; Step B3: If the entity e has not had an attribute conflict, first obtain the result data set where the entity e is located, then perform a breadth-first traversal of the multi-granularity traceability graph starting from the result data set to obtain the traceability record. The termination condition of the breadth-first traversal is the source data set. Finally, query the traceability relationship between the result entity and the entity in the source data set to obtain the final traceability record and return it.
Citation Information
Patent Citations
Data organization method oriented to sea-cloud collaboration network computing network
CN103838847A
Product source tracing method and device
CN104951906A