Multi-source data association fusion method based on metadata management

Through the multi-source data association fusion method based on metadata management, the problem of inefficiency in traditional databases in multi-source data association fusion is solved, dynamic management of metadata and multi-level data association fusion are realized, data processing efficiency and flexibility are improved, and high-performance association retrieval is supported.

CN119938753APending Publication Date: 2025-05-06AEROSPACE SCI & IND INTELLIGENT OPERATION RES & INFORMATION SECURITY RES INST (WUHAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411956521.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Faced with diverse data sources, complex data structures and huge data scales, traditional databases are inefficient in multi-source data association and fusion, and are difficult to dynamically scale and process changing metadata information.

Method used

The multi-source data association fusion method based on metadata management is adopted. By maintaining table and field metadata information, the non-relational document database materializes metadata and establishes indexes, supports visual drag and drop maintenance table association relationships, dynamically constructs entity relationship diagrams, and realizes multi-level data association fusion and high-performance full-base data retrieval.

Benefits of technology

It realizes dynamic definition and automation of metadata, reduces the workload of users, improves the efficiency and flexibility of multi-source data association fusion, supports large-scale data processing, and provides high-performance association retrieval capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938753A_ABST
    Figure CN119938753A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information system data management, in particular to a multi-source data association fusion method based on metadata management, which comprises the following steps of: aiming at a multi-source data fusion scene, firstly maintaining table and field metadata information, and then materializing metadata and establishing an index by using a non-relational document database; the table association relationship is maintained in a visual dragging mode, so that a metadata table entity relationship graph is obtained; multi-source data import is executed, a fusion association task is automatically executed after data import is completed, the metadata table entity relation graph is traversed based on maintained metadata information, all data directly or indirectly associated with main table data is obtained and stored, and multi-level association fusion of the data is achieved; and finally, performing data compression and cleaning escape on the associated and fused data, and pushing the data to a full-text index to realize high-performance associated retrieval of the whole-library data. According to the method, metadata management functions including table and field management are provided, tables and fields can be dynamically created, related attributes can be configured, materialized creation is automatically completed, and dynamic expansion is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information system data management, and specifically to a multi-source data association and fusion method based on metadata management. Background Art

[0002] With the accelerated popularization of Internet applications, the degree of social informatization continues to deepen. People's behaviors based on the network are increasing, and the user behavior data and attribute data in the network are also increasing. The data generated by various information systems show an explosive growth trend. Based on the user's behavior data and attribute data in the network, the user's information in multiple dimensions can be obtained, and the user portrait can be obtained, which has wide applications in many fields such as marketing, content recommendation, financial services and crime prevention. Obtaining the summary data after association fusion is the premise for forming a character portrait and the basis for multi-dimensional personnel identity analysis. However, in the face of increasingly diverse data sources, increasingly complex data structures, and increasingly large data scales, the efficiency of direct structured data retrieval is getting lower and lower, and the constantly changing metadata information also brings thorny problems to development and expansion. The foreign key association performance of traditional relational databases is low and difficult to expand dynamically. The semantic association cost based on knowledge graphs is high and difficult to intervene manually. In order to solve these problems, the present invention proposes a multi-source data association fusion method and system based on metadata management, which can realize the dynamic definition and automatic animalization of metadata, realize the maintenance of visual table association relationships, stably and efficiently complete the multi-source data association fusion, and effectively reduce the workload of users. Summary of the invention

[0003] 1. Technical issues to be resolved

[0004] The technical problem to be solved by the present invention is: how to provide a multi-source data association fusion method based on metadata management.

[0005] (II) Technical solution

[0006] In order to solve the above technical problems, the present invention provides a multi-source data association fusion method based on metadata management, characterized in that, for the multi-source data fusion scenario, the method first maintains table and field metadata information, then uses a non-relational document database to materialize the metadata and establish an index; maintains table association relationships in a visual drag-and-drop manner, thereby obtaining a metadata table entity relationship diagram; performs multi-source data import, and automatically executes the fusion association task after the data import is completed. Based on the maintained metadata information, the metadata table entity relationship diagram is traversed to obtain and store all data directly or indirectly associated with the main table data, thereby realizing multi-level association fusion of data; finally, the data completed with association fusion is compressed and cleaned and escaped, and then pushed to the full-text index, thereby realizing high-performance association retrieval of the entire library data.

[0007] Wherein, the method comprises the following steps:

[0008] Step 1: Maintain table and field metadata information; users can dynamically create tables on the system and define table names, descriptions, types and other attributes. After the table is created, it will automatically bind system fields (ID, etc.); support dynamic creation of fields on the system, and configure the field names, data types and business type attributes in Chinese and English;

[0009] Step 2: Materialize the database table based on the metadata information. After the metadata definition is completed, the system will materialize the metadata in the document database, insert an empty document after creating a collection, and automatically create indexes on some system fields.

[0010] Step 3: Maintain the association relationship. First, define the main table and configure a field in the main table as a unique ID. Taking personnel information as an example, the unique ID can be a field such as ID card number or student ID number. Then, maintain the association relationship between data tables in the system by dragging and dropping lines. It supports multi-field association and AND / OR logical association.

[0011] When associating, the following are checked: (1) Whether the field types match, including that a Boolean type field cannot be associated with a string type field or an array type; (2) Whether there is a self-association, which is prohibited; (3) Whether there is a duplicate association, which is prohibited;

[0012] After the check is passed, an association record is inserted into the database, including the source table ID, the target table ID, the source field ID array, the target field ID array, and the association logic information;

[0013] Step 4: Build an entity relationship diagram, read all data in the association relationship record table, build an entity relationship diagram based on the source table information and target table information in the association relationship record, and cache it;

[0014] Step 5: Import multi-source data. Users upload offline data files or connect to external databases, configure field mapping, and execute data import tasks.

[0015] Step 6: Data association fusion, starting from the main table, obtain each non-repeated unique ID in the main table, traverse the table entity relationship diagram according to the data of each unique ID in the main table, obtain all table data associated with the unique ID, and store them uniformly;

[0016] Step 7: Push to full-text index; obtain the associated fusion results from the associated fusion result table, aggregate the associated fusion result data in separate tables, convert the stored primary key ID into a complete data row, filter out the fields that do not need to be pushed to the full-text index based on the metadata information, and then escape special fields such as Boolean values ​​and dates based on the field type defined in the metadata information. Finally, convert each row of data in each table into text in the form of key-value pairs and write them into the full-text index to form an index structure of table + key-value pairs.

[0017] The data association fusion method in step 6 is specifically as follows:

[0018] The first step is to obtain all non-repeated unique IDs from the main table. Taking personnel identity data as an example, obtain the list of ID cards after deduplication, recorded as set S1. From the actual business point of view, these are all the personnel who need to be associated and fused; obtain the list of ID cards after deduplication from the associated fusion result table, recorded as set S2, which is the result of the last associated fusion execution; the associated fusion result table stores the results of the associated fusion, including the ID number, table metadata ID and the data primary key ID (array) of the person under the corresponding table; divide the data according to S1 and S2, the part where S1 and S2 intersect needs to update the associated fusion result, the part where S1 exceeds S2 needs to add the associated fusion result, and the part where S2 exceeds S1 needs to delete the associated fusion result;

[0019] The second step is to traverse the ID card numbers that need to be updated and added, and obtain the association fusion status of the current ID card from the association fusion status table; the association fusion status table is used to record the status of each unique ID executing the association fusion task, including the ID card number, operation type, status, and execution batch information; it is used to record the status of the current data being updated / deleted; based on the association fusion result table, recover from the breakpoint and calculate the association fusion progress and speed; according to the association fusion status, determine whether the currently traversed ID card number needs to be associated and fused;

[0020] The third step is to process the data that needs to be associated and fused. According to the ID card number, the corresponding association fusion state is obtained, and the association fusion state is changed to RUNNING. If other execution threads read the ID card, no calculation will be performed. First, the corresponding main table data is obtained, and then the adjacency list directly associated with the main table is obtained according to the entity relationship diagram. The adjacency table is traversed, and the details of the association relationship between the main table and the adjacency table are obtained from the metadata information. The associated fields and the associated logic are clarified, and the query conditions are constructed. Among them, the logic requires all field data to match, or the logic only requires any field to match. The array field needs to be queried with in, and the other fields are queried with equality. All the associated data in the adjacency table can be obtained based on the main table data and the constructed query conditions. The above logic is recursively executed to record the calculation status of each data table until all tables are calculated, and all the associated data of all tables can be obtained.

[0021] The fourth step is to store the association fusion results. After obtaining all the associated data of a person, the complete data object is written into the synchronization queue. The main thread will continuously scan the synchronization queue. When the length of the synchronization queue reaches the threshold, the data objects in the queue are taken out and written into the association fusion result table in batches. At the same time, the association fusion status of the corresponding persons in the queue is updated to success in batches. In addition, if abnormal data is encountered in the previous step, the data object containing the error information is written into the abnormal synchronization queue. Similarly, when the abnormal queue reaches the threshold, the system writes the abnormal data into the association fusion status table.

[0022] Step 5: Process the data to be deleted, traverse the part of S2 that exceeds S1 in the first step, delete the relevant records in the association fusion result table, and simultaneously update the association fusion status table;

[0023] The sixth step is to count the execution results. After all the executions are completed, query the associated fusion status table to count the amount of data that was successfully calculated, the amount of data that failed to be calculated, and the time taken to execute the task.

[0024] (III) Beneficial effects

[0025] Compared with the prior art, the key innovations of the present invention are:

[0026] (1) Provides metadata management functions including table and field management, can dynamically create tables and fields and configure related attributes, automatically complete materialized creation, and facilitate dynamic expansion;

[0027] (2) It supports maintaining table association relationships by dragging and dropping connections, which lowers the usage threshold, supports complex association logic, and supports multi-node association;

[0028] (3) When merging associations, traverse the entity relationship diagram, dynamically build data query conditions, and quickly obtain all data related to the unique ID starting from the unique ID of the main table;

[0029] (4) During the association fusion, the association fusion status is stored to avoid repeated calculations, improve calculation efficiency, and realize the statistics of association fusion progress;

[0030] (5) After the associated fusion task is interrupted, it can be resumed from the breakpoint;

[0031] (6) When storing the associated fusion results, the result primary key array is stored in a separate table to save storage space;

[0032] (7) The result data after association fusion is compressed and escaped and stored in the full-text index to support efficient retrieval.

[0033] The advantages of the present invention are as follows:

[0034] (1) Metadata management technology is used to support efficient creation of tables and fields and automatically complete database table materialization, ensuring semantic consistency and avoiding the complexity of direct database operations;

[0035] (2) Supports visual maintenance of table association relationships and interaction through drag-and-drop, which greatly reduces the difficulty of using the system. It supports definition and / or association logic and can handle more complex association scenarios.

[0036] (3) It supports dynamic expansion and has extremely high flexibility. It can create metadata based on the data source to complete data source adaptation;

[0037] (4) The association fusion is highly efficient and has low computational cost, and can be applied to large-scale data. The association fusion process has state records, which can display the association fusion progress in detail and support continued execution after interruption, thus improving the robustness of the system.

[0038] (5) It supports multi-level associations. Based on the entity relationship diagram, it can complete direct associations and multi-level indirect associations. It can cope with complex data association networks. After the association fusion is completed, the primary keys of the associated data of all tables are stored. The actual data acquisition efficiency is high and the storage space is saved.

[0039] (6) Push the fused data to the full-text index to support high-performance retrieval. It also supports configuring which fields do not need to be pushed to the index, which improves the flexibility of the system. It supports special field escape, which reduces the difficulty for users to understand. Convert the fused data into a text format of table-field-value, which not only simplifies the index structure and compresses the data volume, but also enables accurate data traceability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the overall flow chart of the present invention.

[0041] Figure 2 This is a flow chart of the similarity calculation method. DETAILED DESCRIPTION

[0042] In order to make the purpose, content, and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings and examples.

[0043] In order to solve the above technical problems, the present invention provides a multi-source data association fusion method based on metadata management, characterized in that, for the multi-source data fusion scenario, the method first maintains table and field metadata information, then uses a non-relational document database to materialize the metadata and establish an index; maintains table association relationships in a visual drag-and-drop manner, thereby obtaining a metadata table entity relationship diagram; performs multi-source data import, and automatically executes the fusion association task after the data import is completed. Based on the maintained metadata information, the metadata table entity relationship diagram is traversed to obtain and store all data directly or indirectly associated with the main table data, thereby realizing multi-level association fusion of data; finally, the data completed with association fusion is compressed and cleaned and escaped, and then pushed to the full-text index, thereby realizing high-performance association retrieval of the entire library data.

[0044] Wherein, the method comprises the following steps:

[0045] Step 1: Maintain table and field metadata information; users can dynamically create tables on the system and define table names, descriptions, types and other attributes. After the table is created, it will automatically bind system fields (ID, etc.); support dynamic creation of fields on the system, and configure the field names, data types and business type attributes in Chinese and English;

[0046] Step 2: Materialize the database table based on the metadata information. After the metadata definition is completed, the system will materialize the metadata in the document database, insert an empty document after creating a collection, and automatically create indexes on some system fields.

[0047] Step 3: Maintain the association relationship. First, define the main table and configure a field in the main table as a unique ID. Taking personnel information as an example, the unique ID can be a field such as ID card number or student ID number. Then, maintain the association relationship between data tables in the system by dragging and dropping lines. It supports multi-field association and AND / OR logical association.

[0048] When associating, the following are checked: (1) Whether the field types match, including that a Boolean type field cannot be associated with a string type field or an array type; (2) Whether there is a self-association, which is prohibited; (3) Whether there is a duplicate association, which is prohibited;

[0049] After the check is passed, an association record is inserted into the database, including the source table ID, the target table ID, the source field ID array, the target field ID array, and the association logic information;

[0050] Step 4: Build an entity relationship diagram, read all data in the association relationship record table, build an entity relationship diagram based on the source table information and target table information in the association relationship record, and cache it;

[0051] Step 5: Import multi-source data. Users upload offline data files or connect to external databases, configure field mapping, and execute data import tasks.

[0052] Step 6: Data association fusion, starting from the main table, obtain each non-repeated unique ID in the main table, traverse the table entity relationship diagram according to the data of each unique ID in the main table, obtain all table data associated with the unique ID, and store them uniformly;

[0053] Step 7: Push to full-text index; obtain the associated fusion results from the associated fusion result table, aggregate the associated fusion result data in separate tables, convert the stored primary key ID into a complete data row, filter out the fields that do not need to be pushed to the full-text index based on the metadata information, and then escape special fields such as Boolean values ​​and dates based on the field type defined in the metadata information. Finally, convert each row of data in each table into text in the form of key-value pairs and write them into the full-text index to form an index structure of table + key-value pairs.

[0054] The data association fusion method in step 6 is specifically as follows:

[0055] The first step is to obtain all non-repeated unique IDs from the main table. Taking personnel identity data as an example, obtain the list of ID cards after deduplication, recorded as set S1. From the actual business point of view, these are all the personnel who need to be associated and fused; obtain the list of ID cards after deduplication from the associated fusion result table, recorded as set S2, which is the result of the last associated fusion execution; the associated fusion result table stores the results of the associated fusion, including the ID number, table metadata ID and the data primary key ID (array) of the person under the corresponding table; divide the data according to S1 and S2, the part where S1 and S2 intersect needs to update the associated fusion result, the part where S1 exceeds S2 needs to add the associated fusion result, and the part where S2 exceeds S1 needs to delete the associated fusion result;

[0056] The second step is to traverse the ID card numbers that need to be updated and added, and obtain the association fusion status of the current ID card from the association fusion status table; the association fusion status table is used to record the status of each unique ID executing the association fusion task, including the ID card number, operation type, status, and execution batch information; it is used to record the status of the current data being updated / deleted; based on the association fusion result table, recover from the breakpoint and calculate the association fusion progress and speed; according to the association fusion status, determine whether the currently traversed ID card number needs to be associated and fused;

[0057] The third step is to process the data that needs to be associated and fused. According to the ID card number, the corresponding association fusion state is obtained, and the association fusion state is changed to RUNNING. If other execution threads read the ID card, no calculation will be performed. First, the corresponding main table data is obtained, and then the adjacency list directly associated with the main table is obtained according to the entity relationship diagram. The adjacency table is traversed, and the details of the association relationship between the main table and the adjacency table are obtained from the metadata information. The associated fields and the associated logic are clarified, and the query conditions are constructed. Among them, the logic requires all field data to match, or the logic only requires any field to match. The array field needs to be queried with in, and the other fields are queried with equality. All the associated data in the adjacency table can be obtained based on the main table data and the constructed query conditions. The above logic is recursively executed to record the calculation status of each data table until all tables are calculated, and all the associated data of all tables can be obtained.

[0058] The fourth step is to store the association fusion results. After obtaining all the associated data of a person, the complete data object is written into the synchronization queue. The main thread will continuously scan the synchronization queue. When the length of the synchronization queue reaches the threshold, the data objects in the queue are taken out and written into the association fusion result table in batches. At the same time, the association fusion status of the corresponding persons in the queue is updated to success in batches. In addition, if abnormal data is encountered in the previous step, the data object containing the error information is written into the abnormal synchronization queue. Similarly, when the abnormal queue reaches the threshold, the system writes the abnormal data into the association fusion status table.

[0059] Step 5: Process the data to be deleted, traverse the part of S2 that exceeds S1 in the first step, delete the relevant records in the association fusion result table, and simultaneously update the association fusion status table;

[0060] The sixth step is to count the execution results. After all the executions are completed, query the associated fusion status table to count the amount of data that was successfully calculated, the amount of data that failed to be calculated, and the time taken to execute the task.

[0061] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A multi-source data association fusion method based on metadata management, characterized in that: Aiming at the multi-source data fusion scenario, the method first maintains table and field metadata information, then uses a non-relational document database to materialize the metadata and establish an index; maintains table associations in a visual drag-and-drop manner to obtain a metadata table entity relationship diagram; performs multi-source data import, and automatically executes the fusion association task after the data import is completed. Based on the maintained metadata information, the metadata table entity relationship diagram is traversed to obtain and store all data directly or indirectly associated with the main table data, thereby realizing multi-level association fusion of data; finally, the data completed with association fusion is compressed, cleaned and escaped, and then pushed to the full-text index, thereby realizing high-performance association retrieval of the entire database data.

2. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The method comprises the following steps: Step 1: Maintain table and field metadata information; users can dynamically create tables on the system and define table names, descriptions, types and other attributes. After the table is created, it will automatically bind to system fields; the system supports dynamic creation of fields, and can configure the Chinese and English names, data types and business type attributes of the fields; Step 2: Materialize the database table based on the metadata information. After the metadata definition is completed, the system will materialize the metadata in the document database, insert an empty document after creating a collection, and automatically create indexes on some system fields. Step 3: Maintain the association relationship. First, define the main table and configure a field in the main table as a unique ID. Taking personnel information as an example, the unique ID can be a field such as ID card number or student ID number. Then, maintain the association relationship between data tables in the system by dragging and dropping lines. It supports multi-field association and AND / OR logical association. When associating, the following are checked: (1) Whether the field types match, including that a Boolean type field cannot be associated with a string type field or an array type; (2) Whether there is a self-association, which is prohibited; (3) Whether there is a duplicate association, which is prohibited; After the check is passed, an association record is inserted into the database, including the source table ID, the target table ID, the source field ID array, the target field ID array, and the association logic information; Step 4: Build an entity relationship diagram, read all data in the association relationship record table, build an entity relationship diagram based on the source table information and target table information in the association relationship record, and cache it; Step 5: Import multi-source data. Users upload offline data files or connect to external databases, configure field mapping, and execute data import tasks. Step 6: Data association fusion, starting from the main table, obtain each non-repeated unique ID in the main table, traverse the table entity relationship diagram according to the data of each unique ID in the main table, obtain all table data associated with the unique ID, and store them uniformly; Step 7: Push to full-text index; obtain the associated fusion results from the associated fusion result table, aggregate the associated fusion result data in separate tables, convert the stored primary key ID into a complete data row, filter out the fields that do not need to be pushed to the full-text index based on the metadata information, and then escape special fields such as Boolean values ​​and dates based on the field type defined in the metadata information. Finally, convert each row of data in each table into text in the form of key-value pairs and write them into the full-text index to form an index structure of table + key-value pairs.

3. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The data association fusion method in step 6 is specifically as follows: The first step is to obtain all non-repeated unique IDs from the main table. Taking personnel identity data as an example, obtain the list of ID cards after duplication is removed, recorded as set S1. From the actual business point of view, this is all the personnel who need to be associated and fused. Obtain the list of ID cards after duplication is removed from the association and fusion result table, recorded as set S2, which is the result after the last association and fusion execution. The association fusion result table stores the results of association fusion, including the ID number, table metadata ID, and the data primary key ID of the person in the corresponding table; the data is divided according to S1 and S2. The intersection of S1 and S2 needs to update the association fusion result, the part where S1 exceeds S2 needs to add an association fusion result, and the part where S2 exceeds S1 needs to delete the association fusion result; The second step is to traverse the ID card numbers that need to be updated and added, and obtain the association fusion status of the current ID card from the association fusion status table; The association fusion status table is used to record the status of each unique ID executing the association fusion task, including the ID number, operation type, status, and execution batch information; Used to record the status of the current data being updated / deleted; Based on the association fusion result table, recover from the breakpoint and calculate the association fusion progress and speed; according to the association fusion status, determine whether the currently traversed ID card number needs to be associated and fused; The third step is to process the data that needs to be associated and fused; According to the ID card number, obtain the corresponding association fusion state, change the association fusion state to RUNNING, and other execution threads will not perform calculations if they read the ID card; first obtain the corresponding main table data, and then obtain the adjacency list directly associated with the main table according to the entity relationship diagram, traverse the adjacency table, obtain the details of the association relationship between the main table and the adjacency table from the metadata information, clarify the associated fields and the associated logic, and construct query conditions; among them, the logic requires all field data to match, or the logic only requires any field to match, the array field needs to be queried with in, and the other fields are queried with equality. According to the main table data and the constructed query conditions, all the associated data in the adjacency table can be obtained; recursively execute the above logic, record the calculation status of each data table, until all tables are calculated, and all the associated data of all tables can be obtained; The fourth step is to store the association fusion results. After obtaining all the associated data of a person, the complete data object is written into the synchronization queue. The main thread will continuously scan the synchronization queue. When the length of the synchronization queue reaches the threshold, the data objects in the queue are taken out and written into the association fusion result table in batches. At the same time, the association fusion status of the corresponding persons in the queue is updated to success in batches. In addition, if abnormal data is encountered in the previous step, the data object containing the error information is written into the abnormal synchronization queue. Similarly, when the abnormal queue reaches the threshold, the system writes the abnormal data into the association fusion status table. Step 5: Process the data to be deleted, traverse the part of S2 that exceeds S1 in the first step, delete the relevant records in the association fusion result table, and simultaneously update the association fusion status table; The sixth step is to count the execution results. After all the executions are completed, query the associated fusion status table to count the amount of data that was successfully calculated, the amount of data that failed to be calculated, and the time taken to execute the task.

4. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The method provides metadata management functions including table and field management, can dynamically create tables and fields and configure related attributes, automatically complete materialized creation, and facilitate dynamic expansion.

5. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The method supports maintaining table association relationships in a drag-and-drop connection mode, reduces the usage threshold, supports complex association logic, and supports multi-node association.

6. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: When associating and merging, the method traverses the entity relationship diagram, dynamically builds data query conditions, and quickly obtains all data related to the unique ID starting from the unique ID of the main table.

7. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The method stores the association fusion state during association fusion, avoids repeated calculations, improves calculation efficiency, and can realize association fusion progress statistics.

8. The multi-source data association fusion method based on metadata management as claimed in claim 1, characterized in that: The method can resume execution from a breakpoint after the associated fusion task is interrupted.

9. The multi-source data association fusion method based on metadata management as claimed in claim 1, characterized in that: When storing the associated fusion results, the method stores the result primary key array in a separate table to save storage space.

10. The multi-source data association fusion method based on metadata management according to claim 1, characterized in that: The method compresses and escapes the result data after the association fusion and stores it in the full-text index, supporting efficient retrieval.

Citation Information

Cited By

  • Method and system for automatically discovering incidence relation between data tables

    CN121092638A