A method and device for synchronizing heterogeneous model data
By uniformly encapsulating heterogeneous model data into JSON format, and data content conversion and data model conversion are carried out in the data synchronization system, the problems of high complexity of heterogeneous model data synchronization and high user learning cost in the existing technology are solved, and a simplified data conversion process and improved practicality are realized.
Patent Information
- Application Number
- CN202411081925.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-08
AI Technical Summary
When the existing technology realizes data synchronization of heterogeneous model, it is necessary to design different data positioning methods and data conversion rules for each source database and destination database, resulting in high implementation complexity, high user learning cost and poor practicality.
The data objects of heterogeneous models are uniformly packaged in JSON format, and data content conversion and data model conversion are performed within the data synchronization system, simplifying the data conversion process.
Through unified encapsulation into JSON format, the conversion operations between heterogeneous model data are simplified, the implementation complexity and user learning costs are reduced, and the practicality of the data synchronization system is improved.
Smart Images

Figure CN118964486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data synchronization, and in particular to a method and device for synchronizing heterogeneous model data. Background Art
[0002] The data models in the prior art mainly include: relational data model, document data model, graph data model, wide table model (for example, HBase model) and time series model. Based on different data models, different types of database management systems such as relational database, document database, graph database, HBase and time series database have been developed in the database field. Data of different data models are heterogeneous model data.
[0003] In order to meet the data storage needs of users using different data models in different usage scenarios, it is often necessary to synchronize heterogeneous model data between different databases. The existing technology defines a set of data location methods and data conversion rules between the data models of each source database and each target database; after the data object is read from the source database into the data synchronization system, it is first converted at the data content level based on the corresponding data location method and data conversion rules, and then the converted data object is converted at the data model level, and finally written to the target database.
[0004] However, this implementation approach requires that in a data synchronization system that supports heterogeneous data synchronization, data location methods and data conversion rules based on different data models be designed, which increases the complexity and workload of implementing the data synchronization system; and when users use the corresponding data synchronization system, for each data model of the source database and each data model of the destination database, they need to learn the data location method and data conversion rules between the two, which has a high learning cost and is not very practical.
[0005] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in the field of this technology. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a method and device for synchronizing heterogeneous model data. Its purpose is to provide a universal heterogeneous model data synchronization solution. After reading the data object from the source database, the data objects of different models are uniformly encapsulated into JSON format, and the JSON format is uniformly used to convert the data content within the data synchronization system. The data model of the data object in JSON format is then converted and loaded into the destination database, which solves the problem of poor practicality caused by the data synchronization system in the prior art designing different conversion schemes for the data model of each source database and the data model of each destination database.
[0007] The present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a method for synchronizing heterogeneous model data, comprising:
[0009] Read original data from a source database, and convert the original data from a source format into a JSON format to obtain first data to be converted; wherein the source format is the data format of the source database;
[0010] Performing data content conversion on the first data to be converted to obtain second data to be converted in JSON format;
[0011] Performing data model conversion on the second data to be converted to obtain data to be loaded in a target format; writing the data to be loaded into a target database; wherein the target format is the data format of the target database.
[0012] Further, the reading of original data from a source database, converting the original data from a source format to a JSON format, and obtaining first data to be converted includes:
[0013] When the original data is document data, determining the original data as corresponding first data to be converted;
[0014] When the original data is a point in graph data, a point data table is constructed according to the point, and the first data to be converted is constructed according to the point data table; or when the original data is an edge in graph data, an edge data table is constructed according to the edge, and the first data to be converted is constructed according to the edge data table;
[0015] When the original data is relational data or time series data, selectively converting a single source database table corresponding to the original data into single-layer JSON data according to the conversion mode; or converting multiple source database tables corresponding to the original data into multi-layer nested JSON data to obtain the first data to be converted;
[0016] When the original data is wide table data, two layers of JSON data are constructed according to the original data to obtain the first data to be converted.
[0017] Further, when the original data is a point in graph data, a point data table is constructed according to the point, and the first data to be converted is constructed according to the point data table; or when the original data is an edge in graph data, an edge data table is constructed according to the edge, and the first data to be converted is constructed according to the edge data table, including:
[0018] If the original data is a point, the point attribute name and the point attribute value corresponding to the point are stored in the initialized point data table according to the unique ID of the point, so as to obtain a constructed point data table; based on the constructed point data table, the unique ID of the point is used as the key name of the first key-value pair, the point attribute name and the corresponding point attribute value are used as the key value of the first key-value pair, a single-layer JSON data is obtained based on the first key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted;
[0019] If the original data is an edge, the edge attribute name and edge attribute value corresponding to the edge are stored in the initialized edge data table according to the unique ID of the edge, so as to obtain a constructed edge data table; based on the constructed edge data table, the unique ID of the edge is used as the key name of the second key-value pair, and the edge attribute name and the corresponding edge attribute value are used as the key value of the second key-value pair, and a single-layer JSON data is obtained based on the second key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted.
[0020] Further, when the original data is relational data or time series data, selectively converting a single source database table corresponding to the original data into single-layer JSON data according to a conversion mode; or converting multiple source database tables corresponding to the original data into multi-layer nested JSON data to obtain the first data to be converted includes:
[0021] If the conversion mode is a single-table conversion mode, the column name of the source database table is used as the key name of the third key-value pair, and the column value corresponding to the column name of the source database table is used as the key value of the third key-value pair to obtain the third key-value pair of at least one piece of data in the source database table; all the third key-value pairs corresponding to the source database table are constructed into the single-layer JSON data;
[0022] If the conversion mode is a multi-table conversion mode, then based on the single-table conversion mode, a single-layer JSON data corresponding to each source database table is obtained; according to the primary and foreign key relationships between the source database tables in the source database, the parent-child node relationship between the multiple single-layer JSON data is determined, and multi-layer nested JSON data is obtained based on the parent-child node relationship.
[0023] Furthermore, when the original data is wide table data, constructing two layers of JSON data according to the original data to obtain the first data to be converted includes:
[0024] Using the column names of all columns included in the column cluster of the original data as the key names of the fourth key-value pair, using the column values corresponding to the column names in the column cluster as the key values of the fourth key-value pair, and obtaining JSON data of the leaf node based on the fourth key-value pair of the original data;
[0025] The column cluster name of the column cluster of the original data is used as the key name of the fifth key-value pair, the JSON data of the leaf node is used as the key value of the fifth key-value pair, and the JSON data of the root node is obtained based on the fifth key-value pair of the original data.
[0026] Furthermore, the performing data content conversion on the first data to be converted to obtain the second data to be converted in JSON format includes:
[0027] According to the path specified by the user, the key-value pairs to be converted in the first data to be converted for data content conversion are determined in sequence;
[0028] According to the data conversion rules, the key name and / or key value of the key-value pair to be converted is modified until the data content conversion of all the key-value pairs to be converted in the user-specified path is completed.
[0029] Further, the step of sequentially determining the key-value pairs to be converted in the first data to be converted for data content conversion in this time according to the path specified by the user includes:
[0030] Searching for corresponding key-value pairs to be converted in the data synchronization system in sequence according to the database name, the schema name, the table name of the source database table and the column name of the source database table of the original data in the source database;
[0031] and / or, searching in sequence in the data synchronization system for corresponding key-value pairs to be converted according to the node names of the original data in the source database and the attribute names in the node names;
[0032] And / or, according to the table name of the source database table of the original data in the source database, the column cluster name in the source database table and the column name in the corresponding column cluster, the corresponding key-value pairs to be converted are sequentially searched in the data synchronization system.
[0033] Further, the data synchronization system reads original data from the source database, and writes the data to be loaded into the destination database; the data synchronization system includes at least one reading component, at least one conversion component and at least one loading component;
[0034] The method for synchronizing heterogeneous model data further includes:
[0035] When the CPU load of the device running the data synchronization system exceeds a load threshold, a plurality of first data to be converted in JSON format are transferred each time between a reading component and a conversion component connected to the reading component, and / or a plurality of second data to be converted in JSON format are transferred each time between a conversion component and a loading component connected to the conversion component;
[0036] When the CPU load of the device running the data synchronization system does not exceed the load threshold, a single piece of first data to be converted in JSON format is transferred each time between a reading component and a conversion component connected to the reading component, and / or a single piece of second data to be converted in JSON format is transferred each time between a conversion component and a loading component connected to the conversion component;
[0037] Among them, the reading component is used to read the original data from the source database, convert the original data from the source format to the JSON format, and obtain the first data to be converted; the conversion component is used to convert the data content of the first data to be converted, and obtain the second data to be converted in the JSON format; the loading component is used to perform data model conversion on the second data to be converted, obtain the data to be loaded in the target format, and write the data to be loaded to the destination database.
[0038] In a second aspect, the present invention further provides a device for synchronizing heterogeneous model data, which is used to implement the method for synchronizing heterogeneous model data in the first aspect. The device for synchronizing heterogeneous model data includes:
[0039] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the method for synchronizing heterogeneous model data described in the first aspect.
[0040] In a third aspect, the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, which are executed by one or more processors to complete the method for synchronizing heterogeneous model data described in the first aspect.
[0041] In a fourth aspect, a chip is provided, comprising: a processor and an interface, for calling and running a computer program stored in the memory from a memory, and executing the method for synchronizing heterogeneous model data as in the first aspect.
[0042] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer or a processor, enables the computer or the processor to execute the method for synchronizing heterogeneous model data as described in the first to fourth aspects and any one of the aspects.
[0043] In a sixth aspect, a method system for synchronizing heterogeneous model data is provided, comprising a device for synchronizing heterogeneous model data as in the second aspect, and using the method for synchronizing heterogeneous model data as described in the first aspect to complete the interaction of the device for synchronizing heterogeneous model data of the second aspect.
[0044] Different from the prior art, the present invention has at least the following beneficial effects:
[0045] The present invention reads original data from a source database, converts the source format of the original data into JSON format, and obtains first data to be converted; performs data content conversion on the first data to be converted, and obtains second data to be converted in JSON format; performs data model conversion on the second data to be converted, and obtains data to be loaded in the target format used by the destination database, and writes it to the destination database; when the present invention reads the original data into the data synchronization system, it uniformly encapsulates the data into JSON format, and the conversion of the data content only needs to be based on the JSON format, thereby simplifying the conversion operation between heterogeneous model data, and can realize data extraction, conversion and loading tasks between homogeneous databases or heterogeneous databases, solving the problem that the data synchronization system in the prior art designs different conversion schemes for the data model of each source database and the data model of each destination database, and has poor practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 It is a flowchart of a method for synchronizing heterogeneous model data provided by an embodiment of the present invention;
[0048] Figure 2 This is a tree structure diagram of JSON data provided by an embodiment of the present invention;
[0049] Figure 3 It is a specific example of JSON data provided by an embodiment of the present invention;
[0050] Figure 4 is a flow chart of step 10 provided in an embodiment of the present invention;
[0051] Figure 5 is a flowchart of step 102 provided by an embodiment of the present invention;
[0052] Figure 6 is a flowchart of step 103 provided by an embodiment of the present invention;
[0053] Figure 7 is a flowchart of step 104 provided by an embodiment of the present invention;
[0054] Figure 8is a flow chart of step 20 provided in an embodiment of the present invention;
[0055] Fig. 9 It is a functional module diagram of a data synchronization system provided by an embodiment of the present invention;
[0056] Fig.10 is a schematic diagram of a data synchronization operation provided by an embodiment of the present invention;
[0057] Fig.11 It is a flowchart of another method for synchronizing heterogeneous model data provided by an embodiment of the present invention;
[0058] Fig.12 It is a schematic diagram of the architecture of a heterogeneous model data synchronization device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0060] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0061] Unless the context requires otherwise, throughout the specification and claims, the term "including" is to be interpreted as open inclusion, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "examples", "specific examples" or "some examples" and the like are intended to indicate that specific features, structures, materials or characteristics associated with the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representation of the above terms does not necessarily refer to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner, that is, although they may be carried in the embodiments or examples of the above terms due to reasons such as the order and position of appearance, it is not limited to that they can be carried in combination by one embodiment or example.
[0062] In the description of the present invention, it is necessary to understand that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present disclosure.
[0063] In the description of the present invention, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "multiple" is two or more. In addition, for example, the same type of nouns may be described as two independent individuals by adding "A" and "B" at the end. In this case, the corresponding features defined as "A" and "B" are only used to distinguish the same type of individuals for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.
[0064] When describing some embodiments, the expressions "coupling", "coupling" and "connection" and their derivatives may be used. For example, when describing some embodiments, the term "connection" may be used to indicate that two or more components are in direct physical or electrical contact with each other. For another example, when describing some embodiments, the term "coupling" may be used to indicate that two or more components are in direct physical or electrical contact. However, the terms "connection" or "coupling" may also refer to two or more components that are not in direct contact with each other, but still cooperate or interact with each other, such as "optical path coupling", "wireless connection", etc. The embodiments disclosed herein are not necessarily limited to the contents of the present invention.
[0065] In the description of the present invention, the expression "A and / or B" (where A and B are used to formally represent specific characteristic contents) will be involved, and the corresponding expressions include the following three combinations: only A, only B, and a combination of A and B.
[0066] As used herein, "about," "substantially," or "approximately" includes the stated value and an average value that is within an acceptable range of deviation from the particular value as determined by one of ordinary skill in the art taking into account the measurements in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).
[0067] Embodiment 1:
[0068] The basic functions of the data synchronization system are: first read the data objects to be synchronized from the source database, then convert the read data objects according to the user's configuration, or do not convert them, and finally write the data objects to the destination database.
[0069] Due to historical reasons such as the long-term monopoly of relational databases, the data synchronization systems currently on the market are basically developed for relational databases. Data objects flow and process between the processing modules of the data synchronization system in the form of a relational model with a row and column structure. The source database and the destination database are both relational databases, or can be regarded as relational databases. There is no or no need to consider the issue of data model conversion between the two. Existing data synchronization systems generally convert data objects in the representation of relational data model data, locate data through database names, schema names, table names, and column names, and after locating the data object, transform the data object according to the set transformation rules, or transform the database name, schema name, table name, or column name.
[0070] In recent years, in order to meet users' higher demands for database performance and not be limited by the expressive power of relational databases, various non-relational databases developed for specific application types have emerged, such as document databases, time series databases, and graph databases. In their respective application fields, the performance of these non-relational databases in terms of data storage space utilization, data writing, data query, and aggregate computing is superior to that of traditional general-purpose relational databases. In these specific application fields, application systems that were developed based on general-purpose relational databases in the early days now have the need to migrate to specific non-relational databases (for example, real-time monitoring systems). Relational databases were used to store periodically collected time series data in the early days, but now time series databases can be used directly to store these time series data and provide query computing services to the outside world. At the same time, there is also a need to synchronize heterogeneous model data between relational databases and non-relational databases, as well as between various types of non-relational databases.
[0071] Data synchronization between heterogeneous model databases includes two levels of conversion: one is the conversion of data models, and the other is the conversion of data content. For example, when a relational database is migrated to a document database, the relational data of multiple relational tables are merged into one document data. The conversion of data content is to perform content or form conversion operations on the data, for example, desensitizing, cleaning and converting relational data.
[0072] When performing data synchronization between heterogeneous model databases, there is a need for rule-based data conversion. A simple implementation idea in the prior art is: define a set of data location methods and data conversion rules for data objects of each data model. After the data objects are read from the source database of a specific data model into the data synchronization system, the data content of the data objects is first converted based on the data location method and data conversion rules between the source database and the destination database, and then the converted data is converted into a data model, and finally written to the destination database of another data model.
[0073] This implementation idea requires that in the data synchronization system that supports heterogeneous data synchronization, data location methods and data conversion rules based on different data models be designed. In actual production scenarios, according to the specific usage needs of different users, after the data objects to be synchronized are read into the data synchronization system, there are often a large number of data objects that need to be converted at the data content level, such as data desensitization. If the data location method and data conversion rules for each source database data model and each destination database data model are defined according to the specific usage requirements, on the one hand, the complexity and workload of implementing the data synchronization system are huge, which is not conducive to the maintenance and update iteration of the subsequent data synchronization system. On the other hand, there are too many usage methods that users need to learn when using the data synchronization system. In addition, since the data location method and data conversion rules for the same requirement may be different in the process of synchronizing from the source database of the A data model to the destination database of the B data model, compared with the process of synchronizing from the source database of the B data model to the destination database of the A data model, they have to be designed separately, which further increases the development cost of the data synchronization system and the user's usage cost.
[0074] In the prior art, another implementation idea for heterogeneous data synchronization is to define an intermediate representation of data within the data synchronization system, design a data location method and data conversion rules based on this intermediate representation, convert the data object into this intermediate representation when reading it into the data synchronization system, and convert it from this intermediate representation to the model data of the target database when writing it to the target database. However, there are many difficulties in the conversion between the document data of the document data model and the relational data of the relational data model, especially for multi-level document data, which is difficult to represent with a single relational data, and the conversion cost between the two is also very high. When synchronizing document data to a relational database, a solution can be used in which multiple relational data represent a multi-level document data. However, in actual use scenarios, there is also a need to synchronize data between document databases. Due to the limitation of the expression ability of the relational data model, in the synchronization application scenario from document database to document database, the use of this solution faces the problem of how to reconstruct multiple relational data in the intermediate representation into one document data.
[0075] In order to solve the above problems, Figure 1 As shown, an embodiment of the present invention provides a method for synchronizing heterogeneous model data, including:
[0076] Step 10: Read original data from a source database, and convert the original data from a source format into a JSON format to obtain first data to be converted; wherein the source format is the data format of the source database.
[0077] Among them, JavaScript Object Notation (JSON) is a lightweight data exchange format. Figure 2 As shown in the figure, JSON data is a multi-level nested format. The data at each level is a key-value pair set, and the data at the next level constitutes an item of a key-value pair in the previous level. JSON data can be organized into a tree structure in memory. The tree structure consists of a root node, intermediate nodes, and leaf nodes. The root node is a key-value pair set (that is, Figure 2 MAP shown in ), the intermediate node and the leaf node are a set of key-value pairs, or a linked list consisting of a set of key-value pairs (i.e., Figure 2LIST(MAP) shown in ); where the value can be a specific data value or point to the next level tree node. For JSON data, the specified information is extracted from it through the JsonPath path; the industry has defined a set of common JsonPath syntax, which is used to write the JsonPath path and locate the key-value pair or key value based on the JsonPath path; the JsonPath path starts from the key name specified in the root node and is specified to the lower nodes step by step to locate the selected key value. For example, Figure 3 As shown in the figure, the value located by the JsonPath path “$.store.book[0].title” is “Sayings of the Century”; the value located by the JsonPath path “$.store.bicyle.color” is “red”.
[0078] For example, when the source database is a relational database, multiple relational data tables (ie, original data) are read, and the multiple relational data tables are converted into a JSON format to obtain first data to be converted.
[0079] In the embodiment of the present invention, no matter what data model the source format is based on, the original data is converted into JSON format when it is read into the data synchronization system.
[0080] In an optional embodiment, a corresponding reading module can be implemented for each source format according to synchronization requirements, and a default rule for converting the corresponding source format data into JSON data is built into each reading module, so that when the reading module reads the original data into the data synchronization system, it automatically converts it into the JSON format represented within the data synchronization system.
[0081] Step 20: Perform data content conversion on the first data to be converted to obtain second data to be converted in JSON format.
[0082] Data content conversion refers to the transformation of the key name or key value in the first data to be converted specified by the user, or the transformation of the key name and key value at the same time. The transformation of the key includes renaming, deleting and adding key-value pairs; the transformation of the key value includes desensitizing, cleaning and converting the value content, etc., presenting the data to a specific application system in another form.
[0083] For example, for a certain value A in the first data to be converted in JSON format, data desensitization is performed on it to obtain the desensitized value B, that is, the second data to be converted in JSON format is obtained.
[0084] Data content conversion is implemented based on rules. In an optional embodiment, the data synchronization system can predefine a batch of data conversion rules, or support user-defined conversion rules. When configuring a data synchronization task, the user configures the data conversion rules together with the data object, locates the data object to be converted (i.e., the key name or key value in the first data to be converted) through JsonPath path positioning, and applies the specified data conversion rules to the data object to achieve the purpose of data conversion.
[0085] Step 30: Perform data model conversion on the second data to be converted to obtain data to be loaded in a target format; write the data to be loaded into a target database; wherein the target format is the data format of the target database.
[0086] Among them, data model conversion refers to converting the second data to be converted from JSON format to the target format; for example, when the target database is a document-type database, for the second data to be converted that is converted from multiple relational data tables read from a relational database, the data content of the second data to be converted is organized into corresponding document data (i.e., data to be loaded) according to the primary and foreign key relationships between the corresponding multiple relational data tables. Data model conversion only changes the organizational form structure of the corresponding data content. The hierarchical relationship between real things expressed by the multiple relational data tables and the primary and foreign key relationships between them is exactly the same as the hierarchical relationship between real things constructed in the corresponding data to be loaded, which is equivalent to using different data models to express the same real things and the hierarchical relationship between them.
[0087] In an optional embodiment, a corresponding loading module can be implemented for each source format according to synchronization requirements, and a default rule for converting JSON data into the corresponding target format can be built into each loading module, thereby automatically converting JSON data into the target format when it is written from within the data synchronization system to the destination database.
[0088] The present invention reads original data from a source database, converts the source format of the original data into JSON format, and obtains first data to be converted; performs data content conversion on the first data to be converted, and obtains second data to be converted in JSON format; performs data model conversion on the second data to be converted, and obtains data to be loaded in the target format used by the destination database, and writes it to the destination database; when the present invention reads the original data into the data synchronization system, it uniformly encapsulates it into JSON format, and the conversion of the data content only needs to be based on the JSON format, thereby simplifying the conversion operation between heterogeneous model data, and can realize data extraction, conversion and loading tasks between homogeneous databases or heterogeneous databases, solving the problem that the data synchronization system in the prior art designs different conversion schemes for the data model of each source database and the data model of each destination database, and has poor practicality.
[0089] It is worth noting that in actual usage scenarios, users have a large demand for using the data synchronization system to convert data content, so the workload of data content conversion is large, the implementation is complex, and the user learning cost is high; while the conversion of data models, compared with formulating data positioning methods and data conversion rules for conversion needs between different data models, the implementation complexity, workload, and user learning cost are greatly reduced, simplifying the design and implementation of the "rules" themselves when performing data content conversion, as well as the user's usage threshold when using the data synchronization system to configure data synchronization jobs.
[0090] For example, when the data synchronization system needs to synchronize heterogeneous model data between relational databases, document databases, and graph databases, the data synchronization system only needs to formulate: (1) rules for converting relational data into JSON data; (2) rules for converting document data into JSON data; (3) rules for converting graph data into JSON data; (4) rules for data content conversion based on JSON format; (5) rules for converting JSON data into relational data; (6) rules for converting JSON data into document data; (7) rules for converting JSON data into graph data. It is not necessary to design data content conversion rules from each source format to the target format based on different user needs on the basis of designing 6 data model conversion rules for the three heterogeneous databases. This allows the embodiment of the present invention to directly implement the data location method and data conversion rules based on the JSON format when there are a large number of different data content conversion requirements for users. When using the data synchronization system or customizing the data location method and data conversion rules, the user does not need to consider the conversion between heterogeneous data models, which greatly reduces the implementation complexity, workload, related resource costs, and user learning costs of the data synchronization system.
[0091] The following describes the method for synchronizing heterogeneous model data according to an embodiment of the present invention with respect to various specific data models:
[0092] The relational data model is a traditional two-dimensional table structure of rows and columns. From the perspective of expression, each table consists of fixed attribute fields. However, real things are hierarchical, and each table of the relational data model is an abstract modeling of a certain level of real things. When using the relational data model to model things, it is often necessary to use multiple tables to fully model and describe real things and their relationships.
[0093] The document data model is a multi-level nested structure of key-value pairs, which can represent, store, and query the hierarchical relationships of real things together. In terms of logical concepts and even physical concepts, all the contents of document data are organized and stored together. Compared with relational databases that obtain all data information through connection operations, the query performance of document databases is much better. The expressive power of the document data model is also much richer than that of the relational data model.
[0094] The conversion between relational model data and document model data cannot be simply done one-to-one at the data model level, such as one relational data table corresponding to one document model data set. In order to fully utilize the storage performance and query performance of the two databases, the relational model data and document model data should be in a one-to-many relationship, that is, multiple row and column data from different relational data tables are read and merged into a multi-level nested JSON data, or vice versa, a document data is read and decomposed into multiple row and column relational data with different relational data table structures.
[0095] The graph data model (referring to the attribute graph model) uses points and edges to model things and their relationships. Point objects model entities, and edge objects model the connections between entities. Point objects and edge objects have their own attribute sets (the corresponding attribute sets are composed of one or more key-value pairs). The same type of point objects or edge objects can correspond to a relational data table in the relational data model, and the attributes of points or edges correspond to the field columns of the relational data table. The points or edges in the graph data model form a one-to-one correspondence with the relational data tables in the relational data model. In the early days of graph data applications, relational data models were used to store the points and edges of the graph, roughly using one relational table for point objects and another for edge objects. After the emergence of graph databases, they gradually replaced relational databases as the storage and query service carrier for graph application system data.
[0096] For the HBase wide table model, one or more column clusters (column familyu) are specified when the wide table is defined. Each data must have a unique column, and other columns are located under the column cluster. Different rows of data can have different columns in the same column cluster, which has a one-to-one correspondence with the relational data model. From another perspective, the column cluster can be regarded as a node in the document data model, and the "column name-column value pair" is regarded as a key-value pair set. Then the HBase wide table model can be regarded as a two-layer structure of JSON data, a simple document data model.
[0097] The time series data model can be viewed as relational data with timestamps. Each piece of time series data must have a timestamp field and multiple indicator fields. In terms of expression, the time series data can correspond one-to-one with the tables of the relational data model.
[0098] In order to illustrate the process of converting the original data of different data models into JSON format, Figure 4 As shown, the step 10 includes:
[0099] Step 101: When the original data is document data, the original data is determined as corresponding first data to be converted.
[0100] The document data is read from the source database as JSON data, and no additional data processing is required.
[0101] Step 102: When the original data is a point in graph data, a point data table is constructed based on the points, and the first data to be converted is constructed based on the point data table; or, when the original data is an edge in graph data, an edge data table is constructed based on the edge, and the first data to be converted is constructed based on the edge data table.
[0102] Among them, graph data consists of nodes, edges and corresponding attribute objects; in a graph database based on the attribute graph model, nodes and edges can be regarded as two different relational data tables, and nodes and edges are constructed as JSON data with only one hierarchical structure.
[0103] Step 103: When the original data is relational data or time series data, selectively convert a single source database table corresponding to the original data into single-layer JSON data according to a conversion mode; or, convert multiple source database tables corresponding to the original data into multi-layer nested JSON data to obtain the first data to be converted.
[0104] Among them, the relational data model is in the form of a two-dimensional table structure of rows and columns. The data model of time series data is a time series model, which is composed of a two-dimensional table structure of rows and columns with timestamps; single-attribute or multi-attribute time series data must have a time field. The indicator field and time field can be regarded as columns in the relational data table. Time series data is formally a multi-field row data of a relational model. The conversion method of relational data is the same as that of time series data.
[0105] Step 104: When the original data is wide table data, construct two layers of JSON data according to the original data to obtain the first data to be converted.
[0106] Among them, wide table data is a collection of column clusters; the HBase wide table model adds the concept of column clusters on the basis of relational data tables. The wide table model stores all data in a single wide table, so there is no need to consider the primary and foreign key relationships between multiple relational data tables like when processing relational data. Instead, you only need to consider the relationship between column clusters and columns. This can be seen as adding a hierarchical structure on the basis of the relational model, that is, a two-layer JSON data. The first data to be converted is constructed according to this principle.
[0107] From the perspective of representation, in the above five data models, the most basic building blocks of data can be regarded as key-value pairs: the basic data unit in the relational data model is a key-value pair consisting of "column name-column value", but the structure of the relational data table is fixed, so the embodiment of the present invention extracts the column name and stores it separately; the basic data unit of the document data model is a key-value pair organized in a multi-level nested structure; the attributes of point objects and edge objects in the graph data model are represented by key-value pairs; the wide table model can be regarded as a simplified document data model with only two layers of structure; the time series data model itself can be regarded as a relational data model.
[0108] In order to explain in detail the process of converting graph data into JSON data, Figure 5 As shown, step 102 includes:
[0109] Step 1021: If the original data is a point, the point attribute name and the point attribute value corresponding to the point are stored in the initialized point data table according to the unique ID of the point, so as to obtain a constructed point data table; based on the constructed point data table, the unique ID of the point is used as the key name of the first key-value pair, and the point attribute name and the corresponding point attribute value are used as the key value of the first key-value pair, and a single-layer JSON data is obtained based on the first key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted.
[0110] Step 1022: If the original data is an edge, then according to the unique ID of the edge, the edge attribute name and edge attribute value corresponding to the edge are stored in the initialized edge data table to obtain a constructed edge data table; based on the constructed edge data table, the unique ID of the edge is used as the key name of the second key-value pair, and the edge attribute name and the corresponding edge attribute value are used as the key value of the second key-value pair, and a single-layer JSON data is obtained based on the second key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted.
[0111] The graph data is composed of points and edges. Points have unique IDs and other attributes. Edges have unique IDs and other attributes. Other attributes are composed of attribute names and attribute values, which are equivalent to key-value pairs. Edge attribute values include the source point ID and destination point ID corresponding to the edge.
[0112] In an optional embodiment, a single-layer JSON data may be constructed in the format of "(ID: ID value)(attribute name: attribute value)", wherein the key name of the first key-value pair or the second key-value pair is "(ID: ID value)", and the key value is "(attribute name: attribute value)". When the original data is a point, the single-layer JSON data is "(ID: unique ID of the point)(point attribute name: point attribute value)"; when the original data is an edge, the single-layer JSON data is "(ID: unique ID of the edge)(edge attribute name: edge attribute value)".
[0113] In order to explain in detail the process of converting relational data or time series data into JSON data, Figure 6 As shown, step 103 includes:
[0114] Step 1031: If the conversion mode is a single-table conversion mode, the column name of the source database table is used as the key name of the third key-value pair, and the column value corresponding to the column name of the source database table is used as the key value of the third key-value pair to obtain the third key-value pair of at least one data in the source database table; all the third key-value pairs corresponding to the source database table are constructed into the single-layer JSON data.
[0115] The embodiment of the present invention is divided into a single table conversion mode and a multi-table conversion mode to construct JSON data. In the single table conversion mode, the column name and the corresponding column value of the table are combined into a key-value pair, and multiple groups of column value pairs of the data table constitute a key-value pair set, which is represented as a mapping structure (MAP) in the memory to construct JSON data with only one hierarchical structure.
[0116] Step 1032: If the conversion mode is a multi-table conversion mode, then based on the single-table conversion mode, a single-layer JSON data corresponding to each source database table is obtained; according to the primary and foreign key relationships between the source database tables in the source database, the parent-child node relationship between the multiple single-layer JSON data is determined, and multi-layer nested JSON data is obtained based on the parent-child node relationship.
[0117] In the multi-table conversion mode, according to the primary and foreign key relationships between relational data tables, such as the primary and foreign key relationships of the "primary table-supplementary table" mode, a multi-layered JSON data structure is constructed to organize the relational data of different relational data tables together and represent them in memory as a "mapping-linked list mapping nested body" (i.e., Figure 2 MAP-LIST shown <map>). The same is true for time series data.
[0118] In order to explain in detail the process of converting wide table data into JSON data, Figure 7 As shown, step 104 includes:
[0119] Step 1041: Using the column names of all columns included in the column cluster of the original data as the key names of the fourth key-value pair, using the column values corresponding to the column names in the column cluster as the key values of the fourth key-value pair, and obtaining the JSON data of the leaf node based on the fourth key-value pair of the original data.
[0120] Step 1042: Using the column cluster name of the column cluster of the original data as the key name of the fifth key-value pair, using the JSON data of the leaf node as the key value of the fifth key-value pair, and obtaining the JSON data of the root node based on the fifth key-value pair of the original data.
[0121] In the tree structure of JSON data, the key name of the key-value pair of the root node is the column cluster name (key), and the key value points to the child nodes (value) composed of all columns under the column cluster (for example, Figure 2 The valueAi shown in is a key value, whose value is a pointer to the LIST(MAP)1 of the next level); the leaf node is a key-value pair set consisting of all columns and corresponding column values under the column cluster.
[0122] In order to illustrate the process of data content conversion, Figure 8 As shown, the step 20 includes:
[0123] Step 201: according to the path specified by the user, the key-value pairs to be converted in the first data to be converted for data content conversion are determined in sequence.
[0124] Among them, the user-specified path is selected by technical personnel in this field according to the specific usage scenario. Because the relational database is located by database name, schema name, table name and column name; the document database is located by JsonPath; the graph database is located by points and lines to the attribute name; HBase is located by table name, column cluster and column name; the time series database is the same as the relational database. The key-value pair that needs to be converted is found by positioning. Therefore, the embodiment of the present invention searches for the corresponding key-value pairs to be converted in the data synchronization system in sequence according to the database name, schema name, table name of the source database table and column name in the source database table of the original data in the source database; and / or, searches for the corresponding key-value pairs to be converted in sequence in the data synchronization system according to the node name of the original data in the source database and the attribute name in the node name; and / or, searches for the corresponding key-value pairs to be converted in sequence in the data synchronization system according to the table name of the source database table of the original data in the source database, the column cluster name in the source database table and the column name in the corresponding column cluster.
[0125] Step 202: modify the key name and / or key value of the key-value pair to be converted according to the data conversion rule until the data content conversion of all the key-value pairs to be converted in the user-specified path is completed.
[0126] The data conversion rules define how to convert data content, and the data conversion rules are selected by technicians in this field according to specific usage scenarios; for example, the desensitization rule masks and converts the fixed number of digits in the middle of the customer's ID card information with asterisk characters. The data synchronization system can predefine a batch of built-in conversion rules, and also provide a mechanism to support users to customize conversion rules.
[0127] An embodiment of the present invention provides a synchronization requirement list between heterogeneous model data, as shown in the following table. The time series model is classified as a relational data model and is omitted in the table.
[0128]
[0129]
[0130] In the above synchronization requirements list, Relational Database Management System (RDBMS) refers to databases implemented based on relational data models and time series databases; JSON refers to databases implemented based on document models; Graph Database (GDB) refers to databases implemented based on graph models; HBase refers to databases implemented based on wide table data models, such as HBase.
[0131] Among them, "model conversion" refers to data model conversion; for example, the relational data of multiple relational data tables are merged into one document data. "Value conversion" refers to data content conversion; for example, in a relational database, the field values in a row of relational data with multiple fields are converted at the data content level. "JSON value conversion" refers to the data content level conversion of key-value pairs in JSON format, including the conversion of key names and key values. The steps referenced by brackets "[]" indicate non-essential and optional steps in the data synchronization process. "Direct write" means directly writing the original data read from the source database to the destination database. The attribute graph model refers to the use of graphs to represent data. The graph consists of "points" and "edges". To store the graph data modeled by the data graph model in a relational database, it is necessary to create two relational data tables in the relational database. One table is used to store "point" data, called the "point table", and the other table is used to store "edge" data, called the "line table"; "point table, [attribute value conversion]" refers to synchronizing graph data to a relational database. First, the "points" are read out from the graph database as a two-dimensional row and column structure of a relational data table. If there is a configuration requirement, the column values of the two-dimensional row and column structure data are converted accordingly; "[point attribute value conversion]" refers to synchronizing graph data from one graph database to another, and data conversion may be required in the middle.
[0132] The synchronization requirements list shows the steps required for data synchronization between different data models, from columns to rows, which means that after reading the original data from the source database of the specified data model, it is written to the destination database of the specified data model. In the middle, processing steps such as data model conversion and data content conversion may need to be performed; for example, "RDBMS read" and "JSON write" mean that after reading the original data from the relational database, it is written to the document database, which includes one-to-one or one-to-many model conversion and optional data content conversion. Among them, the one-to-one and many-to-one conversion of relational data to document data refers to the conversion of a relational data into a document data and the fusion of multiple relational data into a document data; conversely, the one-to-one and many-to-one conversion of document data to relational data refers to the conversion of a document data into a relational data and the extraction of multiple relational data after decomposition of a document data. The data conversion between other data models is similar.
[0133] In an optional embodiment, the reading modules for different data models construct JSON data, and the conversion module determines the type of the JSON data source and the synchronization requirements (for example, knowing that this JSON data comes from HBase and needs to be written to GDB) according to the different data reading modules (for example, the ID of the data reading module), and then compares the synchronization requirements list to determine what operations need to be performed (for example, point table, attribute value conversion; line table, attribute value conversion), and finally obtains the converted JSON data and gives it to the loading module. Among them, the conversion module is used to convert the data content of the first data to be converted in JSON format to obtain the second data to be converted in JSON format.
[0134] like Fig. 9 As shown, the embodiment of the present invention provides a module division scheme for a data synchronization system: according to basic functions such as data extraction (Extract), transformation (Transformation), loading (Loading), etc., the data synchronization system is divided into three basic functional modules, namely, a reading module, a transformation module and a loading module. The reading module reads the original data from the source database, reads and organizes it into JSON format, and passes it to the subsequent transformation module; the transformation module locates the key-value pair or key value based on the set JsonPath path, and converts the data content of the key-value pair or key value using the set conversion rules. After the conversion is completed, it is transmitted to the subsequent loading module; the converted JSON data is passed to the loading module, and the loading module is responsible for data model conversion, that is, converting the data in JSON format into a script adapted to different destination databases, submitting it to the data source for execution, and completing the data loading work; wherein, the transformation module is optional, and if data content conversion is not required, the transformation module is not required.
[0135] Among them, the reading module implements a functional module for the database of each data model, such as a relational data reading module, a document data reading module, a graph data reading module, and a wide table data reading module, etc.; a data reading module can also be implemented according to each actual database, such as for the relational database Oracle, a functional module that reads Oracle table data and builds it into JSON format is implemented.
[0136] The conversion module converts JSON data into two categories: a. Adjusting the tree structure, such as increasing or decreasing tree branches, so that JSON data from the same data set has the same tree branch structure. b. Modifying the content of key-value pairs, such as changing key names, normalizing key values, or desensitizing key values; transforming and modifying sensitive data under given rules.
[0137] The loading module is responsible for writing data to the specified data source (i.e., the destination database) and converting the JSON data into different scripting languages according to different target formats: for relational databases, it is converted into a Structured Query Language (SQL) script and submitted to the destination database for execution; for graph databases, such as Neo4j, it is converted into a Cypher script and submitted to the Neo4J database server for execution; for document model databases, such as MongoDB, the JSON data can be directly submitted to the MongoDB server for execution.
[0138] Similarly, the reading module and loading module can be further specified according to the data sources of different models. For example, a loading module can be implemented for each data model, such as a relational data loading module, a document data loading module, a graph data loading module, and an HBase data loading module. The loading module can also be implemented for each specific database, such as an Oracle database loading module, which converts JSON data into SQL scripts and submits them to the Oracle server for execution.
[0139] After reading the original data, the reading module constructs it into JSON format, shielding the differences in the representation of different data models; the conversion module converts the data based on the JSON format; the loading module converts the JSON data into the script language of the destination database and submits it to the destination database for execution to achieve the purpose of data synchronization.
[0140] When implementing the data synchronization system, it faces each specific database, such as the relational database Oracle and the document database MongoDB. The aforementioned abstract functional modules must be concretized before they can be used for data synchronization tasks between homogeneous data models or heterogeneous model data sources in reality.
[0141] A relational database is a highly standardized database system that defines a unified standard database operation language, SQL. SQL scripts written in standard SQL syntax can be directly executed on different relational databases that comply with SQL standard specifications. Therefore, for relational databases, a common reading module and a common loading module can be implemented separately.
[0142] There is currently no standard operation script syntax for document model databases and graph databases. Databases belonging to the same model, such as the graph database Neo4j of the property graph model and the graph database TigerGraph, have relatively large differences in their operation scripts. For document model databases and graph databases, when the database types are different, the data synchronization system implements a functional module for each specific database.
[0143] The concrete functional module is called a functional component. The component implements a single function as much as possible, such as the relational database reading component, which reads relational data from the relational database and reconstructs it into JSON format; while the MongoDB reading component can only read relational data from the MongoDB database. Conversion components include structural adjustment components (structural adjustment components can implement branching and pruning operations on tree structures) and key-value pair content adjustment components. Loading components can implement Neo4j point and edge loading components, etc.
[0144] The data synchronization system reads the original data from the source database and writes the data to be loaded to the destination database; the data synchronization system includes at least one reading component, at least one conversion component and at least one loading component. The data synchronization system provides single-function reading components, conversion components and loading components. Users can select appropriate functional components according to the needs of the data synchronization task and combine them to design the data synchronization job. Among them, the reading component can be connected to one or more conversion components, or directly connected to one or more loading components; the conversion component can be connected to one or more conversion components, or directly connected to one or more loading components. Fig.10 As shown, according to this rule, rich data synchronization tasks can be designed.
[0145] To illustrate the process, Fig.11 As shown, the method for synchronizing heterogeneous model data also includes:
[0146] In step 401, when the CPU load of the device running the data synchronization system exceeds a load threshold, multiple first data to be converted in JSON format are transferred each time between a reading component and a conversion component connected to the reading component, and / or multiple second data to be converted in JSON format are transferred each time between a conversion component and a loading component connected to the conversion component.
[0147] In step 402, when the CPU load of the device running the data synchronization system does not exceed the load threshold, a single piece of first data to be converted in JSON format is transferred each time between a reading component and a conversion component connected to the reading component, and / or a single piece of second data to be converted in JSON format is transferred each time between a conversion component and a loading component connected to the conversion component.
[0148] Among them, the reading component is used to read the original data from the source database, convert the original data from the source format to the JSON format, and obtain the first data to be converted; the conversion component is used to convert the data content of the first data to be converted, and obtain the second data to be converted in the JSON format; the loading component is used to perform data model conversion on the second data to be converted, obtain the data to be loaded in the target format, and write the data to be loaded to the destination database.
[0149] When the data synchronization job is executed in the data synchronization system, the data flowing between components is in JSON format. The data is transmitted between two connected components through a blocked synchronization queue. The upstream component writes data to the synchronization queue, and the downstream component reads data from the synchronization queue. Each component has a separate working thread. There are two different ways to transmit data between two adjacent upstream and downstream components: single JSON transmission and multiple JSON batch transmission.
[0150] Data is transferred between two components through a synchronization queue. This process involves one write operation (i.e., writing data to the synchronization queue) and one read operation (i.e., reading data from the synchronization queue). When a single piece of data is transferred between two connected components, each piece of data needs to go through two operations, one write and one read. When batch data is transferred between two connected components, for example, ten pieces of data are a batch. This batch of data only needs to go through two operations, one write and one read, to complete the transfer of the ten pieces of data. Compared with the transfer of a single piece of data, nine rounds of write and one read operations are reduced.
[0151] After a component reads a batch of data from a synchronization queue, the batch of data is processed one by one in sequence inside the component. When the total number of data to be processed remains unchanged, the total processing consumption of the CPU is certain. When the CPU load is high, it is necessary to reduce the CPU workload as much as possible. In the scenario where data is transferred between two components through a synchronization queue, the number of read and write times of the synchronization queue is reduced as much as possible. Therefore, in order to reduce the number of reads and writes to the synchronization queue, the embodiment of the present invention adopts a batch transfer mode when the CPU load is high; in order to increase the timeliness of data processing, a single transfer mode is adopted when the CPU load utilization is low.
[0152] In an optional embodiment, an operation script for generating different models from JSON is added to the loading component to realize data loading from JSON format to different target formats.
[0153] The present invention designs a synchronization method for heterogeneous model data, uses the JSON format as an intermediate data form to perform data conversion operations, and for a data synchronization system, realizes the generalization of data conversion operations across data models.
[0154] Embodiment 2:
[0155] The data synchronization system of the embodiment of the present invention is a device for synchronizing heterogeneous model data. Fig.12 , is a schematic diagram of the architecture of a synchronization device for heterogeneous model data according to an embodiment of the present invention. The synchronization device for heterogeneous model data according to this embodiment includes one or more processors 21 and a memory 22. Fig.12 A processor 21 is taken as an example.
[0156] The processor 21 and the memory 22 may be connected via a bus or other means. Fig.12 The example of connecting through bus is taken in the following.
[0157] The memory 22 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs and non-volatile computer executable programs, such as the method for synchronizing heterogeneous model data in this embodiment. The processor 21 executes the method for synchronizing heterogeneous model data by running the non-volatile software programs and instructions stored in the memory 22.
[0158] The memory 22 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely arranged relative to the processor 21, and these remote memories may be connected to the processor 21 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The program instructions / modules are stored in the memory 22, and when executed by the one or more processors 21, the method for synchronizing heterogeneous model data in the above embodiment is executed, for example, the method described above is executed. Figure 1 , Figure 4-Figure 8 and Fig.11 The steps shown.
[0160] An embodiment of the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors, such as Fig.12 A processor 21 can enable the one or more processors to execute the method for synchronizing heterogeneous model data in a specific embodiment of the present invention, for example, to execute the above-described Figure 1 , Figure 4-Figure 8 and Fig.11 The steps shown can also be implemented Fig.12 The modules and units described above; or the method for synchronizing heterogeneous model data in a specific embodiment of the present invention, for example, executing the method described above Figure 1 , Figure 4-Figure 8 and Fig.11 The steps shown can also be implemented Fig.12 The various modules and units described.
[0161] It is worth noting that the information interaction, execution process, etc. between the modules and units within the above-mentioned devices and systems are based on the same concept as the processing method embodiment of the present invention. The specific contents can be found in the description of the method embodiment of the present invention and will not be repeated here.
[0162] A person skilled in the art may understand that all or part of the steps in the various methods of the embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0163] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.< / map>
Claims
1. A method for synchronizing heterogeneous model data, characterized in that: include: Reading original data from a source database, converting the original data from a source format into a JSON format, and obtaining first data to be converted; wherein the source format is the data format of the source database; the reading original data from a source database, converting the original data from a source format into a JSON format, and obtaining first data to be converted comprises: when the original data is document data, determining the original data as the corresponding first data to be converted; when the original data is a point in graph data, constructing a point data table according to the point, and constructing the first data to be converted according to the point data table; when the original data is an edge in graph data, constructing an edge data table according to the edge, and constructing the first data to be converted according to the edge data table; when the original data is relational data or time series data, selectively converting a single source database table corresponding to the original data into a single-layer JSON data according to a conversion mode; converting multiple source database tables corresponding to the original data into multi-layer nested JSON data to obtain the first data to be converted; when the original data is wide table data, constructing two layers of JSON data according to the original data to obtain the first data to be converted; Performing data content conversion on the first data to be converted to obtain second data to be converted in JSON format; Performing data model conversion on the second data to be converted to obtain data to be loaded in a target format; writing the data to be loaded into a target database; wherein the target format is the data format of the target database; The data synchronization system reads original data from a source database and writes the data to be loaded into a destination database; the data synchronization system includes at least one reading component, at least one conversion component and at least one loading component; the method for synchronizing heterogeneous model data also includes: When the CPU load of the device running the data synchronization system exceeds the load threshold, multiple pieces of first data to be converted in JSON format are transferred each time between the reading component and the conversion component connected to the reading component, and / or, multiple pieces of second data to be converted in JSON format are transferred each time between the conversion component and the loading component connected to the conversion component; when the CPU load of the device running the data synchronization system does not exceed the load threshold, a single piece of first data to be converted in JSON format is transferred each time between the reading component and the conversion component connected to the reading component, and / or, a single piece of second data to be converted in JSON format is transferred each time between the conversion component and the loading component connected to the conversion component; wherein the reading component is used to read original data from a source database, convert the original data from a source format into a JSON format, and obtain first data to be converted; the conversion component is used to perform data content conversion on the first data to be converted, and obtain second data to be converted in JSON format; the loading component is used to perform data model conversion on the second data to be converted, obtain data to be loaded in a target format, and write the data to be loaded to a destination database.
2. The method for synchronizing heterogeneous model data according to claim 1, characterized in that: When the original data is a point in graph data, constructing a point data table according to the point, and constructing the first data to be converted according to the point data table; when the original data is an edge in graph data, constructing an edge data table according to the edge, and constructing the first data to be converted according to the edge data table includes: If the original data is a point, the point attribute name and the point attribute value corresponding to the point are stored in the initialized point data table according to the unique ID of the point, so as to obtain a constructed point data table; based on the constructed point data table, the unique ID of the point is used as the key name of the first key-value pair, the point attribute name and the corresponding point attribute value are used as the key value of the first key-value pair, a single-layer JSON data is obtained based on the first key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted; If the original data is an edge, the edge attribute name and edge attribute value corresponding to the edge are stored in the initialized edge data table according to the unique ID of the edge, so as to obtain a constructed edge data table; based on the constructed edge data table, the unique ID of the edge is used as the key name of the second key-value pair, and the edge attribute name and the corresponding edge attribute value are used as the key value of the second key-value pair, and a single-layer JSON data is obtained based on the second key-value pair, and the single-layer JSON data is used as the corresponding first data to be converted.
3. The method for synchronizing heterogeneous model data according to claim 1, characterized in that: When the original data is relational data or time series data, selectively converting a single source database table corresponding to the original data into single-layer JSON data according to a conversion mode; Converting the multiple source database tables corresponding to the original data into multi-layer nested JSON data to obtain the first data to be converted includes: If the conversion mode is a single-table conversion mode, the column name of the source database table is used as the key name of the third key-value pair, and the column value corresponding to the column name of the source database table is used as the key value of the third key-value pair to obtain the third key-value pair of at least one piece of data in the source database table; all the third key-value pairs corresponding to the source database table are constructed into the single-layer JSON data; If the conversion mode is a multi-table conversion mode, then based on the single-table conversion mode, a single-layer JSON data corresponding to each source database table is obtained; according to the primary and foreign key relationships between the source database tables in the source database, the parent-child node relationship between the multiple single-layer JSON data is determined, and multi-layer nested JSON data is obtained based on the parent-child node relationship.
4. The method for synchronizing heterogeneous model data according to claim 1, characterized in that: When the original data is wide table data, constructing two layers of JSON data according to the original data to obtain the first data to be converted includes: Using the column names of all columns included in the column cluster of the original data as the key names of the fourth key-value pair, using the column values corresponding to the column names in the column cluster as the key values of the fourth key-value pair, and obtaining JSON data of the leaf node based on the fourth key-value pair of the original data; The column cluster name of the column cluster of the original data is used as the key name of the fifth key-value pair, the JSON data of the leaf node is used as the key value of the fifth key-value pair, and the JSON data of the root node is obtained based on the fifth key-value pair of the original data.
5. The method for synchronizing heterogeneous model data according to claim 1, characterized in that: The performing data content conversion on the first data to be converted to obtain the second data to be converted in JSON format comprises: According to the path specified by the user, the key-value pairs to be converted in the first data to be converted for data content conversion are determined in sequence; According to the data conversion rules, the key name and / or key value of the key-value pair to be converted is modified until the data content conversion of all the key-value pairs to be converted in the user-specified path is completed.
6. The method for synchronizing heterogeneous model data according to claim 5, characterized in that: The step of sequentially determining the key-value pairs to be converted in the first data to be converted for data content conversion in this step according to the path specified by the user includes: Searching for corresponding key-value pairs to be converted in the data synchronization system in sequence according to the database name, the schema name, the table name of the source database table and the column name of the source database table of the original data in the source database; and / or, searching in sequence in the data synchronization system for corresponding key-value pairs to be converted according to the node names of the original data in the source database and the attribute names in the node names; And / or, according to the table name of the source database table of the original data in the source database, the column cluster name in the source database table and the column name in the corresponding column cluster, the corresponding key-value pairs to be converted are sequentially searched in the data synchronization system.
7. A synchronization device for heterogeneous model data, characterized in that: The device for synchronizing heterogeneous model data includes at least one processor and a memory, wherein the at least one processor and the memory are connected via a data bus, and the memory stores instructions that can be executed by the at least one processor, and after the instructions are executed by the processor, they are used to implement the method for synchronizing heterogeneous model data described in any one of claims 1-6.
8. A non-volatile computer storage medium, characterized in that: The computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors to complete the method for synchronizing heterogeneous model data as described in any one of claims 1-6.
Citation Information
Patent Citations
Data synchronization method and device
CN111209344A
Method and device for converting document model data into relational model data
CN118152353A