Construction method and system of distributed graph database storage layer based on TiDB

By building a distributed graph database storage layer on TiDB and accelerating data access using TiKV and TiFlash, the constraint and access efficiency problems of the distributed graph database management system in the existing technology are solved, and efficient graph data storage and analysis are achieved.

CN120144552APending Publication Date: 2025-06-13GLOBAL ENERGY INTERCONNECTION RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311713814.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, a distributed graph database management system based on KV storage system cannot provide distributed constraints efficiently, while a distributed graph database system based on a relational database is limited by a relational database query mechanism, resulting in inefficient access to graph database points and edge elements.

Method used

The distributed graph database storage layer based on TiDB is adopted to store graph data through TiDB, and the underlying direct point and edge object data access mechanism and TiFlash columnar storage are used to accelerate data access.

Benefits of technology

It realizes efficient distributed constraint management, improves the access efficiency of point and edge data of graph database, and is suitable for the data storage and analysis needs of graph databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144552A_ABST
    Figure CN120144552A_ABST
Patent Text Reader

Abstract

The invention provides a TiDB-based distributed graph database storage layer construction method and system. The method comprises the following steps: storing point data and edge data in graph data based on TiDB; point and edge data access is carried out on the TiDB storing point data and edge data in the graph data based on a TiKV bottom layer direct point and edge object data access mechanism; and carrying out large-scale point and edge data access on the TiDB which stores the point data and the edge data in the graph data based on TiFlash column type storage. According to the method, the distributed graph database of the TiDB is adopted to efficiently provide distributed constraints, and the data access process is accelerated through a bottom layer direct point and edge object data access mechanism of the TiKV and TiFlash column type storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph database management, and specifically to a method and system for constructing a distributed graph database storage layer based on TiDB. Background Art

[0002] With the explosive development of the Internet, mobile Internet, social networks, Internet of Things, and industrial field-related networks such as power networks, there is a great demand for the storage of relationship graphs and applications such as network topology analysis and functional analysis based on relationship graphs, which has also contributed to the research and development boom of graph databases.

[0003] A graph database is a data management system that uses points and edges as basic storage units and is designed to efficiently store and query graph data.

[0004] In a distributed database management system, it means that data is stored in different local databases respectively, managed by different local database management systems, run on different machines, connected by different communication networks, and presented as an overall database management system in terms of product form. A distributed database is logically a unified whole and physically stored on different physical nodes respectively. From the user's perspective, a distributed database system is the same as a centralized database system logically, and users can execute global applications at any site. It seems that the data is stored on the same computer and managed by a single database management system (DBMS), and users don't feel any difference.

[0005] Common techniques used in the storage layer of conventional distributed graph database management systems are:

[0006] Implementing distributed graph data storage based on a single-machine version of KV (key-value) data storage;

[0007] Implementing a distributed graph database system storage based on a relational database system (such as MySQL);

[0008] Implementing the storage of a distributed graph database in a native graph storage manner from the bottom layer based on a file system;

[0009] Analyzing the problems existing in the above common distributed graph database management systems at the graph database storage layer are:

[0010] For a distributed graph database management system implemented based on a KV storage system (such as RocksDB / HBase), due to the characteristics of the KV storage system, it is often unable to efficiently provide distributed constraints (such as unique value constraints, whether it can be null, etc.);

[0011] The distributed graph database system implemented based on a relational database (such as MySQL) is restricted by the query mechanism of the relational database. When accessing specific points or edges, it is necessary to locate row data through an index, thus slowing down the process of accessing common point and edge elements in the graph database. Summary of the Invention

[0012] In order to solve the problems of the distributed graph database management system implemented based on the KV storage system in the prior art, which is restricted by the characteristics of the KV storage system and cannot efficiently provide distributed constraints, and the distributed graph database system implemented based on the relational database, which is restricted by the query mechanism of the relational database and slows down the process of accessing common point and edge elements in the graph database, the present invention proposes a construction method for the storage layer of a distributed graph database based on TiDB, including:

[0013] Storing the point data and edge data in the graph data based on TiDB;

[0014] Accessing the point and edge data of TiDB storing the point data and edge data in the graph data based on the underlying direct point and edge object data access mechanism of TiKV;

[0015] Performing large-scale point and edge data access to TiDB storing the point data and edge data in the graph data based on TiFlash columnar storage.

[0016] Optionally, the storing the point data and edge data in the graph data based on TiDB includes:

[0017] Introducing all points of each point type in the graph data into a table in TiDB for storage;

[0018] Introducing all edges of each edge type in the graph data into a table in TiDB for storage.

[0019] Optionally, the introducing all points of each point type in the graph data into a table in TiDB for storage includes:

[0020] Using the row ID of the TiDB table as the unique ID of each stored point data;

[0021] Using the columns corresponding to the graph data type in the TiDB table as the attributes of the point type, and keeping the attributes of the point type, data constraints consistent with the column data types and data constraints in the TiDB table.

[0022] Optionally, the introducing all edges of each edge type in the graph data into a table in TiDB for storage includes:

[0023] Using the row ID of the TiDB table as the ID of the stored edge;

[0024] Use the columns of the TiDB table as the attributes of the edge types in the graph data, and keep the attributes of the edge types in the graph data, as well as the data constraints, consistent with the column data types and data constraints in the TiDB table.

[0025] Optionally, the underlying direct vertex and edge object data access mechanism based on TiKV accesses the vertex and edge data in the TiDB storing the graph data, including:

[0026] For the row-level data stored in TiDB, the Key in TiKV has: tablePrefix{TableID}_recordPrefixSep{RowID}, where tablePrefix and recordPrefixSep are specific string constants;

[0027] Obtain the TableID according to the table information corresponding to the vertex type or edge type, combine the ID of the vertex or edge, know the Key of the vertex or edge in the underlying KV storage of TiDB, and directly access the corresponding vertex and edge data based on the Key of the vertex or edge.

[0028] Optionally, the large-scale vertex and edge data access to the TiDB storing the graph data based on TiFlash columnar storage includes:

[0029] When deploying the TiDB cluster, configure TiFlash as a RAFT Learner. For the update of the vertex and edge data in TiDB, TiFlash will obtain a copy of the vertex and edge data and store it in TiFlash in a columnar storage manner;

[0030] When processing a large number of vertices and edges of the whole graph or subgraph, perform large-scale vertex and edge data access through TiFlash.

[0031] On the other hand, the present invention also provides a construction system for the distributed graph database storage layer based on TiDB, including:

[0032] A storage module for storing the vertex data and edge data in the graph data based on TiDB;

[0033] A direct access module for accessing the vertex and edge data in the TiDB storing the graph data based on the underlying direct vertex and edge object data access mechanism of TiKV;

[0034] A big data access module for performing large-scale vertex and edge data access to the TiDB storing the graph data based on TiFlash columnar storage.

[0035] Optionally, the storage module includes:

[0036] A point data storage sub-module, configured to introduce all points in each point type in the graph data into a table in TiDB for storage;

[0037] An edge data storage sub-module, configured to introduce all edges in each edge type in the graph data into a table in TiDB for storage.

[0038] Optionally, the point data storage sub-module is specifically configured to:

[0039] Use the row ID of the TiDB table as the unique ID for each stored point data;

[0040] Use the columns corresponding to the graph data type in the TiDB table as the attributes of the point type, and keep the attributes of the point type, data constraints consistent with the column data types and data constraints in the TiDB table.

[0041] Optionally, the edge data storage sub-module is specifically configured to:

[0042] Use the row ID of the TiDB table as the ID of the stored edge;

[0043] Use the columns of the TiDB table as the attributes of the edge type in the graph data, and keep the attributes of the edge type of the graph data, data constraints consistent with the column data types and data constraints in the TiDB table.

[0044] Optionally, the direct access module is specifically configured to:

[0045] For the row-level data stored in TiDB, the Key in TiKV has: tablePrefix{TableID}_recordPrefixSep{RowID}, where tablePrefix and recordPrefixSep are specific string constants;

[0046] Obtain the TableID according to the table information corresponding to the point type or edge type, combine with the ID of the point or edge, know the Key of the point or edge in the underlying KV storage of TiDB, and directly access the corresponding point and edge data based on the Key of the point or edge.

[0047] Optionally, the big data access module is specifically configured to:

[0048] When deploying the TiDB cluster, configure TiFlash as a RAFT Learner. For the update of point and edge data in TiDB, TiFlash will obtain a copy of the point and edge data and store it in TiFlash in a columnar storage manner;

[0049] When processing a large number of points and edges in the whole graph or sub - graphs, large - scale point and edge data access is performed through TiFlash.

[0050] On the other hand, the present application also provides a computing device, including: one or more processors;

[0051] The processor is used to execute one or more programs;

[0052] When the one or more programs are executed by the one or more processors, the construction method of the distributed graph database storage layer based on TiDB as described above is implemented.

[0053] On the other hand, the present application also provides a computer - readable storage medium, on which there is a computer program, and when the computer program is executed, the construction method of the distributed graph database storage layer based on TiDB as described above is implemented.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0055] The present invention provides a construction method of a distributed graph database storage layer based on TiDB, including: storing point data and edge data in the graph data based on TiDB; performing point and edge data access on TiDB storing point data and edge data in the graph data through the underlying direct point and edge object data access mechanism of TiKV; performing large - scale point and edge data access on TiDB storing point data and edge data in the graph data based on TiFlash columnar storage. The present invention efficiently provides distributed constraints by using the distributed graph database of TiDB, and accelerates the data access process through the underlying direct point and edge object data access mechanism of TiKV and TiFlash columnar storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flowchart of the construction method of the distributed graph database storage layer based on TiDB of the present invention;

[0057] Figure 2 It is a schematic diagram of the construction principle of the distributed graph database storage layer based on TiDB of the present invention. DETAILED DESCRIPTION

[0058] The present invention proposes a construction method of a distributed graph database storage layer based on TiDB, which introduces support for distributed storage of graph databases based on the TiDB distributed database management system, bringing the following conveniences:

[0059] 1) TiDB provides constraint features of relational databases in a distributed environment, such as non - null constraints and unique value constraints;

[0060] 2) TiDB has good online scaling and multiple redundancy functions, thus bringing distributed basic characteristics to the distributed graph storage system built on it;

[0061] 3) In addition to the characteristics of relational data, TiDB stores data in KV mode at the bottom layer. Therefore, the graph database built on it can apply this characteristic to directly access point and edge elements through KV instead of relational data;

[0062] 4) In addition to providing the characteristics based on row storage, TiDB also provides the ability based on column storage at the bottom layer, bringing advantages in column storage performance for large-scale graph analysis;

[0063] Combining the above items, the distributed + relational data constraint + KV access ability + large-scale row storage access ability provided by TiDB proposed by the present invention is suitable for graph database data storage.

[0064] Example 1:

[0065] A construction method for the storage layer of a distributed graph database based on TiDB, as Figure 1 shown, includes:

[0066] Step S1: Store the point data and edge data in the graph data based on TiDB;

[0067] Step S2: Access the point and edge data of TiDB that stores the point data and edge data in the graph data through the direct point and edge object data access mechanism of the underlying TiKV;

[0068] Step S3: Perform large-scale point and edge data access on TiDB that stores the point data and edge data in the graph data based on TiFlash columnar storage.

[0069] The present invention is a construction method for the storage layer of a distributed graph database based on TiDB. On top of the distributed system based on TiDB, it innovatively realizes the implementation of the data layer storage technology for the distributed graph database. The following combines Figure 2 to specifically introduce the present invention:

[0070] Step S1: Store the point data and edge data in the graph data based on TiDB, specifically including:

[0071] Introduce all the points in each point type in the graph data into a table in TiDB for storage;

[0072] Introduce all the edges in each edge type in the graph data into a table in TiDB for storage.

[0073] Further, all points of each point type in the graph data are introduced into a table in TiDB for storage, including:

[0074] Use the row ID of the TiDB table as the unique ID for each point data stored;

[0075] Use the columns corresponding to the graph data type in the TiDB table as the attributes of the point type, and keep the attributes of the point type, data constraints consistent with the column data types and data constraints in the TiDB table.

[0076] Further, all edges of each edge type in the graph data are introduced into a table in TiDB for storage, including:

[0077] Use the row ID of the TiDB table as the ID of the stored edge;

[0078] Use the columns of the TiDB table as the attributes of the edge type in the graph data, and keep the attributes of the edge type of the graph data and data constraints consistent with the column data types and data constraints in the TiDB table.

[0079] The graph database of the graph database is implemented with the distributed storage characteristics of TiDB, and brings global distributed characteristics such as uniqueness constraints and non-null constraints;

[0080] The access to points and edges of the graph database depends on TiKV for access, which accelerates the data access process compared with the access to row-level data in relational databases;

[0081] For the graph storage based on TiDB, the present invention introduces the column storage characteristics based on TiFlash for storing the attributes of points and edges. For specific graph analysis (such as graph computing, etc.), read-only access to graph data is performed through TiFlash storage, and its column storage characteristics are used to accelerate the graph analysis process.

[0082] This step S1 is to introduce a storage mechanism of a distributed graph database with strong constraints on TiDB. The storage of point data and edge data will be introduced separately below:

[0083] 1.1 Storage of point data

[0084] For each point type in the graph data, introduce a table in TiDB to store all points of this type, that is, use a single table to store the point type in a single graph database.

[0085] Since the point data in the graph database is stored in tables in TiDB, it is necessary to reference individual point objects more efficiently in some way. Therefore, the row ID (i.e., RowID) of TiDB is used as the unique ID for each piece of stored point data. This ID remains unique and immutable, that is, it does not change with the change of the values of any user-visible attributes.

[0086] At the same time, columns corresponding to the graph data types are introduced into the TiDB table for the attributes of the point type, and the attribute data types and data constraints of the graph data are kept consistent with the column data types and data constraints in the TiDB table. The distributed characteristics of TiDB bring global data constraint guarantees, such as the uniqueness of attributes, non-null constraints, etc.

[0087] Based on the distributed storage of TiDB, its built-in partitioning characteristics (such as the partitioning method based on hash values) can be relied on to provide partitioning characteristics for the storage of point data. At the same time, the multiple redundant backup mechanism of TiDB also provides a backup mechanism for the storage of point data.

[0088] 1.2 Edge data storage

[0089] For each edge type in the graph data, a table in TiDB is introduced in the present invention to store all edges of this type, that is, a single table is also used to store the edge types in a single graph database. Similarly, the row ID (i.e., RowID) of TiDB is used as the ID of the stored edge for each edge in the edge type. This ID is unique and immutable, and can globally and uniquely reference each specific edge.

[0090] At the same time, columns corresponding to the graph data types are introduced into the TiDB table for the attributes of the edge type, and the attribute data types and data constraints of the edges in the graph data are kept consistent with the column data types and data constraints in the TiDB table. Similar to the aforementioned point attribute storage mechanism, the global strong consistency constraint of TiDB ensures the consistency characteristics of edge attributes, and the edge attributes can be accessed more quickly through columnar storage subsequently.

[0091] Based on the distributed storage of TiDB, its built-in partitioning characteristics (such as the partitioning method based on hash values) can be relied on to provide partitioning characteristics for the storage of edge data. At the same time, the multiple redundant backup mechanism of TiDB also provides a backup mechanism for the storage of edge data.

[0092] Through step 1 here, the technical solution based on TiDB provides an effective mechanism for the storage of points and edges in graph data:

[0093] Based on the distributed consistency characteristics of TiDB, it provides distributed characteristics for the storage of point and edge data;

[0094] Based on the constraint features of TiDB, better consistency features are brought to the point and edge data.

[0095] Step S2: Access the point and edge data in TiDB that stores the point data and edge data in the graph data based on the underlying direct point and edge object data access mechanism of TiKV. Specifically, it includes:

[0096] For the row-level data stored in TiDB, the Key in TiKV has: tablePrefix{TableID}_recordPrefixSep{RowID}, where tablePrefix and recordPrefixSep are specific string constants;

[0097] Obtain the TableID according to the table information corresponding to the point type or edge type, combine it with the ID of the point or edge, know the Key of the point or edge in the underlying KV storage of TiDB, and directly access the corresponding point and edge data based on the Key of the point or edge.

[0098] This step S2 introduces the underlying direct point and edge object data access mechanism based on TiKV. The specific content is as follows.

[0099] TiDB is a distributed relational database, but its underlying layer is based on TiKV (a distributed KV storage engine). Although TiKV does not directly provide an interface at the SQL level for external applications of TiDB, the characteristics of its open-source database also bring convenience for secondary development. Based on this feature, the present invention introduces the direct access mechanism of TiKV for specific point and edge data (i.e., point and edge data facing a specified ID) to accelerate the read-only access process of graph data.

[0100] Specifically, for the row-level data stored in TiDB, the Key in TiKV has the following format: tablePrefix{TableID}_recordPrefixSep{RowID}, where both tablePrefix and recordPrefixSep are specific string constants. Therefore, by obtaining the TableID according to the table information (table information corresponding to the point type or edge type) and combining it with the ID of the point or edge (i.e., RowID), the Key of the point or edge in the underlying KV storage of TiDB can be known, and thus the corresponding point and edge data can be directly accessed.

[0101] The method of direct access to vertices and edges through TiKV introduced by the present invention avoids the problem of potential row positioning through the primary key when performing row-level data access based on IDs (or RowIDs) in the construction of a graph database usually based on a relational database (such as MySQL). Instead, it directly uses KV storage for the positioning of vertex and edge data, thus improving the efficiency of vertex and edge data access.

[0102] Step S3: Perform large-scale vertex and edge data access to TiDB storing vertex data and edge data in the graph data based on TiFlash columnar storage, specifically including:

[0103] When deploying the TiDB cluster, configure TiFlash as a RAFT Learner. For the update of vertex and edge data in TiDB, TiFlash will obtain a copy of the vertex and edge data and store it in TiFlash in a columnar storage manner;

[0104] When processing a large number of vertices and edges of the whole graph or a subgraph, perform large-scale vertex and edge data access through TiFlash.

[0105] This step S3 introduces a large-scale graph data analysis process based on TiFlash columnar storage, and the specific content is as follows.

[0106] In the analysis scenario of the graph database, there is the following characteristic data access:

[0107] Processing large-scale data, such as processing all vertices and edges of the whole graph;

[0108] Only a limited number of attributes are read during the processing, and it is not necessary to use all the attributes of vertices and edges.

[0109] Considering such characteristics of graph analysis data, the present invention combines the storage of vertex and edge data in the foregoing steps and introduces TiFlash based on TiDB for large-scale vertex and edge data access:

[0110] When deploying the TiDB cluster, configure TiFlash as a RAFT Learner. For the update of vertex and edge data in TiDB, TiFlash will also obtain a copy of the vertex and edge data and store it in TiFlash in a columnar storage manner;

[0111] When performing large-scale graph data analysis (such as PageRank graph calculation, shortest path calculation, etc.), a large number of vertices and edges in the whole graph or subgraph are processed, and only some attributes are involved (for example, PageRank only involves the weight attribute of edges, and the shortest path only involves the path length value attribute). The calculated data is accessed through TiFlash to apply the access efficiency brought by TiFlash column storage to improve the large-scale calculation process on the graph database solution introduced by the present invention.

[0112] Embodiment 2:

[0113] Based on the same inventive concept, the present invention also provides a construction system for a distributed graph database storage layer based on TiDB, including:

[0114] A storage module for storing vertex data and edge data in graph data based on TiDB;

[0115] A direct access module for accessing vertex and edge data in TiDB that stores vertex data and edge data in graph data based on the underlying direct vertex and edge object data access mechanism of TiKV;

[0116] A big data access module for accessing large-scale vertex and edge data in TiDB that stores vertex data and edge data in graph data based on TiFlash columnar storage.

[0117] Optionally, the storage module includes:

[0118] A vertex data storage sub-module for introducing all vertices in each vertex type in graph data into a table in TiDB for storage;

[0119] An edge data storage sub-module for introducing all edges in each edge type in graph data into a table in TiDB for storage.

[0120] Optionally, the vertex data storage sub-module is specifically used for:

[0121] Using the row ID of the TiDB table as the unique ID of each stored vertex data;

[0122] Using the column corresponding to the graph data type in the TiDB table as the attribute of the vertex type, and keeping the attributes of the vertex type, data constraints consistent with the column data types and data constraints in the TiDB table.

[0123] Optionally, the edge data storage sub-module is specifically used for:

[0124] Using the row ID of the TiDB table as the ID of the stored edge;

[0125] Use the columns of the TiDB table as the attributes of the edge types in the graph data, and keep the attributes of the edge types of the graph data, the data constraints, the column data types in the TiDB table, and the data constraints consistent.

[0126] Optionally, the direct access module is specifically used for:

[0127] For the row-level data stored in TiDB, the Key in TiKV has: tablePrefix{TableID}_recordPrefixSep{RowID}, where tablePrefix and recordPrefixSep are specific string constants;

[0128] Obtain the TableID according to the table information corresponding to the point type or edge type, combine the ID of the point or edge, know the Key of the point or edge in the underlying KV storage of TiDB, and directly access the corresponding point and edge data based on the Key of the point or edge.

[0129] Optionally, the big data access module is specifically used for:

[0130] When deploying the TiDB cluster, configure TiFlash as a RAFT Learner. For the update of the point and edge data in TiDB, TiFlash will obtain a copy of the point and edge data and store it in TiFlash in a columnar storage manner;

[0131] When processing a large number of points and edges of the whole graph or subgraph, perform large-scale point and edge data access through TiFlash.

[0132] Embodiment 3:

[0133] Based on the same inventive concept, the present invention further provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of the method for constructing a distributed graph database storage layer based on TiDB in the above embodiments.

[0134] Embodiment 4:

[0135] Based on the same inventive concept, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of the method for constructing a distributed graph database storage layer based on TiDB in the above embodiments.

[0136] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0137] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0138] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0140] The above are only embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention pending approval.

Claims

1. Method for constructing storage layer of distributed graph database based on TiDB, Characterized in that, Comprising: Storing vertex data and edge data in graph data based on TiDB; Performing vertex and edge data access on TiDB storing vertex data and edge data in graph data based on underlying direct vertex and edge object data access mechanism of TiKV; Performing large-scale vertex and edge data access on TiDB storing vertex data and edge data in graph data based on TiFlash columnar storage.

2. The method according to claim 1, Characterized in that, The storing vertex data and edge data in graph data based on TiDB comprises: Introducing all vertices in each vertex type in graph data into a table in TiDB for storage; Introducing all edges in each edge type in graph data into a table in TiDB for storage.

3. The method according to claim 2, Characterized in that, The introducing all vertices in each vertex type in graph data into a table in TiDB for storage comprises: Using row ID of TiDB table as unique ID of each stored vertex data; Using columns corresponding to graph data types in TiDB table as attributes of vertex type, and keeping attributes of vertex type, data constraints consistent with column data types and data constraints in TiDB table.

4. The method according to claim 2, Characterized in that, The introducing all edges in each edge type in graph data into a table in TiDB for storage comprises: Using row ID of TiDB table as ID of stored edges; Using columns of TiDB table as attributes of edge type in graph data, and keeping attributes of edge type of graph data, data constraints consistent with column data types and data constraints in TiDB table.

5. The method according to claim 1, Characterized in that, The performing vertex and edge data access on TiDB storing vertex data and edge data in graph data based on underlying direct vertex and edge object data access mechanism of TiKV comprises: For row-level data stored in TiDB, Key in TiKV has: tablePrefix{TableID}_recordPrefixSep{RowID}, where tablePrefix and recordPrefixSep are specific string constants; Obtaining TableID according to table information corresponding to vertex type or edge type, combining with ID of vertex or edge, knowing Key of vertex or edge in underlying KV storage of TiDB, and directly accessing corresponding vertex and edge data based on the Key of the vertex or edge.

6. The method according to claim 1, Characterized in that, The performing large-scale vertex and edge data access on TiDB storing vertex data and edge data in graph data based on TiFlash columnar storage comprises: When deploying TiDB cluster, configuring TiFlash as RAFT Learner, for updates of vertex and edge data in TiDB, TiFlash will obtain a copy of vertex and edge data and store it in TiFlash in columnar storage manner; When processing a large number of points and edges in the whole graph or sub-graph, large-scale point and edge data access is performed through TiFlash.

7. A system for constructing a storage layer of a distributed graph database based on TiDB Characterized in that It includes: A storage module for storing point data and edge data in the graph data based on TiDB; A direct access module for performing point and edge data access on TiDB that stores point data and edge data in the graph data based on the underlying direct point and edge object data access mechanism of TiKV; A big data access module for performing large-scale point and edge data access on TiDB that stores point data and edge data in the graph data based on TiFlash columnar storage.

8. The system according to claim 7 Characterized in that The storage module includes: A point data storage sub-module for introducing all points in each point type in the graph data into a table in TiDB for storage; An edge data storage sub-module for introducing all edges in each edge type in the graph data into a table in TiDB for storage.

9. A computer device Characterized in that It includes: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, the method for constructing a storage layer of a distributed graph database based on TiDB according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium Characterized in that A computer program is stored thereon, and when the computer program is executed, the method for constructing a storage layer of a distributed graph database based on TiDB according to any one of claims 1 to 6 is implemented.