Adjacent edge multiple storage method and system of distributed graph database
By introducing one-fold redundant storage of edges and centralized storage of adjacent edges into the distributed graph database, the problem of low query efficiency caused by cross-partition join operation is solved, and fast adjacent data acquisition of edges and points is achieved, which improves the query performance of graph database.
Patent Information
- Application Number
- CN202311873595.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In a distributed graph database system based on relational data storage, the join operation of table data may span multiple partition servers, affecting the query traversal efficiency of graph data, especially when there are a large number of adjacent edges at some points.
In the distributed graph data environment on the embedded relational data engine, a one-fold redundant storage of edges and a centralized overall storage field of adjacent edges are introduced. The edge information is quickly queried through one-fold redundant storage of edges, and a centralized overall storage field of adjacent edges is introduced for each point to quickly query the adjacent edge data of the point.
Through redundant edge data storage processing, the cross-partition storage problem of table data corresponding to each other is accelerated by the dual adjacency edge data storage mechanism, the adjacency edge acquisition operation of points is improved, and the query performance of the graph database is improved.
Smart Images

Figure CN120234342A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graph databases, and particularly to a method and system for multi-storing adjacent edges of a distributed graph database. Background Art
[0002] With the explosive development of the Internet, mobile Internet, social networks, Internet of Things, and industrial field-related networks such as power networks, there is a great demand for the storage of relationship graphs and applications such as network topology analysis and functional analysis based on relationship graphs, which has also contributed to the research and development boom of graph databases.
[0003] A graph database is a data management system based on vertices and edges as basic storage units, designed to efficiently store and query graph data. The vertices and edges in a graph database have attributes, that is, data associated with a certain vertex or edge in the form of key-value pairs. When representing the attributes of a graph database, the attribute values of vertices or edges can exist in a certain data type (such as integer, string, double precision number, etc.).
[0004] Based on a relational database system (such as MySQL, PostgreSQL, Oracle, etc.), it is a technical solution for the storage system of a graph database, that is, the point or edge Schema in the graph database is implemented through the tables of the relational database, and a specific point or edge in the graph database is represented by a row of data in the relational database.
[0005] However, in this actual technology, when performing some conventional operations on a graph database, such as querying adjacent vertices of a specific type through a certain vertex or performing a BFS (Breath First Search) operation, there may be performance problems, because this operation of "finding adjacent edges through a certain vertex" and then finding adjacent vertices, or performing a BFS traversal, is actually a "Join operation" from a relational table corresponding to a certain type of vertex and another table corresponding to edges.
[0006] Generally, compared with the number of vertex data, the number of edges is larger, and there may be a situation where some vertices have a large number of adjacent edges. One scenario is: in some social network environments (such as Weibo), the relationship of one person "following" another person is represented by an edge. Some vertices (such as big Vs in Weibo - that is, VIP users with a large number of followers) will have a huge number of edges (some big Vs may have up to millions or more followers). At the same time, the operation of finding adjacent edges through vertices and then finding adjacent vertices in a graph database is a basic operation of graph traversal operations. Therefore, the high cost of such adjacent data access operations will ultimately affect the overall performance of the graph database.
[0007] Furthermore, in a distributed graph database system based on relational data storage, this problem may be more obvious because the relational database tables that support the node and edge data of the graph may be stored on different partitions. Therefore, the Join operation of the table data may span multiple partition servers, further affecting the query traversal efficiency of the graph data. Summary of the Invention
[0008] To solve the problem in the prior art that in a distributed graph database system based on relational data storage, the Join operation of the table data may span multiple partition servers, further affecting the query traversal efficiency of the graph data, the present invention proposes a method for storing adjacent edges multiple times in a distributed graph database, including:
[0009] Introduce a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly query edge information through the single redundancy storage of edges;
[0010] In a distributed graph data environment on an embedded relational data engine, introduce a centralized overall storage field for adjacent edges for each node, and quickly query the adjacent edge data of the node through the centralized overall storage field of the adjacent edges.
[0011] Optionally, the introducing a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine includes:
[0012] On the partition server where the end point of the edge is located, introduce row data storage for the same edge;
[0013] For the edge data stored in the embedded relational engine in the start point partition and the redundant edge data in the end point partition, establish a B+ tree index facing the node ID.
[0014] Optionally, the establishing a B+ tree index facing the node ID for the edge data stored in the embedded relational engine in the start point partition and the redundant edge data in the end point partition includes:
[0015] Establish a B+ tree index corresponding to the start point ID for the edge data stored in the start point partition;
[0016] Establish a B+ tree index corresponding to the end point ID for the redundant edge data stored in the end point partition.
[0017] Optionally, the quickly querying edge information through the single redundancy storage of edges includes:
[0018] According to the start point ID of the edge to be queried, query the edge data corresponding to the start point ID in the B+ tree index corresponding to the start point ID;
[0019] Query the edge data corresponding to the end point ID in the B+ tree index corresponding to the end point ID according to the end point ID of the edge to be queried.
[0020] Optionally, introducing a centralized overall storage field for adjacent edges for each vertex in the distributed graph data environment on the embedded relational data engine includes:
[0021] Introduce an internal column of the graph database in the row-level storage of each vertex, and store the edge ID and the ID of the adjacent vertex connected to the vertex by the edge in binary form;
[0022] Introduce two other internal columns in the row-level storage of each vertex. One internal column records the last update timestamp of the "all adjacent information" internal column, and the other internal column records the timestamp of the change of the "adjacent edges" of the vertex.
[0023] Optionally, the fast query of the adjacent edge data of the vertex through the centralized overall storage field of the adjacent edges includes:
[0024] First, check the update timestamp and the changed timestamp. When the timestamps match, query the adjacent information internal column of the vertex to obtain the data information of all adjacent edges of the vertex;
[0025] If the time does not match, obtain the adjacent edges and adjacent vertex information of the edge through the edge data of the current partition and / or the redundant edge data table through indexing;
[0026] Judge whether the obtained adjacent edge and adjacent vertex information of the edge meet the query requirements, and determine the query result based on the judgment result.
[0027] Optionally, the judging whether the obtained adjacent edge and adjacent vertex information of the edge meet the query requirements, and determining the query result based on the judgment result includes:
[0028] When the obtained adjacent edge information and adjacent vertex information already meet the query requirements, use the adjacent edge information and adjacent vertex information as the query result;
[0029] When the obtained adjacent edge information and adjacent vertex information do not meet the query requirements, connect to the corresponding partition server for further query according to the obtained adjacent vertices and edge data, and use the adjacent edge information, adjacent vertex information and the further query result as the query result.
[0030] On the other hand, the present application also provides an adjacent edge multiple storage system for a distributed graph database, including:
[0031] An edge redundant storage module for introducing a first-level redundant storage of edges in the distributed graph data environment on the embedded relational data engine, and quickly querying edge information through the first-level redundant storage of edges;
[0032] A point redundancy storage module, which is used to introduce a centralized overall storage field for adjacent edges for each point in a distributed graph data environment on an embedded relational data engine, and quickly query the adjacent edge data of the point through the centralized overall storage field of the adjacent edges.
[0033] Optionally, the edge redundancy storage module includes:
[0034] An edge redundancy introduction sub-module, which is used to introduce row data storage for the same edge on the partition server where the end point of the edge is located; and establish a B+ tree index for the point ID for the edge data stored in the embedded relational engine of the start partition and the edge redundancy data of the end partition.
[0035] An edge query sub-module, which is used to query the edge data corresponding to the start ID in the B+ tree index corresponding to the start ID according to the start ID of the edge to be queried; or query the edge data corresponding to the end ID in the B+ tree index corresponding to the end ID according to the end ID of the edge to be queried.
[0036] Optionally, the specific implementation steps of establishing a B+ tree index for the point ID for the edge data stored in the embedded relational engine of the start partition and the edge redundancy data of the end partition in the edge redundancy introduction sub-module include:
[0037] Establish a B+ tree index corresponding to the start ID for the edge data stored in the start partition.
[0038] Establish a B+ tree index corresponding to the end ID for the redundant edge data stored in the end partition.
[0039] Optionally, the point redundancy storage module includes:
[0040] A point redundancy introduction sub-module, which is used to introduce internal columns of the graph database in the row-level storage of each point, store the edge ID and the ID of the adjacent point connected by the edge in binary form; and introduce two other internal columns in the row-level storage of each point, one of which records the last update timestamp of the "all adjacent information" internal column, and the other records the timestamp of the change of the "adjacent edges" of the point.
[0041] A point redundancy query sub-module, which is used to first check the update timestamp and the changed timestamp. When the timestamps match, query the adjacent information internal column of the point to obtain the data information of all adjacent edges of the point; if the time does not match, obtain the adjacent edges and adjacent point information of the edge through the edge data and / or redundant edge data table of the current partition through the index; judge whether the obtained adjacent edges and adjacent point information of the edge meet the query requirements, and determine the return result based on the judgment result.
[0042] Optionally, the specific implementation steps of the point redundancy query sub-module for determining whether the obtained adjacent edge and adjacent point information of the edge meet the query requirements and determining the query result based on the judgment result include:
[0043] When the obtained adjacent edge information and adjacent point information already meet the query requirements, use the adjacent edge information and adjacent point information as the query result;
[0044] When the obtained adjacent edge information and adjacent point information do not meet the query requirements, connect to the corresponding partition server for further query according to the data of the adjacent points and edges obtained by the query, and use the adjacent edge information, adjacent point information and the further query result as the query result.
[0045] On the other hand, the present application also provides a computing device, including: one or more processors;
[0046] The processor is used to execute one or more programs;
[0047] When the one or more programs are executed by the one or more processors, the method for storing adjacent edges multiple times in a distributed graph database as described above is implemented.
[0048] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the method for storing adjacent edges multiple times in a distributed graph database as described above is implemented.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] The present invention provides a method for storing adjacent edges multiple times in a distributed graph database, including: introducing a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly querying edge information through the single redundancy storage of edges; introducing a centralized overall storage field for adjacent edges for each point in the distributed graph data environment on the embedded relational data engine, and quickly querying the adjacent edge data of the point through the centralized overall storage field of adjacent edges. The present invention introduces redundant stored edge data so that each point can obtain the stored data of the edge from its corresponding partition. The present invention introduces a centralized storage of adjacent edges for each point so that the ID information of all edges of this point can be directly obtained from each relational data row corresponding to the point.
[0051] The present invention processes the problem of cross-partition storage of table data corresponding to edges through redundant edge data storage, and accelerates the operation of obtaining adjacent edges of points through a dual adjacent edge data storage mechanism. Description of the Drawings
[0052] Figure 1Flowchart of an adjacent edge multiple storage method for a distributed graph database of the present invention;
[0053] Figure 2 Schematic diagram of the adjacent edge multiple storage method for a distributed graph database in the embodiment of the present invention. Detailed implementation manners
[0054] The present invention proposes an adjacent edge multiple storage method for a distributed graph database, which processes the problem of cross-partition storage of table data corresponding to edges through redundant edge data storage, and accelerates the operation of obtaining adjacent edges of a point through a dual adjacent edge data storage mechanism.
[0055] Embodiment 1:
[0056] An adjacent edge multiple storage method for a distributed graph database, as Figure 1 shown, includes:
[0057] Step S1: Introduce a first-level redundant storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly query edge information through the first-level redundant storage of edges;
[0058] Step S2: Introduce a centralized overall storage field for adjacent edges for each point in a distributed graph data environment on an embedded relational data engine, and quickly query the adjacent edge data of the point through the centralized overall storage field for adjacent edges.
[0059] The adjacent edge multiple storage method introduced by the present invention is based on an embedded relational engine, and introduces a first-level redundant data storage for edge data, that is, redundant stored edge data is respectively introduced corresponding to the starting point and the ending point of the edge stored in different partitions, so that each point can obtain the stored data of the edge from the partition where it is located. In addition to edge data, the present invention introduces a centralized storage of adjacent edges for each point, that is, the "dual edge data" centralized for each point in addition to edge data, so that the ID information of all edges of this point can be directly obtained from each relational data row corresponding to the point. This combination of the first-level adjacent edge data storage and the dual edge data of the point accelerates the operation of obtaining adjacent edges in a distributed graph environment.
[0060] The present invention stores graph data based on an embedded relational data engine. First, the storage background of the embedded relational engine is introduced, and its technical characteristics are as follows:
[0061] Distributed graph data is partitioned for point data through a partitioning strategy (such as through common Hash partitioning), and the data of points is distributed and stored on different partition servers according to this strategy. Generally, the data of an edge and the starting point of this edge are stored on the same partition, but this is only part of the data, and the further storage of the edge data will be further described in the implementation steps of the subsequent technical solution.
[0062] Corresponding to the partition service area in the graph database instance, the underlying storage of each partition provides the storage capacity of graph data through an independent relational data engine running in the kernel of the graph data partition server.
[0063] Each relational data engine only serves a single partition server, and there is no interaction between these relational storage engines, which are independent of each other. This is different from the common way of using a single integrated relational data engine to serve a graph server.
[0064] These relational engines only provide the storage and retrieval (i.e., writing and reading) operations of the graph data in the partition where they are located. Data synchronization, backup, query logic, etc. across partitions are all handled by the graph database kernel.
[0065] All these embedded relational data engines as a whole form the storage support layer of the graph data.
[0066] In such a storage architecture based on embedded relational engines, a single embedded relational engine provides the storage function of the graph data (vertices, edges, attributes) on this partition in the form of a native table.
[0067] Generally speaking, the vertices and edges within the partition are stored through the native tables of the embedded relational engine:
[0068] Vertices are stored in the table of the embedded relational engine, and their ID is the global ID of the distributed graph database, which also serves as the primary key of the table in the embedded relational engine. The attributes of the vertices are stored in the columns of the table in the embedded relational engine.
[0069] Edges are also stored through the table of the embedded relational engine. The ID of the edge is the global ID of the distributed graph database and exists as the primary key of the embedded relational engine. At the same time, the data of the table in the embedded relational engine corresponding to each edge also stores the starting ID and ending ID of the edge. The attributes of the edge are stored in the columns of the table in the embedded relational engine.
[0070] Based on the above storage background of the embedded relational engine, the following Figure 2 describes the adjacent edge multiple storage method of a distributed graph database of the present invention step by step.
[0071] Step S1: Introduce a single redundancy storage of edges in the distributed graph data environment on the embedded relational data engine, and quickly query edge information through the single redundancy storage of edges, specifically including:
[0072] Introduce a single redundancy storage of edges in the distributed graph data environment based on the embedded relational data engine.
[0073] Generally speaking, in an analytical database such as a graph database, the operations of edge data have the following characteristics:
[0074] The topological structure information of the edge (i.e., the information of the connected start point and end point) belongs to the operation of writing less and reading more. After the edge is created, this connection information of the edge is rarely changed. Therefore, the graph database has operations of creating and deleting edges, but there is no operation of modifying the "connection" of the edge (i.e., modifying the two points connected by a certain edge). However, at the same time, the connection information of the edge is widely used in topological traversal operations (such as the aforementioned BFS traversal).
[0075] The attributes of the edge belong to the data that can be changed. Therefore, "attribute change" is also one of the basic operations of the graph database.
[0076] As mentioned above, when storing point data, the point storage is partitioned according to the corresponding partitioning strategy, and the edge data is stored in the start point partition. Considering the operation characteristics of the edge connection information and attribute information, the present invention introduces a redundant storage for the edge topological information and introduces a point adjacent edge index on top of the embedded relationship engine:
[0077] On the partition server where the end point of the edge is located, row data storage is also introduced for the same edge. However, this row-level storage of "end point edge data" is a simplified storage of edge data, which only stores the ID of the edge and the start point ID and end point ID of the edge, without storing the attribute data of the edge, so as to reduce the data space required for storage. Therefore, the topological structure of the edge (i.e., the edge ID, the start point ID and end point ID of the edge) is redundantly stored;
[0078] For the edge data stored in the embedded relationship engine of the start point partition and the edge redundant data of the end point partition, a B+ tree index is established for the edge data facing the point ID:
[0079] A B+ tree index corresponding to the start point ID is established for the edge data stored in the start point partition. That is, through the index of the embedded relationship engine, the corresponding edge data can be quickly obtained according to the start point ID, avoiding the global scan operation of the edge data during the adjacent edge query;
[0080] Similarly, a B+ tree index corresponding to the end point ID is established for the redundant edge data stored in the end point partition;
[0081] It can be seen that through the above operations:
[0082] For the point data, in the partition where the point data is located, all the adjacent edge information of this point, including the edge ID and the ID of the adjacent point of the point corresponding to this edge, can always be obtained at one time through the embedded relationship engine inside the partition;
[0083] And this operation of the point obtaining the adjacent edge is all carried out through the index of the partition embedded engine;
[0084] The above operations avoid cross-partition access to the adjacent edge data of points and also avoid global scans of the relationship table for acquisition.
[0085] Step S2: In a distributed graph data environment on an embedded relational data engine, introduce a centralized overall storage field for adjacent edges for each point, and quickly query the adjacent edge data of the point through the centralized overall storage field of the adjacent edges, specifically as follows:
[0086] Introduce asynchronous dual-redundant adjacent edge direct storage for point data, that is, introduce a centralized overall storage field for adjacent edges for each individual point.
[0087] Based on the row-based edge data redundant storage introduced in Step S1, the present invention introduces centralized adjacent edge data storage for each point, that is, introduce an independent field in the row data where each point is located, specifically for storing the adjacent edge information of this point. Specifically:
[0088] Introduce an internal column in the row-level storage of each point (i.e., a column invisible to user applications), and store it in binary format (i.e., a BLOB type column in the relational table);
[0089] Therefore, each point is a column field that stores all adjacent edge data, including the edge ID and the IDs of adjacent points connected to this point through the edge.
[0090] Through the edge adjacency information stored in the internal column of the point here, all adjacent edge data of a specific point can be obtained at one time. Compared with the data in the first-level edge table, the internal column method of the point here avoids the process of obtaining multiple rows of data in the sub-table (even through the index in Step 1, it is actually multiple rows of data), and obtains the adjacent edge data of the edge at one time.
[0091] However, this "internal column" method stores all adjacent information of a point in one column, which is extremely inconvenient in the case of changes to the association relationship of the point (such as a point deleting an adjacent edge or adding a new adjacent edge). Therefore, when there are additions or deletions of edges, only the data table of the edge (and the redundant data table of the edge) is maintained, and the data in the dual-adjacent data internal column is not synchronously maintained. The dual-adjacent data column is maintained asynchronously:
[0092] In addition to the internal column of all adjacent information of the point, introduce two other internal columns: one internal column records the last update timestamp of the "all adjacent information" internal column (referred to as the adjacent information update timestamp), and the second internal column records the timestamp of changes to the "adjacent edges" of the point (such as adjacent edge addition, deletion) (referred to as the adjacent edge change timestamp);
[0093] When performing edge addition and deletion operations, update the adjacent edge change timestamps of the two points (starting point and ending point) corresponding to the edge;
[0094] When querying the adjacency information of a point, if the two timestamps are the same, read the internal columns of the adjacency information to obtain the adjacency information of the point;
[0095] If the two timestamps are the same, read the edge data table (and the edge redundancy data table) to obtain the adjacency data of the point;
[0096] While obtaining the adjacent edges of a certain point through the edge data table, or updating the adjacency data in the "adjacency information" internal column of this point through a background thread / process.
[0097] Through the above process, the data in the internal column of the adjacency data of the point introduced in this step is maintained asynchronously. When the double adjacency data is not updated in time (i.e., the above timestamps do not match), obtain the adjacency information of the point through the data in the edge table instead of the double adjacency data.
[0098] Based on the above process of querying the adjacent edges of a point under the double adjacency edge storage mechanism.
[0099] When introducing a multiple adjacency edge data storage mechanism, the process of querying the adjacency information of a certain point is as follows:
[0100] First, check the two timestamps. When the timestamps match, query the internal column of the adjacency information of the point to obtain the data information of all the adjacent edges of the point;
[0101] If the time does not match, then according to the need, through the edge data of the current partition (for the edges starting from the points in this partition) and / or the redundant edge data (for the edges ending at the points stored in this partition) table, obtain the adjacent edges and adjacent point information of this edge through indexing;
[0102] If the obtained adjacent edge information and adjacent point information already meet the query requirements (such as only querying the adjacent edges or adjacent point IDs of the point), then directly return the query result;
[0103] If there are further query and traversal requirements (such as needing the attribute information of the adjacent edges or points, or entering the next layer of adjacent points and edge queries), then accordingly, based on the data of the adjacent points and edges obtained from the query, connect to the corresponding partition server for query.
[0104] It can be seen that according to the adjacency data storage and query methods introduced by these steps:
[0105] For the points in the graph database, the direct query of the adjacent data points and edges only needs to be performed on the current partition service, and no cross-server operations are required;
[0106] The redundant storage of the first-level adjacency information of the edges in the page partition ensures that the query of the adjacency data of points does not require crossing the partition server.
[0107] For a graph database with more reads than writes, the second-level adjacency information of points will be directly queried through the internal adjacency information column, which is an operation to obtain all data at once and has performance advantages. At the same time, the existence of the row-level data of the edges in the edge table brings convenience to the maintenance of edge data (operations of adding and deleting edges).
[0108] Generally speaking, the multiple adjacency information storage method introduced by the embedded partition of the present invention for storage partitions, with the first-level adjacency information cooperating with the asynchronously maintained second-level adjacency information, provides the query of point adjacency information for the partition server, and also provides a balance between the support for edge maintenance operations and the query of point adjacency information. Generally, it brings a performance improvement to the query of the common point adjacency information in graph data.
[0109] Embodiment 2:
[0110] Based on the same inventive concept, the present invention also provides an adjacency edge multiple storage system for a distributed graph database, including:
[0111] An edge redundant storage module, used to introduce a first-level redundant storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly query edge information through the first-level redundant storage of edges;
[0112] A point redundant storage module, used to introduce a centralized overall storage field of adjacent edges for each point in a distributed graph data environment on an embedded relational data engine, and quickly query the adjacent edge data of points through the centralized overall storage field of adjacent edges.
[0113] Furthermore, the edge redundant storage module includes:
[0114] An edge redundancy introduction sub-module, used to introduce row data storage for the same edge on the partition server where the end point of the edge is located; and establish a B+ tree index for the edge data stored in the embedded relational engine of the start partition and the edge redundant data of the end partition, facing the point ID;
[0115] An edge query sub-module, used to query the edge data corresponding to the start ID in the B+ tree index corresponding to the start ID of the edge to be queried according to the start ID of the edge to be queried; or query the edge data corresponding to the end ID in the B+ tree index corresponding to the end ID of the edge to be queried according to the end ID of the edge to be queried.
[0116] Furthermore, the specific implementation steps of establishing a B+ tree index for the edge data stored in the embedded relational engine of the start partition and the edge redundant data of the end partition in the edge redundancy introduction sub-module include:
[0117] Create a B+ tree index corresponding to the starting point ID for the edge data stored in the starting point partition;
[0118] Create a B+ tree index corresponding to the ending point ID for the redundant edge data stored in the ending point partition.
[0119] Furthermore, the point redundancy storage module includes:
[0120] A point redundancy introduction sub-module, which is used to introduce internal columns of the graph database in the row-level storage of each point, store the edge ID and the ID of the adjacent point connected to the point by the edge in binary form; and introduce two other internal columns in the row-level storage of each point, where one internal column records the last update timestamp of the "all adjacent information" internal column, and the other internal column records the timestamp of the change of the "adjacent edges" of the point;
[0121] A point redundancy query sub-module, which is used to first check the update timestamp and the changed timestamp. When the timestamps match, query the adjacent information internal column of the point to obtain the data information of all adjacent edges of the point; if the time does not match, obtain the adjacent edges and adjacent point information of the edge through the edge data and / or redundant edge data table of the current partition through the index; determine whether the obtained adjacent edge and adjacent point information of the edge meet the query requirements, and determine the return result based on the judgment result.
[0122] Furthermore, the specific implementation steps of the point redundancy query sub-module for determining whether the obtained adjacent edge and adjacent point information meet the query requirements and determining the query result based on the judgment result include:
[0123] When the obtained adjacent edge information and adjacent point information already meet the query requirements, use the adjacent edge information and adjacent point information as the query result;
[0124] When the obtained adjacent edge information and adjacent point information do not meet the query requirements, connect to the corresponding partition server for further query according to the adjacent points and edge data obtained by the query, and use the adjacent edge information, adjacent point information and further query result as the query result.
[0125] Example 3:
[0126] Based on the same inventive concept, the present invention further provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for adjacent edge multiple storage of a distributed graph database in the above embodiments.
[0127] Embodiment 4:
[0128] Based on the same inventive concept, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of a method for adjacent edge multiple storage of a distributed graph database in the above embodiments.
[0129] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0130] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0133] The above are only embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the scope of the claims of the present invention pending approval.
Claims
1. A method for multiple storage of adjacent edges in a distributed graph database, characterized in that, Including: Introduce a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly query edge information through the single redundancy storage of edges; In a distributed graph data environment on an embedded relational data engine, introduce a centralized overall storage field for adjacent edges for each vertex, and quickly query the adjacent edge data of the vertex through the centralized overall storage field of adjacent edges.
2. The method according to claim 1, wherein The introducing of a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine includes: On the partition server where the end point of the edge is located, introduce row data storage for the same edge; For the edge data stored in the embedded relational engine in the start partition and the edge redundancy data in the end partition, establish a B+ tree index facing the vertex ID.
3. The method according to claim 2, wherein The establishing of a B+ tree index facing the vertex ID for the edge data stored in the embedded relational engine in the start partition and the edge redundancy data in the end partition includes: Establish a B+ tree index corresponding to the start ID for the edge data stored in the start partition; Establish a B+ tree index corresponding to the end ID for the redundant edge data stored in the end partition.
4. The method according to claim 1, wherein The quickly querying of edge information through the single redundancy storage of edges includes: According to the start ID of the edge to be queried, query the edge data corresponding to the start ID in the B+ tree index corresponding to the start ID; According to the end ID of the edge to be queried, query the edge data corresponding to the end ID in the B+ tree index corresponding to the end ID.
5. The method according to claim 1, wherein The introducing of a centralized overall storage field for adjacent edges for each vertex in a distributed graph data environment on an embedded relational data engine includes: Introduce internal columns of the graph database in the row-level storage of each vertex, and store the edge ID and the ID of the adjacent vertex connected by the edge to the vertex in binary form; Introduce two other internal columns in the row-level storage of each vertex, where one internal column records the last update timestamp of the "all adjacent information" internal column, and the other internal column records the timestamp of the change of the "adjacent edges" of the vertex.
6. The method according to claim 1, wherein The quickly querying of the adjacent edge data of the vertex through the centralized overall storage field of adjacent edges includes: First, check the update timestamp and the change timestamp. When the timestamps match, query the adjacent information internal column of the vertex to obtain the data information of all adjacent edges of the vertex; If the time does not match, then through the edge data and / or redundant edge data tables in the current partition, obtain the adjacent edges and adjacent vertex information of the edge through the index; Judge whether the obtained adjacent edges and adjacent vertex information of the edge meet the query requirements, and determine the query result based on the judgment result.
7. The method according to claim 6, wherein The judging whether the obtained adjacent edges and adjacent vertex information of the edge meet the query requirements and determining the query result based on the judgment result includes: When the obtained adjacent edge information and adjacent vertex information already meet the query requirements, use the adjacent edge information and adjacent vertex information as the query result; When the obtained adjacent edge information and adjacent vertex information do not meet the query requirements, connect to the corresponding partition server for further query according to the adjacent vertices and edge data obtained by the query, and use the adjacent edge information and adjacent vertex information and the further query result as the query result.
8. An adjacency edge multiple storage system for a distributed graph database, characterized in that, Including: An edge redundant storage module is used to introduce a single redundancy storage of edges in a distributed graph data environment on an embedded relational data engine, and quickly query edge information through the single redundancy storage of edges; A vertex redundant storage module is used to introduce a centralized overall storage field of adjacent edges for each vertex in a distributed graph data environment on an embedded relational data engine, and quickly query the adjacent edge data of the vertex through the centralized overall storage field of adjacent edges.
9. A computer device, characterized in that, Comprising: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, a method for multiple storage of adjacent edges of a distributed graph database as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, There is a computer program stored thereon, and when the computer program is executed, a method for multiple storage of adjacent edges of a distributed graph database as described in any one of claims 1 to 7.