Graph database system and graph query method
By adjusting the SSTable file format and index method in the key-value storage system, the perceived storage of graph data is realized, the problem of low graph query performance in the prior art is solved, and the efficiency of graph query is improved.
Patent Information
- Application Number
- CN202510278070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing graph storage system built on key-value storage systems is not perceived by graph data, resulting in low graph query performance.
By adjusting and improving the SSTable file format and indexing method in the key-value storage system, it can be perceived by graphs, including organizing nodes and edges into key-value pairs and aggregating storage, and optimizing the storage structure using graph topology and query characteristics.
It reduces the complexity of graph query operations, improves graph query efficiency, and ensures that the graph storage engine has high query performance.
Smart Images

Figure CN120296208A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present application relate to the technical field of storage systems, and in particular, to a graph database system and a graph query method. Background Art
[0002] A graph is a data structure that stores and manages data in a graphical structure. In a graph, nodes (Node), edges (Edge), and properties (Property) are used to store data. This storage method is very suitable for expressing complex relationships between entities. In a graph, nodes represent entities such as people, places, events, etc., and each node can have multiple properties to describe the specific information of the entity. Edges are used to represent the relationships between nodes, such as "know", "belong to", "be located at", etc., and edges can also contain properties to describe the characteristics of the relationships, such as the strength of the relationship, the establishment time, etc. Properties are data fields attached to nodes or edges to store specific information, such as a person's name, age, or the strength of the relationship, the establishment time, etc.
[0003] Graphs are very suitable for expressing and processing data with complex relationships and can therefore usually provide relatively high query performance. Correspondingly, for files used to store graph data such as nodes, edges, and properties, how to design the format of such files so that these files are graph-aware, thereby being able to retain the relatively high query performance of the graph and thus improving the efficiency of graph queries in such files has become a much-concerned issue. Summary of the Invention
[0004] One or more embodiments of the present application provide the following technical solutions:
[0005] The present application provides a graph database system for storing graph data, where the graph data includes nodes and edges; wherein:
[0006] An SSTable file for storing sorted string tables is stored in the storage space of the graph database system; the SSTable file includes at least one first data block for storing key-value pairs corresponding to the nodes and the edges;
[0007] The key-value pairs include first key-value pairs and second key-value pairs; the key in the first key-value pair includes the node identifier of the node, and the value includes the node attributes of the node; the key in the second key-value pair includes the start point identifier and the end point identifier of the edge, and the value includes the edge attributes of the edge;
[0008] The first data block aggregately stores each target first key-value pair and a target second key-value pair whose start point identifier is the same as the node identifier in the target first key-value pair.
[0009] The present application also provides a graph query method; wherein, nodes and edges in the graph are stored in a sorted string table (SSTable) file; the SSTable file includes at least one first data block for storing key-value pairs corresponding to the nodes and the edges; the nodes and the edges are respectively organized into a first key-value pair and a second key-value pair and stored in the first data block; the key in the first key-value pair includes the node identifier of the node, and the value includes the node attributes of the node; the key in the second key-value pair includes the start identifier and the end identifier of the edge, and the value includes the edge attributes of the edge; each target first key-value pair, and a target second key-value pair whose start identifier is the same as the node identifier in the target first key-value pair, are aggregated and stored in the first data block;
[0010] The method includes:
[0011] Obtaining a target key for querying the graph;
[0012] Querying in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier;
[0013] Querying in the target first data block based on the target key to determine a target value corresponding to the target key, and determining a graph query result based on the target value.
[0014] The present application also provides an electronic device, including:
[0015] A processor;
[0016] A memory for storing instructions executable by the processor;
[0017] Wherein, the processor realizes the steps of the method as described in any one of the above by running the executable instructions.
[0018] The present application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above are realized.
[0019] In the above technical solution, the nodes and edges in the graph can be stored in an SSTable file, which can include at least one first data block for storing key-value pairs corresponding to the nodes and edges; the nodes and edges can be respectively organized into a first key-value pair and a second key-value pair and stored in the first data block. Among them, the key in the first key-value pair corresponding to a node can include the node identifier of the node, and the value can include the node attributes of the node. The key in the second key-value pair corresponding to an edge can include the start point identifier and the end point identifier of the edge, and the value can include the edge attributes of the edge; each target first key-value pair and the target second key-value pair whose start point identifier is the same as the node identifier in the target first key-value pair can be aggregated and stored in the first data block. Accordingly, when obtaining a target key for querying the graph, first, based on the target node identifier in the target key, query in the SSTable file to determine the target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier, and then, based on the target key, query in the target first data block to determine the target value corresponding to the target key, and determine the graph query result based on the target value.
[0020] By adopting the above method, the optimization of the key-value storage system is realized by using the characteristics of graph topology and graph query, specifically including adjusting and improving aspects such as the SSTable file format and indexing method in the key-value storage system, so that the SSTable file for storing graph data in the key-value storage system is graph-aware, thereby reducing the complexity of the graph query operation of the graph storage engine based on the key-value storage system in the key-value storage system, improving the graph query efficiency, and thus ensuring that the graph storage engine has relatively high graph query performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The following will describe the drawings required for the description of the exemplary embodiments, where:
[0022] Figure 1 is a schematic diagram of a key-value storage system shown in an exemplary embodiment of the present application.
[0023] Figure 2 is a schematic diagram of a graph storage file shown in an exemplary embodiment of the present application.
[0024] Figure 3 is a flowchart of a graph query method shown in an exemplary embodiment of the present application.
[0025] Figure 4 is a schematic diagram of the structure of a device shown in an exemplary embodiment of the present application.
[0026] Figure 5 is a block diagram of a graph query device shown in an exemplary embodiment of the present application. Detailed Implementation Modes
[0027] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation modes described in the following exemplary embodiments do not represent all implementation modes consistent with one or more embodiments of the present application. On the contrary, they are merely examples consistent with some aspects of one or more embodiments of the present application.
[0028] It should be noted that in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in the present application. In some other embodiments, the steps included in the method may be more or less than those described in the present application. In addition, a single step described in the present application may be decomposed into multiple steps for description in other embodiments; and multiple steps described in the present application may also be combined into a single step for description in other embodiments.
[0029] A storage system generally refers to the combination of hardware and software in the entire computer system for storing data. It includes not only physical storage media such as hard disk drives (HDDs), solid-state drives (SSDs), tape libraries, etc., but also file systems for managing these media, storage services at the operating system level, etc. In a database environment, the storage system is responsible for providing persistent storage space, ensuring that data can be safely and reliably stored on physical media and can be accessed efficiently.
[0030] A storage engine is a component in a database management system (DBMS) that is responsible for managing how data is stored, retrieved, and updated. It determines the physical storage structure of data on disk and supports functions such as transaction processing and concurrency control.
[0031] The storage engine depends on the underlying storage system to implement actual data read and write operations. Without the physical storage capabilities provided by the storage system, the storage engine cannot complete the persistent storage of data. Moreover, the storage engine performs performance optimization by understanding or controlling the characteristics of the storage system; for example, some storage engines may adjust their index structures or caching policies according to the characteristics of the storage medium to improve data access efficiency.
[0032] The storage system can be regarded as the underlying storage infrastructure of the storage engine. The storage system is responsible for providing reliable storage capabilities, ensuring data security and availability, while the storage engine implements more advanced data management and access functions on this basis.
[0033] Graph Storage refers to a data storage solution specifically designed for graph data structures, aiming to efficiently store and manage large amounts of graph data such as interconnected nodes, edges, and attributes.
[0034] Graph storage systems usually optimize operations on nodes, edges, and their attributes to support fast queries, traversals, and other graph-related computational tasks. In a graph storage system, graph data is typically stored in the form of files, for example, persistently stored on disk in file form; these files constitute the basic storage units of graph data in the graph storage system.
[0035] The graph storage engine built on top of the graph storage system, that is, the graph storage engine with the graph storage system as the underlying storage, is responsible for the specific data access and storage logic corresponding to the graph storage system, including storage structure (such as file format), indexing method, transaction processing, concurrency control, etc.
[0036] In practical applications, graph storage systems can generally be divided into two categories: Native Graph Storage systems and Non-Native Graph Storage systems.
[0037] A Native Graph Storage system refers to a storage system specifically designed for graphs, which directly organizes data in the form of nodes, edges, and attributes. The core features of this storage model include: Nodes and edges are directly represented as independent objects at the physical storage level, with their own identifiers (IDs), sets of attributes, and references to other related objects, rather than being indirectly represented through relational tables or key-value pairs; Nodes usually contain a list of references to all their adjacent nodes (i.e., the target nodes pointed to by the edges connected to them), and edges can also directly link to the two nodes they connect, enabling fast graph traversal operations, that is, accessing or processing all nodes and edges in the graph according to a certain strategy, such as jumping from one node to another along the edges; Providing consistency and persistence guarantees for graph data.
[0038] For Native Graph Storage systems, they are usually optimized by leveraging the characteristics of graph topology and graph queries; among them, graph topology refers to the connection method and structural characteristics between nodes and edges in a graph, and graph query refers to the process of querying graph data in the storage system. That is, according to the characteristics of the graph data structure and common graph query patterns, adjustments and improvements are made to aspects such as the file format and indexing method corresponding to the storage system, making the files used to store graph data in the Native Graph Storage system graph-aware. In this case, for the graph storage engine built on top of the Native Graph Storage system, it can reduce the complexity of graph query operations in the Native Graph Storage system and improve graph query efficiency, thus ensuring that the graph storage engine has relatively high graph query performance.
[0039] Non-graph-native storage systems usually convert graphs into other forms of data structures for storage, such as tables in relational databases or key-value pairs in key-value storage systems.
[0040] Non-graph-native storage systems generally do not have the ability to directly map nodes and edges, nor do they specifically optimize graph traversal operations. That is to say, the files used to store graph data in non-graph-native storage systems are usually unaware of the graph, resulting in poor query performance of the graph storage engine built on non-graph-native storage systems, that is, the complexity of graph query operations in the graph storage engine in non-graph-native storage systems is high and the graph query efficiency is low.
[0041] A key-value store is a simple and efficient database model that uses key-value pairs to store data. Each record in a key-value store has a unique key and the value associated with that key. The key is a unique identifier used to locate a specific value stored in the database. Usually, the key is in the form of a string, but it can also be other types of data. The value is the actual data item associated with the key. It can be any form of data, such as a string, a number, a binary object, etc. In some cases, the value can be a complex data structure, such as a JSON document or a serialized object.
[0042] The design of key-value stores is very simple. Usually, it only provides basic operations such as Put, Get / Exact Lookup, Range Lookup or Seek, Delete, etc., without complex schema definitions or relationship constraints, which makes them easy to implement, deploy, and maintain. Due to the simple data structure, key-value stores usually have high read and write performance, capable of providing extremely high read and write throughput and low-latency response times.
[0043] Key-value stores have excellent scalability and flexibility. Specifically, most key-value stores are designed as distributed systems, supporting horizontal scaling to adapt to larger data volumes and higher throughput requirements. In addition, since the format and content of the value in the key-value pair are usually not restricted, it is possible to freely choose how to organize the data according to actual needs to form a storage data structure. Key-value stores can also provide different consistency guarantees, ranging from strong consistency to eventual consistency, depending on the specific application scenario and technical implementation.
[0044] Due to the good scalability and flexibility of key-value stores, in practical applications, graph data can be converted into a key-value pair data structure and stored in a key-value store, thus transforming the key-value store into a graph storage system. Building a graph storage engine based on a key-value store can utilize the existing complete key-value store ecosystem and infrastructure.
[0045] However, a graph storage system built on a key-value storage system belongs to a non-native graph storage system. This means that the files used to store graph data in such storage systems are actually unaware of the graph. Therefore, the graph query performance of a graph storage engine built on a key-value storage system will be greatly affected.
[0046] In this application, the key-value storage system is optimized by leveraging the characteristics of graph topology and graph query. A graph-aware SSTable file format is proposed, which is obtained by adjusting and improving the original SSTable file format in the key-value storage system. Accordingly, adjustments and improvements are made to aspects such as the indexing method, enabling the graph storage engine built on the optimized key-value storage system to have relatively high query performance.
[0047] Specifically, one or more embodiments of this application provide a technical solution for implementing graph queries. In this technical solution, the nodes and edges in a graph can be stored in an SSTable file, which can include at least one first data block for storing key-value pairs corresponding to the nodes and edges. The nodes and edges can be organized into first key-value pairs and second key-value pairs respectively and stored in the first data block. Among them, the key in the first key-value pair corresponding to a node can include the node identifier of the node, and the value can include the node attributes of the node. The key in the second key-value pair corresponding to an edge can include the start point identifier and end point identifier of the edge, and the value can include the edge attributes of the edge. Each target first key-value pair and the target second key-value pair whose start point identifier is the same as the node identifier in the target first key-value pair can be aggregated and stored in the first data block. Accordingly, when obtaining a target key for querying the graph, first, based on the target node identifier in the target key, query in the SSTable file to determine the target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier. Then, based on the target key, query in the target first data block to determine the target value corresponding to the target key, and determine the graph query result based on the target value.
[0048] By adopting the above method, the optimization of the key-value storage system by leveraging the characteristics of graph topology and graph query is realized. Specifically, it includes adjusting and improving aspects such as the SSTable file format and indexing method in the key-value storage system, enabling the SSTable file for storing graph data in the key-value storage system to be graph-aware, thereby reducing the complexity of the graph query operation of the graph storage engine built on the key-value storage system in the key-value storage system, improving the graph query efficiency, and ensuring that the graph storage engine has relatively high graph query performance.
[0049] In this application, a graph database system can refer to a system for storing, managing, and querying graph data structures.
[0050] In the storage space of the above graph database system, SSTable (Sorted String Table) files can be stored, and these SSTable files are used to store the nodes and edges in the graph. Among them, the SSTable file can include at least one data block (which can be called the first data block) for storing key-value pairs corresponding to the nodes and edges.
[0051] The nodes in the above graph can be organized into first key-value pairs corresponding to the nodes, and the edges in the graph can be organized into second key-value pairs corresponding to the edges; the first key-value pairs and the second key-value pairs can be stored in the above first data block. Among them, the key in the first key-value pair can include the node identifier of the node, and the value can include the node attributes of the node; the key in the second key-value pair can include the start point identifier and the end point identifier of the edge, and the value can include the edge attributes of the edge.
[0052] Each first key-value pair can be used as a target first key-value pair in turn, and the second key-value pair with the start point identifier the same as the node identifier in the target first key-value pair can be used as a target second key-value pair; the target first key-value pair and the target second key-value pair can be aggregated and stored in the above first data block.
[0053] It should be noted that the above graph database system can specifically include a key-value storage system for storing graph data in the data structure of key-value pairs, that is, the key-value storage system provides the storage space of the graph database system. The graph database system can also include a storage engine for converting graph data into the data structure of key-value pairs and making it meet specific storage requirements. Among them, the storage engine can be a graph storage engine built based on the key-value storage system, that is, a graph storage engine with the key-value storage system as the underlying storage.
[0054] The following takes the key-value storage system and the graph storage engine built based on the key-value storage system as examples for detailed description.
[0055] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a key-value storage system shown in an exemplary embodiment of the present application.
[0056] As Figure 1 shown, the graph storage engine can use the data structure of LSM-Tree (Log Structured Merge Tree) to organize data in the key-value storage system.
[0057] The LSM-Tree is a data structure optimized for handling a large number of write operations. It is designed to append new data to the end of a file rather than update it in-place. The LSM-Tree first appends all new write operations to one or more sorted data structures in memory and then, at an appropriate time, writes this data to persistent storage on disk in batches and sequentially. This data structure can reduce random write operations on disk, thereby reducing the overhead of random writes, and can maintain query efficiency through a background merging process.
[0058] The LSM-Tree uses a multi-level structure to manage data. These levels start with the MemTable in memory and then gradually migrate to multiple levels on the hard disk. Each level contains one or more SSTable files. As the number of levels increases, the amount of data in each level also gradually increases, and the size of each level is usually several times that of the previous level.
[0059] Specifically, when data is first written, it is first stored in a memory data structure called the MemTable. The MemTable is in memory and is a collection of sorted key-value pairs. All newly inserted data accumulates here first. Since it is memory-based operation, the write speed is very fast and lookup operations can be performed quickly.
[0060] After the current MemTable reaches a certain size, it is frozen and converted to an immutable state, i.e., the Immutable MemTable, waiting to be written to disk as an immutable file, i.e., the SSTable file. This approach reduces frequent disk write operations, thereby improving write performance.
[0061] The SSTable file is written to disk to form the first level (L0). The SSTable is a read-only file format that contains key-value pairs arranged in order. To speed up lookups, each SSTable can have associated metadata information (e.g., Bloom Filter and index), which can be used to quickly locate the block that may contain the target key.
[0062] Over time, the L0 level will accumulate more and more SSTable files. To keep the system running efficiently and with good query performance, the system will periodically perform a compaction operation to merge the data in the L0 level to the next level (L1), while removing duplicate or expired data items. The compaction process will continue downwards until the data reaches the bottom level. The compaction operation is usually asynchronous and does not block the write path.
[0063] Note that the L0 layer is the layer closest to the memory and usually contains SSTable files that have just been flushed from the MemTable in the memory to the disk. The SSTable files in the L0 layer may not be sorted by key because they are generated independently from different MemTables. When the number of SSTable files in the L0 layer exceeds a certain threshold, a Minor Compaction is triggered to merge and sort these files and then migrate the results to the L1 layer.
[0064] Starting from the L1 layer, the SSTable files in each layer are sorted by key. The number and size of the SSTable files in each layer usually increase layer by layer. For example, the SSTable files in the L1 layer may be larger than those in the L0 layer, and the SSTable files in the L2 layer are larger than those in the L1 layer. When the number or total size of the SSTable files in a certain layer exceeds a preset threshold, a Major Compaction is triggered to merge a part of the files in this layer with the files in the next layer to generate larger SSTable files and then migrate the results to the next layer.
[0065] When the graph storage engine performs a query on a key-value storage system using an LSM-Tree, it needs to start from the MemTable in the memory and then sequentially check the SSTable files at all levels. Since the SSTable files in each layer are sorted by key, efficient algorithms such as binary search can be used to quickly locate the data.
[0066] Specifically, when performing a query, it will first look for the data in the MemTable in the memory.
[0067] If the required data cannot be found in the MemTable, it will continue to look for it sequentially in the SSTable files at all levels that already exist on the disk (if there are Immutable MemTables, it will first look in the Immutable MemTables and then in the SSTable files). The impact of compaction is also considered during the query process, that is, all relevant SSTable files need to be traversed to ensure that the latest data version is obtained.
[0068] Before accessing an SSTable file, its associated Bloom Filter can be used to determine whether the SSTable file is likely to contain the target key. If the Bloom Filter indicates that the target key does not exist in the SSTable file, the SSTable file can be skipped, thus avoiding unnecessary disk I / O operations.
[0069] Finally, the query process combines the results from the MemTable, Immutable MemTables, and individual SSTables, and returns the latest and unique key-value pairs.
[0070] For example, assume in a key-value storage system using an LSM-Tree, there is one MemTable, one SSTable file in level L0, and two SSTable files in level L1 (in the order of SSTable file 1 and SSTable file 2); the key space ranges from key_0 to key_9. Then, when querying for the key key_5, first look in the MemTable, assume key_5 is not found; continue to look in the SSTable file in level L0, assume key_5 is found in the SSTable file in level L0, and its corresponding value is value_old; continue to look in SSTable file 1 in level L1, assume key_5 is not found; continue to look in SSTable file 2 in level L1, assume key_5 is also found in SSTable file 2 in level L1, and its corresponding value is value_new, and it is determined to be the updated value according to the timestamp or other mechanisms. In this case, the finally returned value is the latest value value_new corresponding to key_5.
[0071] As can be seen from the above, the key-value pairs stored in the original SSTable file are arranged in the order of keys. Therefore, when using the original SSTable file to store graph data, the graph data needs to be converted into key-value pair form data for storage. That is, the original SSTable file is graph-agnostic.
[0072] In a Figure 1 key-value storage system as shown, the original graph-agnostic SSTable file can be replaced with an adjusted and improved graph-aware SSTable file (which can be called a Graph-Aware SSTable file).
[0073] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a graph storage file shown in an exemplary embodiment of the present application.
[0074] As Figure 2 shown, the above graph storage file (i.e., the file for storing graph data) can be a Graph-Aware SSTable file obtained by adjusting and improving on the basis of the original SSTable file. Among them, the Graph-Aware SSTable file is graph-aware.
[0075] In the original SSTable file, it usually contains a Data Block and an Index Block. The Data Block is the most basic data storage unit in the SSTable file, used to store data in the form of actual key-value pairs, and the stored key-value pairs are arranged in the order of the keys. The Index Block is used to store the meta-information of each Data Block, such as the key range (the range of the minimum key and the maximum key) of the Data Block and the offset of the Data Block in the SSTable file (which can indicate the starting position). It can accelerate the search process and enable quick positioning to the Data Block where the target key is located.
[0076] In addition to the Data Block and the Index Block, a complete SSTable file may also contain other auxiliary components, such as a Footer, a Bloom Filter, and Metadata. The Footer is located at the end of the file and contains the offsets of the Index Block and other metadata information (such as the Bloom Filter) in the SSTable file, which is used to quickly find the Index Block. The Bloom Filter is a probabilistic data structure used to quickly determine whether a certain key exists in the SSTable file, reducing unnecessary disk I / O operations. Metadata contains other auxiliary information, such as the size of the file, the creation time, the version number, etc.
[0077] For a Graph-Aware SSTable file as Figure 2 shown, first of all, it should be noted that the Graph-Aware SSTable file stores key-value pairs (which can be respectively called the first key-value pairs and the second key-value pairs) organized by the nodes and edges in the graph. Specifically, the nodes in the graph can be organized into the first key-value pairs corresponding to the nodes, and the edges in the graph can be organized into the second key-value pairs corresponding to the edges.
[0078] In practical applications, in order to facilitate the search for specific key-value pairs in the above-mentioned first data block, all the key-value pairs stored in the first data block can be arranged in the order of the keys (usually in Lexicographical Order, that is, the dictionary order). When there are multiple first data blocks, the key-value pairs stored in two adjacent first data blocks are also arranged in the order of the keys.
[0079] Taking a node as an example, the key in the first key-value pair organized by this node may include the node identifier of this node, and the value may include the node attributes of this node (specifically, it may be the encoded node attribute data). Taking an edge as an example, the key in the second key-value pair organized by this edge may include the node identifier of the Source Node (source node / starting node) of this edge (abbreviated as the starting point identifier) and the node identifier of the Destination Node (destination node / ending node) of this edge (abbreviated as the ending point identifier), and the value includes the edge attributes of this edge (specifically, it may be the encoded edge attribute data).
[0080] In practical applications, graphs can be divided into directed graphs (Directed Graph / Digraph) and undirected graphs (Undirected Graph). In a directed graph, each edge has a specific direction, pointing from one node to another node. Therefore, each edge in a directed graph can be organized into a second key-value pair. In an undirected graph, the edges have no direction, which means that the two nodes connected by each edge are reachable in both directions. Therefore, for an undirected graph, each undirected edge in it can be regarded as consisting of two directed edges. For example, assuming that in an undirected graph, an edge connects node A and node B, then this edge can be regarded as an edge pointing from node A to node B, and an edge pointing from node B to node A. Thus, this edge can be organized into two second key-value pairs. In one second key-value pair, the key contains the starting point identifier as the node identifier of node A and the ending point identifier as the node identifier of node B. In the other second key-value pair, the key contains the starting point identifier as the node identifier of node B and the ending point identifier as the node identifier of node A.
[0081] The above first key-value pair can be as shown in Table 1 below:
[0082] Table 1
[0083]
[0084] In some embodiments, the key in the first key-value pair organized by a node, in addition to including the node identifier of this node, may also include the node type of this node. The specific contents of the key and value in the first key-value pair can be set according to actual needs, and this application does not impose special restrictions on this.
[0085] The above second key-value pair can be as shown in Table 2 below:
[0086] Table 2
[0087]
[0088] In some embodiments, the key in the second key-value pair organized by an edge may include, in addition to the start identifier and the end identifier of the edge, the node type of the Source Node of the edge (referred to as the start type for short), and / or the creation timestamp of the edge. The specific content of the key and value in the second key-value pair can be set according to actual needs, and this application does not impose special restrictions on this.
[0089] In the above-mentioned Graph-Aware SSTable file, it may include at least one data block (which can be called the first data block) for storing the above-mentioned first key-value pair and the above-mentioned second key-value pair. Among them, the data block is similar to the Data Block in the original SSTable file.
[0090] It should be noted that each of the first key-value pairs in all the first key-value pairs (which can be called the target first key-value pairs), and the second key-value pair whose start identifier is the same as the node identifier in the target first key-value pair (which can be called the target second key-value pair), can be aggregated and stored in the above-mentioned first data block. For example, assuming that the node identifier included in the key of the target first key-value pair is node identifier A, then the start identifier included in the key of the target second key-value pair should also be node identifier A, and these target first key-value pairs and target second key-value pairs can be aggregated and stored in the first data block.
[0091] The node identifier in the above-mentioned target first key-value pair is the same as the start identifier in the above-mentioned target second key-value pair, which means that the edge corresponding to the target second key-value pair starts from the node corresponding to the target first key-value pair. Therefore, by aggregating and storing the target first key-value pair and the target second key-value pair, the common access to the node and its connected edges can be realized, thereby accelerating the query of the node and its connection points and improving the efficiency of graph query and graph traversal.
[0092] Correspondingly, in the above-mentioned Graph-Aware SSTable file, the above-mentioned first key-value pair and the above-mentioned second key-value pair can be specifically arranged in the order of the smallest key in each group of aggregated first key-value pairs and second key-value pairs. For example, the above-mentioned target first key-value pair and the above-mentioned target second key-value pair are a group of aggregated first key-value pairs and second key-value pairs. Since the data content included in the key of the first key-value pair is less than the data content included in the key of the second key-value pair, based on the lexicographical order, the smallest key in this group of aggregated target first key-value pairs and target second key-value pairs is usually the key in the target first key-value pair.
[0093] Aggregating and storing the above-mentioned target first key-value pair and the above-mentioned target second key-value pair in the above-mentioned first data block can specifically be aggregating and storing the target first key-value pair and the target second key-value pair in the same first data block.
[0094] For example, assume that the node identifier included in the key of the first key-value pair 1 is node identifier A, the node identifier included in the key of the first key-value pair 2 is node identifier B, the start identifiers included in the keys of the second key-value pairs 1-1 and 1-2 are both node identifier A, and the start identifiers included in the keys of the second key-value pairs 2-1, 2-2, 2-3, 2-4, 2-5, and 2-6 are all node identifier B. Also assume that the lexicographical order of the keys is the key of the first key-value pair 1 < the key of the second key-value pair 1-1 < the key of the second key-value pair 1-2, the key of the first key-value pair 2 < the key of the second key-value pair 2-1 < the key of the second key-value pair 2-2 < the key of the second key-value pair 2-3 < the key of the second key-value pair 2-4 < the key of the second key-value pair 2-5 < the key of the second key-value pair 2-6, and the key of the first key-value pair 1 < the key of the first key-value pair 2. Then, (the first key-value pair 1, the second key-value pair 1-1, the second key-value pair 1-2) can be stored in one first data block, and (the first key-value pair 2, the second key-value pair 2-1, the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, the second key-value pair 2-6) can be stored in another first data block.
[0095] Similar to the Data Block in the original SSTable file, the size of each first data block may have a certain limit to improve the query efficiency of the first data block; when this limit is reached, new key-value pairs will be written to the next first data block. That is, the amount of data of the first key-value pairs and the second key-value pairs stored in each first data block is not greater than a preset threshold. Here, the amount of data can refer to the number of key-value pairs or the data size of the key-value pairs, and this application does not impose special restrictions on this. Correspondingly, aggregating and storing the above-mentioned target first key-value pairs and the above-mentioned target second key-value pairs in the first data block can specifically be aggregating and storing the target first key-value pairs and the target second key-value pairs in one or consecutive multiple first data blocks.
[0096] The first key-value pair and the second key-value pair containing different node identifiers can be stored in different first data blocks. For example, assume that the node identifier contained in the key of the first key-value pair 1 is node identifier A, the node identifier contained in the key of the first key-value pair 2 is node identifier B, the starting point identifiers contained in the keys of the second key-value pair 1-1 and the second key-value pair 1-2 are both node identifier A, the starting point identifiers contained in the keys of the second key-value pair 2-1, the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, and the second key-value pair 2-6 are all node identifier B, and assume that the number of key-value pairs stored in a first data block is not greater than 5, and assume that the lexicographical order of the keys is the key of the first key-value pair 1 < the key of the second key-value pair 1-1 < the key of the second key-value pair 1-2, the key of the first key-value pair 2 < the key of the second key-value pair 2-1 < the key of the second key-value pair 2-2 < the key of the second key-value pair 2-3 < the key of the second key-value pair 2-4 < the key of the second key-value pair 2-5 < the key of the second key-value pair 2-6, and the key of the first key-value pair 1 < the key of the first key-value pair 2. Then, (the first key-value pair 1, the second key-value pair 1-1, the second key-value pair 1-2) can be stored in one first data block, (the first key-value pair 2, the second key-value pair 2-1, the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4) can be stored in another first data block, and (the second key-value pair 2-5, the second key-value pair 2-6) can be stored in yet another first data block.
[0097] The first key-value pair and the second key-value pair containing different node identifiers can also be stored in the same first data block. Continuing with the above example, (the first key-value pair 1, the second key-value pair 1-1, the second key-value pair 1-2, the first key-value pair 2, the second key-value pair 2-1) can be stored in one first data block, and (the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, the second key-value pair 2-6) can be stored in another first data block.
[0098] In addition, in the above Graph-Aware SSTable file, an index (which can be called a first-level index) created based on the node identifier in the above first key-value pair can also be included to facilitate determining the first key-value pair and the second key-value pair for aggregated storage that contain a specific node identifier.
[0099] It should be noted that the above first-level index can be used to indicate the first data block where the first key-value pair and the second key-value pair containing each node identifier are located. Among them, the first key-value pair and the second key-value pair containing the same node identifier are the first key-value pair and the second key-value pair for aggregated storage.
[0100] Specifically, the above first-level index may include: the correspondence between each node identifier and the storage location information of the first data block where the first key-value pair and the second key-value pair containing the node identifier are located in the above Graph-Aware SSTable file. Among them, the storage location information of the first data block in the Graph-Aware SSTable file may include the offset of the first data block in the Graph-Aware SSTable file (which can indicate the starting position); in addition, the storage location information may further include other auxiliary information, such as the size of the first data block, so as to facilitate determining the termination position of the first data block.
[0101] For example, assume that the first data block 1 stores (the first key-value pair 1, the second key-value pair 1-1, the second key-value pair 1-2, the first key-value pair 2, the second key-value pair 2-1), and the first data block 2 stores (the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, the second key-value pair 2-6). Then, since the first data block where the first key-value pair 1, the second key-value pair 1-1, and the second key-value pair 1-2 containing the node identifier A are located is the first data block 1, and the first data block where the first key-value pair 2, the second key-value pair 2-1, the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, and the second key-value pair 2-6 containing the node identifier B are located includes the first data block 1 and the first data block 2, the first-level index created based on the node identifiers in the first key-value pair 1 and the first key-value pair 2 may be as shown in Table 3 below:
[0102] Table 3
[0103]
[0104] The storage location information of the first data block 1 in Table 3 above is the storage location information of the first data block 1 in the above Graph-Aware SSTable file, and the storage location information of the first data block 2 is the storage location information of the first data block 2 in the Graph-Aware SSTable file.
[0105] Furthermore, the above first-level index may include: the correspondence between each node identifier and the storage location information of the first data block where the first key-value pair and the second key-value pair containing the node identifier are located in the above Graph-Aware SSTable file. Correspondingly, the correspondence between the node identifier and the storage location information in the first-level index can be arranged in the order of the first key-value pairs containing each node identifier in the first data block. Specifically, when the key-value pairs stored in the first data block are arranged in the order of keys, the correspondence between the node identifier and the storage location information in the first-level index can be arranged in the order of the smallest keys in the first key-value pairs and the second key-value pairs containing each node identifier.
[0106] For example, assume that the node identifier contained in the key of the first key-value pair 1 is node identifier A, the node identifier contained in the key of the first key-value pair 2 is node identifier B, and the node identifier contained in the key of the first key-value pair 3 is node identifier C. Also assume that the key of the first key-value pair 1 < the key of the first key-value pair 2 < the key of the first key-value pair 3, and assume that the first key-value pair and the second key-value pair containing node identifier A are aggregated and stored in the first data block 1, the first key-value pair and the second key-value pair containing node identifier B are aggregated and stored in the first data block 1 and the first data block 2, and the first key-value pair and the second key-value pair containing node identifier C are aggregated and stored in the first data block 3. Then, in the above first-level index, the arrangement order of the correspondence between the node identifier and the storage location information can be: the correspondence between node identifier A and the storage location information of the first data block 1, the correspondence between node identifier B and the storage location information of the first data block 1, the correspondence between node identifier C and the storage location information of the first data block 3. Since the first data block where the first key-value pair and the second key-value pair containing node identifier A are located is the first data block 1, the first data block where the first key-value pair and the second key-value pair containing node identifier C are located is also the first data block 1, and the first data block where the first key-value pair and the second key-value pair containing node identifier C are located is the first data block 3, the first-level index created based on the node identifiers in the first key-value pair 1, the first key-value pair 2, and the first key-value pair 3 can be as shown in Table 4 below:
[0107] Table 4
[0108]
[0109] In the above case, the first data block in the above Graph-Aware SSTable file for aggregating and storing the first key-value pair and the second key-value pair containing the same node identifier can be determined according to the arrangement order of the correspondence between the node identifier and the storage location information in the above first index. Specifically, for a node identifier, according to the corresponding relationship containing the node identifier and the next corresponding relationship of the corresponding relationship, in the Graph-Aware SSTable file, the first data blocks in the range starting from the first data block indicated by the storage location information contained in the corresponding relationship and ending at the first data block indicated by the storage location information contained in the next corresponding relationship of the corresponding relationship can be determined as the first data block in the Graph-Aware SSTable file for aggregating and storing the first key-value pair and the second key-value pair containing the node identifier. Continuing with the above example, since the next corresponding relationship of the corresponding relationship between the node identifier B and the storage location information of the first data block 1 is the corresponding relationship between the node identifier C and the storage location information of the first data block 3, the first data blocks in the range from the first data block 1 to the first data block 3 can be determined as the first data block for aggregating and storing the first key-value pair and the second key-value pair containing the node identifier B, that is, the first key-value pair and the second key-value pair containing the node identifier B must be stored in the first data block 1 and the first data block 2, and the first key-value pair and the second key-value pair containing the node identifier B may be stored in the first data block 3.
[0110] To reduce the lookup time complexity of the above first-level index, thereby accelerating the query of nodes and their connection points and improving the efficiency of graph query and graph traversal, the first-level index can be a hash index. That is, the first-level index can specifically include: the correspondence between the hash value of each node identifier and the storage location information of the first data block where the first key-value pair and the second key-value pair containing the node identifier are located in the above Graph-Aware SSTable file.
[0111] Alternatively, the above first-level index can also follow the index block in the original SSTable file. At this time, the first-level index can specifically include: the correspondence between the range of node identifiers in the first key-value pairs contained in each first data block and the storage location information of the first data block in the above Graph-Aware SSTable file.
[0112] It should be noted that for the graph composed of nodes and edges corresponding to the key-value pairs stored in the above-mentioned Graph-Aware SSTable file, if the graph is a sparse graph (i.e., the number of edges is relatively small), it indicates that a first data block in the Graph-Aware SSTable file may contain multiple first key-value pairs. Therefore, the first-level index in the Graph-Aware SSTable file can be the index block in the SSTable file; if the graph is a dense graph (i.e., the number of edges is relatively large, for example, the number of edges is close to the square of the number of nodes), it indicates that the first key-value pairs in the Graph-Aware SSTable file may be scattered in discontinuous first data blocks. Therefore, the first-level index in the Graph-Aware SSTable file can be a hash index.
[0113] In practical applications, Cuckoo Hashing can be used to create the above first-level index based on the node identifiers in the above first key-value pairs.
[0114] To facilitate determining the first data block for aggregating and storing the first key-value pairs and second key-value pairs containing the same node identifier, thereby accelerating the query of nodes and their connection points and improving the efficiency of graph query and graph traversal, when the above target first key-value pairs and the above target second key-value pairs are aggregated and stored in multiple consecutive first data blocks, an index (which can be called a second-level index) created based on the key in the target second key-value pair can also be stored in the first first data block among these multiple first data blocks.
[0115] It should be noted that the above second-level index can be used to indicate the first data block where the above target second key-value pair is located. In this case, the first-level index can specifically be used to indicate the first first data block where the first key-value pairs and second key-value pairs containing each node identifier are located. That is, at this time, only the first first data block where the first key-value pairs and second key-value pairs containing the node identifier in the target key are located needs to be found using the first-level index, and the remaining first data blocks can be found through the second-level index in the first first data block.
[0116] Specifically, the second-level index can include: the correspondence between the key ranges of the target second key-value pairs in each of the first data blocks included in the above multiple consecutive first data blocks and the storage location information of this first data block in the Graph-Aware SSTable file. Among them, the storage location information can include the offset (which can indicate the starting position) of this first data block in the Graph-Aware SSTable file; in addition, the storage location information can also include other auxiliary information, such as the size of this first data block, to facilitate determining the termination position of this first data block.
[0117] In practical applications, for the Nth (N is a natural number) of the above-mentioned consecutive first data blocks, the key range of the target second key-value pairs in the Nth first data block can be represented by the minimum key and the maximum key among these target second key-value pairs, or can be represented by the maximum key in the target second key-value pairs in the Nth first data block and the maximum key in the target second key-value pairs in the (N + 1)th first data block. That is to say, the key range of the target second key-value pairs in a first data block in the above-mentioned secondary index can specifically be the range of the minimum key and the maximum key in the target second key-value pairs in this first data block, or can only be the maximum key in the target second key-value pairs in this first data block.
[0118] For example, assume that the first data block 1 stores (first key-value pair 1, second key-value pair 1-1, second key-value pair 1-2, first key-value pair 2, second key-value pair 2-1), and the first data block 2 stores (second key-value pair 2-2, second key-value pair 2-3, second key-value pair 2-4, second key-value pair 2-5, second key-value pair 2-6). It indicates that when the target first key-value pair is the first key-value pair 2 and the target second key-value pairs include the second key-value pair 2-1, the second key-value pair 2-2, the second key-value pair 2-3, the second key-value pair 2-4, the second key-value pair 2-5, and the second key-value pair 2-6, the target first key-value pair and the target second key-value pairs are aggregated and stored in two consecutive first data blocks, namely the first data block 1 and the first data block 2. In this case, since the maximum key in the target second key-value pairs in the first data block 1 is the key in the second key-value pair 2-1, and the maximum key in the target second key-value pairs in the first data block 2 is the key in the second key-value pair 2-6, the secondary index created based on the keys in the target second key-value pairs can be as shown in Table 5 below:
[0119] Table 5
[0120]
[0121] Moreover, the data stored in the first data block 1 can specifically be (first key-value pair 1, second key-value pair 1-1, second key-value pair 1-2, first key-value pair 2, second key-value pair 2-1, the secondary index corresponding to the first key-value pair 2).
[0122] Please, on the basis of Figure 1 and Figure 2 , refer to Figure 3 , Figure 3 which is a flowchart of a graph query method shown in an exemplary embodiment of the present application.
[0123] In this embodiment, as mentioned above, the nodes and edges in the graph can be stored in an SSTable file (specifically, such as Figure 2The Graph-Aware SSTable file shown), and the SSTable file may include at least one first data block for storing key-value pairs corresponding to nodes and edges in the graph.
[0124] Nodes in the above graph can be organized into first key-value pairs corresponding to the nodes, and edges in the graph can be organized into second key-value pairs corresponding to the edges, and stored in the above first data block. Among them, the key in a first key-value pair may include the node identifier of a node, and the value may include the node attributes of this node; the key in a second key-value pair may include the start identifier and end identifier of an edge, and the value may include the edge attributes of this edge.
[0125] For a first key-value pair among all first key-value pairs, this first key-value pair can be used as the target first key-value pair. The target first key-value pair, and the second key-value pair whose start identifier is the same as the node identifier in the target first key-value pair (which can be called the target second key-value pair), are aggregated and stored in the above first data block.
[0126] It should be noted that the above graph query method can be implemented by a graph storage engine built based on a key-value storage system, and is used for graph query in the key-value storage system.
[0127] As Figure 3 shown, the above graph query method may include the following steps:
[0128] Step 302: Obtain a target key for querying the graph.
[0129] In this embodiment, the graph can be queried based on a key (which can be called the target key) through Get operation and Seek operation, that is, execute Get(Key) operation or Seek(Key) operation, where Key is the target key.
[0130] Step 304: Query in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier.
[0131] In this embodiment, as described above, whether it is the above first key-value pair or the above second key-value pair, the key therein contains a node identifier. Therefore, first, the node identifier (which can be called the target node identifier) can be extracted from the above target key. Subsequently, based on the target node identifier, query can be performed in the above SSTable file, so as to determine the first data block (which can be called the target first data block) for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier.
[0132] In some embodiments, the above SSTable file may further include a first-level index created based on the node identifier in the first key-value pair. The first-level index can be used to indicate the first data block where the first key-value pair and the second key-value pair containing each node identifier are located, that is, the first data block that aggregately stores the corresponding first key-value pair and second key-value pair.
[0133] Since the first-level index is an index created based on the node identifier and used to indicate the first data block where the first key-value pair and the second key-value pair containing each node identifier are located, by querying the first-level index, it is possible to determine the above-mentioned first data block for aggregately storing the first key-value pair and the second key-value pair containing the target node identifier.
[0134] In some embodiments, when the above target key is obtained, it is possible to first query in the corresponding LSM-Tree based on the target key to determine the SSTable file in the LSM-Tree for storing the key-value pair containing the target key, and then query in the above-mentioned first-level index included in the determined SSTable file based on the above target node identifier in the target key.
[0135] As described above, it is possible to determine whether the SSTable file may contain the target key through its associated Bloom Filter. That is, it is possible to determine the above-mentioned SSTable file for storing the key-value pair containing the above target key through the Bloom Filter associated with each SSTable file in the LSM-Tree included in the above key-value storage system.
[0136] Step 306: Query in the target first data block based on the target key to determine the target value corresponding to the target key, and determine the graph query result based on the target value.
[0137] In this embodiment, it is possible to query in the above target first data block based on the above target key. Since the first data block is used to actually store the above first key-value pair and the above second key-value pair, by querying the target first data block, it is possible to determine the target value corresponding to the target key, and thus it is possible to determine the graph query result based on the target value.
[0138] Specifically, if the above target value is the value in the above first key-value pair, that is, the node attribute, then the node attribute can be determined as the graph query result; if the above target value is the value in the above second key-value pair, that is, the edge attribute, then the edge attribute can be determined as the graph query result.
[0139] In some embodiments, as described above, the data volumes of the first key-value pair and the second key-value pair stored in the first data block do not exceed a preset threshold. Accordingly, the target first key-value pair and the target second key-value pair may be aggregated and stored in one first data block, or may be aggregated and stored in multiple consecutive first data blocks.
[0140] In some embodiments, when the target first key-value pair and the target second key-value pair are aggregated and stored in multiple consecutive first data blocks, the first first data block among the multiple consecutive first data blocks may also be used to store a secondary index created based on the key in the target second key-value pair.
[0141] The secondary index may be used to indicate the first data block where the target second key-value pair is located. Accordingly, the primary index may specifically be used to indicate the first first data block where the first key-value pair and the second key-value pair containing each node identifier are located.
[0142] When determining the graph query result based on the above target value, if the target value is the secondary index, a further query may be performed in the secondary index based on the target key. Since the secondary index is an index created based on the key in the target second key-value pair and used to indicate the first data block where the target second key-value pair is located, by querying the secondary index, the first data block corresponding to the target key can be determined. Subsequently, based on the target key, a query may be performed in the first data block corresponding to the target key to determine the value corresponding to the target key, and the graph query result may be determined based on the value corresponding to the target key.
[0143] Similarly to the foregoing content, when determining the graph query result based on the value corresponding to the target key, since the secondary index hit by the target key at this time is created based on the key in a specific second key-value pair, the value corresponding to the target key should be an edge attribute, and thus the edge attribute may be determined as the graph query result.
[0144] It should be noted that the above secondary index may exist independently in the first first data block, that is, it is not included as the value in a key-value pair in the first data block. In this case, when the target key does not match the keys in each key-value pair included in the first data block, the secondary index in the first data block may be directly determined as the target value corresponding to the target key.
[0145] Alternatively, the above secondary index may be included as the value in a key-value pair in the above first first data block, and the key in this key-value pair may be the key range of the above target second key-value pair. In this case, when the above target key hits the key range in this first data block (at this time, the target key does not match the keys in each key-value pair included in this first data block), the secondary index in this first data block may be determined as the above target value corresponding to the target key.
[0146] In some embodiments, as described above, the above secondary index may include: the correspondence between the key ranges of the target second key-value pairs in each of the above consecutive first data blocks and the storage location information of this first data block in the above SSTable file.
[0147] Therefore, when further querying in the above secondary index based on the above target key to determine the first data block corresponding to the target key, specifically, a binary search may be further performed in the secondary index based on the target key to determine the storage location information of the first data block corresponding to the target key in the above SSTable file, and based on this storage location information, the first data block corresponding to the target key may be determined in this SSTable file.
[0148] In some embodiments, as described above, the above primary index may include: the correspondence between each node identifier and the storage location information of the first first data block where the first key-value pair and the second key-value pair containing this node identifier are located in the above SSTable file. Correspondingly, the correspondence between the node identifier and the storage location information in the primary index may be arranged in the order of the first key-value pairs containing each node identifier in the first data block.
[0149] In the above case, the above target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the above target node identifier in the above SSTable file may be determined according to the arrangement order of the correspondence between the node identifier and the storage location information in the above first index. Specifically, first, a query may be performed in the above primary index included in this SSTable file based on the target node identifier to determine the corresponding relationship (which may be referred to as the target corresponding relationship) containing the target node identifier, and the next corresponding relationship of the target corresponding relationship, and then the first data blocks within the range from the first data block indicated by the storage location information included in the target corresponding relationship to the first data block indicated by the storage location information included in the next corresponding relationship of the target corresponding relationship in this SSTable file may be determined as the target first data block.
[0150] In some embodiments, for the above key-value storage system, the specific format of the SSTable files included in the LSM-Tree can be selected according to the actual application scenarios and requirements.
[0151] Taking the above SSTable file as an example, when the subgraph composed of the nodes and edges corresponding to the key-value pairs stored in the SSTable file is a sparse subgraph, the above first-level index in the SSTable file can be the index block in the SSTable file. When the subgraph composed of the nodes and edges corresponding to the key-value pairs stored in the SSTable file is a dense subgraph, the above first-level index in the SSTable file can be a hash index.
[0152] For another example, in the case of graph analysis scenarios, storing edge attributes column-wise is more convenient for access. Therefore, the edge attributes can be separately stored in the above second data block.
[0153] It should be noted that the specific formats of the SSTable files included in the same LSM-Tree can also be different. In this case, when merging the SSTable files, the SSTable files with the same format can be merged.
[0154] In the above technical solution, the nodes and edges in the graph can be stored in the SSTable file, and the SSTable file can include at least one first data block for storing the key-value pairs corresponding to the nodes and edges; the nodes and edges can be respectively organized into a first key-value pair and a second key-value pair and stored in the first data block. Among them, the key in the first key-value pair corresponding to a node can include the node identifier of the node, and the value can include the node attributes of the node. The key in the second key-value pair corresponding to an edge can include the start point identifier and the end point identifier of the edge, and the value can include the edge attributes of the edge; each target first key-value pair and the target second key-value pair whose start point identifier is the same as the node identifier in the target first key-value pair can be aggregated and stored in the first data block. Correspondingly, when obtaining the target key for querying the graph, first, based on the target node identifier in the target key, query in the SSTable file to determine the target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier, then, based on the target key, query in the target first data block to determine the target value corresponding to the target key, and determine the graph query result based on the target value.
[0155] In the above - mentioned manner, the optimization of the key - value storage system is achieved by utilizing the characteristics of graph topology and graph query. Specifically, it includes adjusting and improving aspects such as the SSTable file format and indexing method in the key - value storage system, making the SSTable file for storing graph data in the key - value storage system graph - aware. Thereby, it can reduce the complexity of graph query operations of the graph storage engine based on the key - value storage system in the key - value storage system, improve the graph query efficiency, and thus ensure that the graph storage engine has relatively high - efficient graph query performance.
[0156] Corresponding to the embodiments of the foregoing method, the present application also provides embodiments of an apparatus.
[0157] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non - volatile memory 410. Of course, other required hardware may also be included. One or more embodiments of the present application can be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non - volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of the present application do not exclude other implementation manners, such as a logic device or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic module and can also be hardware or a logic device.
[0158] Please refer to Figure 5 , Figure 5 which is a block diagram of a graph query device shown in an exemplary embodiment of the present application.
[0159] The above - mentioned graph query device can be applied to Figure 5 the device shown in
[0160] to implement the technical solution of the present application. Among them, the nodes and edges in the figure are stored in a sorted string table SSTable file; the SSTable file includes at least one first data block for storing key - value pairs corresponding to the nodes and the edges; the nodes and the edges are respectively organized into a first key - value pair and a second key - value pair and stored in the first data block; the key in the first key - value pair includes the node identifier of the node, and the value includes the node attributes of the node; the key in the second key - value pair includes the start - point identifier and the end - point identifier of the edge, and the value includes the edge attributes of the edge; each target first key - value pair, and the target second key - value pair whose start - point identifier is the same as the node identifier in the target first key - value pair, are aggregated and stored in the first data block.
[0161] An acquisition module 502 acquires a target key for querying the graph.
[0162] A first query module 504 queries the SSTable file based on a target node identifier in the target key to determine a target first data block for aggregating and storing a first key-value pair and a second key-value pair containing the target node identifier.
[0163] A second query module 506 queries the target first data block based on the target key to determine a target value corresponding to the target key, and determines a graph query result based on the target value.
[0164] In some embodiments, the SSTable file further includes a first-level index created based on node identifiers in the first key-value pair; the first-level index is used to indicate the first data blocks where the first key-value pair and the second key-value pair containing each node identifier are located.
[0165] The querying in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing a first key-value pair and a second key-value pair containing the target node identifier includes:
[0166] Querying in the first-level index included in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing a first key-value pair and a second key-value pair containing the target node identifier.
[0167] In some embodiments, the SSTable file is an SSTable file in a Log-Structured Merge Tree (LSM-Tree).
[0168] The apparatus further includes:
[0169] A third query module queries the LSM-Tree based on the target key to determine the SSTable file for storing the key-value pair containing the target key before querying the SSTable file based on the target node identifier in the target key.
[0170] In some embodiments, the data volume of the first key-value pair and the second key-value pair stored in each first data block is not greater than a preset threshold; the target first key-value pair and the target second key-value pair are aggregated and stored in one or a continuous plurality of first data blocks.
[0171] In some embodiments, when the target first key-value pair and the target second key-value pair are aggregately stored in a plurality of consecutive first data blocks, the first first data block among the plurality of consecutive first data blocks is further used to store a secondary index created based on the key in the target second key-value pair; the secondary index is used to indicate the first data block where the target second key-value pair is located; the primary index is used to indicate the first first data block where the first key-value pair and the second key-value pair containing each node identifier are located;
[0172] Determining the graph query result based on the target value includes:
[0173] If the target value is the secondary index, query in the secondary index based on the target key to determine the first data block corresponding to the target key;
[0174] Query in the first data block corresponding to the target key based on the target key to determine the value corresponding to the target key, and determine the graph query result based on the value corresponding to the target key.
[0175] In some embodiments, the secondary index includes: the correspondence between the key range of the target second key-value pair in each of the first data blocks included in the plurality of consecutive first data blocks and the storage location information of this first data block in the SSTable file;
[0176] Querying in the secondary index based on the target key to determine the first data block corresponding to the target key includes:
[0177] Perform a binary search in the secondary index based on the target key to determine the storage location information of the first data block corresponding to the target key in the SSTable file, and determine the first data block corresponding to the target key in the SSTable file according to this storage location information.
[0178] In some embodiments, the primary index includes: the correspondence between each node identifier and the storage location information of the first first data block where the first key-value pair and the second key-value pair containing this node identifier are located in the SSTable file; the correspondence in the primary index is arranged in the order of the first key-value pair in the first data block;
[0179] Querying in the primary index included in the SSTable file based on the target node identifier in the target key to determine the target first data block for aggregately storing the first key-value pair and the second key-value pair containing the target node identifier includes:
[0180] Query in the first-level index included in the SSTable file based on the target node identifier in the target key to determine a target correspondence relationship that includes the target node identifier and the next correspondence relationship of the target correspondence relationship;
[0181] In the SSTable file, determine the first data blocks within the range from the first data block indicated by the storage location information included in the target correspondence relationship to the first data block indicated by the storage location information included in the next correspondence relationship as target first data blocks for aggregating and storing a first key-value pair and a second key-value pair that include the target node identifier.
[0182] In some embodiments, when the subgraph composed of nodes and edges corresponding to the key-value pairs stored in the SSTable file is a sparse subgraph, the first-level index in the SSTable file is the index block in the SSTable file; when the subgraph composed of nodes and edges corresponding to the key-value pairs stored in the SSTable file is a dense subgraph, the first-level index in the SSTable file is a hash index.
[0183] In some embodiments, the key in the first key-value pair further includes the node type of the node.
[0184] In some embodiments, the key in the second key-value pair further includes the start point type of the edge and / or the establishment timestamp of the edge.
[0185] For the device embodiments, they basically correspond to the method embodiments. Therefore, for the relevant parts, refer to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the technical solution of this application.
[0186] The systems, devices, modules, or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer may be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email receiving and sending device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0187] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0188] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0189] Computer-readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0190] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0191] The above specific embodiments of the present application are described. Other embodiments are within the scope of the present application. In some cases, the actions or steps recorded in the present application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0192] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" are also intended to include the plural forms unless the context clearly dictates otherwise. The term "and / or" means any and all possible combinations of one or more of the associated listed items.
[0193] The description of the terms "one embodiment", "some embodiments", "example", "specific example", or "a mode of implementation" used in one or more embodiments of the present application means that the specific features or characteristics described in connection with the embodiment are included in at least one embodiment of the present application. The schematic descriptions of these terms do not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described may be combined in a suitable manner in one or more embodiments of the present application. In addition, different embodiments and the specific features or characteristics in different embodiments may be combined without contradiction.
[0194] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of the present application to describe various information, the information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0195] The above description is only the preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of the present application shall be included within the scope of protection of one or more embodiments of the present application.
[0196] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
Claims
1. A graph database system for storing graph data, the graph data including nodes and edges; wherein: An SSTable file of a sorted string table is stored in the storage space of the graph database system; the SSTable file includes at least one first data block for storing key-value pairs corresponding to the nodes and the edges; The key-value pairs include a first key-value pair and a second key-value pair; the key in the first key-value pair includes a node identifier of the node, and the value includes node attributes of the node; the key in the second key-value pair includes a start point identifier and an end point identifier of the edge, and the value includes edge attributes of the edge; In the first data block, various target first key-value pairs, and target second key-value pairs whose start point identifiers are the same as the node identifiers in the target first key-value pairs are aggregately stored.
2. The graph database system according to claim 1, wherein the SSTable file further includes a first-level index created based on the node identifiers in the first key-value pairs; the first-level index is used to indicate the first data blocks where the first key-value pairs and the second key-value pairs containing each node identifier are located.
3. The graph database system according to claim 1, wherein a log-structured merge tree (LSM-Tree) is stored in the storage space of the graph database system; the SSTable file is the SSTable file in the LSM-Tree.
4. The graph database system according to claim 2, wherein the data amount of the first key-value pairs and the second key-value pairs stored in each first data block is not greater than a preset threshold; the target first key-value pairs and the target second key-value pairs are aggregately stored in one or a continuous plurality of first data blocks.
5. The graph database system according to claim 4, when the target first key-value pair and the target second key-value pair are aggregately stored in a plurality of consecutive first data blocks, a secondary index created based on the key in the target second key-value pair is further stored in the first first data block among the plurality of consecutive first data blocks; wherein, The second-level index is used to indicate the first data block where the target second key-value pair is located; the first-level index is used to indicate the first first data block where the first key-value pairs and the second key-value pairs containing each node identifier are located.
6. The graph database system according to claim 5, wherein the secondary index comprises: The corresponding relationship between the key ranges of the target second key-value pairs in each of the continuous plurality of first data blocks and the storage location information of the first data block in the SSTable file.
7. The graph database system according to claim 2, wherein the primary index includes: The corresponding relationship between each node identifier and the storage location information of the first first data block where the first key-value pair and the second key-value pair containing the node identifier are located in the SSTable file; The corresponding relationship in the first-level index is arranged in the order of the first key-value pairs in the first data block.
8. The graph database system according to claim 2, when the subgraph composed of the nodes and edges corresponding to the key-value pairs stored in the SSTable file is a sparse subgraph, the first-level index in the SSTable file is an index block in the SSTable file; when the subgraph composed of the nodes and edges corresponding to the key-value pairs stored in the SSTable file is a dense subgraph, the first-level index in the SSTable file is a hash index.
9. A method for querying a graph; wherein, The nodes and edges in the figure are stored in a sorted string table SSTable file; the SSTable file includes at least one first data block for storing key-value pairs corresponding to the nodes and the edges; the nodes and the edges are respectively organized into a first key-value pair and a second key-value pair and stored in the first data block; The key in the first key-value pair includes the node identifier of the node, and the value includes the node attributes of the node; the key in the second key-value pair includes the start point identifier and the end point identifier of the edge, and the value includes the edge attributes of the edge; each target first key-value pair and a target second key-value pair whose start point identifier is the same as the node identifier in the target first key-value pair are aggregated and stored in the first data block; The method includes: Obtaining a target key for querying the figure; Querying in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier; Querying in the target first data block based on the target key to determine a target value corresponding to the target key, and determining a figure query result based on the target value.
10. The method according to claim 9, wherein the SSTable file further includes a first-level index created based on the node identifiers in the first key-value pairs; the first-level index is used to indicate the first data block where the first key-value pairs and the second key-value pairs containing each node identifier are located; The querying in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier includes: Querying in the first-level index included in the SSTable file based on the target node identifier in the target key to determine a target first data block for aggregating and storing the first key-value pair and the second key-value pair containing the target node identifier.
11. The method according to claim 9, wherein the SSTable file is the SSTable file in a log-structured merge tree LSM-Tree; Before querying in the SSTable file based on the target node identifier in the target key, the method further includes: Querying in the LSM-Tree based on the target key to determine the SSTable file for storing the key-value pair containing the target key.
12. The method according to claim 10, wherein the data volume of the first key-value pairs and the second key-value pairs stored in each first data block is not greater than a preset threshold; the target first key-value pairs and the target second key-value pairs are aggregated and stored in one or a continuous plurality of first data blocks.
13. According to the method described in claim 12, when the target first key-value pair and the target second key-value pair are aggregately stored in a plurality of consecutive first data blocks, the first first data block among the plurality of consecutive first data blocks is further used to store a secondary index created based on the key in the target second key-value pair; the secondary index is used to indicate the first data block where the target second key-value pair is located; the primary index is used to indicate the first first data block where the first key-value pair and the second key-value pair containing each node identifier are located; Determining the graph query result based on the target value includes: If the target value is the secondary index, query in the secondary index based on the target key to determine the first data block corresponding to the target key; Query in the first data block corresponding to the target key based on the target key to determine the value corresponding to the target key, and determine the graph query result based on the value corresponding to the target key.
14. According to the method described in claim 13, the secondary index includes: The correspondence between the key ranges of the target second key-value pairs in each of the plurality of consecutive first data blocks and the storage location information of the first data block in the SSTable file; Querying in the secondary index based on the target key to determine the first data block corresponding to the target key includes: Perform a binary search in the secondary index based on the target key to determine the storage location information of the first data block corresponding to the target key in the SSTable file, and determine the first data block corresponding to the target key in the SSTable file according to the storage location information.
15. The method according to claim 10, wherein the primary index includes: The correspondence between each node identifier and the storage location information of the first first data block where the first key-value pair and the second key-value pair containing the node identifier are located in the SSTable file; The correspondence in the primary index is arranged in the order of the first key-value pair in the first data block; Querying in the primary index included in the SSTable file based on the target node identifier in the target key to determine the target first data block for aggregately storing the first key-value pair and the second key-value pair containing the target node identifier includes: Query in the primary index included in the SSTable file based on the target node identifier in the target key to determine the target correspondence relationship containing the target node identifier and the next correspondence relationship of the target correspondence relationship; Determine the first data blocks within the range from the first data block indicated by the storage location information contained in the target correspondence relationship to the first data block indicated by the storage location information contained in the next correspondence relationship in the SSTable file as the target first data blocks for aggregately storing the first key-value pair and the second key-value pair containing the target node identifier.
16. According to the method described in claim 10, when the subgraph composed of nodes and edges corresponding to the key-value pairs stored in the SSTable file is a sparse subgraph, the primary index in the SSTable file is the index block in the SSTable file; when the subgraph composed of nodes and edges corresponding to the key-value pairs stored in the SSTable file is a dense subgraph, the primary index in the SSTable file is a hash index.
17. An electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein, the processor realizes the method according to any one of claims 9 to 16 by running the executable instructions.
18. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method according to any one of claims 9 to 16 is realized.
Citation Information
Cited By
Data updating method and device, storage medium and computer equipment
CN121433571A
Graph database storage method and device based on Hash mapping, equipment and medium
CN121502043A