Query optimization method, system and equipment for graph database and storage medium
By pushing the Join operator down to the storage engine for execution in the graph database, the problem of underutilized computing power of the storage layer is solved, query speed and efficiency are improved, and data transmission and resource waste are reduced.
Patent Information
- Application Number
- CN202511087281.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing graph database query methods have the problem of underutilized computing power of the storage layer and slow query speed. Especially when processing large-scale graph data, a large amount of raw data needs to be transferred to the query engine layer, resulting in high memory pressure, low CPU utilization and slow query speed.
By identifying Join operators in the query engine and pushing down Join operators that meet the preset push-down conditions to the storage engine for execution, the storage engine is used to perform Join operations on vertex and edge data, reducing data transmission and improving computing efficiency.
By pushing the Join operator down to the storage engine, data transmission and computing resource waste are reduced, query speed and efficiency are improved, and the computing power of the storage layer is fully utilized.
Smart Images

Figure CN120596516A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of graph databases, and in particular to query optimization methods, systems, devices, and storage media for graph databases. Background Art
[0002] A graph database is a non-relational database specifically designed for processing highly connected data. It uses a graph structure (nodes, edges, and attributes) to represent and store data. Compared to traditional relational databases, graph databases exhibit significant advantages in scenarios such as multi-hop queries and complex relational network analysis. This data model is well-suited for processing highly interconnected data. Currently, graph databases (such as NebulaGraph) typically adopt a separate compute-storage architecture, with the query engine layer responsible for query parsing, optimization, and execution coordination, and the storage engine layer responsible for data persistence and local computation. A common query pattern currently is two-stage query processing: basic scanning operations are performed first at the storage engine layer, followed by data joins, filtering, and aggregations at the query engine layer.
[0003] However, as the scale of graph data expands, existing graph database query methods need to transfer a large amount of raw data to the query engine layer, the vertex and edge data need to be serialized / deserialized multiple times, and intermediate data that is not needed for the final result is transmitted. In addition, the connection efficiency is low, and the traditional Join algorithm is not suitable for the characteristics of graph data. The query engine needs to cache a large number of intermediate results, and the computing power of the storage layer is not fully utilized. Therefore, the existing graph data query method has problems such as high memory pressure, low CPU utilization, and slow query speed. Summary of the Invention
[0004] The embodiments of the present application provide a query optimization method, system, device and storage medium for a graph database, so as to at least solve the problem in the related art that the computing power of the storage layer is not fully utilized and the query speed is slow in the existing graph data query method.
[0005] In a first aspect, embodiments of the present application provide a query optimization method for a graph database, which is applied to a graph database query system having a query engine and a storage engine. The method includes: Obtaining a graph data query statement through the query engine, disassembling and analyzing the graph data query statement, and identifying a Join operator; Obtain the preset push-down condition, determine whether the Join operator meets the push-down condition, and if so, push the Join operator down to the storage engine; Querying the point-edge scanning operator upstream of the Join operator through the storage engine to obtain first point-edge data, and passing the first point-edge data to the Join operator; The storage engine performs a Join operation on the first point-edge data according to the Join operator to obtain first connection data, and returns the first connection data to the query engine.
[0006] In one embodiment, the method further comprises: When the Join operator does not meet the push-down condition, query the point-edge scanning operator upstream of the Join operator through the storage engine to obtain the second point-edge data; After returning the second point edge data to the query engine, it is passed to the Join operator; The query engine performs a Join operation on the second point-edge data according to the Join operator to obtain second connection data, and passes the second connection data to a downstream operator.
[0007] In one embodiment, after returning the first connection data to the query engine, the method further includes: If there is no query operator downstream of the Join operator, the first connection data is fed back to the user; If there is a query operator downstream of the Join operator, then other data upstream of the query operator is obtained, and the query engine performs corresponding query processing on the first connection data and other data according to the query operator; If the storage engine returns multiple first connection data to the query engine, the query engine merges the multiple second connection data.
[0008] In one embodiment, the storage format of point data is: graph identifier + point type identifier + point identifier; the storage format of edge data is: graph identifier + edge type identifier + edge start identifier + edge end identifier; The storage engine includes multiple storage partitions, and the push-down condition is: The point data and edge data obtained by the upstream scanning of the Join operator are stored in the same storage partition, and the execution logic of the Join operator is that the starting point identifier of the edge data is equal to the identifier of the point data.
[0009] In one embodiment, the decomposing and analyzing the graph data query statement includes: Parsing the graph data query statement into an abstract syntax tree; The abstract syntax tree is converted into a logical execution plan to obtain query operators, which include a point scan operator, an edge scan operator and a Join operator.
[0010] In one embodiment, the storage engine queries the vertex-edge scan operator upstream of the Join operator, including: Querying a point identifier in the storage partition by using a primary key index, and obtaining corresponding point data according to the point identifier; Scan the same storage partition according to the point identifier to obtain edge data starting from the point identifier.
[0011] In one embodiment, the storage engine includes multiple storage partitions, wherein: Multiple storage partitions can execute Join operators that meet the push-down conditions in parallel. During the execution of the Join operator, each storage partition performs a Join operation on the vertex-edge data it holds.
[0012] In a second aspect, an embodiment of the present application provides a graph database query system, the system comprising a query engine and a storage engine, and the system implements the query optimization method for the graph database as described in any one of the above embodiments when running; wherein, The query engine is used to receive graph data query statements, read graph data from the storage engine and execute corresponding query operators, and feedback the final query results to the user; wherein, the query engine can execute all query operators except the vertex-edge scanning operator and the write operator; The storage engine is used to store graph data and execute vertex-edge scanning operators, write operators, and join operators that meet the push-down conditions.
[0013] In a third aspect, an embodiment of the present application provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the query optimization method for the graph database as described in the first aspect above is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the query optimization method for the graph database as described in the first aspect above is implemented.
[0015] The query optimization method, system, device, and storage medium for graph databases provided by the embodiments of the present application have at least the following technical effects: In this application, the query engine obtains a graph data query statement, disassembles and analyzes the graph data query statement, and identifies the Join operator; obtains a preset push-down condition, determines whether the Join operator meets the push-down condition, and if the judgment result is yes, pushes the Join operator down to the storage engine; queries the point-edge scanning operator upstream of the Join operator through the storage engine to obtain the first point-edge data, and passes the first point-edge data to the Join operator; the storage engine performs a Join operation on the first point-edge data according to the Join operator to obtain the first connection data, and returns the first connection data to the query engine. This application pushes the qualified Join operator down to the storage engine for execution, thereby eliminating the need to transfer a large amount of original data to the query engine layer, reducing the transmission of intermediate data that is not required for the final result, and thus improving the query speed; and by pushing down the Join operator, the computing power of the storage layer is fully utilized, reducing the waste of computing resources.
[0016] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 This is a flowchart of a query optimization method for a graph database in one embodiment of the present application; Figure 2 It is a logical execution and data flow diagram of a map data query statement in one embodiment of the present application; Figure 3 It is a structural block diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.
[0019] Obviously, the drawings described below are merely examples or embodiments of the present application. Those skilled in the art can, without inventive effort, apply the present application to other similar scenarios based on these drawings. Furthermore, it is also understood that, although the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, changes in design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as an insufficiency of the content disclosed in this application.
[0020] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.
[0021] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meaning as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "an," "the," and similar expressions used herein do not denote quantitative limitations and may refer to either the singular or the plural. The terms "comprise," "include," "have," and any variations thereof, used herein, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules (units) is not limited to the listed steps or units but may also include steps or units not listed, or may include other steps or units inherent to the process, method, product, or apparatus. The terms "connected," "connected," "coupled," and similar expressions used herein are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used herein, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone; A and B exist simultaneously; or B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0022] Graph database: A database management system specifically designed to store and query graph-structured data, which consists of nodes and edges.
[0023] Vertex: A graph database represents an entity. For example, Zhang San is a person. A vertex is used to store an entity in the graph database. Each vertex has a unique identifier (a globally unique integer) in the graph database.
[0024] Edges: Graph databases represent relationships between entities. For example, if Zhang San and Li Si know each other, the relationship between entities is stored in the form of edges. Each edge actually stores the identifiers of the starting and ending points.
[0025] In addition, the Join operation in a graph database is significantly different from that in a traditional relational database. The Join operation in a graph database is the process of matching and connecting vertices with edges or other vertices based on a specific relationship. This connection is not based on value matching, but on the topological relationship of the graph structure itself.
[0026] Currently, joins are primarily implemented using the following methods: In-memory joins pull vertex-edge data into graphd memory for join execution; distributed joins shuffle data across nodes before joining; and indexed nested loops use index lookups to connect vertex-edges. Developing new optimization solutions to push join operators down to storage engines, leveraging the locality of graph data and the computational power of modern storage engines, remains a key technical area of research.
[0027] Based on the above reasons, this application pushes pre-defined joins (joins that meet pre-defined pushdown conditions in the graph database model) to stored execution, enabling parallel execution of these joins while reducing the amount of data required for each join, significantly improving graph database query performance. This application is widely applicable to various graph database query optimization scenarios, including social network analysis and recommendation systems.
[0028] Based on the above situation, on the first aspect, an embodiment of the present application provides a graph database management system, in which there are two main components: a query engine (graphd) and a storage engine (storaged). The query engine is used to receive graph data query statements, read graph data from the storage engine and execute corresponding query operators, and feedback the final query results to the user; for example, the query engine executes all query operators except the node-edge scanning operator and the write operator. The storage engine is used to store graph data (i.e., node-edge data) and execute node-edge scanning operators, write operators, and Join operators that meet the push-down conditions. Specifically, storaged is only responsible for executing the node scanning operator (NodeScan), the edge scanning operator (EdgeScan), and the write operator (the write operator refers to the execution of data modification operations, such as data insertion, update, deletion, etc.), and all operators except the scan operator (Scan) and the write operator are executed in graphd.
[0029] In the embodiment of the present application, the point data (referred to as points) and edge data (referred to as edges) in the graph data are stored in the storage based on key-value. The relevant key format is as follows: Point: GraphID (unique identifier of the graph) + NodeType (unique identifier of the point type) + NodeID (unique identifier of the point); Edge: GraphID (unique identifier of the graph) + EdgeType (unique identifier of the edge type) + SrcID (unique identifier of the starting point of the edge) + DstID (unique identifier of the end point of the edge).
[0030] Based on the above storage format, this embodiment of stored can complete the following query operations: 1. In a graph, to get a point of a given type NodeType with an identifier of X, you only need to specify GraphID+NodeType+X to read, that is, you can read a point with an identifier of X; 2. In a graph, to get all edges of a given edge type EdgeType, you only need to specify GraphID+EdgeType to perform a prefix scan to read multiple edges; 3. In a graph, to obtain all edges with a given edge type EdgeType and a starting identifier of X, you only need to specify GraphID+EdgeType+X for prefix scanning. This will read multiple edges (because there may be multiple edges with X as the starting point).
[0031] In addition to the aforementioned vertex edges, for each vertex, storage also stores a primary key index, which maps the primary key attribute to the vertex identifier, also in key-value format. For example, suppose there is a vertex type Person, whose primary key attribute is name. Then, for each Person-type vertex, a key-value pair is stored, where the key is the vertex's name attribute and the value is the corresponding vertex identifier.
[0032] In the second aspect, the embodiment of the present application provides a query optimization method for a graph database, which can be run in the above-mentioned graph database management system, such as Figure 1 As shown, the query optimization method of the present application is implemented by the following contents.
[0033] Step S1: Obtain a graph data query statement through the query engine, analyze the query statement, and identify join operators. Specifically, the query statement is parsed into an abstract syntax tree (AST). The AST is then converted into a logical execution plan to obtain query operators, which include a point scan operator, an edge scan operator, and a join operator.
[0034] In a specific embodiment, a commonly used graph database query statement is given in this embodiment. This embodiment uses this query statement to explain the query optimization method of this application in detail. The query statement is: MATCH (p1:Person{name:"Alice"})-[e1:KNOWS]->(p2:Person) RETURN p2 The query statement above is parsed into an abstract syntax tree (AST) to identify key components in the query: vertex patterns (e.g., (p1:Person{name:"Alice"})), edge patterns (e.g., [e1:KNOWS]), relation patterns (e.g., (p1)-[e1]->(p2)), and return results (e.g., RETURN p2). The query statement in this example searches for all data in the graph that meets the following conditions: a vertex named p1 of type Person with a name attribute of Alice; a vertex named p2 of type Person; and an edge of type KNOWS between p1 and p2, named e1. The query ultimately returns p2, which represents all people Alice knows.
[0035] This query will be translated into the following logical execution plan and physical execution plan in the graph database, refer to Figure 2 , the arrows represent the data flow, in the logical plan: NodeScan(p1): finds p1, that is, all points of type Person and named Alice; EdgeScan(e1): finds e1, that is, all edges of type KNOWS; Join1: The condition is left_node_id(e1) = element_id(p1), that is, find all data that satisfies the starting point identifier of edge e1 equal to the point identifier of p1; NodeScan(p2): finds p2, i.e. all points with the label Person; Join2: The condition is element_id(p2) = right_node_id(e1), that is, find all data that satisfies the end point identifier of the e1 edge equal to the midpoint identifier of p2.
[0036] Step S2: Obtain the preset pushdown conditions and determine whether the Join operator meets the pushdown conditions. If so, push the Join operator to the storage engine. In this embodiment, the storage format of point data is: graph identifier + point type identifier + point identifier; the storage format of edge data is: graph identifier + edge type identifier + edge start identifier + edge end identifier.
[0037] In this embodiment, the storage engine includes multiple storage partitions, and NebulaGraph adopts a hash sharding strategy. First, for each point, each point has a corresponding primary key attribute. The hash value of the primary key attribute of each point is calculated to obtain an integer X, and then the partition ID is obtained modulo the number of partitions. That is, partition ID = Hash(primary key attribute) % number of partitions. In addition, the unique identifier of each point will also contain the partition ID information. Secondly, for edges, the partition in which each edge is stored is determined by the unique identifier of the starting point of the edge. That is, this edge is stored in the partition where the starting point is located. Through the above operations, it is guaranteed that "a point and all edges starting from this point will only be stored in a given storage partition."
[0038] In the graph database of the embodiment of the present application, storaged has multiple instances (i.e., storage partitions), each of which is responsible for storing a portion of the data in the graph. For a point, assuming its identifier is X, then the point with identifier X and all edges starting from X are stored in the same storaged in the graph database. Therefore, as long as any identifier X is known, the point p with identifier X and all edges e starting from X can be queried in the same storaged, that is, the Join operator with the Join condition left_node_id(e) = element_id(p) can be executed in this storaged. Such a Join is also called a Pre-defined Join.
[0039] Based on the storage sharding strategy of this embodiment, the push-down condition of this application is: the point data and edge data obtained by the upstream scan of the Join operator are stored in the same storage partition, and the execution logic of the Join operator is that the starting point identifier of the edge data is equal to the identifier of the point data. In other words, the point edge data upstream of the Join, that is, the data read by NodeScan and the data read by EdgeScan, are stored in the same storage, and the Join condition is that the starting point identifier of the edge = the point identifier.
[0040] exist Figure 2 In the example, Join1 satisfies both conditions and is therefore a pre-defined join, whereas Join2 does not match the join conditions and is therefore not a pre-defined join.
[0041] Step S3: Query the point-edge scanning operator upstream of the Join operator through the storage engine to obtain first point-edge data, and pass the first point-edge data to the Join operator. Specifically, during the point-edge scanning process, query the point identifier in the storage partition using the primary key index, and obtain the corresponding point data based on the point identifier; then scan the same storage partition based on the point identifier to obtain edge data starting from the point identifier.
[0042] Taking the above query statement as an example, when executing NodeScan(p1), storaged queries the primary key index and uses Alice to find the corresponding vertex identifier, assuming it is X. It then reads the vertex using GraphID + NodeType + X to obtain p1, which is the vertex of type Person and named Alice. Note that unlike the original solution, storaged passes the identifier of each vertex in p1 to EdgeScan(e1), ultimately passing the corresponding data to the downstream Join1 and not returning it to graphd.
[0043] When executing EdgeScan(e1), storaged will perform a prefix scan on each point identifier received from NodeScan(p1) using GraphID + EdgeType + point identifier, and pass the corresponding data to the downstream Join1 without returning it to graphd.
[0044] Step S4: The storage engine performs a Join operation on the first edge data according to the Join operator to obtain first connection data, and returns the first connection data to the query engine. Figure 2 After receiving the data from NodeScan(p1) and EdgeScan(e1), storaged's Join1 can perform a Join based on the condition left_node_id(e1) = element_id(p1), and then return the corresponding data to graphd.
[0045] In a preferred embodiment, if the storage engine returns multiple first-join data points to the query engine, the query engine merges the multiple second-join data points. Multiple storage partitions can execute Join operators that meet the push-down conditions in parallel. During the execution of the Join operators, each storage partition performs a Join operation on its own vertex-edge data. Graphd merges the Join1 results from each storage partition and passes them to the downstream operator.
[0046] In one embodiment of the present application, after the first connection data is returned to the query engine, if there is no query operator downstream of the Join operator, the first connection data is fed back to the user; if there is a query operator downstream of the Join operator, Figure 2 Join2 in the query operator, then obtain other data upstream of the query operator such as Figure 2 The query engine then performs corresponding query processing on the first connection data and other data according to the query operator, such as performing a Join2 operation.
[0047] In an excellent embodiment of the present application, when the Join operator does not meet the push-down condition, the storage engine queries the point-edge scanning operator upstream of the Join operator to obtain second point-edge data; the second point-edge data is returned to the query engine and passed to the Join operator; the query engine performs a Join operation on the second point-edge data according to the Join operator to obtain second connection data, and passes the second connection data to the downstream operator.
[0048] Specifically, taking the above query statement as an example, and referring to Figure 2 If Join1 does not meet the push-down conditions, the query is performed as follows: When executing NodeScan(p1), storaged queries the primary key index and finds the corresponding vertex identifier through Alice, assuming it is X. It then reads the vertex using GraphID + NodeType + X to read p1, which is all the vertices of type Person and named Alice. Finally, the read data is returned to graphd. Then, when executing EdgeScan(e1), storaged performs a prefix scan based on GraphID + EdgeType to read e1, that is, all edges of type KNOWS, and finally returns the read data to graphd. After receiving the data from NodeScan(p1) and EdgeScan(e1), graphd performs Join1, which is a join based on the condition left_node_id(e1) = element_id(p1). It then outputs the corresponding data to the downstream. The rest of the process is similar to the push-down to the storage engine described above and is not described in detail here.
[0049] In summary, the query optimization method for graph databases provided by the embodiment of the present application pushes down the Pre-defined Join that meets the push-down conditions provided by the present application to the storage engine, thereby enabling the Pre-defined Join to be completed in parallel in multiple storaged (each storaged joins the part of the data it holds, and ultimately merges the data results), which can greatly improve execution efficiency. Executing Pre-defined Join in storaged greatly reduces the amount of data to be joined, while avoiding frequent data transfer between the storage engine and the query engine (such as returning a large amount of vertex-edge data to the query engine), thereby improving query speed and efficiency. In addition, by pushing down some Joins to the storage engine, the present application can fully utilize the computing memory of the storage engine and avoid wasting computing resources.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, Figure 3 FIG is a block diagram of an electronic device according to an exemplary embodiment. Figure 3 As shown, the electronic device may include a processor 11 and a memory 12 storing computer program instructions.
[0051] Specifically, the processor 11 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0052] The memory 12 may include a large-capacity memory for data or instructions. By way of example, and not limitation, the memory 12 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 12 may include removable or non-removable (or fixed) media. Where appropriate, the memory 12 may be internal or external to the data processing device. In certain embodiments, the memory 12 is non-volatile memory. In certain embodiments, the memory 12 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0053] The memory 12 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 11 .
[0054] The processor 11 implements any one of the query optimization methods for the graph database in the above embodiments by reading and executing computer program instructions stored in the memory 12.
[0055] In one embodiment, the electronic device may further include a communication interface 13 and a bus 10. Figure 3 As shown, the processor 11 , the memory 12 , and the communication interface 13 are connected via a bus 10 and communicate with each other.
[0056] The communication interface 13 is used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application. The communication interface 13 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0057] The bus 10 includes hardware, software, or both, and couples the components of the electronic device to each other. The bus 10 includes, but is not limited to, at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 10 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Bus 10 may include one or more buses, where appropriate. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.
[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the query optimization method for the graph database provided in the first aspect is implemented.
[0059] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0060] In a possible implementation, the present invention can also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of the query optimization method for the graph database provided in the first aspect.
[0061] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0062] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0063] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A query optimization method for a graph database, characterized in that: Applied in a graph database query system having a query engine and a storage engine, the method includes: Obtaining a graph data query statement through the query engine, disassembling and analyzing the graph data query statement, and identifying a Join operator; Obtain the preset push-down conditions and push the Join operators that meet the push-down conditions to the storage engine; Querying the point-edge scanning operator upstream of the Join operator through the storage engine to obtain first point-edge data, and passing the first point-edge data to the Join operator; The storage engine performs a Join operation on the first point-edge data according to the Join operator to obtain first connection data, and returns the first connection data to the query engine.
2. The query optimization method according to claim 1, characterized in that: The method further comprises: When the Join operator does not meet the push-down condition, query the point-edge scanning operator upstream of the Join operator through the storage engine to obtain the second point-edge data; After returning the second point edge data to the query engine, it is passed to the Join operator; The query engine performs a Join operation on the second point-edge data according to the Join operator to obtain second connection data, and passes the second connection data to a downstream operator.
3. The query optimization method according to claim 1, characterized in that: After returning the first connection data to the query engine, the method further includes: If there is no query operator downstream of the Join operator, the first connection data is fed back to the user; If there is a query operator downstream of the Join operator, then other data upstream of the query operator is obtained, and the query engine performs corresponding query processing on the first connection data and other data according to the query operator; If the storage engine returns multiple first connection data to the query engine, the query engine merges the multiple second connection data.
4. The query optimization method according to claim 1, wherein: The storage format of point data is: graph identifier + point type identifier + point identifier; the storage format of edge data is: graph identifier + edge type identifier + edge start identifier + edge end identifier; The storage engine includes multiple storage partitions, and the push-down condition is: The point data and edge data obtained by the upstream scanning of the Join operator are stored in the same storage partition, and the execution logic of the Join operator is that the starting point identifier of the edge data is equal to the identifier of the point data.
5. The query optimization method according to claim 1, characterized in that: The disassembling and analyzing the graph data query statement includes: Parsing the graph data query statement into an abstract syntax tree; The abstract syntax tree is converted into a logical execution plan to obtain query operators, which include a point scan operator, an edge scan operator, and a Join operator.
6. The query optimization method according to claim 4, characterized in that: The storage engine queries the point-edge scanning operator upstream of the Join operator, including: Querying a point identifier in the storage partition by using a primary key index, and obtaining corresponding point data according to the point identifier; Scan the same storage partition according to the point identifier to obtain edge data starting from the point identifier.
7. The query optimization method according to claim 1, characterized in that: The storage engine includes multiple storage partitions, wherein: Multiple storage partitions can execute Join operators that meet the push-down conditions in parallel. During the execution of the Join operator, each storage partition performs a Join operation on the vertex-edge data it holds.
8. A graph database query system, characterized in that: The system includes a query engine and a storage engine, and when the system is running, the query optimization method for a graph database according to any one of claims 1 to 7 is implemented; wherein, The query engine is used to receive graph data query statements, read graph data from the storage engine and execute corresponding query operators, and feedback the final query results to the user; wherein, the query engine executes all query operators except the vertex-edge scanning operator and the write operator; The storage engine is used to store graph data and execute vertex-edge scanning operators, write operators, and join operators that meet the push-down conditions.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the query optimization method for a graph database as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the query optimization method for a graph database according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for database query language based on SQL (Structured Query Language) extension graph
CN115858872A
Data query method, electronic equipment and computer readable storage medium
CN117112613A
Graph data query method for graph database and related equipment
CN117591564A
Graph database query method based on runtime filtering
CN120371870A
Full-text indexing method and system based on graph database
US20220335086A1