Graph database processing method and device, storage medium and program product

By dynamically selecting the storage format of adjacent tables in the graph database and combining the cost indicators of the historical query task set, the problem of low performance of the graph database in precise query and column scanning tasks is solved, and the optimization of storage costs and efficient utilization of resources is achieved.

CN120541270APending Publication Date: 2025-08-26BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510629544.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing graph databases have low query performance when performing precise query tasks and column scanning tasks, and traditional dual-format storage solutions lead to increased storage costs.

Method used

By predicting the cumulative query cost indicators of adjacent tables under row format and column format storage, dynamically select the target storage format, optimize resource utilization, and avoid data redundancy.

Benefits of technology

Improve the query performance of query tasks, reduce storage costs, and optimize resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541270A_ABST
    Figure CN120541270A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a graph database processing method and device, a storage medium and a program product. Predicting a first accumulated query cost index when the adjacency list is stored in a row format and a second accumulated query cost index stored in a column format; determining a target storage format of the adjacency list according to the first accumulated query cost index and the second accumulated query cost index; in leaf nodes of the index tree of the graph database, the adjacency list is stored in the target storage format. According to the embodiment of the invention, the target storage format of the adjacency list is dynamically selected by measuring the query cost of different storage formats according to the adjacency list of the target graph data in combination with the condition of the historical query task set, so that the storage cost can be reduced, unnecessary data redundancy is avoided, resource utilization is optimized, and the query performance of the query task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a graph database processing method, device, storage medium, and program product. Background Art

[0002] A graph database is a database specifically designed for efficiently storing, querying, and processing relational data. It uses a graph structure (nodes, edges, attributes) to directly represent data and its relationships.

[0003] Currently, graph databases typically use adjacency lists to efficiently record the connection relationships between vertices in the graph. However, in the existing technology, the adjacency lists of graph databases are usually stored in row format, which results in poor query performance when facing query tasks such as precise query tasks and column scanning tasks. Summary of the Invention

[0004] The embodiments of the present disclosure provide a graph database processing method, device, storage medium, and program product to dynamically select the target storage format of the adjacency table, optimize resource utilization, and improve the query performance of query tasks.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for processing a graph database, including:

[0006] Predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format;

[0007] determining a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator;

[0008] In a leaf node of an index tree of the graph database, the adjacency list is stored in the target storage format.

[0009] In a second aspect, an embodiment of the present disclosure provides a graph database processing device, including:

[0010] A prediction unit is configured to predict, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format;

[0011] a format determining unit, configured to determine a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator;

[0012] A storage unit is used to store the adjacency list in the leaf node of the index tree of the graph database using the target storage format.

[0013] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0014] The memory stores computer-executable instructions;

[0015] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the graph database processing method described in the first aspect and various possible designs of the first aspect.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the graph database processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0017] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the graph database processing method described in the first aspect and various possible designs of the first aspect.

[0018] The processing method, device, storage medium and program product of the graph database provided by the embodiments of the present disclosure predict the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when it is stored in column format based on the historical query task set of the adjacency table of the target graph data in the graph database; determine the target storage format of the adjacency table based on the first cumulative query cost index and the second cumulative query cost index; and store the adjacency table in the target storage format in the leaf node of the index tree of the graph database. In the embodiments of the present disclosure, for the adjacency table of the target graph data, combined with the situation of the historical query task set, the target storage format of the adjacency table is dynamically selected by weighing the query costs of different storage formats, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of query tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 A diagram illustrating a scenario of a method for processing a graph database provided in one embodiment of the present disclosure;

[0021] Figure 2 A flowchart of a method for processing a graph database provided in one embodiment of the present disclosure;

[0022] Figure 3 A schematic flow chart of a method for processing a graph database provided in another embodiment of the present disclosure;

[0023] Figure 4 A schematic diagram of an adjacency table provided in an embodiment of the present disclosure being stored in row format;

[0024] Figure 5 A schematic diagram of an adjacency table provided in an embodiment of the present disclosure being stored in a column format;

[0025] Figure 6 A structural block diagram of a graph database processing device provided in one embodiment of the present disclosure;

[0026] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0028] First, the technical terms in this disclosure are explained:

[0029] Graph database: A database specifically designed for efficient storage, query, and processing of relational data, which uses a graph structure (nodes, edges, attributes) to directly represent data and its relationships.

[0030] Adjacency list: It is one of the most commonly used storage representation methods for graph data structures. It efficiently records the connection relationships between vertices in the graph.

[0031] Row format storage: the entire row of data is stored continuously on disk / memory;

[0032] Column format storage: Each column of data is stored separately.

[0033] Exact Query: This means that the query conditions must completely match the nodes, edges, or attributes in the graph, without involving fuzzy matching or complex conditions. For example, you can retrieve all the attributes of the 10 nearest edges of a specific node's direct neighbor nodes.

[0034] Wide Column Scan: This is a data access mode for wide-column databases that allows efficient scanning of datasets with a large number of columns (possibly sparsely distributed). For example, it can scan a single column with excessive data. It can also filter data based on specific conditions when scanning columns.

[0035] In the prior art, graph databases typically use adjacency lists to efficiently record the connection relationships between vertices in the graph. However, the adjacency lists of graph databases in the prior art are typically stored in row format, which results in excellent performance when executing precise query tasks such as filter, limit, and property, but poor query performance when facing column scanning tasks.

[0036] Considering that databases typically store data in both row and column formats, we evaluated the performance of precise query and column scan tasks and found that when the query per second (QPS) was low, the difference in CPU consumption between different data formats was not significant. However, as the QPS increased, the difference in CPU consumption gradually became apparent. Specifically, using the same number of CPU cores, in the precise query task scenario, the maximum throughput of the row format was 1.65 times that of the column format. Conversely, in the column scan scenario, the maximum throughput using the column format was 2.02 times that of the row format.

[0037] Traditional databases have established practices for meeting the diverse needs of precise query tasks and column scan tasks. It's generally believed that row storage is more suitable for precise query tasks, while column storage is more suitable for column scan tasks. The aforementioned test results also confirm that this approach remains applicable and effective for graph database workloads. Given the extremely high throughput of OLTP workloads, with single-machine clusters reaching millions of queries per second (QPS), performance optimization becomes our top priority. To this end, we need to design a dual-format storage solution for adjacency tables, optimizing for precise query tasks (row scans) and column scans, respectively, to fully meet the high-performance requirements of different workloads.

[0038] The cost increases associated with multiple storage formats need to be effectively controlled. To address the dual challenges of massive data volumes and highly variable workloads, traditional HTAP databases generally employ data redundancy, maintaining both row- and column-based storage formats to optimize performance. However, in scenarios dealing with massive data volumes, this approach inevitably leads to a sharp increase in storage costs. While dual-format storage improves data processing efficiency to accommodate diverse business scenarios, it also presents the challenge of data redundancy. Given that graph data and its access patterns generally follow a power-law distribution, fully redundant storage in multiple formats is clearly not cost-effective for adjacency tables with smaller data volumes and lower access throughput. Therefore, it is essential to thoroughly analyze the distribution characteristics of graph data and design an adaptable data storage solution to meet business needs for reducing storage costs.

[0039] In order to solve the above technical problems, the embodiment of the present disclosure provides a method for processing a graph database, which predicts a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in the graph database; determines the target storage format of the adjacency table based on the first cumulative query cost index and the second cumulative query cost index; and stores the adjacency table in the target storage format in the leaf node of the index tree of the graph database. In this embodiment, for the adjacency table of the target graph data, combined with the situation of the historical query task set, the target storage format of the adjacency table is dynamically selected by weighing the query costs of different storage formats, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of the query task.

[0040] The application scenarios of the graph database processing method of the embodiment of the present disclosure are as follows: Figure 1 As shown, based on the set of historical query tasks for the adjacency table of target graph data in the graph database, the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when it is stored in column format are predicted; based on the first cumulative query cost index and the second cumulative query cost index, the target storage format of the adjacency table is determined; in the leaf node of the index tree of the graph database, the adjacency table is stored in the target storage format.

[0041] It should be noted that the activation of relevant functions of the embodiments of the present disclosure, the data obtained, the processing and storage methods of the data, etc., should all be authorized in advance by the user and other rights holders associated with the user, and should comply with the relevant laws and regulations and the agreement rules between rights holders.

[0042] The processing method of the graph database disclosed in the present invention will be introduced in detail below in conjunction with specific embodiments.

[0043] refer to Figure 2 , Figure 2 This is a flow chart of a graph database processing method provided in one embodiment of the present disclosure. The method of this embodiment can be applied to electronic devices such as terminal devices or servers. The graph database processing method includes:

[0044] S201. Based on a set of historical query tasks for an adjacency table of target graph data in a graph database, predict a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format.

[0045] In this embodiment, a graph database may include multiple graph data, where graph data consists of nodes and edges connecting nodes. Nodes can represent entities, such as people, places, products, and events, while edges represent connections between nodes. Attributes are information attached to nodes or edges, used to store entity characteristics or metadata about relationships. For any graph data in a graph database, an adjacency list can be used to represent the graph data structure, efficiently describing the connections and attributes between nodes. In an adjacency table, a row of data represents the connection relationship and attributes between a node and a neighboring node. If a node has multiple neighboring nodes, the connection relationship and attributes between the node and each neighboring node occupy a row of data.

[0046] The adjacency table of a traditional graph database is usually stored in row format, storing the entire row data of the adjacency table continuously together. However, the adjacency table stored in a single row format cannot meet different query tasks, especially column scanning tasks, and has poor workload performance. It is possible to consider using a dual format storage for the adjacency table, which is both row format and column format. The column format stores each column data of the neighbor node separately, which can better meet the needs of column scanning tasks. However, if the adjacency tables of all graph data in the graph database are stored in dual format, although it can improve performance, it will face the problem of increased cost due to data redundancy. In particular, it is not cost-effective for adjacency tables with small data volume and low access throughput. For adjacency tables with relatively few column scanning tasks, the column format storage in the dual format storage will lead to increased costs. For adjacency tables with relatively few precise query tasks, the row format storage in the dual format storage will lead to increased costs, etc. In order to avoid unnecessary data redundancy and reduce storage costs, in this embodiment, the appropriate storage format can be dynamically selected for the adjacency table of the target graph data based on the historical query tasks of the adjacency table of the target graph data, rather than fixedly using a single format storage or fixed dual format storage, to optimize resource utilization.

[0047] Specifically, in this embodiment, a historical query task set of the adjacency table of the target graph data can be obtained, which may include historical precise query tasks and historical column scanning tasks. The historical query task set may be a historical query task set within a predetermined time period in the past. Furthermore, based on the historical query task set, the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when the adjacency table is stored in column format can be predicted. For any precise query task, the query cost index when the adjacency table is stored in row format is usually lower than the query cost index when the adjacency table is stored in column format, and for any column scanning task, the query cost index when the adjacency table is stored in row format is usually higher than the query cost index when the adjacency table is stored in column format. Therefore, when the adjacency table is stored in row format, the query cost index of each precise query task and column scanning task can be predicted and accumulated to obtain the first cumulative query cost index. When the adjacency table is stored in column format, the query cost index of each precise query task and column scanning task can be predicted and accumulated to obtain the second cumulative query cost index. This is so that by comparing the first cumulative query cost index and the second cumulative query cost index, the target storage format of the adjacency table can be determined, thereby reducing storage costs by unnecessary data redundancy.

[0048] Optionally, in this embodiment, the query cost indicator may be measured by query time consumption, or may be measured by resource consumption, or may be measured by other consumption, or may be measured by comprehensively considering multiple consumptions.

[0049] S202: Determine a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator.

[0050] In this embodiment, after the first cumulative query cost indicator and the second cumulative query cost indicator are predicted, the target storage format of the adjacency table may be determined based on the first cumulative query cost indicator and the second cumulative query cost indicator.

[0051] Optionally, the size of the first cumulative query cost indicator and the second cumulative query cost indicator can be simply considered. If the first cumulative query cost indicator is greater than the second cumulative query cost indicator, the target storage format adopts the row format. If the first cumulative query cost indicator is less than the second cumulative query cost indicator, the target storage format adopts the column format. If the first cumulative query cost indicator is equal to the second cumulative query cost indicator, either the row format or the column format can be adopted. There is no need to convert the storage format, and the current storage format can be maintained.

[0052] Optionally, the following logic can be used to determine the target storage format:

[0053] If both the first cumulative query cost indicator and the second cumulative query cost indicator exceed the preset cost threshold, determining that the target storage format is a dual format, where the dual format is a coexistence of a row format and a column format; or

[0054] If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed a preset cost threshold, determining a difference between the first cumulative query cost indicator and the second cumulative query cost indicator;

[0055] If the difference exceeds the preset difference, determining the target storage format to be the storage format corresponding to the smaller one of the first cumulative query cost indicator and the second cumulative query cost indicator;

[0056] If the difference does not exceed the preset difference, the target storage format is determined to be the current storage format.

[0057] In this embodiment, if both the first cumulative query cost indicator and the second cumulative query cost indicator exceed the preset cost threshold, it means that the first cumulative query cost indicator and the second cumulative query cost indicator are both very high, that is, the query cost of using only one storage format is very high. Therefore, the target storage format is determined to be a dual format in which row format and column format coexist. Then, the precise query task can apply to the data in row format, and the column scan task can apply to the data in column format, which can effectively reduce the query cost.

[0058] If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed the preset cost threshold, it means that the cost of using the storage format corresponding to the smaller of the cumulative query cost indicators is relatively low. However, since the storage format conversion also requires a certain amount of resources and time, it is also necessary to consider whether it is necessary to convert the storage format. The size of the gap between the first cumulative query cost indicator and the second cumulative query cost indicator can be compared. If the difference (absolute value) between the first cumulative query cost indicator and the second cumulative query cost indicator exceeds the preset difference, it means that the gap between the first cumulative query cost indicator and the second cumulative query cost indicator is relatively large. The query is executed after the storage format conversion. The benefit of the task (that is, the reduced query cost) is greater than the consumption of executing the storage format conversion process, so the storage format corresponding to the smaller of the cumulative query cost indicators can be selected as the target storage format; if the difference (absolute value) between the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed the preset difference, it means that the gap between the first cumulative query cost indicator and the second cumulative query cost indicator is not large, and the benefit of executing the query task after the storage format conversion (that is, the reduced query cost) is not greater than the consumption of executing the storage format conversion process, so the storage format conversion can be omitted and the current storage format can be maintained. The current storage format is the storage format currently used, which can be a row format or a column format.

[0059] Of course, logic of other target storage formats may also be used in this embodiment, which is not limited in this embodiment.

[0060] S203. In a leaf node of the index tree of the graph database, the adjacency list is stored in the target storage format.

[0061] In this embodiment, after determining the target storage format of the adjacency table of the target graph data, the adjacency table can be stored in the target storage format. In this embodiment, the adjacency table can be stored in the target storage format in the leaf node of the index tree of the graph database, wherein the index tree is an index structure specially optimized for the graph data model, which is used to accelerate the search, traversal and relationship query of nodes, edges and attributes. The optional index tree can be a BW index tree. The BW index tree is divided into three layers, from top to bottom: the BW tree index layer (root), the buffer layer (internal), and the storage layer. The BW tree index layer provides an API for operating this index tree structure. The buffer layer connects the index layer and the storage layer, and records the mapping of logical page numbers to physical pointers through a mapping table. The storage layer may include multiple leaf nodes (LeafPage). The data of an adjacency table is usually stored in the same leaf node. The leaf node can provide two different storage formats for the adjacency table. Therefore, the adjacency table can be stored in the leaf node in the target storage format.

[0062] Regardless of whether the row format or the column format is used, the basic structure in the leaf node is similar, mainly consisting of a pointer array and data blocks (which may include row blocks or column blocks).

[0063] It should be noted that the above-mentioned graph database processing method can be executed according to manual triggering instructions, or it can be executed periodically, or it can be executed when changes in the situation of precise query tasks and column scanning tasks, or changes in query costs (such as extended time consumption, etc.) are detected. In this way, the target storage format of the adjacency table can be dynamically adjusted to adapt to changes in various situations, optimize performance and control costs.

[0064] The graph database processing method provided in this embodiment predicts a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in the graph database; determines the target storage format of the adjacency table based on the first cumulative query cost index and the second cumulative query cost index; and stores the adjacency table in the target storage format in the leaf nodes of the index tree of the graph database. In this embodiment, for the adjacency table of target graph data, combined with the situation of the historical query task set, the target storage format of the adjacency table is dynamically selected by measuring the query cost of the adjacency table in row format and column format, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of query tasks.

[0065] Based on the above-mentioned graph database processing method, different proportions of real query tasks of OLTP (Online Transaction Processing) and OLAP (Online Analytical Processing) were selected for testing. The extreme throughput performance was evaluated by adopting single row format storage, single column format storage, dual format storage, and the dynamic selection of target storage format of the embodiment of the present disclosure for the adjacency table. In the three scenarios of "more OLTP", "OLTP and OLAP are equal", and "more OLAP", the dynamic selection of target storage format of the embodiment of the present disclosure improved the performance by approximately 31.8%, 53.5%, and 70.5% respectively compared with single row format storage; compared with single column format storage, the dynamic selection of target storage format of the embodiment of the present disclosure improved the performance by approximately 31.8%, 53.5%, and 70.5% respectively. Compared with single-row format storage, the performance is improved by approximately 40%, 19.1% and 11.7% respectively; dual-format storage uses redundant row and column data to serve OLTP and OLAP respectively, thereby achieving the most extreme performance, but it brings about twice the idle overhead compared to single row format storage or single column format storage. The dynamic selection of the target storage format in the embodiment of the present disclosure saves 41.25%, 45.75% and 47.25% of storage space respectively compared with dual-format storage, while the optimal performance only decreases by approximately 2.2%, 5.3% and 3.7%.

[0066] Based on any of the above embodiments, the method of predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format, includes:

[0067] Obtaining a task scale feature of each query task included in the historical query task set; wherein the task scale feature includes the number of rows of data queried by the query task and the number of attributes queried in the row of data;

[0068] According to the task scale characteristics of each query task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format are predicted.

[0069] In this embodiment, when predicting the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when it is stored in column format, taking into account the different task scale characteristics of each precise query task and column scanning task, the corresponding query costs are also different. If a query task queries more row data, the query cost will be greater, and the more attributes queried in the row data, the query cost will also be greater. Therefore, the task scale characteristics of the precise query tasks and column scanning tasks included in the historical query task set can be obtained, and then based on the task scale characteristics of the precise query tasks and the column scanning tasks, the query costs of each precise query task and each column scanning task when the adjacency table is stored in row format and column format are predicted respectively. Then, by accumulation, the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when it is stored in column format are obtained, thereby improving the accuracy of predicting the first cumulative query cost index and the second cumulative query cost index.

[0070] Furthermore, when predicting the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when the adjacency table is stored in column format based on the task scale characteristics of each query task, the following steps may be specifically performed:

[0071] Determining a unit cost index of each query behavior in various types of query tasks under different storage formats of the adjacency table, wherein the types of query tasks include exact query tasks and column scan tasks;

[0072] Predicting, based on the task scale characteristics of each query task and the unit cost index, a query cost index for each query task when the adjacency table is stored in a row format and a column format;

[0073] The first cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in row format, and the second cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in column format.

[0074] In this embodiment, the unit cost indicators of each query behavior in various types of query tasks under different storage formats of the adjacency table are different. The unit cost indicators of each query behavior include but are not limited to the unit cost of querying a single leaf node where the adjacency table is located, the unit cost of querying a row of data in the adjacency table, and the unit cost of querying a certain attribute in a certain row. The unit cost indicators of each query behavior in various types of query tasks under different storage formats of the adjacency table can be determined in combination with historical query tasks. Furthermore, based on the task scale characteristics of the precise query task and the column scanning task and the above-mentioned unit costs, the query cost of each query task when the adjacency table is stored in row format can be more accurately and conveniently predicted, and the first cumulative query cost indicator can be accumulated. The query cost of each query task when the adjacency table is stored in column format can also be predicted, and the second cumulative query cost indicator can be accumulated. Among them, when the adjacency table is stored in row format or column format, the cost of any query task locating the leaf node of the adjacency table can be determined according to the above-mentioned unit cost index, and the cost of the query task querying all target row target attributes can be determined according to the task scale characteristics and the above-mentioned unit cost index, and accumulated to obtain the query cost index of the query task.

[0075] Optionally, considering that the prediction cost index has a high requirement for timeliness, in order to simplify the prediction process, in this embodiment, the unit cost index of some query behaviors of the adjacency table is represented by an average cost index. Specifically, the unit cost index of the query behavior in various query tasks of the adjacency table in different storage formats may include a first cost constant, a second cost constant, a first unit cost, a second unit cost, and a third unit cost.

[0076] The first cost constant (denoted as C) is the average cost index of the single leaf node location, cache overhead and read / write overhead of each query task of the adjacency table. The second cost constant (denoted as B) is the average cost index of the single row data location and row data version query of each query task of the adjacency table. The first unit cost (denoted as a row ) is the average cost index of a single attribute query in any row of data for any query task when the adjacency table is stored in row format. The second unit cost (denoted as a ’ col,eq ) and the third unit cost (denoted as a ’ col,wcs ) are the average cost indicators of the precise query task and the column scan task for single attribute queries in any column data when the adjacency table is stored in column format.

[0077] By using the corresponding average cost index to represent the above cost characteristics, the requirement for timeliness can be reduced, and the short-term fluctuation of cost characteristics over time that increases the difficulty of prediction can be avoided.

[0078] Based on the above embodiment, predicting the first cumulative query cost indicator when the adjacency table is stored in row format may specifically include:

[0079] For any query task among the precise query task and the column scan task, determining a query cost index for any row of data according to the number of attributes queried in the row of data, the first unit cost, and the second cost constant; determining a query cost index for any query task according to the query cost index for any row of data, the number of row data queried, and the first cost constant;

[0080] The query cost index of each query task is accumulated to obtain a first accumulated query cost index when the adjacency table is stored in a row format.

[0081] In this embodiment, when the adjacency table is stored in row format, for any precise query task or any column scan task, leaf node location must be performed based on the query conditions, i.e., the leaf node where the adjacency table is located must be located. This also incurs cache overhead and read / write overhead, which is the first cost constant C mentioned above.

[0082] Secondly, for any precise query or column scan task, it is necessary to locate the row data in the adjacency table according to the query conditions. This is usually done through a pointer array. If different versions are involved, the row data of the required version must also be located. This part of the query cost is also the second cost constant B mentioned above.

[0083] Then, for any precise query task or any column scan task, it is necessary to query the required attributes from the row data, and the query cost is positively correlated with the number of queried attributes. The first unit cost a row Cost is the average cost index of a single attribute query in any row of data for any query task when the adjacency table is stored in row format. Therefore, the query cost index for querying each row of data is also Cost. single_row =a row ×N attr +B, where N attr The number of attributes being queried in the row data;

[0084] If a query task needs to query multiple rows of data, the query cost index of multiple rows of data is also Cost multi_row =N rows ×Cost single_row , where N rows The number of rows of data queried;

[0085] The query cost index of each query task is accumulated to obtain the first cumulative query cost index:

[0086]

[0087] Where i represents the i-th query task, i ranges from 1 to Q, and Q represents the number of query tasks in the historical query task set.

[0088] Based on the above embodiment, predicting the second cumulative query cost indicator of the adjacency table stored in a column format may specifically include:

[0089] For any precise query task, determining a query cost index for any row of data based on the number of attributes being queried in the row of data, the second unit cost, and the second cost constant; determining a query cost index for any precise query task based on the query cost index for any row of data, the number of rows of data being queried, and the first cost constant;

[0090] For any column scanning task, determining a query cost index for any column data according to the number of row data being queried and the second unit cost; determining a query cost index for any column scanning task according to the query cost index for any column data, the number of attributes being queried in the row data, and the first cost constant;

[0091] The query cost index of each precise query task and the query cost index of each column scanning task are accumulated to obtain a second accumulated query cost index when the adjacency table is stored in a column format.

[0092] In this embodiment, when the adjacency table is stored in a column format, there is a certain difference in the query cost between the precise query task and the column scan task.

[0093] First, when the adjacency table is stored in column format, for any precise query task or any column scan task, it is necessary to locate the leaf node according to the query conditions, that is, to locate the leaf node where the adjacency table is located. At the same time, there will also be cache overhead and read and write overhead. This part of the query cost is also the first cost constant mentioned above. It should be noted that the first cost constant may be different for precise query tasks or column scan tasks, respectively expressed as C ’ col,eq and C ’ col,wcs express;

[0094] Secondly, for any precise query task, the query process when the adjacency table is stored in column format is very similar to that in row format. However, due to the difference in row and column storage methods, the cost of accessing an attribute is also different. In this embodiment, the second unit cost (denoted as a ’ col,eq) represents the average cost index of the precise query task for a single attribute query in any column of data when the adjacency table is stored in column format. Therefore, any precise query task needs to query N of any row of data. attr The number of attributes requires access to N attr Column data, that is, the query cost index of the precise query task to query any row of data is (a ’ col,eq ×N attr +B), and if an accurate query task requires querying multiple rows of data, the query cost index of multiple rows of data is also Cost col,eq,multi_row =N rows ×(a ’ col,eq ×N attr +B), where N rows The number of rows of data queried;

[0095] For any column scanning task, when the adjacency table is stored in column format, it is necessary to scan each column of data to obtain data that meets the screening conditions. The access cost of each column of data is proportional to the number of rows that need to be scanned. In this embodiment, the third unit cost a is used. ’ col,wcs It represents the average cost index of a column scan task for a single attribute query in any column data when the adjacency table is stored in column format, that is, the average cost index of scanning a row in the column data. Therefore, for any column scan task, the query cost index of scanning each column is a ’ col,wcs ×N rows , and if a column scan task needs to access multiple columns of data, the query cost index of multiple columns of data is also Cost col,wcs,multi_row =a ’ col,eq ×N rows ×N attr ;

[0096] The query cost index of each precise query task and the query cost index of each column scan task are accumulated to obtain the second cumulative query cost index:

[0097]

[0098] Among them, i represents the i-th query task, Q wcs Indicates the number of column scan tasks in the historical query task set, Q eq Indicates the exact number of query tasks in the historical query task set.

[0099] On the basis of any of the above embodiments, the unit cost indicators of various query behaviors in various types of query tasks in the adjacency table under different storage formats, including the first cost constant, the second cost constant, the first unit cost, the second unit cost and the third unit cost, etc., can be determined based on historical query tasks. For example, assuming that the query time is used as the cost indicator, the task scale characteristics and time consumption of the sample adjacency table stored in the row format in executing various precise query tasks and column scanning tasks can be obtained, and the task scale characteristics and time consumption of the sample adjacency table stored in the column format in executing various precise query tasks and column scanning tasks can be obtained, and they can be respectively substituted into the above formulas, and each unit cost indicator is used as the unknown number to be solved to solve each unit cost indicator.

[0100] Based on any of the above embodiments, when the target storage format is used to store the adjacency table, the following steps may be specifically included:

[0101] If the current storage format of the adjacency list is the dual format and the target storage format is the row format or the column format, clearing the data of the adjacency list that is not in the target storage format by garbage collection; or

[0102] If the current storage format of the adjacency list is not the dual format, and the target storage format is different from the current storage format, the data in the current storage format of the adjacency list is converted into data in the target storage format.

[0103] In this embodiment, if the current storage format of the adjacency list is dual format, and the target storage format is a single storage format of row format or column format, then no real storage format conversion actually occurs, and only redundant data is cleaned up. Therefore, it is only necessary to clean up the non-target storage format data of the adjacency list through garbage collection. For example, if the current storage format is dual format and the target storage format is row format, then it is only necessary to retain the row format and clean up the column format.

[0104] If the current storage format of the adjacency list is not dual format, and the target storage format is different from the current storage format, storage format conversion is involved, that is, conversion from row format to column format, or from column format to row format. The process of converting from a single storage format to dual format also involves conversion from row format to column format, or from column format to row format.

[0105] In the above embodiment, when it comes to storage format conversion, considering that the adjacency table may still be updated in real time, in order to avoid the storage format conversion affecting the online real-time update of the adjacency table and not losing the updated data, the storage format conversion process can be as follows: Figure 3 As shown, specifically including:

[0106] S301, storing the incremental data of the adjacency table in an additional storage space;

[0107] S302: Obtain a snapshot of the existing data in the current storage format of the adjacency list, and obtain data in the target storage format of the adjacency list based on the snapshot;

[0108] S303. Write the incremental data in the additional storage space into the data in the target storage format, stop storing the subsequent incremental data of the adjacency table in the additional storage space, and directly write the subsequent incremental data of the adjacency table into the data in the target storage format.

[0109] In this embodiment, when it is necessary to write incremental data to the adjacency table when the adjacency table is updated, the incremental data can be intercepted and stored in an additional storage space instead of being directly written to the adjacency table. Then, a snapshot of the existing data in the current storage format of the adjacency table is obtained, and the data is converted into the target storage format based on the snapshot, that is, the row format is converted into the column format, or the column format is converted into the row format based on the snapshot, and then the incremental data in the additional storage space is written into the data in the target storage format, and the subsequent incremental data of the adjacency table is stopped from being stored in the additional storage space, but is directly written into the data in the target storage format. In this way, the storage format conversion is completed and the impact on the adjacency table update is reduced.

[0110] Optionally, in S303, in the process of writing the incremental data in the additional storage space into the data in the target storage format, the incremental data of the adjacency table will continue to be stored in the additional storage space. This may result in the inability to complete the writing of all the incremental data in the additional storage space into the data in the target storage format. Therefore, when the incremental data in the additional storage space is written into the data in the target storage format, when the incremental data in the additional storage space is less than the preset data volume, the subsequent incremental data of the adjacency table can be stopped from being stored in the additional storage space. After waiting for all the incremental data in the additional storage space to be written into the data in the target storage format, the subsequent incremental data of the adjacency table can be started to be directly written into the data in the target storage format, thereby reducing the time affecting the update of the adjacency table and eliminating the need for offline storage format conversion.

[0111] It should be noted that if the current storage format of the adjacency list is a single storage format and the target storage format is another single storage format, then when the incremental data in the additional storage space is written to the data in the target storage format, the incremental data in the additional storage space is also synchronously written to the data in the current storage format. When the subsequent incremental data of the adjacency list can be directly written to the data in the target storage format and the data in the current storage format, the data in the current storage format of the adjacency list is cleared through garbage collection, thereby ensuring the security of the storage format conversion process and avoiding data loss due to conversion failure. Of course, if the target storage format is a dual format, there is no need to clear the data in the current storage format of the adjacency list through garbage collection.

[0112] Based on any of the above embodiments, the method further includes:

[0113] Receiving a target precise query task for an adjacency table of target graph data; searching for a leaf node where the adjacency table of the target graph data is located in an index tree of the graph database according to the target precise query task;

[0114] If the target storage format is a row format or a column format, searching for target row data in the leaf node according to the target precise query task, and obtaining a query result based on the target row data; or

[0115] If the target storage format is dual format, target row data is searched in the row format data in the leaf node according to the target precise query task, and query results are obtained according to the target row data.

[0116] In this embodiment, subsequent query tasks can be performed based on the data in the target storage format of the adjacency table.

[0117] The data query interface for LeafPage, the leaf node of the index tree, offers two key features: row scanning and filtering capabilities, and column scanning of large amounts of data. Furthermore, it supports a multi-version storage mechanism to meet the multi-version concurrency control requirements of transactions. The row scanning feature allows users to quickly locate a specific row of data within the index tree and, based on this, traverse adjacent rows in a specified direction. During the traversal process, data that meets the requirements is filtered based on attribute filtering conditions, and the version chain is searched to retrieve the required specific version. For example, when querying the 10 most recently liked videos of a user node (i.e., the user's neighbor nodes), the system will locate a video record starting from the current time point and retrieve 10 edges in reverse order. Secondly, LeafPage provides a large-scale column scanning capability. This query can scan all data in the LeafPage, including neighbor vertices and required attributes, making it particularly suitable for complex multi-degree analysis scenarios. During the scanning process, the system uses the data's timestamp columns to ensure that the data version that meets the time requirements is retrieved.

[0118] The adjacency table of the target graph data is stored in row format as follows Figure 4 As shown, using column format storage can be as follows Figure 5 shown.

[0119] Furthermore, for target precise query tasks, such as query Figure 4 and Figure 5 For all attributes of K3 in the adjacency table, regardless of whether the target graph data's adjacency table is in row or column format, the index tree must first be searched for the leaf node (K1-K5) where the target graph data's adjacency table is located. The next step, for row-stored and column-stored queries, is similar. The target row data can be searched for in the leaf node according to the target precise query task, and query results can be obtained based on the target row data. In specific implementation, after locating the leaf node where the adjacency table is located, the target row data is searched for using a pointer array. A determination is then made as to whether the version of the target row data meets the required version requirements, such as MVCC (Multi-Version Concurrency Control) requirements, which require that the version creation time match the creation time of the target precise query task. If the required version requirements are not met, the target row data's version chain is searched for a version that meets the required version requirements, and the query results are then obtained from that version. The version chain is formed by linking different versions of the target row data via pointers, with each version including a pointer to the previous version.

[0120] It should be noted that the difference between the column format storage and row format storage when performing target precise query tasks is that each attribute of a row of data in the adjacency table is not continuous in memory and needs to be queried separately.

[0121] If the target storage format is dual format, considering that the query cost of performing precise query tasks on row-format data is relatively lower, the target row data can be searched in the row-format data in the leaf node according to the target precise query task, and the query results can be obtained based on the target row data.

[0122] Based on any of the above embodiments, the method further includes:

[0123] Receive a target column scanning task for an adjacency table of target graph data; search an index tree of the graph database for a leaf node where the adjacency table of the target graph data is located according to the target column scanning task;

[0124] If the target storage format is a row format, traverse each row of data in the leaf node and obtain the target attribute queried by the target column scanning task from each row of data; or

[0125] If the target storage format is a column format, then the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained according to the attribute column data, the creation time column data, and the deletion time column data; or

[0126] If the target storage format is dual format, the attribute column data, creation time column data and deletion time column data queried by the target column scanning task are obtained from the column format data in the leaf node, and the target attributes queried by the target column scanning task are obtained based on the attribute column data, the creation time column data and the deletion time column data.

[0127] In this embodiment, for a target column scanning task, such as query Figure 4 and Figure 5 For an attribute in an attribute column of the adjacency table that meets the preset conditions, no matter whether the adjacency table of the target graph data adopts row format or column format, it is necessary to first search the leaf node leaf (K1-K5) where the adjacency table of the target graph data is located in the index tree;

[0128] Furthermore, if the target storage format is a row format, each row of data is traversed in the leaf node, and the target attribute queried by the target column scanning task is obtained from each row of data. In the specific implementation, after locating the leaf node where the adjacency list is located, each row of data is traversed according to the pointer array (Pointer Array), and then it is determined whether the version of the row of data meets the required version requirements. For example, MVCC requires that the version creation time must match the creation time of the target precise query task. If the required version requirements are not met, the version chain of the row of data is searched for a version that meets the required version requirements, and then the target attribute to be queried is obtained from the version, and then it is determined whether the target attribute of the row of data meets the preset conditions.

[0129] If the target storage format is a column format, the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained based on the attribute column data, creation time column data, and deletion time column data. In this embodiment, considering that the column format is not easy to link various versions in the form of a version chain, the creation time column data and the deletion time column data can be used to judge the version of the target attribute, so as to find the target attribute that meets the preset conditions and meets the required version requirements, and meets the required version requirements, that is, the time of the required version falls between the creation time and the deletion time.

[0130] If the target storage format is dual format, considering that the query cost of performing a column scan task on column format data is relatively lower, the target column scan task can be performed on the column format data in the leaf node.

[0131] Based on any of the above embodiments, the data not pointed to by the pointer array in the leaf node can be cleaned up through garbage collection, that is, Figure 4 and Figure 5 The compaction process in can realize the recovery of storage resources.

[0132] Corresponding to the processing method of the graph database in the above embodiment, Figure 6 A structural block diagram of a graph database processing device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 6 The graph database processing device 600 includes: a prediction unit 601, a format determination unit 602, and a storage unit 603.

[0133] The prediction unit 601 is configured to predict, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format;

[0134] a format determining unit 602, configured to determine a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator;

[0135] The storage unit 603 is used to store the adjacency list in the leaf node of the index tree of the graph database using the target storage format.

[0136] The processing device of the graph database provided by the embodiment of the present disclosure predicts a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in the graph database; determines the target storage format of the adjacency table based on the first cumulative query cost index and the second cumulative query cost index; and stores the adjacency table in the target storage format in the leaf node of the index tree of the graph database. In this embodiment, for the adjacency table of the target graph data, the target storage format of the adjacency table is dynamically selected by weighing the query costs of different storage formats in combination with the situation of the historical query task set, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of query tasks.

[0137] In one or more embodiments of the present disclosure, the prediction unit 601, when predicting a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in a graph database, is configured to:

[0138] Obtaining task scale characteristics of the precise query task and the column scan task included in the historical query task set; wherein the task scale characteristics include the number of row data queried by the query task and the number of attributes queried in the row data;

[0139] According to the task scale characteristics of the precise query task and the column scan task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format are predicted.

[0140] In one or more embodiments of the present disclosure, the prediction unit 601, when predicting the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when the adjacency table is stored in column format based on the task scale characteristics of the precise query task and the column scan task, is configured to:

[0141] Determining cost characteristics of various query tasks for the adjacency table in different storage formats;

[0142] According to the task scale characteristics and cost characteristics of the precise query task and the column scan task, a first cumulative query cost index when the adjacency table is stored in row format and a second cumulative query cost index when the adjacency table is stored in column format are predicted.

[0143] In one or more embodiments of the present disclosure, the cost feature includes a first cost constant, a second cost constant, a first unit cost, a second unit cost, and a third unit cost;

[0144] The first cost constant is the average cost index of leaf node positioning, cache overhead and read-write overhead for each query task of the adjacency table; the second cost constant is the average cost index of row data positioning and row data version query for each query task of the adjacency table; the first unit cost is the average cost index of single attribute query in any row data by any query task when the adjacency table is stored in row format; the second unit cost and the third unit cost are respectively the average cost index of single attribute query in any column data by precise query task and column scanning task when the adjacency table is stored in column format.

[0145] In one or more embodiments of the present disclosure, when predicting the first cumulative query cost indicator when the adjacency table is stored in row format based on the task scale characteristics and cost characteristics of the precise query task and the column scan task, the prediction unit 601 is configured to:

[0146] For any query task among the precise query task and the column scan task, determining a query cost index for any row of data according to the number of attributes queried in the row of data, the first unit cost, and the second cost constant; determining a query cost index for any query task according to the query cost index for any row of data, the number of row data queried, and the first cost constant;

[0147] The query cost index of each query task is accumulated to obtain a first accumulated query cost index when the adjacency table is stored in a row format.

[0148] In one or more embodiments of the present disclosure, when predicting the second cumulative query cost indicator of the adjacency table stored in a column format based on the task scale characteristics and cost characteristics of the precise query task and the column scan task, the prediction unit 601 is configured to:

[0149] For any precise query task, determining a query cost index for any row of data based on the number of attributes being queried in the row of data, the second unit cost, and the second cost constant; determining a query cost index for any precise query task based on the query cost index for any row of data, the number of rows of data being queried, and the first cost constant;

[0150] For any column scanning task, determining a query cost index for any column data according to the number of row data being queried and the second unit cost; determining a query cost index for any column scanning task according to the query cost index for any column data, the number of attributes being queried in the row data, and the first cost constant;

[0151] The query cost index of each precise query task and the query cost index of each column scanning task are accumulated to obtain a second accumulated query cost index when the adjacency table is stored in a column format.

[0152] In one or more embodiments of the present disclosure, when determining the target storage format of the adjacency table based on the first cumulative query cost indicator and the second cumulative query cost indicator, the format determination unit 602 is configured to:

[0153] If both the first cumulative query cost indicator and the second cumulative query cost indicator exceed a preset cost threshold, determining that the target storage format is a dual format, wherein the dual format is a coexistence of a row format and a column format; or

[0154] If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed a preset cost threshold, determining a difference between the first cumulative query cost indicator and the second cumulative query cost indicator;

[0155] If the difference exceeds a preset difference, determining the target storage format to be the storage format corresponding to the smaller one of the first cumulative query cost indicator and the second cumulative query cost indicator;

[0156] If the difference does not exceed the preset difference, the target storage format is determined to be the current storage format.

[0157] In one or more embodiments of the present disclosure, when the storage unit 603 stores the adjacency table in the target storage format, it is configured to:

[0158] If the current storage format of the adjacency list is the dual format and the target storage format is the row format or the column format, clearing the data of the adjacency list that is not in the target storage format by garbage collection; or

[0159] If the current storage format of the adjacency list is not the dual format, and the target storage format is different from the current storage format, the data in the current storage format of the adjacency list is converted into data in the target storage format.

[0160] In one or more embodiments of the present disclosure, when the storage unit 603 converts the data in the current storage format of the adjacency table into the data in the target storage format, it is configured to:

[0161] Storing the incremental data of the adjacency table in an additional storage space;

[0162] Obtaining a snapshot of the existing data in the current storage format of the adjacency list, and obtaining data in the target storage format of the adjacency list based on the snapshot;

[0163] The incremental data in the additional storage space is written into the data in the target storage format, and the subsequent incremental data of the adjacency table is stopped from being stored in the additional storage space, and the subsequent incremental data of the adjacency table is directly written into the data in the target storage format.

[0164] In one or more embodiments of the present disclosure, after converting the data in the current storage format of the adjacency table into the data in the target storage format, the storage unit 603 is further configured to:

[0165] If the target storage format is not the dual format, the data in the current storage format of the adjacency table is cleared by garbage collection.

[0166] In one or more embodiments of the present disclosure, the device further includes a query unit 604 configured to:

[0167] Receiving a target precise query task for an adjacency table of target graph data; searching for a leaf node where the adjacency table of the target graph data is located in an index tree of the graph database according to the target precise query task;

[0168] If the target storage format is a row format or a column format, searching for target row data in the leaf node according to the target precise query task, and obtaining a query result based on the target row data; or

[0169] If the target storage format is dual format, target row data is searched in the row format data in the leaf node according to the target precise query task, and query results are obtained according to the target row data.

[0170] In one or more embodiments of the present disclosure, the query unit 604 is further configured to:

[0171] Receive a target column scanning task for an adjacency table of target graph data; search an index tree of the graph database for a leaf node where the adjacency table of the target graph data is located according to the target column scanning task;

[0172] If the target storage format is a row format, traverse each row of data in the leaf node and obtain the target attribute queried by the target column scanning task from each row of data; or

[0173] If the target storage format is a column format, then the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained according to the attribute column data, the creation time column data, and the deletion time column data; or

[0174] If the target storage format is dual format, the attribute column data, creation time column data and deletion time column data queried by the target column scanning task are obtained from the column format data in the leaf node, and the target attributes queried by the target column scanning task are obtained based on the attribute column data, the creation time column data and the deletion time column data.

[0175] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0176] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.

[0177] refer to Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The electronic device 700 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0178] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0179] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0180] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0181] It should be noted that the computer-readable storage medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium that can transmit, propagate, or convey a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be conveyed using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0182] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0183] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0184] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0186] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0187] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0188] The electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present disclosure predict a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in a graph database; determine the target storage format of the adjacency table based on the first cumulative query cost index and the second cumulative query cost index; and store the adjacency table in the target storage format in the leaf nodes of the index tree of the graph database. In this embodiment, for the adjacency table of target graph data, combined with the situation of the historical query task set, the target storage format of the adjacency table is dynamically selected by weighing the query costs of different storage formats, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of query tasks.

[0189] In a first aspect, according to one or more embodiments of the present disclosure, a method for processing a graph database is provided, comprising:

[0190] Predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format;

[0191] determining a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator;

[0192] In a leaf node of an index tree of the graph database, the adjacency list is stored in the target storage format.

[0193] According to one or more embodiments of the present disclosure, the method of predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format includes:

[0194] Obtaining a task scale feature of each query task included in the historical query task set; wherein the task scale feature includes the number of rows of data queried by the query task and the number of attributes queried in the row of data;

[0195] According to the task scale characteristics of each query task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format are predicted.

[0196] According to one or more embodiments of the present disclosure, predicting, based on the task scale characteristics of each query task, a first cumulative query cost indicator when the adjacency table is stored in a row format and a second cumulative query cost indicator when the adjacency table is stored in a column format includes:

[0197] Determining a unit cost index of each query behavior in various types of query tasks under different storage formats of the adjacency table, wherein the types of query tasks include exact query tasks and column scan tasks;

[0198] Predicting, based on the task scale characteristics of each query task and the unit cost index, a query cost index for each query task when the adjacency table is stored in a row format and a column format;

[0199] The first cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in row format, and the second cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in column format.

[0200] According to one or more embodiments of the present disclosure, determining the target storage format of the adjacency table based on the first cumulative query cost indicator and the second cumulative query cost indicator includes:

[0201] If both the first cumulative query cost indicator and the second cumulative query cost indicator exceed a preset cost threshold, determining that the target storage format is a dual format, wherein the dual format is a coexistence of a row format and a column format; or

[0202] If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed a preset cost threshold, determining a difference between the first cumulative query cost indicator and the second cumulative query cost indicator;

[0203] If the difference exceeds a preset difference, determining the target storage format to be the storage format corresponding to the smaller one of the first cumulative query cost indicator and the second cumulative query cost indicator;

[0204] If the difference does not exceed the preset difference, the target storage format is determined to be the current storage format.

[0205] According to one or more embodiments of the present disclosure, storing the adjacency table in the target storage format includes:

[0206] If the current storage format of the adjacency list is the dual format and the target storage format is the row format or the column format, clearing the data of the adjacency list that is not in the target storage format by garbage collection; or

[0207] If the current storage format of the adjacency list is not the dual format, and the target storage format is different from the current storage format, the data in the current storage format of the adjacency list is converted into data in the target storage format.

[0208] According to one or more embodiments of the present disclosure, converting the data in the current storage format of the adjacency table into the data in the target storage format includes:

[0209] Storing the incremental data of the adjacency table in an additional storage space;

[0210] Obtaining a snapshot of the existing data in the current storage format of the adjacency list, and obtaining data in the target storage format of the adjacency list based on the snapshot;

[0211] The incremental data in the additional storage space is written into the data in the target storage format, and the subsequent incremental data of the adjacency table is stopped from being stored in the additional storage space, and the subsequent incremental data of the adjacency table is directly written into the data in the target storage format.

[0212] According to one or more embodiments of the present disclosure, after converting the data in the current storage format of the adjacency table into the data in the target storage format, the method further includes:

[0213] If the target storage format is not the dual format, the data in the current storage format of the adjacency table is cleared by garbage collection.

[0214] According to one or more embodiments of the present disclosure, the method further includes:

[0215] Receive a target precise query task for the adjacency list of target graph data;

[0216] According to the target precise query task, searching for a leaf node where the adjacency table of the target graph data is located in the index tree of the graph database;

[0217] If the target storage format is a row format or a column format, searching for target row data in the leaf node according to the target precise query task, and obtaining a query result based on the target row data; or

[0218] If the target storage format is dual format, target row data is searched in the row format data in the leaf node according to the target precise query task, and query results are obtained according to the target row data.

[0219] According to one or more embodiments of the present disclosure, the method further includes:

[0220] Receive a target column scanning task for the adjacency list of the target graph data;

[0221] Searching, according to the target column scanning task, for a leaf node where the adjacency table of the target graph data is located in the index tree of the graph database;

[0222] If the target storage format is a row format, traverse each row of data in the leaf node and obtain the target attribute queried by the target column scanning task from each row of data; or

[0223] If the target storage format is a column format, then the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained according to the attribute column data, the creation time column data, and the deletion time column data; or

[0224] If the target storage format is dual format, the attribute column data, creation time column data and deletion time column data queried by the target column scanning task are obtained from the column format data in the leaf node, and the target attributes queried by the target column scanning task are obtained based on the attribute column data, the creation time column data and the deletion time column data.

[0225] In a second aspect, according to one or more embodiments of the present disclosure, a graph database processing device is provided, including:

[0226] A prediction unit is configured to predict, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format;

[0227] a format determining unit, configured to determine a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator;

[0228] A storage unit is used to store the adjacency list in the leaf node of the index tree of the graph database using the target storage format.

[0229] According to one or more embodiments of the present disclosure, the prediction unit, when predicting a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format based on a set of historical query tasks for the adjacency table of target graph data in a graph database, is configured to:

[0230] Obtaining a task scale feature of each query task included in the historical query task set; wherein the task scale feature includes the number of rows of data queried by the query task and the number of attributes queried in the row of data;

[0231] According to the task scale characteristics of each query task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format are predicted.

[0232] According to one or more embodiments of the present disclosure, the prediction unit, when predicting the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when the adjacency table is stored in column format based on the task scale characteristics of each query task, is configured to:

[0233] Determining a unit cost index of each query behavior in various types of query tasks under different storage formats of the adjacency table, wherein the types of query tasks include exact query tasks and column scan tasks;

[0234] Predicting, based on the task scale characteristics of each query task and the unit cost index, a query cost index for each query task when the adjacency table is stored in a row format and a column format;

[0235] The first cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in row format, and the second cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in column format.

[0236] According to one or more embodiments of the present disclosure, when determining the target storage format of the adjacency table based on the first cumulative query cost indicator and the second cumulative query cost indicator, the format determination unit is configured to:

[0237] If both the first cumulative query cost indicator and the second cumulative query cost indicator exceed a preset cost threshold, determining that the target storage format is a dual format, wherein the dual format is a coexistence of a row format and a column format; or

[0238] If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed a preset cost threshold, determining a difference between the first cumulative query cost indicator and the second cumulative query cost indicator;

[0239] If the difference exceeds a preset difference, determining the target storage format to be the storage format corresponding to the smaller one of the first cumulative query cost indicator and the second cumulative query cost indicator;

[0240] If the difference does not exceed the preset difference, the target storage format is determined to be the current storage format.

[0241] According to one or more embodiments of the present disclosure, when the storage unit stores the adjacency table in the target storage format, it is configured to:

[0242] If the current storage format of the adjacency list is the dual format and the target storage format is the row format or the column format, clearing the data of the adjacency list that is not in the target storage format by garbage collection; or

[0243] If the current storage format of the adjacency list is not the dual format, and the target storage format is different from the current storage format, the data in the current storage format of the adjacency list is converted into data in the target storage format.

[0244] According to one or more embodiments of the present disclosure, when the storage unit converts the data in the current storage format of the adjacency table into the data in the target storage format, it is configured to:

[0245] Storing the incremental data of the adjacency table in an additional storage space;

[0246] Obtaining a snapshot of the existing data in the current storage format of the adjacency list, and obtaining data in the target storage format of the adjacency list based on the snapshot;

[0247] The incremental data in the additional storage space is written into the data in the target storage format, and the subsequent incremental data of the adjacency table is stopped from being stored in the additional storage space, and the subsequent incremental data of the adjacency table is directly written into the data in the target storage format.

[0248] According to one or more embodiments of the present disclosure, after converting the data in the current storage format of the adjacency table into data in the target storage format, the storage unit is further configured to:

[0249] If the target storage format is not the dual format, the data in the current storage format of the adjacency table is cleared by garbage collection.

[0250] According to one or more embodiments of the present disclosure, the device further includes a query unit configured to:

[0251] Receiving a target precise query task for an adjacency table of target graph data; searching for a leaf node where the adjacency table of the target graph data is located in an index tree of the graph database according to the target precise query task;

[0252] If the target storage format is a row format or a column format, searching for target row data in the leaf node according to the target precise query task, and obtaining a query result based on the target row data; or

[0253] If the target storage format is dual format, target row data is searched in the row format data in the leaf node according to the target precise query task, and query results are obtained according to the target row data.

[0254] According to one or more embodiments of the present disclosure, the query unit is further configured to:

[0255] Receive a target column scanning task for an adjacency table of target graph data; search an index tree of the graph database for a leaf node where the adjacency table of the target graph data is located according to the target column scanning task;

[0256] If the target storage format is a row format, traverse each row of data in the leaf node and obtain the target attribute queried by the target column scanning task from each row of data; or

[0257] If the target storage format is a column format, then the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained according to the attribute column data, the creation time column data, and the deletion time column data; or

[0258] If the target storage format is dual format, the attribute column data, creation time column data and deletion time column data queried by the target column scanning task are obtained from the column format data in the leaf node, and the target attributes queried by the target column scanning task are obtained based on the attribute column data, the creation time column data and the deletion time column data.

[0259] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0260] The memory stores computer-executable instructions;

[0261] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the graph database processing method described in the first aspect and various possible designs of the first aspect.

[0262] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method for processing the graph database as described in the first aspect and various possible designs of the first aspect is implemented.

[0263] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the graph database processing method described in the first aspect and various possible designs of the first aspect.

[0264] In summary, based on the historical query task set for the adjacency table of the target graph data in the graph database, the first cumulative query cost index when the adjacency table is stored in row format and the second cumulative query cost index when stored in column format are predicted; based on the first cumulative query cost index and the second cumulative query cost index, the target storage format of the adjacency table is determined; and in the leaf nodes of the index tree of the graph database, the adjacency table is stored in the target storage format. In this embodiment, for the adjacency table of the target graph data, combined with the situation of the historical query task set, the target storage format of the adjacency table is dynamically selected by weighing the query costs of different storage formats, which can reduce storage costs, avoid unnecessary data redundancy, optimize resource utilization, and improve the query performance of the query task.

[0265] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0266] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0267] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for processing a graph database, characterized in that: include: Predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format; determining a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator; In a leaf node of an index tree of the graph database, the adjacency list is stored in the target storage format.

2. The method according to claim 1, characterized in that The method of predicting, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format, includes: Obtaining a task scale feature of each query task included in the historical query task set; wherein the task scale feature includes the number of rows of data queried by the query task and the number of attributes queried in the row of data; According to the task scale characteristics of each query task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format are predicted.

3. The method according to claim 2, characterized in that The predicting, based on the task scale characteristics of each query task, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format includes: Determining a unit cost index of each query behavior in various types of query tasks under different storage formats of the adjacency table, wherein the types of query tasks include exact query tasks and column scan tasks; Predicting, based on the task scale characteristics of each query task and the unit cost index, a query cost index for each query task when the adjacency table is stored in a row format and a column format; The first cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in row format, and the second cumulative query cost index is obtained by accumulating the query cost indicators of each query task when the adjacency table is stored in column format.

4. The method according to claim 1, wherein The determining, according to the first cumulative query cost indicator and the second cumulative query cost indicator, a target storage format of the adjacency table includes: If both the first cumulative query cost indicator and the second cumulative query cost indicator exceed a preset cost threshold, determining that the target storage format is a dual format, wherein the dual format is a coexistence of a row format and a column format; or If at least one of the first cumulative query cost indicator and the second cumulative query cost indicator does not exceed a preset cost threshold, determining a difference between the first cumulative query cost indicator and the second cumulative query cost indicator; If the difference exceeds a preset difference, determining the target storage format to be the storage format corresponding to the smaller one of the first cumulative query cost indicator and the second cumulative query cost indicator; If the difference does not exceed the preset difference, the target storage format is determined to be the current storage format.

5. The method according to claim 1, wherein The storing the adjacency table in the target storage format includes: If the current storage format of the adjacency list is the dual format and the target storage format is the row format or the column format, clearing the data of the adjacency list that is not in the target storage format by garbage collection; or If the current storage format of the adjacency list is not the dual format, and the target storage format is different from the current storage format, the data in the current storage format of the adjacency list is converted into data in the target storage format.

6. The method according to claim 5, characterized in that The converting the data in the current storage format of the adjacency table into the data in the target storage format includes: Storing the incremental data of the adjacency table in an additional storage space; Obtaining a snapshot of the existing data in the current storage format of the adjacency list, and obtaining data in the target storage format of the adjacency list based on the snapshot; The incremental data in the additional storage space is written into the data in the target storage format, and the subsequent incremental data of the adjacency table is stopped from being stored in the additional storage space, and the subsequent incremental data of the adjacency table is directly written into the data in the target storage format.

7. The method according to claim 5 or 6, characterized in that After converting the data in the current storage format of the adjacency table into the data in the target storage format, the method further includes: If the target storage format is not the dual format, the data in the current storage format of the adjacency table is cleared by garbage collection.

8. The method according to claim 1, characterized in that The method further comprises: Receive a target precise query task for the adjacency list of target graph data; According to the target precise query task, searching for a leaf node where the adjacency table of the target graph data is located in the index tree of the graph database; If the target storage format is a row format or a column format, searching for target row data in the leaf node according to the target precise query task, and obtaining a query result based on the target row data; or If the target storage format is dual format, target row data is searched in the row format data in the leaf node according to the target precise query task, and query results are obtained according to the target row data.

9. The method according to claim 1, characterized in that The method further comprises: Receive a target column scanning task for the adjacency list of the target graph data; Searching, according to the target column scanning task, for a leaf node where the adjacency table of the target graph data is located in the index tree of the graph database; If the target storage format is a row format, traverse each row of data in the leaf node and obtain the target attribute queried by the target column scanning task from each row of data; or If the target storage format is a column format, then the attribute column data, creation time column data, and deletion time column data queried by the target column scanning task are obtained in the leaf node, and the target attribute queried by the target column scanning task is obtained according to the attribute column data, the creation time column data, and the deletion time column data; or If the target storage format is dual format, the attribute column data, creation time column data and deletion time column data queried by the target column scanning task are obtained from the column format data in the leaf node, and the target attributes queried by the target column scanning task are obtained based on the attribute column data, the creation time column data and the deletion time column data.

10. A graph database processing device, characterized in that: include: A prediction unit is configured to predict, based on a set of historical query tasks for an adjacency table of target graph data in a graph database, a first cumulative query cost index when the adjacency table is stored in a row format and a second cumulative query cost index when the adjacency table is stored in a column format; a format determining unit, configured to determine a target storage format of the adjacency table according to the first cumulative query cost indicator and the second cumulative query cost indicator; A storage unit is used to store the adjacency list in the leaf node of the index tree of the graph database using the target storage format.

11. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • A method and device for accessing graph data based on a grouping association table

    CN109255055A

  • Graph database query method and device, electronic equipment and storage medium

    CN116521956A

  • Virtual experiment simulation data-oriented database architecture and data query method

    CN119739744A

  • Methods and system for recommending storage format for migrating a rdbms

    US20240311350A1

  • Method and apparatus for conversion of data storage formats

    WO2015139193A1