Multi-version data caching device and method based on graph database

Through multi-version data caching devices and methods, the operator cache strategy is dynamically adjusted, and cache is enabled only for high-value operators. Combined with read requests, it quickly locates the optimal version cache, which solves the performance degradation of graph databases under frequent updates, and achieves efficient query and high throughput.

CN120234353AActive Publication Date: 2025-07-01启元实验室
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510716079.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The cache scheme of the existing graph database has a sharp decline in performance when data is updated frequently, and traditional cache schemes are difficult to effectively utilize the reuse value of intermediate results, resulting in inefficient query.

Method used

The multi-version data cache device is adopted. By adding only the storage edge and node relationship data, dynamically adjusting the operator cache strategy in combination with the statistical analysis module, only cache is enabled for high-value operators, and quickly locate the optimal version cache in combination with read requests, reducing duplicate calculations, and recovering the cache when the service is restarted through the materialized loading module.

Benefits of technology

In frequently updated graph data scenarios, maintain high query throughput and data consistency, avoid cache invalidation caused by data changes, and significantly accelerate complex queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234353A_ABST
    Figure CN120234353A_ABST
Patent Text Reader

Abstract

The invention provides a multi-version data caching device and method based on a graph database, and relates to the technical field of data processing. According to the device, a data storage module continuously stores relational data corresponding to edges and nodes based on an only addition mode, and determines a physical offset value of a scanning end position corresponding to a read request; the statistical analysis module is used for pre-adjusting a preset default operator according to actual business requirements and collecting statistical data; the cache operator selection module determines a target operator from the pre-adjusted default operators according to the statistical data, and enables or disables the cache function of the target operator; the cache module performs independent storage, version maintenance and elimination processing on each node of the target operator to generate overall cache information; and the query execution module determines the optimal version cache corresponding to the target node in the target operator from the overall cache information based on the relational data and the read request so as to adjust the preliminary execution plan based on the optimal version cache and the physical deviation value and execute the corresponding query action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, for example, to a caching device and method for multi-version data based on a graph database. Background Art

[0002] A graph database is a database management system that uses nodes and edges as basic storage units, directly storing and processing relationship data between entities. The theoretical basis of graph databases comes from graph theory, which is particularly suitable for processing datasets with complex association relationships and supports efficient graph traversal operations. With the rapid development of information technology, the generation and storage volume of data have increased exponentially. Moreover, as the complexity and scale of graph data continue to expand, the computational complexity of queries has increased significantly. The query process of a graph database usually requires a large number of node traversals and edge scans, which have large operation overheads, but the result set generated during the query process may not be large. For example, in social network analysis, although a large number of user relationships need to be traversed, ultimately only a few qualified nodes closest to a specific user may need to be returned, and the result sets of different queries in a graph scenario often overlap. This characteristic makes caching results particularly valuable. By caching query results or phased results of complex queries with large computational overheads, small data volumes, and high reuse probabilities, once the subsequent query hits the cache, not only can data access be reduced, but also repeated intensive calculations can be avoided, significantly improving query performance. Especially in a distributed environment, data transmission between nodes will bring additional query latency and processor performance occupancy. Through a reasonable caching strategy, the network communication overhead in subsequent queries can also be significantly reduced.

[0003] In related technologies, generally, read-only query results are cached, that is, the return of the entire query request is cached. However, in subsequent use, only exactly the same query can hit the cached read-only query results, and once there is data update, it may cause the cache to become invalid, and a series of calculation processes need to be performed again for querying, resulting in low query efficiency. Summary of the Invention

[0004] This application aims to provide a caching device and method for multi-version data based on a graph database.

[0005] According to one aspect of this application, a caching device for multi-version data based on a graph database is proposed, including: A data storage module, configured to continuously store relationship data corresponding to edges and nodes based on the append-only method, and determine the physical offset value of the scan end position corresponding to the read request according to the read request and the relationship data; A statistical analysis module, configured to pre-adjust a preset default operator according to actual business requirements and collect statistical data; A cache operator selection module, configured to determine a target operator from pre-adjusted default operators according to statistical data, and turn on or off the cache function of the target operator, where the target operator is an operator that supports caching intermediate results; A cache module, configured to independently save, maintain versions, and perform elimination processing on each node of the target operator to generate overall cache information; A query execution module, configured to determine an optimal version cache of the target operator from the overall cache information based on relational data and a read request, adjust a preset preliminary execution plan based on the optimal version cache and a physical offset value, and perform a corresponding query action.

[0006] According to one aspect of the present application, a cache method for multi-version data based on a graph database is provided, including: Based on the append-only method, continuously store relational data corresponding to edges and nodes, and determine a physical offset value of a scan end position corresponding to the read request according to the read request and the relational data; Pre-adjust a preset default operator according to actual business requirements, and collect statistical data; Determine a target operator from the pre-adjusted default operators according to the statistical data, and turn on or off the cache function of the target operator, where the target operator is an operator that supports caching intermediate results; Independently save, maintain versions, and perform elimination processing on each node of the target operator to generate overall cache information; Based on the relational data and the read request, determine an optimal version cache corresponding to a target node in the target operator from the overall cache information, adjust a preset preliminary execution plan based on the optimal version cache and the physical offset value, and perform a corresponding query action.

[0007] According to one aspect of the present application, an electronic device is provided, the electronic device includes: a processor; a memory storing a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method as described above.

[0008] According to one aspect of the present application, a non-transitory computer-readable medium is provided, on which readable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method as described above.

[0009] It should be understood that the above general description and subsequent detailed description are exemplary and should not limit the present application.

[0010] Beneficial effects: Through the above embodiments provided by this application, the statistical analysis module dynamically adjusts the operator caching strategy, enables caching only for high-value operators, and avoids resource occupation by invalid caching. Combining with read requests ensures that the optimal version cache can be quickly located during query execution, reduces duplicate calculations, and significantly accelerates complex queries. Based on append-only storage, historical version relational data is retained, enabling the intermediate result cache to be bound to the intermediate result version and read requests. Together with the physical offset value corresponding to the optimal version cache, only incremental updates are required instead of full-scale reconstruction. The query execution module determines the optimal version cache based on relational data and read requests, and then adjusts the preliminary execution plan to avoid cache invalidation problems caused by data changes. In scenarios of frequently updated graph data, traditional caching schemes experience a significant drop in performance due to frequent data changes. This application maintains high query throughput while ensuring data consistency through append-only storage, intermediate result caching, and multi-version caching. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without exceeding the scope protected by this application.

[0012] Figure 1 It is a block diagram of a caching device for multi-version data based on a graph database provided by an embodiment of this application; Figure 2 It is a block diagram of another caching device for multi-version data based on a graph database provided by an embodiment of this application; Figure 3 It is a flowchart of a caching method for multi-version data based on a graph database provided by an embodiment of this application; Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted.

[0014] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0015] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0016] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0017] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of the concept of the present application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0018] Specific implementation manners may refer to the following embodiments.

[0019] Figure 1 It is a block diagram of a cache device for multi-version data based on a graph database provided for the embodiments of the present application. As Figure 1 shown, the cache device 10 for multi-version data based on a graph database includes a data storage module 11, a statistical analysis module 12, a cache operator selection module 13, a cache module 14, and a query execution module 15.

[0020] The data storage module 11 is used to continuously store the relationship data corresponding to edges and nodes in an append-only manner, and determine the physical offset value of the scan end position corresponding to the read request according to the read request and the relationship data.

[0021] In this device, the data storage module 11 can be a graph database storage engine, and the storage of its edges is based on the append-only mechanism. The graph database stores edges as first-class citizens, and the edges of the same node are generally stored physically or logically continuously through pointers. The stored content can be regarded as relational data. The append-only method is a data storage or processing strategy, and its core principle is that all new data can only be appended to the end of the existing data, and the previous data cannot be modified or deleted.

[0022] In a database system, a read request is an operation to access data of a specific version of the database, and the version of the data read can be precisely controlled through a timestamp or a transaction identifier (which can be abbreviated as transaction ID). Timestamps play a key role in concurrency control and multi-version data management, ensuring transaction isolation and consistent reads. To support incremental updates of the cache, the same type of edges of the same node are stored in a compact and append-only manner. In scenarios such as social network graph databases, person_likes_post related to the Person node can be used to represent a type of edge, which is used to connect the node of the post it likes; person_workAt_organisation represents another type of edge, which is used to connect the node representing the work organization. The person_likes_post edges of the same Person are stored continuously, and subsequent updates will only append information at the end. Deletion is also represented by appending a deletion message. The person_workAt_organisation of each Person is stored continuously separately in the same way.

[0023] When reading a certain type of edge of a node, in addition to returning normal data for each read request, it can also return the physical offset value corresponding to the end position of the data scan, that is, the end position of the visible edge data of the latest transaction at the time of initiating the read.

[0024] The statistical analysis module 12 is used to pre-adjust the preset default operators according to actual business requirements and collect statistical data.

[0025] The graph database in this application has many operators, such as breadth-first search, subgraph or path matching, attribute aggregation, n-degree neighbor query, shortest path operator, time range aggregation operator, group counting operator, etc., which can be used as preset default operators. The statistical analysis module 12 can pre-transform the default operators in advance to make them support reading and writing intermediate result caches according to the granularity of the traversed nodes. This process can be pre-adjustment. In some implementation manners, an operator can be used to represent a fixed type of computational sub-process in the graph query process.

[0026] In some implementations, since the implementation and invocation methods of different default operators in the graph database are greatly affected by the actual business usage and access methods, the usage logic of different default operators in this device for caching intermediate results can be completed during the graph database program development stage, or an interface can be reserved and later dynamically loaded, added, or modified in a plug-in manner.

[0027] The statistical data can include the invocation frequencies of different default operators, the scale of data input and output, and the execution cost, etc. These statistical data can be collected by the statistical analysis module 12.

[0028] The cache operator selection module 13 is used to determine a target operator from the pre-adjusted default operators according to the statistical data, and enable or disable the cache function of the target operator, where the target operator is an operator that supports caching intermediate results.

[0029] The cache operator selection module 13 can specifically select the intermediate result cache that can be cached to optimize the cache. In this application, the satisfaction conditions of the operators that support caching intermediate results can be set in advance. By comparing the satisfaction conditions with the statistical data, it can be determined which default operators are the target operators. Since the query processes corresponding to different read requests are different and the involved target operators are different, the cache functions of the target operators can be enabled or disabled.

[0030] In some implementations, four constants l, m, n, and t can be pre-configured for each target operator to represent weights, and a, b, and c are used to represent the operator invocation frequency, the ratio of the estimated scanned item scale of the query plan and the intermediate result item scale, and the histogram median of the estimated execution cost of each operator output by the optimizer in the statistical data respectively. When la + mb + nc > t, the cache function of the operator is enabled.

[0031] In some other implementations, the target operator can be randomly selected multiple times by Monte Carlo or all possible selections can be enumerated, and the cache function is enabled preferentially according to the execution cost benefit effect after enabling the cache. The specific process is as follows: First, the user can input queries that are considered to be frequently accessed subsequently, or allow the query intermediate result cache selection and optimization module to read several queries with the highest request frequencies by itself. Then, confirm the operators that support caching in the queries. The cache functions of each operator are enabled and disabled independently. For complex queries with a large number of operators that support caching, Monte Carlo can be used to randomly select and enable multiple times to reduce the number of enumerated combinations. Finally, compare all combinations from the perspective of the cost benefit before and after the query is enabled and disabled, select the optimal solution, and set the operators in the selected solution to enable the cache function.

[0032] In some other implementation manners, after the caching function of the target operator is enabled, the caching scale and cache hit rate of the target operator can be continuously counted. Three configurable constants p, q, and r can be preset, and a, b, and c are used to represent the call frequency of the target operator, cache hit rate, and the total number of cache entries of the operator in the statistical data respectively. When pab - qc < r, it is marked that the target operator can be preferentially closed, that is, if the total cache scale reaches the preset upper limit subsequently, the caching function of this target operator is preferentially closed, and its existing cache is cleared.

[0033] The cache module 14 is used to independently save, maintain the version, and perform elimination processing on each node of the target operator, so as to generate overall cache information.

[0034] In this application, the operations within a target operator involve multiple nodes in the graph, and each node can cache multiple versions of the intermediate results of the node-related calculations. A new read request may trigger an update of the version, so version maintenance in different situations is required. As the amount of stored version data increases, some versions may no longer be used, so it may be necessary to perform elimination processing on the versions. In some implementation manners, the old version elimination processing is not to directly delete, but may be to merge inappropriate versions with other versions.

[0035] The cache module 14 can cache the multi-version intermediate results of each node in the target operator, respond to read requests and update requests, and provide a cache elimination function. In some implementation manners, a maintenance period and elimination conditions can be preset, and the cache module 14 performs maintenance and elimination.

[0036] In some implementation manners, MVCC (Multi-Version Concurrency Control) can be used to support the incremental update and multi-version maintenance of the cache, and ensure sufficient concurrency of reading and writing. The version read by the read request is based on the timestamp obtained during the request processing, and the data updated by the write transaction after this timestamp is invisible to this read request.

[0037] For example, three read requests successively obtain transaction timestamps t1, t2, and t4. And there is an update transaction t3 before t4. Eventually, the order in which the transactions execute to the operator that supports the cache function is t1, t3, t4, t2. When t1 executes, the cache version R1 is created. After that, when t3 updates the edge data related to the cached node and commits, when t4 executes, it reads the new data and updates the cache version to R2. When t2 executes, it can only read the R1 version because t2 should not see the data updated by t3. When implementing MVCC in this application, the timestamp can be represented by simple consecutive numbers. Only the write transaction increments this timestamp value by one each time, and the read transaction timestamp remains unchanged. Each piece of data stored is marked with the transaction timestamp at the time of writing. The read operation can only see the data whose marked value is not greater than its own timestamp. In the above example, the read transaction timestamps t1 = t2, t3 is the incremented write transaction timestamp, so t3 = t2 + 1. After that, t4 is the read transaction timestamp that remains unchanged, so t4 = t3. Therefore, in order to enable the query transaction to read the correct cache, the final modification timestamp of the corresponding data can be marked for the cache version. The timestamp value of the cache version R1 mentioned above is t1 (= t2), and the timestamp of R2 is t4 (= t3).

[0038] In the process of maintaining a multi-version cache of the intermediate results of each node item of a single target operator, there are two different ways. One is based on the latest version T4, that is, maintaining the latest version and a series of rollback logs (undo). Since most read requests read the latest version, usually the first way is used to maintain the latest version. Each time the cache is updated, the latest cache is directly calculated, and the modification is recorded as an undo log. When an old version needs to be read, the version is rolled back with the undo log. That is, if a request wants to obtain the cache data of version T4, it can directly call the latest version. If it wants to obtain version T1, it needs to reverse from version T4 to get version T1.

[0039] The other is based on the initial version T1, that is, maintaining a relatively old version and a series of incremental data (todo). If you want to obtain version T1, you can directly call the basic initial version.

[0040] For example, in a social network scenario, when searching for the M most active forums among the friends within N degrees of a person A, the minimum caching granularity is the sorting of the number of posts made by each individual friend on various forums. In the examples of t1, t2, t4, and t3 above, assume that after t1 is executed, the number of posts made by a friend on the football, badminton, and basketball forums are 6, 7, and 8 respectively. Then, in the t3 transaction, the user adds two posts each on the football and basketball forums and one post on the bowling forum. After t4, the cached baseline data in the item corresponding to this friend in the HashMap is {Basketball 10, Football 8, Badminton 7, Bowling 1}. The differential data corresponds to the modification of the intermediate cache results of the nodes caused by scanning the edges related to this node and performing calculations on this basis during the incremental calculation of this operator, that is: {undo: -2 for basketball after target scanning object splicing and sending, -2 for football after target scanning object splicing and sending, -1 for bowling after target scanning object splicing and sending; delete the bowling forum item}. When the read request created at time t2 accesses the cache, based on the baseline data and the above undo, a playback operation is executed to obtain the old version {Basketball 8, Badminton 7, Football 6}. If it is cumbersome to record the undo operations for each node's cache in individual operators or the cost of executing undo is high, the second method can be used to maintain a relatively old cache version and retain incremental information to generate todo after target scanning object splicing and sending during updates, delaying the direct modification of the baseline version. In the above example, if the second method is used, the updated cache record after t4 target scanning object splicing and sending is {Basketball after target scanning object splicing and sending 8, Badminton after target scanning object splicing and sending 7, Football after target scanning object splicing and sending 6} after target scanning object splicing and sending and the differential item after target scanning object splicing and sending is {todo: add the bowling forum item; +2 for basketball after target scanning object splicing and sending, +2 for football after target scanning object splicing and sending, +1 for bowling after target scanning object splicing and sending}. In this example, the first method is more optimal. The second method requires periodically merging the baseline version and some old redo after target scanning object splicing and sending to ensure that the baseline cache does not always stay in the oldest version, while the first method only needs to periodically directly delete the obsolete undo after target scanning object splicing and sending.

[0041] The query execution module 15 is used to determine the optimal version cache corresponding to the target node in the target operator from the overall cache information based on the relational data and the read request, so as to adjust the preset preliminary execution plan based on the optimal version cache and the physical offset value and execute the corresponding query action.

[0042] In this application, there are multiple nodes on the graph involved in the calculation process of the target operator. The read request may include specific read tasks, and the number and range of nodes corresponding to different tasks are different. The nodes of the target operator corresponding to the read request can be used as target nodes. A preliminary execution plan can be generated in advance, that is, the process of querying according to the original query logic without considering node caching. The query execution module 15 can select the cached information of the target operator that meets the conditions from the overall cached information based on the relational data corresponding to the read request and the current transaction ID carried by the read request, and can select the optimal version cache based on the preset optimal selection conditions. It should be noted that the "optimal" in this application refers to the version that meets the conditions and is the most suitable, and this version may be the latest version or other versions.

[0043] According to the preset execution plan adjustment method, add the execution logic of using the cache to the original plan, which includes processing the optimal version cache and the physical offset value, and then execute the query action corresponding to the adjusted plan. In some ways, an offset vector can be calculated based on the physical offset value, and the optimal version cache and the offset vector are processed, and then the query action corresponding to the adjusted plan is executed.

[0044] Traditional cache solutions are difficult to predict the reuse value of intermediate results in graph queries, resulting in a low cache hit rate. The device of this application dynamically adjusts the operator cache policy through the statistical analysis module, enables caching only for high-value operators, and avoids resource occupation by invalid caches. Combining with the read request ensures that the optimal version cache can be quickly located during query execution, reduces repeated calculations, and significantly accelerates complex queries. Based on append-only storage and retaining historical versions of relational data, the intermediate result cache is bound to the intermediate result version and the read request. Cooperating with the physical offset value corresponding to the optimal version cache, only incremental updates are required instead of full reconstruction. The query execution module determines the optimal version cache through relational data and the read request, and then adjusts the preliminary execution plan to avoid the problem of cache invalidation caused by data changes. In scenarios of frequently updated graph data, the performance of traditional cache solutions drops sharply due to frequent data changes. This application maintains high query throughput while ensuring data consistency through append-only storage, intermediate result caching, and multi-version caching.

[0045] According to some embodiments, referring to Figure 2 , the cache device 10 for multi-version data of the graph database further includes a materialized loading module 16, which is used to materialize the overall cached information of the target operator at a preset materialization interval, so as to reload the materialized overall cached information when the service restarts.

[0046] The materialized loading module 16 can be a module for materializing and reloading the cached intermediate results. To prevent cache loss caused by process exit or hardware restart, the cache can be materialized regularly, either onto this module or onto other modules or devices. After the service restarts, the cached materialized data can be loaded and reused. The materialization interval can be preset by the user according to actual needs.

[0047] In the device of the present application, by periodically materializing the cache information of the target operator, the intermediate results in the memory are persistently stored (such as on disk or distributed storage), ensuring rapid loading after the service restarts and avoiding performance degradation caused by cold start. In a production environment that requires high availability, the service may restart due to failures or upgrades. By materializing the cache information, the device of the present application enables the system to directly reuse historical calculation results after recovery, reducing manual intervention and operation and maintenance pressure.

[0048] According to some embodiments, when the scan operation corresponding to the read request is to scan multiple types of edges, the data storage module 11 can obtain the physical offset values corresponding to each type of edge, so as to generate an offset vector according to the physical offset values, where the multiple types of edges are the edges corresponding to the same target operator.

[0049] In the present application, since the target operator corresponds to multiple nodes and each node in the graph database is connected with multiple types of edges, the read request may scan multiple types of edges simultaneously. For example, when obtaining the number of posts a person has made and the posts they like, it is necessary to scan both the edges recording the posts they like and the edges recording their posting information. Therefore, multiple offsets need to be obtained to form an offset vector.

[0050] The data storage module 11 can maintain an append-only log classified by edge type, and each edge write is accompanied by a globally incrementing logical timestamp (or version number). For each type of edge (such as edge types E1, E2... En), the module records its current maximum physical offset value (i.e., the disk file offset or log sequence number of the last write). In response to the read request, the incremental scan range is obtained. The physical offset ranges of multiple types of edges can be encoded as structured data as the physical offset values, and then the physical offset values can be input into a preset structure to output the corresponding offset vector.

[0051] The device of the present application accurately records the data update positions of each type of edge through the physical offset values, enabling the query to only scan the newly added or changed data segments, avoiding redundant I / O operations, and significantly improving the scan efficiency.

[0052] According to some embodiments, the cache module 14 caches the intermediate results according to a preset granularity of the nodes within the operator.

[0053] In this application, graph queries usually start from nodes and traverse nodes through the adjacent edges of the nodes. When the operators in the query traverse the nodes, they often calculate by reading a large amount of information around the nodes involved in the graph, and the calculation results of these nodes are very likely to be reused by other queries later. Therefore, the intermediate results can be cached at the node granularity to improve the reuse rate of the cache. Usually, a HashMap (Hash-Map, hash map) or other data structures that are easy to retrieve are used to organize and manage the specific intermediate result content corresponding to each node within the same operator.

[0054] For example, in a social network scenario, a common query is to search for the M most active forums among the friends within N degrees of a person A. If only the entire query result is cached, the cache is only valid for specific parameters A, N, and M, and the efficiency is very low. During the first query traversal process in this application, the posting situations of all friends within N degrees of A will be cached independently and saved using a HashMap. When subsequently querying the posting situations of the friends within K degrees of a person B, the HashMap can be queried for each friend to try to directly obtain the cached information. If a certain friend of B is not in the cache, the posting information of this friend is obtained through the conventional method and then added to the cache. At this time, each item in the HashMap is the posting situation of each friend. Each item can use a BTreeMap (B-Tree Map, B-tree map) or a red-black tree to sort and save the number of posts of the friends in different forums. This cache can also be used for other types of queries. For example, when querying for the person with the highest overlap with the top ten forums where a person C likes to post among the friends within three degrees of C, because the posting situations of the friends within three degrees of C are very likely to have been cached.

[0055] The device of this application caches according to the node-level granularity (such as single node, neighbor set, subgraph), enabling different queries to reuse the intermediate results of the same node and significantly improving the hit rate.

[0056] According to some embodiments, when the cache operator selection module 13 determines the target operator from the pre-adjusted default operators according to the statistical data, it is specifically used to: determine the cache hit rate and the benefit of enabling the cache of the pre-adjusted default operators according to the statistical data; and determine the default operators whose cache hit rate and benefit of enabling the cache meet the preset cache enabling conditions as the target operators.

[0057] In this application, the cache hit rate refers to the proportion of the points that can directly read information from the cache among all the points involved in the calculation process of the target operator with the cache function enabled, where the cache information is read around the points involved. For example, if an operator is to obtain the Post information of the latest comments of a certain type of Person, the id of each Person that meets the conditions will be used to query the corresponding cache. The proportion of the number of successful cache reads to the total number of queries is the cache hit rate for this time.

[0058] In some implementations, the query count and the successful query count can be cumulatively calculated to obtain the cumulative actual hit rate. In other implementations, the hit rate can also be simply estimated using the cache size and the data size. That is, if the data size of Person is 10,000 and the cache size is 1,000, then the hit rate can be estimated to be 10%.

[0059] The cache opening benefit refers to the query execution cost saved after the cache function is enabled, because the target operator can skip some original executions by using the cache. The generation process of the query execution plan constructs a tree plan using the dynamic programming algorithm, and the entire generation and optimization process is based on a cost model. This model estimates the resource consumption of different execution paths (such as I / O operations, CPU time, etc.) and selects the execution plan with the lowest estimated cost. In this application, in the optimization using the cache, for any target operator, once its cache function is enabled, in the subsequent calculations around the node, the part that hits the cache can be skipped and the intermediate result of the corresponding point can be directly read from the cache, thus saving resource consumption (such as I / O operations, CPU time, etc.). For this operator, all the resources saved minus the cache maintenance cost is the cache opening benefit.

[0060] In this application, the calculation of the cache opening benefit is utilized in two parts. One is applied to evaluate the effectiveness of enabling the cache for the target operator. To conveniently evaluate whether a target operator should enable or disable the cache function, the cumulative hit benefit of the cache function used by the target operator each time can be statistically calculated. The cumulative value is affected by various factors, including the call count, the cache hit rate in each call, and the cost that can be saved by actually skipping steps once a hit occurs in each call. For operators with a low cumulative hit benefit, the query intermediate result cache selection and optimization module will give priority to closing and eliminating operations for them. The other is applied to the opening judgment of the target operator's cache function, that is, for high-frequency queries, Monte Carlo multiple random selections or enumerating all possible selections are used, and the cache function of the target operator is preferentially enabled according to the execution cost-benefit effect after enabling the cache.

[0061] When calculating the cache opening benefit, the historical hit rate or the cache hit rate can be estimated according to the data size, and the additional time cost brought by updating the cache is also considered. The calculation method for the benefit of a single operator enabled in a query is as follows: The original execution cost without enabling the cache - the execution cost after skipping some steps after enabling the cache * the cache hit rate - the cost of updating and maintaining the cache.

[0062] In the specific implementation process, the above parameters can be obtained in advance.

[0063] If the caching function is enabled for multiple operators within a query, the benefit of enabling caching is not the sum of the individual benefits of the multiple operators. Because if several operators are on the same pipeline, after the lower-layer operator enables caching, the original execution cost of the upper-layer operator not enabling caching or missing the cache has already been reduced. Therefore, it is necessary to correct the benefit obtained by the above method. The correction method is to subtract the benefit of the downstream operator with the caching function enabled from the benefit of the upstream operator.

[0064] The device of the present application only enables caching for target operators that meet the conditions based on the cached hit rate and benefit data statistically in real time, avoiding resource occupation by invalid caching. By continuously monitoring the cached hit rate and the benefit of enabling caching, the caching policy is automatically adjusted to adapt to changes in the query pattern.

[0065] According to some embodiments, when the caching module 14 independently stores the target operator and its internal nodes, it is specifically used for: storing the target operator based on the preset calculation type of the target operator.

[0066] In the present application, the caching for a class of the same calculations can be stored uniformly, that is, the target operators are stored independently. During the storage process, the unit storing the target operator can be the node ID. Therefore, for each target operator, it can be stored separately according to its calculation type. Further, the target operator can be stored in detail, that is, the versions corresponding to each node are divided based on the ID of each node.

[0067] In some implementation manners, during the process of independently storing the node versions in the target operator, the append-only mechanism can also be used. In some implementation manners, a globally monotonically increasing transaction timestamp can be adopted to ensure uniqueness throughout the system. When the target operator (such as a path search operator) generates intermediate results, the current timestamp T_i (i.e., the timestamp of the current read request) is bound to the generated functional key result, and a unique exclusive transaction timestamp is set for the corresponding version, and then multiple versions of the intermediate results are stored.

[0068] The device of the present application supports independently storing each node of the target operator based on the calculation type, independently maintaining version information, avoiding resource waste caused by unified storage, and improving the memory / cache utilization rate. In addition, the operators of the same calculation type are stored centrally, which is convenient for quickly retrieving and reusing the calculated results through the cache locality principle, reducing repeated calculations, and is especially suitable for loop structures or frequently called operators.

[0069] According to some embodiments, when the cache module 14 performs an elimination process on each node of the target operator, it is specifically configured to: obtain the number of versions corresponding to each node; when the number of versions is greater than a preset cache limit number, obtain the transaction timestamps corresponding to the multiple versions of each node; and merge the versions whose transaction timestamps meet the preset merging conditions.

[0070] In the present application, the cache module 14 stores multiple target operators, and each node corresponding to each target operator has multiple versions corresponding thereto, and the number of versions of the node can be directly called. The cache limit number can be preset based on historical traffic and the usage conditions of different versions. The merging conditions can also be preset, and the merging conditions can be used to represent the data merging method.

[0071] For any one of the nodes, compare the number of versions corresponding to the node with the cache limit number. If the number of versions is less than or equal to the cache limit number, no operation is required. If the number of versions is greater than the cache limit number, the transaction timestamps of different versions can be detected, and the versions whose transaction timestamps meet the merging conditions are merged.

[0072] In some implementation manners, there are existing versions 0, 1, 2, and 3, and 3 is the version closest to the current moment. The merging condition can be that when 0 is to be eliminated, 0 can be preferentially merged with 1 or 1&2.

[0073] In some other implementation manners, the timestamps of each version can be directly detected, the elimination conditions are preset, and the versions whose timestamps meet the elimination conditions are merged with the latest version or the adjacent version.

[0074] In some other implementation manners, it can be additionally supplemented that each time the operator updates the cache, it will check the transaction ID (transaction timestamp) of the smallest version in the currently active transactions, and then check whether the oldest batch of cache versions in the cache module lag behind the above smallest version transaction ID. When there are multiple lagging versions, they can be merged into one lagging version. For example, if the currently active transaction is t30, and it is found that the cache has t9, t19, t29, and t39, then the t9 and t19 versions can be eliminated. The reason for retaining the version at the t29 moment is to confirm that if the transaction of t30 accesses the cache, the appropriate version can be obtained based on the version at the t29 moment for processing.

[0075] The cache module can passively (e.g., triggered when the preset cache memory is insufficient) or actively eliminate operator caches with low cache value and outdated cache versions (among the n caches older than the time version corresponding to the oldest active transaction, the oldest n - 1 caches will not be accessed and should be eliminated. If incremental / delta storage is used for different versions of the cache, the information in the eliminated version should be merged with the remaining versions; if the device strictly limits the number of cache versions, when the number of cache versions reaches the limit, it will also trigger the cleaning of the old versions. After that, if there are still old transactions accessing the cache and all existing cache versions are newer than the timestamp corresponding to the access, the cache cannot be used).

[0076] The device of the present application can perform overall elimination processing on the target operator cache, avoiding the occupation of resources by target operators that are not used for a long time and improving the query efficiency.

[0077] According to some embodiments, the query execution module 15 is specifically configured to: when the optimal version cache is not empty and the current timestamp corresponding to the read request is less than the transaction timestamp of the oldest version of the target node, determine that the cache read is empty and execute the query according to the preliminary execution plan; when the optimal version cache is not empty and the current timestamp is equal to the transaction timestamp of the version of the target node, execute the query based on the optimal version cache; when the optimal version cache is not empty and the current timestamp is greater than the transaction timestamp of the latest version of the target node, or the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target node, select the cache information of the optimal version of the node with a transaction timestamp less than the current timestamp from the optimal version cache; obtain the physical offset value of the version with a timestamp less than the current timestamp, and perform a supplementary query according to the physical offset value to determine the cache information difference, and execute the query based on the cache information difference and the cache information; when the optimal version cache is empty, execute the query according to the preliminary execution plan.

[0078] In the present application, since the overall cache data corresponding to the target operator includes each node within the operator and each version corresponding to each node, the optimal version cache can be extracted from multiple versions of the target node first, and targeted analysis can be performed on the operator node cache and corresponding query actions can be executed accordingly.

[0079] First, when reading the intermediate result cache of a specified node of a specified target operator, the read request needs to carry the current timestamp T_j of the current read query, which is returned by the cache module 14 in response. The cache module 14 internally maintains multiple versions of the corresponding intermediate results, and each version is bound to the transaction timestamp when it is generated, that is, the transaction timestamp of the corresponding version is set based on the current timestamp T_j. For the corresponding read request, the query execution module 15 has four types of processing methods: The first case: If the optimal version cache is not empty and the current timestamp corresponding to the read request is less than the transaction timestamp of the oldest version of the target operator, the cache read is empty, and the query needs to be executed according to the preliminary execution plan.

[0080] The second case: If the optimal version cache is not empty and the current timestamp is equal to the transaction timestamp of the version of the target node, directly return the cache information of this version.

[0081] The third case: If the optimal version cache is not empty and the current timestamp is greater than the transaction timestamp of the latest version of the target operator, or the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target operator, then return the cache of the latest version T_i less than the timestamp T_j, and the offset value D_i corresponding to the T_i version in the append-only file system. The subsequent query starts reading all visible data (i.e., the cache information difference) from the append-only file D_i, and combines the cache and the newly obtained data to complete the query. In some implementation manners, a supplementary query is performed according to the physical offset vector to determine the cache information difference.

[0082] The fourth case: In the case where the new cache information is empty, directly execute the query according to the preliminary execution plan.

[0083] This application performs targeted queries for different situations of the latest cache information, the current timestamp, and the timestamps of multiple versions of the target operator, improving the query efficiency.

[0084] According to some embodiments, when the cache module 14 independently stores each node of the target operator, it is specifically used for: when the optimal version cache is not empty and the current timestamp is greater than the transaction timestamp of the latest version of the target node, storing the new intermediate result version during the query process of executing the query according to the preliminary execution plan to update multiple versions of the target node; when the optimal version cache is not empty and the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target node, or the current timestamp is greater than or equal to the timestamp of the latest version of the target node, determining that multiple versions of the target node do not need to be updated; when the optimal version cache is empty, storing the intermediate result version during the query process of executing the query according to the preliminary execution plan.

[0085] In this application, the query execution module 15 has different query methods, and the cache module 14 corresponds to different data update methods.

[0086] If the optimal version cache is not empty and the current timestamp is greater than the transaction timestamp of the latest version, the query process can be directly executed according to the preliminary execution plan, and then the newly generated intermediate result version is stored to update multiple versions of the corresponding target node.

[0087] If the optimal version cache is not empty, and the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target node, or the current timestamp is greater than or equal to the transaction timestamp of the latest version of the target node, there is no need to update the version at this time.

[0088] If the latest cache information is empty, the query execution module 15 executes the query according to the preliminary execution plan, and the cache module 14 can directly store the intermediate result versions generated during the query process.

[0089] In different query modes of the query execution module of this application, the version of the target node is updated specifically, and only the incremental part of the data is recalculated, reducing the occupation of computing resources.

[0090] Figure 3 It is a flowchart of the cache method for multi-version data based on a graph database provided by an embodiment of this application. The method of this embodiment can be applied to the above-mentioned cache device for multi-version data based on a graph database. As Figure 3 shown, the method includes: step S30, step S31, step S32, step S33 and step S34.

[0091] In step S30, based on the append-only method, the relationship data corresponding to the edges and nodes is continuously stored, and according to the read request and the relationship data, the physical offset value of the scan end position corresponding to the read request is determined.

[0092] In step S31, according to the actual business requirements, the preset default operator is pre-adjusted, and statistical data is collected.

[0093] In step S32, according to the statistical data, the target operator is determined from the pre-adjusted default operators, and the cache function of the target operator is enabled or disabled, where the target operator is an operator that supports caching intermediate results.

[0094] In step S33, each node of the target operator is independently saved, version maintained, and eliminated to generate overall cache information.

[0095] In step S34, based on the relationship data and the read request, the optimal version cache corresponding to the target node in the target operator is determined from the overall cache information, so as to adjust the preset preliminary execution plan based on the optimal version cache and the physical offset value, and execute the corresponding query action.

[0096] In this application, the data storage module can pre-store the relationship data between edges and nodes and the physical offset value. The user can send a read request through the user request module. The statistical analysis module can obtain relevant statistical data, pre-adjust the current default operator, and then send the adjusted default operator and the statistical data to the cached operator selection module. The cached operator selection module can select a target operator that supports caching intermediate results based on the statistical data, and control the opening or closing of these target operators.

[0097] The query execution module can attempt to obtain the latest cache information in the cache module, perform a query execution action to achieve fast query. If no query result is found, then it can perform a query without using the cached intermediate results according to the preliminary execution plan, send the query result to the user, and at the same time return the target operator and the corresponding intermediate result version used in the query process to the cached operator selection module for update, thereby updating the cache version / the intermediate result of the target operator for the first time in the cache module.

[0098] The materialized loading module can periodically materialize the intermediate results cached in the cache module and reload them at startup.

[0099] The method performs similar functions to the device provided above. For other functions, please refer to the previous description and will not be elaborated here.

[0100] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of this application, as Figure 4 shown, the electronic device 400 of this embodiment may include: a memory 401 and a processor 402.

[0101] A computer program is stored on the memory 401. When the computer program is executed by the processor 402, the foregoing processor 402 executes the method in the above embodiment.

[0102] Among them, the processor 402 and the memory 401 are connected, such as through a bus.

[0103] Optionally, the electronic device 400 may further include a transceiver. It should be noted that in actual applications, there is more than one transceiver, and the structure of the electronic device 400 does not limit the embodiments of this application.

[0104] The processor 402 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 402 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0105] The bus may include a path for transmitting information between the above components. The bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0106] The memory 401 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0107] The memory 401 is used to store the application program code for executing the solution of this application and is controlled by the processor 402 for execution. The processor 402 is used to execute the application program code stored in the memory 401 to implement the content shown in the foregoing method embodiments.

[0108] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 4 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0109] The electronic device of this embodiment can be used to execute the method of any of the above embodiments. The implementation principles and technical effects are similar and will not be elaborated here.

[0110] The present application also provides a non-transitory computer-readable storage medium, on which computer-readable instructions are stored. When the foregoing instructions are executed by a processor, the processor is caused to execute the method in the above embodiments.

[0111] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a non-transitory computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0112] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, changes or deformations made by those skilled in the art based on the idea of the present application in terms of the specific implementation manners and application scope of the present application all fall within the protection scope of the present application. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A cache device for multi-version data based on a graph database, characterized in that, Including: A data storage module, which is used to continuously store the relationship data corresponding to edges and nodes based on the append-only mode, and determine the physical offset value of the scan end position corresponding to the read request according to the read request and the relationship data; A statistical analysis module, which is used to pre-adjust the preset default operators according to actual business requirements and collect statistical data; A cache operator selection module, which is used to determine a target operator from the pre-adjusted default operators according to the statistical data, and enable or disable the cache function of the target operator, where the target operator is an operator that supports caching intermediate results; A cache module, which is used to independently save, maintain versions, and eliminate each node of the target operator to generate overall cache information; A query execution module, which is used to determine the optimal version cache corresponding to the target node in the target operator from the overall cache information based on the relationship data and the read request, and adjust the preset preliminary execution plan based on the optimal version cache and the physical offset value, and execute the corresponding query action.

2. The device according to claim 1, characterized in that It further includes a materialized loading module, which is used to materialize the overall cache information at a preset materialized interval, so as to reload the materialized overall cache information when the service restarts.

3. The device according to claim 1, characterized in that, When the scan operation corresponding to the read request is to scan multiple types of edges, the data storage module obtains the physical offset values corresponding to each of the multiple types of edges, and generates an offset vector according to the physical offset values, where the multiple types of edges are the edges corresponding to the same target operator.

4. The device according to claim 1, characterized in that The cache module caches the intermediate results according to the preset granularity of the nodes.

5. The device according to claim 1, characterized in that, When the cache operator selection module determines a target operator from the pre-adjusted default operators according to the statistical data, it is further used for: Determining the cache hit rate and the benefit of enabling the cache of the pre-adjusted default operators according to the statistical data; Determining the default operators whose cache hit rate and benefit of enabling the cache meet the preset cache enabling conditions as the target operators.

6. The device according to claim 1, characterized in that When the cache module independently saves each node of the target operator, it is specifically used for: Storing the target operator based on the preset calculation type of the target operator.

7. The device according to claim 6, characterized in that, When the cache module eliminates each operator of the target operator, it is specifically used for: Obtaining the number of versions corresponding to each operator; When the number of versions is greater than the preset cache limit number, obtaining the transaction timestamps corresponding to the multiple versions of each operator; Merging the versions whose transaction timestamps meet the preset merging conditions.

8. The device according to claim 1, characterized in that, The query execution module is specifically used for: When the optimal version cache is not empty and the current timestamp corresponding to the read request is less than the transaction timestamp of the oldest version of the target node, determining that the cache read is empty and executing the query according to the preliminary execution plan; When the optimal version cache is not empty and the current timestamp is equal to the transaction timestamp of the version of the target node, executing the query based on the optimal version cache; When the optimal version cache is not empty, and the current timestamp is greater than the transaction timestamp of the latest version of the target node, or the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target node, select the cache information of the optimal version of the node whose transaction timestamp is less than the current timestamp from the optimal version cache; Obtain the physical offset value of the version with a timestamp less than the current timestamp, and perform a supplementary query based on the physical offset value to determine the cache information difference, and execute the query based on the cache information difference and the cache information; When the optimal version cache is empty, execute the query according to the preliminary execution plan.

9. The device according to claim 8, characterized in that, When the cache module performs version maintenance on each node of the target operator, it is specifically used for: When the optimal version cache is not empty, and the current timestamp is greater than the transaction timestamp of the latest version of the target node, store the new intermediate result version in the query process of executing the query according to the preliminary execution plan to update the multiple versions of the target node; When the optimal version cache is not empty, and the current timestamp corresponding to the read request is between the transaction timestamps corresponding to different versions of the target node, or the current timestamp is greater than or equal to the transaction timestamp of the latest version of the target node, determine that the multiple versions of the target node do not need to be updated; When the optimal version cache is empty, store the intermediate result version in the query process of executing the query according to the preliminary execution plan.

10. A caching method for multi-version data based on a graph database, characterized in that, Include: Based on the append-only method, continuously store the relationship data corresponding to the edges and nodes, and determine the physical offset value of the scan end position corresponding to the read request according to the read request and the relationship data; Pre-adjust the preset default operator according to the actual business requirements, and collect statistical data; According to the statistical data, determine the target operator from the pre-adjusted default operator, and enable or disable the cache function of the target operator, where the target operator is an operator that supports caching intermediate results; Independently save, perform version maintenance, and perform elimination processing on each node of the target operator to generate overall cache information; Based on the relationship data and the read request, determine the optimal version cache corresponding to the target node in the target operator from the overall cache information, so as to adjust the preset preliminary execution plan based on the optimal version cache and the physical offset value, and execute the corresponding query action.

Citation Information

Patent Citations

  • Cache control method and device, equipment and storage medium

    CN114547093A

  • Data processing method and device for graph database, equipment and storage medium

    CN116561382A

  • Query result set caching method for database

    CN119474154A

  • Multilayer nested query cache multiplexing method and system based on dynamic materialization strategy

    CN119759976A

  • Multi-rotation type armament apparatus

    KR1020230012204A