A method, system, device, and medium supporting direct update of rule compression graph
Patent Information
- Application Number
- CN202511907466.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-12-17
AI Technical Summary
首先,基于规则的压缩结构本质上是一种层级依赖的语法树,这种紧凑的结构天然缺乏随机访问能力,使得在不解压的情况下快速定位特定的边或节点变得异常困难
1、查询性能保证:本发明通过构建基于路径的索引列表和维护压缩结构的完整性,保证了在压缩图上直接进行图算法分析的高效性。
Smart Images

Figure CN121833717B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of graph data processing and database management technology, specifically to a method, system, device, and medium that supports direct updating of rule-compressed graphs, and is particularly suitable for the storage and querying of large-scale dynamic graph data. Background Technology
[0002] With the rapid development of social networks, bioinformatics, and transportation networks, graph data is exploding at an unprecedented rate, posing severe storage and computational challenges to existing data management systems. To alleviate the storage pressure brought by massive amounts of data, graph compression technology has been widely used. Among them, rule-based compression has attracted much attention because it can achieve high compression rates by utilizing the repetitive patterns of graph structures and allows direct analysis on compressed data. However, traditional graph compression methods are mainly designed for static graphs, and their core assumption is that graph data will not change once compressed. This static assumption is fundamentally in conflict with the dynamic characteristics of graph data in the real world, which is constantly evolving and frequently updated (such as the insertion and deletion of edges). When dealing with dynamic graph data, existing static compression methods usually have to adopt a "global decompression-update-recompression" strategy. This process not only introduces extremely high time latency, often several orders of magnitude slower than direct operation, but also causes memory usage to surge instantaneously during decompression, often exceeding system memory limits, thus making compression technology almost ineffective in dynamic large-scale graph scenarios.
[0003] While some research has been conducted on dynamic graph management, existing solutions often struggle to strike a balance between space efficiency, update response speed, and query performance. On one hand, uncompressed dynamic graph systems, although achieving high update throughput through optimized data structures, suffer from significant memory overhead, limiting their ability to handle large-scale graph data, especially in environments with limited single-machine memory. On the other hand, existing compressed graph systems that support updates, while reducing storage requirements to some extent, often come at the expense of other performance aspects. For example, some systems experience fragmentation of the underlying storage structure during update processing, severely slowing down subsequent graph traversal and query speeds; others rely on local decompression and recompression operations to handle updates, limiting system throughput under high-frequency updates; still others employ excessively complex compression algorithms (such as those based on clique mining), resulting in update performance far inferior to uncompressed systems, making it difficult to meet real-time requirements.
[0004] Specifically, rule-based compression technology faces unique challenges in achieving efficient direct updates. First, the structure of rule-based compression is essentially a hierarchical syntax tree. This compact structure inherently lacks random access capabilities, making it extremely difficult to quickly locate specific edges or nodes without decompression. Second, dynamically inserted data often struggles to immediately integrate with existing compression rules. Improper handling can rapidly disrupt the original compression structure, causing the compression ratio to degrade drastically with each update. Finally, deletion operations are particularly complex because compression rules are often shared across multiple structures in the graph. Directly modifying a rule can trigger a chain reaction, incorrectly affecting other graph data that hasn't been deleted. Furthermore, the data redundancy introduced for safe deletion further diminishes the system's spatial advantages. Summary of the Invention
[0005] To address the aforementioned problems, the purpose of this invention is to provide a method, system, device, and medium that supports direct updating of rule-compressed graphs, which can simultaneously achieve high compression ratio, low-latency updates, and efficient query performance.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for supporting direct updating of rule-compressed graphs, comprising: The original graph data is converted into regular compressed graph data, and an index list based on hierarchical paths is constructed. Based on the theoretical framework of direct update based on rule compression, the rule-compressed graph data is directly updated using an index list to obtain the updated rule-compressed graph data. By using an index list and a rule traversal algorithm, graph analysis query operations are performed directly on the rule-compressed graph to obtain graph analysis query results; Return the updated rule-compressed graph data or graph analysis query results.
[0007] Furthermore, the process of converting the original graph data into regular compressed graph data and constructing an index list based on hierarchical paths includes: The input raw graph data is processed using a preset rule compression algorithm to generate a rule-compressed graph; In the rule compression graph, neighboring nodes are sampled at preset intervals, and the access path from the root rule to the target rule is recorded as an index record to build an index list.
[0008] Furthermore, based on the theoretical framework of direct update using rule compression, the method utilizes an index list to directly update the rule-compressed graph data to obtain the updated rule-compressed graph data, including: A theoretical framework for direct update based on rule compression is established, and the direct update operation is defined. Based on the index list, a two-stage indexing mechanism is used to determine the target edge; Based on the defined target edge, a dual-track architecture of "direct front-end update and asynchronous back-end cleanup" is adopted to directly update the rule-compressed graph data. The index list of the updated rule-compressed graph data is revised to ensure that the rule-compressed graph data is always in a queryable and valid state.
[0009] Furthermore, the direct update operation includes direct insert operation and direct delete operation, defined as follows: If the conditions are met: and These operations are called direct update operations on the rule compression graph; in, This is a direct insert operation; This is a direct delete operation; This is an uncompressed image. For vertex set, It is an edge set; For a regular compressed graph, where, A set of symbols for vertices or rules. For a set of rules; For decompression mapping, .
[0010] Furthermore, based on the determined target edge, a dual-track architecture of "direct front-end update and asynchronous back-end cleanup" is adopted to directly update the rule-compressed graph data, including: Receive the request to insert an edge, and based on the index list, perform a direct insertion operation using a hierarchical buffering strategy based on the LSM-Tree concept to obtain the updated rule compression graph; Receive requests to delete edges, and based on the index list, perform direct deletion operations using a path-based recursive copy-on-write strategy to obtain an updated rule-compressed graph.
[0011] Furthermore, the process of receiving edge insertion requests and performing direct insertion operations based on an index list and a hierarchical buffering strategy based on the LSM-Tree concept to obtain an updated rule compression graph includes: Level 0, Level 1, and Level 2 storage areas are set up respectively to store uncompressed incremental edges, compressed sub-image segments, and the main compressed image; The newly inserted edges are written to the level 0 storage area and stored in the incremental adjacency list as an uncompressed ordered data structure. When the data size in the Level 0 storage area exceeds a preset threshold, the data in the Level 0 storage area is compressed into an immutable compressed sub-image segment and stored in the Level 1 storage area. When the number of compressed sub-image segments in the first-level storage area exceeds the threshold, a background merging process is triggered. The compressed sub-image segments in the first-level storage area are merged with the main compressed image in the second-level storage area through decompression and recompression to obtain an updated compression rule image.
[0012] Furthermore, the process of receiving edge deletion requests and performing direct deletion operations based on the index list using a path-based recursive copy-on-write strategy to obtain an updated rule compression graph includes: Based on the index list, the unique path of the edge to be deleted is obtained, and the edge to be deleted is directly deleted using a recursive copy-on-write strategy. After performing a direct deletion, the index list is dynamically maintained. A two-stage background cleanup mechanism is used to clean up the space expansion caused by recursive write operations.
[0013] Secondly, the present invention provides a system that supports direct updating of rule-compressed graphs, comprising: The initialization module is used to convert the original graph data into regular compressed graph data and build an index list based on hierarchical paths; The direct update module is used to perform direct update operations on the rule-compressed graph data based on the rule-compressed direct update theoretical framework and using an index list to obtain the updated rule-compressed graph data. The query processing module is used to directly perform graph analysis queries on the rule-compressed graph using an index list and a rule traversal algorithm to obtain graph analysis query results. The output module is used to return the updated rule-compressed graph data or graph analysis query results.
[0014] Thirdly, the present invention provides a computer-readable storage medium for storing one or more programs, said one or more programs including instructions that, when executed by a computing device, cause the computing device to perform any method.
[0015] Fourthly, the present invention provides a computing device comprising: one or more processors and a memory, wherein the memory stores one or more programs and is configured to be executed by the one or more processors, the one or more programs including instructions for performing any method.
[0016] The present invention has the following advantages due to the adoption of the above technical solutions: 1. Guarantee of query performance: This invention ensures the efficiency of graph algorithm analysis directly on the compressed graph by constructing a path-based index list and maintaining the integrity of the compressed structure.
[0017] 2. High update responsiveness: By separating the update and compression modules and utilizing buffers and local copy-on-write mechanisms, this invention avoids costly global decompression and recompression operations on the critical update path.
[0018] 3. Excellent space efficiency: This invention effectively eliminates fragmentation and redundancy generated during the update process through regular merging and a two-stage cleanup strategy (rule folding and secondary compression) in the background, so that the compressed image after dynamic update can still maintain a high compression rate.
[0019] Therefore, this invention can be widely applied in the fields of graph data processing and database management technology. Attached Figure Description
[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. In the drawings: Figure 1 This is a flowchart of a method for directly updating rule-compressed graphs provided in an embodiment of the present invention; Figure 2 This is an example diagram illustrating the support for direct updating of rule-compressed graphs provided in this embodiment of the invention; Figure 3 This is a schematic diagram of the index list and path location provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the hierarchical merging process in the direct insertion stage provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of recursive copying and background cleanup during the direct deletion stage provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0022] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0023] In some embodiments of the present invention, a method for supporting direct updates of rule-compressed graphs is provided. This method resolves the contradiction between compressed storage and real-time updates of dynamic graph data by establishing a theoretical framework for direct updates of rule-compressed graphs. Based on this theoretical framework, a separate direct update is performed. In the real-time operation path, locality is prioritized to achieve high throughput. In the background maintenance path, compression efficiency is achieved through periodic merging and recompression, while ensuring update closure throughout the process. This allows the present invention to achieve efficient dynamic graph management without sacrificing query performance and storage advantages.
[0024] Correspondingly, in other embodiments of the present invention, a system, device, and medium are provided that support direct updating of rule-compressed graphs.
[0025] Example 1 like Figure 1 , Figure 2 As shown, the present invention provides a method for supporting direct updating of rule-compressed graphs, which includes the following steps: 1) Convert the original graph data into regular compressed graph data and construct an index list based on hierarchical paths; 2) Based on the theoretical framework of direct update based on rule compression, the rule-compressed graph data is directly updated using an index list to obtain the updated rule-compressed graph data. 3) By using an index list and rule traversal algorithm, graph analysis queries are performed directly on the rule-compressed graph to obtain the graph analysis query results; 4) Return the updated rule-compressed graph data or graph analysis query results.
[0026] Furthermore, in step 1) above, such as Figure 3 As shown, since the rule-compressed graph is essentially a directed acyclic graph (DAG) composed of nested rules, the traditional array index-based access method fails. To achieve fast edge location without decompression, this embodiment designs a path-based indexing mechanism.
[0027] Specifically, it includes the following steps: 1.1) The input original graph data is processed using a preset rule compression algorithm to generate a rule-compressed graph.
[0028] Specifically, the rule compression algorithm can be a variant of the Re-Pair algorithm, or other rule compression algorithms can be selected according to actual needs. This invention does not impose any restrictions on this.
[0029] 1.2) In the rule compression graph, neighbor nodes are sampled at preset intervals, and the access path from the root rule to the target rule is recorded as an index record to build an index list.
[0030] In a regular compressed graph, an edge cannot be located directly using a simple memory address. In this embodiment, a "path" is defined as an ordered sequence of offsets. Each element in this offset sequence Representative at the Offsets within the layer rule body. This sequence acts as navigation instructions: starting from the root rule, based on elements... Find the next level sub-rule reference, enter the sub-rule, and then proceed according to... The addressing continues until a specific edge is finally located. This path description method guarantees the unique indexing capability of any edge in a hierarchical compressed structure.
[0031] Meanwhile, to balance the space overhead of the index with retrieval speed, this embodiment does not create an index for every edge, but instead adopts a "skip-sampling" strategy. The index sampling interval is specified as follows: For each vertex's adjacency list, every... neighboring nodes (e.g.) An index is created, with each index entry recording a tuple (neighbor vertex ID, access path). During compression, the system scans the generated rule stream. When it detects that the distance between the currently processed symbol and the previous index point reaches an interval... And the current rule depth does not exceed the preset maximum depth limit. When an index record is generated, it is inserted into the index list.
[0032] In this embodiment, an index list is used to map vertices to the access paths of their neighboring nodes in the compression rules, thereby supporting random access. Only a rule-compression graph that has been initialized and indexed can be efficiently processed by subsequent modules.
[0033] Furthermore, step 2) above includes the following steps: 2.1) Establish a theoretical framework for direct update based on rule compression, which includes direct insertion and direct deletion operations; 2.2) Based on the index list, a two-stage indexing mechanism is used to determine the target edge; 2.3) Based on the defined target edge, a dual-track architecture of "direct front-end update and asynchronous back-end cleanup" is adopted to directly update the rule-compressed graph data; 2.3) Correct the index list of the updated rule-compressed graph data to ensure that the rule-compressed graph data is always in a queryable and valid state, and avoid subsequent queries from failing due to structural damage.
[0034] Furthermore, in step 2.1) above, the direct update theoretical framework based on rule compression proposed in this invention aims to resolve the contradiction between compressed storage and real-time updates of dynamic graph data. In traditional graph compression technology, compression is usually regarded as a static archiving process, and update operations often require an expensive "decompression-update-recompression" cycle. To break this limitation, this invention introduces the concept of direct update and establishes an axiomatic system to regulate dynamic operations on compressed graphs.
[0035] In terms of form, if the uncompressed image is represented as ,in For vertex set, Let it be an edge set; the rule-compressed graph is represented as... ,in A set of symbols for vertices or rules. For a set of rules. Given a compression mapping. and decompression mapping This invention defines a direct insertion operation on a regular compressed graph. And direct deletion operation If the conditions are met: and These operations are called direct updates on the regular compressed graph. This means that the result of an operation performed directly on the structure of the regular compressed graph must be strictly equivalent to the result of the corresponding operation performed on the original uncompressed graph, that is, it must satisfy semantic consistency.
[0036] To ensure the feasibility and efficiency of direct updates in practical applications, this invention establishes three fundamental axioms that the direct update paradigm must follow: Locality. This is a prerequisite for ensuring update efficiency. Given a rule-compressed graph. Direct insertion and direct deletion operations should only modify the rule compression diagram. rule set in and symbol set A subset of the graph. This means that the computational cost of an update should be proportional to the size of the updated content, and independent of the overall size of the graph. This invention, through copy-on-write and incremental buffering mechanisms, ensures that operations primarily occur within local paths or buffers, thus satisfying the locality requirement.
[0037] Compression efficiency. This is the fundamental reason for the existence of compressed graphs. For any edge set... The size of the updated compressed image must always be smaller than the size of the updated uncompressed image, i.e. However, simply pursuing locality (such as directly generating new rules) often leads to a large number of inefficient rules (i.e., rules with low reference counts), thereby compressing efficiency. Therefore, this invention theoretically introduces a metric for "rule efficiency." Based on this, a background cleanup mechanism was designed to maintain compression efficiency during long-term operation.
[0038] Updating closure properties is crucial for ensuring system availability. (Compressed graph set) The graph should remain closed under direct update operations; that is, the updated compressed graph structure must still be accessible to the originally defined set of query algorithms. This requires that update operations not only modify the data, but also maintain the integrity of the indexes and structure to ensure that the graph is queryable at all times.
[0039] Furthermore, in step 2.2) above, this embodiment establishes a two-stage retrieval mechanism: based on the index list established in step 1), a retrieval algorithm of "binary search + last-mile scanning" is implemented.
[0040] Specifically, it includes the following steps: 2.2.1) First stage (coarse-grained localization): based on the given target edge At the vertex In the index list, perform a binary search on the neighbor vertex IDs to find those less than or equal to... The largest index entry (i.e., the nearest predecessor anchor); 2.2.2) Second stage (fine-grained scanning): Using the paths recorded in the largest index item, quickly locate the specific rule positions inside the rule compression graph; 2.2.3) Starting from the determined specific rule location, perform a linear local decoding scan in the compressed rule stream until the target edge is found. .
[0041] Through the index design in this embodiment, the time complexity of edge retrieval is significantly reduced, and users can flexibly balance memory usage and query performance through adjustable sampling interval parameters.
[0042] Furthermore, based on the aforementioned theoretical framework of rule-based compression for direct updates, this invention designs a direct update scheme, the core principles of which are read-write separation and static / dynamic separation: to simultaneously satisfy high update throughput and high compression rate, this embodiment adopts a dual-track architecture of "foreground direct update and background asynchronous cleanup": Front-end: Responsible for handling real-time insert and delete requests. This path is designed strictly according to locality of reference; that is, any update operation only modifies the rule set. and symbol set A very small subset of each ensures that the time complexity of the operation is independent of the overall size of the graph, and only depends on the length of the local paths involved in the update. Therefore, foreground paths resolutely avoid full graph decompression or expensive global rule reconstruction.
[0043] Backend: Responsible for maintaining compression efficiency. Since partial updates to the frontend (such as copy-on-write or incremental buffering) inevitably introduce data redundancy or fragmentation, the system is designed with an independent background process that periodically performs "subgraph merging," "rule folding," and "secondary compression" operations. Although these operations are costly, they do not block high-frequency trading on the frontend due to asynchronous execution.
[0044] Therefore, in step 2.3) above, when performing the direct update operation, the following steps are included: 2.3.1) Receive the request to insert an edge, and based on the index list, perform a direct insertion operation using a hierarchical buffering strategy based on the LSM-Tree concept to obtain the updated rule compression graph; 2.3.2) Receive the request to delete an edge, and based on the index list, perform a direct deletion operation using a path-based recursive copy-on-write strategy to obtain the updated rule compression graph.
[0045] Furthermore, in step 2.3.1) above, as Figure 4 As shown, to address the difficulty of integrating new data into the rule-compressed graph, this embodiment adopts a hierarchical storage and merging architecture, similar to a variant of the log structure merge tree (LSM-Tree), to support high-throughput direct insertion operations.
[0046] 2.3.1.1) Set up storage areas of level 0 (memory buffer), level 1 (compressed sub-layer), and level 2 (main layer) respectively to store uncompressed incremental edges, compressed sub-image segments, and the main compressed image; 2.3.1.2) Write the newly inserted edge into the level 0 storage area and store the incremental adjacency list in an uncompressed ordered data structure.
[0047] In this embodiment, all newly inserted edges are not directly written to the compressed structure, but are first placed into the level 0 buffer. Within the buffer, an uncompressed ordered data structure (such as a red-black tree or skip list) is used to store the incremental adjacency list. Leveraging the high speed of memory, the insertion operation is ensured to have extremely low latency, and the ordered structure supports fast deduplication and retrieval via binary search.
[0048] 2.3.1.3) When the data size of the Level 0 storage area exceeds the preset threshold, the data in the Level 0 storage area is compressed into an immutable compressed sub-image segment and stored in the Level 1 storage area.
[0049] When the size of the level 0 buffer reaches a preset threshold (e.g., 100,000 edges), a "small merge" is triggered. The system compresses the data in the buffer into an independent, immutable subgraph using a rule-based compression algorithm and promotes it to level 1 storage. At this point, the graph logically consists of a "main compressed graph" and several "incrementally compressed subgraphs." During queries, the system logically combines the adjacency lists of these parts.
[0050] 2.3.1.4) When the number of compressed sub-image segments in the first-level storage area exceeds the threshold, a background merging process is triggered. The compressed sub-image segments in the first-level storage area are merged with the main compressed image in the second-level storage area through decompression and recompression to obtain an updated compression rule image.
[0051] As time progresses, the number of compressed sub-image segments in Level 1 will continuously increase, requiring the traversal of too many compressed sub-image segments during queries, thus reducing read performance. Therefore, when the number of compressed sub-image segments in Level 1 exceeds a threshold (e.g., 10), a "large merge" is triggered. This embodiment explicitly abandons the scheme of directly merging two rule sets because theoretical analysis shows that directly merging rule sets has extremely high complexity. Instead, this embodiment adopts a "decompression-recompression" strategy: First, in the background, all compressed sub-image segments of level 1 and the main compressed image of level 2 are partially decompressed asynchronously; Secondly, the decompressed data stream is merged and sorted; Finally, the rule compression algorithm is re-run on the merged data to generate a brand new Level 2 master compressed graph.
[0052] This process runs entirely in the background, without blocking read and write operations in the foreground, and can periodically eliminate fragmentation caused by incremental updates, ensuring a high compression ratio over the long term. This embodiment, through a layered design, effectively distributes the expensive compression computations to background batch processing, thereby achieving near-native in-memory database insertion performance in the foreground.
[0053] Furthermore, in step 2.3.2 above, as... Figure 5 As shown, this embodiment details a method for direct deletion and background cleanup while maintaining the integrity of the compressed structure. The difficulty in deletion lies in the fact that compression rules are often referenced (shared) in multiple places, and directly modifying the rules can incorrectly affect non-target areas. This embodiment solves this problem by combining "local modification" with "delayed optimization."
[0054] Specifically, it includes the following steps: 2.3.2.1) Obtain the unique path of the edge to be deleted based on the index list, and use a recursive copy-on-write strategy to directly delete the edge to be deleted; 2.3.2.2) After performing a direct deletion, the index list is dynamically maintained; 2.3.2.3) A two-stage background cleanup mechanism is adopted to clean up the space expansion caused by recursive write operations, so that the system can maintain an excellent compression ratio even under long-term frequent updates.
[0055] Furthermore, in step 2.3.2.1) above, the deletion process includes the following steps: ① Pathfinding: Based on the index list, obtain the edge to be deleted. The unique path in the rule tree .
[0056] ② Recursive copying: from path Starting from the lowest-level rule and moving upwards to the root rule, each rule node along the path is copied layer by layer, creating a new rule copy. The content of this new rule copy is identical to the original rule, but the target edge to be deleted or the pointer to the target edge has been removed. The copy of the parent rule will point to the new copy of the child rule, not the original rule.
[0057] ③ Atomic Replacement: At the root rule level, pointers to the original rule chain are replaced with pointers to the newly copied rule chain. The original rule remains physically unchanged (only its reference count decreases) because it may be referenced from other locations in the graph. The newly generated rule chain only serves the currently modified adjacency list. This ensures that modifications are limited to a local scope and do not affect global data consistency.
[0058] Furthermore, in step 2.3.2.2 above, the dynamic maintenance of the index includes the following steps: ① After performing the deletion, the system immediately checks the index list of the affected vertices.
[0059] ② Remove invalid indexes: If the anchor point of an index entry is within the scope of the rules to be deleted, then the index is removed directly.
[0060] ③ Offset Correction: If an index item is located after the deletion point, the system calculates the length change caused by the deletion operation and updates the offset parameter of the path in the index item to ensure that subsequent queries can navigate correctly.
[0061] Furthermore, in step 2.3.2.3 above, the copy-on-write mechanism generates a large number of new rules that are only referenced once, which violates the efficiency principle of rule compression (i.e., rules should be reused multiple times), leading to space expansion. This embodiment designs a background two-stage cleanup process, including: Phase 1: Rule Collapsing. The background scanning process identifies all rules with a reference count of 1. Since these rules no longer offer reuse value, the system directly "inlines" and flattens them into their parent rules, freeing up space previously occupied by those rules. This step eliminates invalid intermediate layers.
[0062] Phase Two: Secondary Compression. After rule folding, the data in the affected regions reverts to a flat sequence. The system then runs the compression algorithm again on these local regions to uncover new potential repeating patterns and generate new, highly efficient rules (reference counting). ).
[0063] This embodiment ensures the logical correctness and isolation of deletion operations through copy-on-write, and effectively recovers the "space debt" generated by updates through two-stage cleanup in the background, enabling the system to maintain an excellent compression ratio even under long-term and frequent updates.
[0064] Example 2 The above-described embodiment 1 provides a method for supporting direct updating of rule-compressed graphs. Correspondingly, this embodiment provides a system for supporting direct updating of rule-compressed graphs. The system provided in this embodiment can implement the method for supporting direct updating of rule-compressed graphs in embodiment 1. The system can be implemented through software, hardware, or a combination of both. For example, the system may include integrated or separate functional modules or units to execute the corresponding steps in the methods of embodiment 1. Since the system in this embodiment is basically similar to the method embodiment, the description process in this embodiment is relatively simple. Relevant details can be found in the description of embodiment 1. The system embodiment provided in this embodiment is merely illustrative.
[0065] The system provided in this embodiment that supports direct updating of rule-compressed graphs includes: The initialization module is used to convert the original graph data into regular compressed graph data and build an index list based on hierarchical paths; The direct update module is used to perform direct update operations on the rule-compressed graph data based on the rule-compressed direct update theoretical framework and using an index list to obtain the updated rule-compressed graph data. The query processing module is used to directly perform graph analysis queries on the rule-compressed graph using an index list and a rule traversal algorithm. The output module is used to return the updated rule-compressed graph data or graph analysis query results.
[0066] Furthermore, directly update the module, including: The definition module is used to establish a theoretical framework for direct updates based on rule compression, which includes direct insertion and direct deletion operations; The index positioning module is used to determine the target edge based on the index list and using a two-stage indexing mechanism. The direct update operation module is used to directly update the rule-compressed graph data based on a defined target edge, using a dual-track architecture of "direct front-end update and asynchronous back-end cleanup". The maintenance module is used to correct the index list of the updated rule-compressed graph data to ensure that the rule-compressed graph data is always in a queryable and valid state, and to avoid subsequent query failures due to structural damage.
[0067] Furthermore, directly update the operation module, including: The direct insertion module receives edge insertion requests and performs direct insertion operations based on the index list and a hierarchical buffering strategy based on the LSM-Tree concept to obtain the updated rule compression graph. The direct deletion module receives requests to delete edges and performs direct deletion operations based on the index list using a path-based recursive copy-on-write strategy, resulting in an updated rule compression graph.
[0068] In this embodiment, the direct update module employs a hierarchical storage architecture (similar to an LSM-Tree design) to balance write speed and query performance. It includes a memory buffer (Level 0) for temporarily storing uncompressed incremental edges; an intermediate layer (Level 1) for storing compressed sub-image segments; and a main storage layer (Level 2) for storing the main compressed image. This module is responsible for monitoring the thresholds of each layer and triggering compression and merging tasks between layers.
[0069] The direct deletion module is responsible for removing edges while maintaining the integrity of the compressed structure. This module implements path-based recursive copy-on-write logic and integrates a background cleanup mechanism. When a deletion request is received, it does not immediately reconstruct the entire graph, but instead performs local rule copying and pointer replacement. To prevent space bloat, this module includes a background garbage collector to scan for and collapse inefficient rules.
[0070] Example 3 This embodiment provides a processing device corresponding to the method for directly updating the rule compression graph provided in Embodiment 1. The processing device can be a client-side processing device, such as a mobile phone, laptop, tablet computer, desktop computer, etc., to execute the method of Embodiment 1.
[0071] The processing device includes a processor, a memory, a communication interface, and a bus. The processor, memory, and communication interface are connected via the bus to enable communication between them. The memory stores a computer program that can run on the processor. When the processor runs the computer program, it executes the method for directly updating the rule-compressed graph provided in Embodiment 1.
[0072] Preferably, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.
[0073] Preferably, the processor can be any type of general-purpose processor such as a central processing unit (CPU) or a digital signal processor (DSP), and there is no limitation herein.
[0074] Example 4 The method for supporting direct updating of rule-compressed graphs in Embodiment 1 can be specifically implemented as a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing the method for supporting direct updating of rule-compressed graphs as described in Embodiment 1 are loaded.
[0075] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.
[0076] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for supporting direct updating of rule-compressed graphs, characterized in that, include: The original graph data is converted into regular compressed graph data, and an index list based on hierarchical paths is constructed. The updated rule-compressed graph data is obtained by directly updating the rule-compressed graph data using the index list. By using an index list and a rule traversal algorithm, graph analysis query operations are performed directly on the rule-compressed graph to obtain graph analysis query results; Return the updated rule-compressed graph data or graph analysis query results; The step of directly updating the rule-compressed graph data using an index list to obtain the updated rule-compressed graph data includes: The direct update operation is defined as including direct insert and direct delete operations. If the conditions are met: and These operations are called direct update operations on the rule compression graph; in, This is a direct insert operation; This is a direct delete operation; This is an uncompressed image. For vertex set, It is an edge set; For a regular compressed graph, where, A set of symbols for vertices or rules. For a set of rules; For decompression mapping, ; Based on the index list, determine the position of the target edge in the rule-compressed graph; Based on a defined target edge, a dual-track architecture of "direct front-end update and asynchronous back-end cleanup" is adopted to directly update the rule-compressed graph data, including: Receive the request to insert an edge, and based on the index list, perform a direct insertion operation using a hierarchical buffering strategy based on the LSM-Tree concept to obtain the updated rule compression graph; Receive requests to delete edges, and based on the index list, perform direct deletion operations using a path-based recursive copy-on-write strategy to obtain an updated rule-compressed graph; The index list of the updated rule-compressed graph data is revised to ensure that the rule-compressed graph data is always in a queryable and valid state.
2. The method for supporting direct updating of rule-compressed graphs as described in claim 1, characterized in that, The process of converting the original graph data into regular compressed graph data and constructing an index list based on hierarchical paths includes: The input raw graph data is processed using a preset rule compression algorithm to generate a rule-compressed graph; In the rule compression graph, neighboring nodes are sampled at preset intervals, and the access path from the root rule to the target rule is recorded as an index record to build an index list.
3. The method for supporting direct updating of rule-compressed graphs as described in claim 1, characterized in that, The process involves receiving edge insertion requests and, based on an index list, performing direct insertion operations using a hierarchical buffering strategy based on the LSM-Tree concept to obtain an updated rule-compressed graph, including: Level 0, Level 1, and Level 2 storage areas are set up respectively to store uncompressed incremental edges, compressed sub-image segments, and the main compressed image; The newly inserted edges are written to the level 0 storage area and stored in the incremental adjacency list as an uncompressed ordered data structure. When the data size in the Level 0 storage area exceeds a preset threshold, the data in the Level 0 storage area is compressed into an immutable compressed sub-image segment and stored in the Level 1 storage area. When the number of compressed sub-image segments in the first-level storage area exceeds the threshold, a background merging process is triggered. The compressed sub-image segments in the first-level storage area are merged with the main compressed image in the second-level storage area through decompression and recompression to obtain an updated compression rule image.
4. The method for supporting direct updating of rule-compressed graphs as described in claim 1, characterized in that, The process involves receiving requests to delete edges and, based on an index list, performing direct deletion operations using a path-based recursive copy-on-write strategy to obtain an updated rule-compressed graph, including: Based on the index list, the unique path of the edge to be deleted is obtained, and the edge to be deleted is directly deleted using a recursive copy-on-write strategy. After performing a direct deletion, the index list is dynamically maintained. A two-stage background cleanup mechanism is used to clean up the space expansion caused by recursive write operations.
5. A system that supports direct updating of rule-compressed graphs, characterized in that, include: The initialization module is used to convert the original graph data into regular compressed graph data and build an index list based on hierarchical paths; The direct update module is used to directly update the rule compression graph data using the index list, and obtain the updated rule compression graph data. The query processing module is used to directly perform graph analysis queries on the rule-compressed graph using an index list and a rule traversal algorithm to obtain graph analysis query results. The output module is used to return the updated rule-compressed graph data or graph analysis query results; The step of directly updating the rule-compressed graph data using an index list to obtain the updated rule-compressed graph data includes: Define the direct update operation: The direct update operation includes direct insert operation and direct delete operation. If the conditions are met: and These operations are called direct update operations on the rule compression graph; in, This is a direct insert operation; This is a direct delete operation; This is an uncompressed image. For vertex set, It is an edge set; For a regular compressed graph, where, A set of symbols for vertices or rules. For a set of rules; For decompression mapping, ; Based on the index list, determine the position of the target edge in the rule-compressed graph; Based on a defined target edge, a dual-track architecture of "direct front-end update and asynchronous back-end cleanup" is adopted to directly update the rule-compressed graph data, including: Receive the request to insert an edge, and based on the index list, perform a direct insertion operation using a hierarchical buffering strategy based on the LSM-Tree concept to obtain the updated rule compression graph; Receive requests to delete edges, and based on the index list, perform direct deletion operations using a path-based recursive copy-on-write strategy to obtain an updated rule-compressed graph; The index list of the updated rule-compressed graph data is revised to ensure that the rule-compressed graph data is always in a queryable and valid state.
6. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 4.
7. A computing device, characterized in that, include: One or more processors and a memory, wherein the memory stores one or more programs and is configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 4.
Citation Information
Patent Citations
Compressed data direct calculation method and system based on nonvolatile storage
CN118363528A
LSM-Tree key value storage system for establishing query index by using underlying information
CN119127867A