A segment-based aggregation hash table reading and writing method and system based on a graph database

By separating the hash table into a hash entry table and a data payload block, and using linear probing and hash prefix filtering to optimize memory access, the low performance problem of traditional hash tables in large data volume scenarios is solved, and the read and write performance and aggregation computing efficiency are improved.

CN120561129BActive Publication Date: 2025-10-17杭州悦数科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511086668.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-17
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

In large data volume scenarios, traditional hash tables have low read and write performance, becoming a performance bottleneck, especially in aggregate computing scenarios, mainly due to frequent hash conflicts and performance degradation caused by expansion operations.

Method used

The hash table is separated into a hash entry table and a data load block. A linear probing method is used to locate hash entries and store data. Hash prefix filtering and continuous memory layout are used to optimize memory access, reduce cache misses and capacity expansion overhead.

Benefits of technology

This improves the read and write performance of hash tables, reduces the overhead of cache misses and capacity expansion operations, and improves the efficiency of aggregate calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561129B_ABST
    Figure CN120561129B_ABST
Patent Text Reader

Abstract

The application relates to a segmented aggregation hash table reading and writing method and system based on a graph database, wherein the method comprises the following steps: separating an original hash table into a hash entry table and a data load block, wherein the hash entry table is used for storing hash entry data, and the data load block is used for storing aggregation data; in response to receiving a write data instruction, locating a hash entry in the hash entry table according to a hash value of the write data; based on a data storage state of the hash entry, storing the write data in the data load block through linear probing, and updating the hash entry; and in response to receiving a read data instruction, obtaining target aggregation data corresponding to the read data through linear probing according to a hash value of the read data. The hash table reading and writing performance is improved, and the problem that the traditional hash table reading and writing process has low performance in a large data amount scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of databases, and in particular to a segmented aggregate hash table reading and writing method and system based on a graph database. BACKGROUND

[0002] In the field of databases, hash tables are widely used in the efficient processing of grouping queries and aggregate functions (such as avg, count) as the core data structure for implementing aggregate calculations. Aggregate functions have higher computational complexity than scalar functions, and queries involving aggregate calculations in large data scenarios often have poor performance, becoming a performance bottleneck for the entire business program.

[0003] Traditional hash tables achieve fast reading and writing through hash mapping of key-value pairs, and their performance depends on the uniform distribution of hash functions and conflict resolution strategies. However, in large data scenarios, frequent hash collisions can significantly reduce performance. For example, the unordered_map of the C++ Standard Template Library (Standard Template Library) uses a linked list method to handle collisions, hanging conflict entries in hash buckets through pointers. This design can solve collisions, but the scattered storage of linked list nodes destroys the spatial locality of data, resulting in low CPU cache prefetch efficiency, and further causing frequent cache misses and CPU stalls (which further reduces pipeline efficiency and has a more serious impact on performance).

[0004] Therefore, in large data scenarios, the traditional hash table reading and writing process has the problem of low performance. SUMMARY

[0005] The embodiments of the present application provide a segmented aggregate hash table reading and writing method and system based on a graph database to at least solve the problem of low performance of the traditional hash table reading and writing process in large data scenarios.

[0006] In a first aspect, the embodiments of the present application provide a segmented aggregate hash table reading and writing method based on a graph database, which comprises:

[0007] Separating the original hash table into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregate data;

[0008] In response to receiving a write data instruction, locating the hash entry in the hash entry table according to the hash value of the write data,

[0009] storing the write data in the data load block by linear probing and updating the hash entry based on a data storage state of the hash entry;

[0010] in response to receiving a read data instruction, obtaining target aggregated data corresponding to the read data by linear probing according to a hash value of the read data.

[0011] In an embodiment, the hash entry table is a list structure in a continuous memory layout manner, wherein the hash entry table includes multiple rows of hash entries, and each row of the hash entries includes:

[0012] a hash value prefix for storing high-order bytes of a hash value;

[0013] a block number for locating a target data load table in the data load block;

[0014] an intra-block address for locating a data load entry in the target data load table.

[0015] In an embodiment, the data load block is a list structure in a continuous memory layout manner, and the data load block includes multiple data load tables with consistent structures, and each data load table includes multiple rows of data load entries, and each row of the data load entries includes:

[0016] a hash value for storing a hash value of a group column, the hash value of the group column being calculated when write data is written and being stored in advance;

[0017] a group column for storing group column data;

[0018] an aggregated column for storing a calculation result of an aggregated function.

[0019] In an embodiment, the locating the hash entry in the hash entry table according to the hash value of the write data includes:

[0020] calculating the hash value of the write data;

[0021] locating a row number of the write data in the hash entry table by taking a remainder of the hash value of the write data divided by a number of rows of the hash entry table;

[0022] locating the hash entry in the hash entry table based on the row number.

[0023] In an embodiment, the storing the write data in the data load block by linear probing and updating the hash entry based on the data storage state of the hash entry includes the following steps:

[0024] determining whether the hash entry stores data;

[0025] If the hash entry does not store data, the write data is stored in the data load block and the data of the hash entry is updated. If the hash entry stores data, the write data is stored in the data load block hash entry through linear probing.

[0026] In an embodiment, the storing the write data in the data load block through linear probing includes:

[0027] determining whether a prefix of a hash value of the write data matches a prefix of a hash value in the hash entry table;

[0028] If not, determining whether a prefix of a hash value of the write data matches a prefix of a hash value in the hash entry table according to a continuous memory access rule;

[0029] If yes, reading a corresponding data load entry according to a block number and an in-block address in the hash entry table, determining whether a hash value in the data load entry is the same as the hash value of the write data and whether a group column key value is the same as a key value of the write data, and updating aggregate column data corresponding to the group column key value of the data load block in the case of being the same.

[0030] In an embodiment, after the data of the hash entry is updated, when a number of rows of the hash entry table reaches a capacity expansion threshold, the method further includes:

[0031] synchronously expanding a capacity of the hash entry table and the data load block to N times of an initial size, wherein a physical storage location in the data load block is unchanged, and data in the hash entry table is updated through the hash value.

[0032] In an embodiment, in response to receiving a read data instruction, target aggregate data corresponding to the read data is obtained through linear probing according to a hash value of the read data, including:

[0033] calculating the hash value of the read data and performing linear probing in the hash entry;

[0034] determining whether a prefix of the hash value of the read data matches a prefix of a hash value in the hash entry table;

[0035] If not matched, according to the continuous memory access rule, it is judged whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the read data, if matched, the corresponding data load entry is read according to the block number and the intra-block address in the hash entry table, it is judged whether the hash value in the data load entry is the same as the hash value of the read data, if the same, it is judged whether the grouping column key value is the same as the key value of the read data, in the case of the same, the target aggregation data corresponding to the read data is obtained.

[0036] In an embodiment, after obtaining the target aggregation data corresponding to the read data, the method further comprises:

[0037] According to the data type, the segmented aggregation hash table is optimized, wherein according to the fixed-length data type, the grouping column data of the data load entry is stored in the hash prefix position of the hash entry, and according to the variable-length string type, the grouping column is stored through the dictionary tree;

[0038] And / or, the segmented aggregation hash table is optimized through a parallel computing strategy, the input data is partitioned according to a hash range, each of the partitions is processed through an independent thread, each of the independent threads processes the hash entry table and the data load block, a local hash table is constructed through address reference, and the target aggregation data corresponding to the read data is obtained based on the local hash table.

[0039] In a second aspect, the embodiments of the present application provide a segmented aggregation hash table read-write system based on a graph database, the system comprising a separation module, a positioning module, a writing module and a reading module; wherein:

[0040] The separation module is configured to separate an original hash table into a hash entry table and a data load block, wherein the hash entry table is configured to store hash entry data, and the data load block is configured to store aggregation data;

[0041] The positioning module is configured to, in response to receiving a write data instruction, locate a hash entry in the hash entry table according to a hash value of the write data,

[0042] The writing module is configured to, based on a data storage state of the hash entry, store the write data in the data load block through linear probing, and update the hash entry;

[0043] The reading module is configured to, in response to receiving a read data instruction, obtain target aggregation data corresponding to the read data through linear probing according to a hash value of the read data.

[0044] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method for reading and writing a segmented aggregate hash table based on a graph database according to the first aspect.

[0045] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, having a computer program stored thereon, and the program is executable by a processor to implement the method for reading and writing a segmented aggregate hash table based on a graph database according to the first aspect.

[0046] The method and system for reading and writing a segmented aggregate hash table based on a graph database provided by the embodiments of the present application have at least the following technical effects.

[0047] The original hash table is separated into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregate data. The hash entry table is decoupled from the data load block, the memory access locality is improved, and the Cache Miss is reduced. In response to receiving a write data instruction, a hash entry in the hash entry table is located according to a hash value of the write data. The target entry is directly located through the hash value, the traversal process is avoided, and the calculation overhead is reduced. Based on the data storage state of the hash entry, the write data is stored in the data load block through linear probing, and the hash entry is updated. The conflict is solved in the continuous space of the hash entry table, and the probing process is accelerated. In response to receiving a read data instruction, target aggregate data corresponding to the read data is obtained through linear probing according to a hash value of the read data. Invalid entries are quickly filtered through the hash prefix, the number of data load block access times is reduced, and the aggregate data reading efficiency is improved. In the present application, the linear probing of the continuous memory layout reduces the cache invalidation when data is written, and the hash value is pre-stored to avoid repeated calculation. When data is queried, the hash prefix is used for hierarchical filtering to skip most non-matching items, only a small amount of data load block data needs to be accessed, the hash table reading and writing performance is improved, and the problem of low performance in the traditional hash table reading and writing process in a large data amount scenario is solved.

[0048] Details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects, and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS

[0049] The drawings described herein are intended to provide further understanding of the present application, form a part of the present application, and serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0050] Figure 1is a flow chart of a graph database based segmented aggregated hash table read-write method;

[0051] Figure 2 is a structural diagram of a hash entry table and data load block according to an exemplary embodiment;

[0052] Figure 3 is a flow chart of step S103 according to an exemplary embodiment;

[0053] Figure 4 is a structural diagram of a parallel segmented aggregated hash table according to an exemplary embodiment;

[0054] Figure 5 is a system structure block diagram of a graph database based segmented aggregated hash table read-write according to an exemplary embodiment;

[0055] Figure 6 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0056] For the purpose of making the purpose, technical scheme and advantages of the present application more clear, the present application is described and explained in the following with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.

[0057] Obviously, the drawings in the following description are only some examples or embodiments of the present application, and for those of ordinary skill in the art, the present application can be applied to other similar scenarios without making creative efforts based on the drawings. In addition, it can be understood that although the efforts made in the development process can be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacture or production changes based on the technical content disclosed in the present application are only routine technical means, and should not be understood as insufficient disclosure of the present application.

[0058] In the present application, "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it independent or alternative to other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in the present application can be combined with other embodiments without conflict.

[0059] Unless otherwise defined, technical terms and scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terms "a", "an", "one", "this", and the like, as used in the specification herein, do not imply quantity of one, but rather denote the presence of at least one of something or unit. The terms "include", "comprise", "have", and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product or device that comprises a list of steps or units (elements) is not limited to the listed steps or units, but can further include other steps or units not listed or can further include other steps or units inherent to such process, method, product or device. The terms "connect", "connected", "coupling" and the like, are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The term "multiple" refers to two or more. The term "and / or" describes the associated relationship of associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects. The terms "first", "second", "third" and the like are merely used to distinguish similar objects, and do not represent a specific order of the objects.

[0060] In the field of databases, aggregate functions are widely used in computing applications, and their computing complexity is higher than that of scalar functions. In large data scenarios, queries containing aggregate calculations often have poor performance and are more likely to become performance bottlenecks of the entire business program. For example, a query statement first queries all teachers and their students through a graph topology, then groups and aggregates each teacher by class, and calculates the score statistics of each teacher's class. In this query statement, there are two aggregate functions avg (average score of students in the class) and count (number of students passing the class in the class). Therefore, the query performance is poor.

[0061] In engineering practice, the performance of queries containing aggregate functions in large data scenarios is often poor, and the query overhead is mainly concentrated in the aggregation operator. Through performance analysis, it is found that the performance bottleneck of the aggregation operator is mainly concentrated in the overhead of the hash table. For example:

[0062] In the existing data writing process, the traditional hash table first calculates the hash value of the key of the key-value pair to be written and locates a specific slot (bucket). If the slot is empty, the key-value pair is written to the slot. If the slot is not empty, a linked list node is created and connected to the tail of the current slot linked list. Hash collision occurs when there are multiple key-value pairs in the slot, which causes a serious decline in performance. In the worst case, the time complexity decreases from O(1) to O(n). The solution to the problem of too high hash collision rate is to expand the capacity, which means that once the amount of data written exceeds a certain threshold, the hash table will be expanded, which will cause the following overheads: a. Large amount of data movement and copying b. Rehash overhead.

[0063] In the existing reading process, the unordered mapping first calculates the hash value of the key to be queried and locates its slot in the hash table. Each slot in the traditional hash table stores all key-value pairs with the same hash value through a linked list. After locating the specified slot, it is still necessary to traverse each key-value pair in the slot and find the key-value pair with the same key. This process has the following overheads: after locating the slot, all nodes in the slot need to be traversed and compared with the key to be searched. In the worst case, all nodes in the slot need to be traversed. The linked list data structure is implemented through a pointer, and during the traversal process, random memory access is generated, which reduces data locality and causes CPU stalls, which seriously reduces computing performance. In addition, when the load factor of the hash table exceeds a threshold, the resize operation needs to recalculate the hash value of all data and migrate the memory, which involves a large amount of data copying and destruction operations (used to automatically execute resource cleaning work when the object life cycle ends), further exacerbating the write performance bottleneck.

[0064] Therefore, in the large data volume scenario, the traditional hash table has the problem of low read-write performance. Improving the read-write performance of the hash table is crucial to the aggregation computing performance.

[0065] Based on the above, the embodiments of the present application provide a segmented aggregation hash table reading and writing method and system based on a graph database.

[0066] In this article, it should be understood that the terms involved can be technical means or other summary technical terms for implementing part of the invention, for example, the terms can include:

[0067] Database: basic software for managing data. The database provides data persistence and functions such as adding, deleting, modifying, and querying data. Users interact with the database through database languages (such as sql, gql) or API interfaces to complete data processing business logic.

[0068] Aggregate function: A type of function in database, which aggregates multiple values into one value, such as avg function calculates the average of multiple values, count function calculates the number of multiple values.

[0069] Aggregate operator: A type of operator in database, which produces the query result by grouping the input data and calculating the aggregate function. The implementation of aggregate operator has a great impact on the query performance.

[0070] Hash table: A type of key-value data structure, which stores data in the form of key-value pairs. When querying, the corresponding key value can be obtained in O(1) time complexity through the key. Each hash value in the hash table corresponds to a slot. If there are multiple key-value pairs in the slot, a hash collision will occur, which will cause a serious performance decline. In the worst case, the time complexity will decrease from O(1) to O(n). When the data volume in the hash table exceeds a certain threshold, frequent hash collisions will occur, which requires resizing (expanding to twice the size). Resizing requires re-computing the hash value of all data and re-arranging their positions in the hash table. The resizing operation involves a large amount of memory data movement, which has a great impact on write performance. The unordered map of C++ standard template library is a common implementation of hash table. In the database field, hash table is often used to implement aggregate or join operators. Different scenarios may have their own specific optimization schemes. The present application is suitable for aggregate computing scenarios.

[0071] Linear probing: A method for handling hash collisions, the main idea is to find the nearest free unit to the collision unit in the hash table and write the data into it.

[0072] Cache: CPU internal cache, which has faster access speed than memory but higher cost. Cache is used to solve the memory wall problem of von Neumann architecture computers by utilizing data locality. If cache is hit, CPU does not need to execute high-overhead memory access instructions and only needs to access cache, thereby improving performance. Conversely, if cache is not hit, not only is the cache not hit, but the CPU also stops (causing pipeline bubbles), resulting in performance degradation.

[0073] CPU stall: CPU stall generally occurs in the case of cache miss, after cache miss, CPU needs to execute memory access instruction, this process will consume hundreds of CPU clock cycles, and CPU does not perform any actual work in this process. On modern multi-core CPU, CPU stall will lead to pipeline efficiency decline, and the impact on performance is more serious.

[0074] In a first aspect, the embodiments of the present application provide a segmented aggregate hash table read-write method based on a graph database, Figure 1 is a flowchart of a segmented aggregate hash table read-write method based on a graph database, as Figure 1 shown, the method comprises:

[0075] Step S101, separate the original hash table into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregate data.

[0076] Step S102, in response to receiving a write data instruction, locate the hash entry in the hash entry table according to the hash value of the write data.

[0077] Step S103, based on the data storage state of the hash entry, store the write data in the data load block through linear probing, and update the hash entry.

[0078] Step S104, in response to receiving a read data instruction, obtain target aggregate data corresponding to the read data through linear probing according to the hash value of the read data.

[0079] In summary, the embodiment of the present application provides a segmented aggregation hash table read-write method based on a graph database. The original hash table is separated into a hash entry table and a data load block. The hash entry table is used to store hash entry data, and the data load block is used to store aggregated data. The hash entry table and the data load block are decoupled, the memory access locality is improved, and the frequent cache miss is reduced. In response to receiving a write data instruction, the hash entry in the hash entry table is located according to the hash value of the write data. The target entry is directly located through the hash value, the traversal process is avoided, and the calculation overhead is reduced. Based on the data storage state of the hash entry, the write data is stored in the data load block through linear probing, and the hash entry is updated. The conflict is solved in the continuous space of the hash entry table, and the probing process is accelerated. In response to receiving a read data instruction, the target aggregated data corresponding to the read data is obtained through linear probing according to the hash value of the read data. Invalid entries are quickly filtered through the hash prefix, the number of data load block access times is reduced, and the aggregation data read efficiency is improved. In the present application, the linear probing of the continuous memory layout reduces the cache invalidation when writing, the hash value is pre-stored to avoid repeated calculation, and when querying, the hash prefix is used for hierarchical filtering to skip most non-matching items. Only a small amount of data load block data needs to be accessed, the hash table read-write performance is improved, and the problem of low performance in the read-write process of the traditional hash table in a large data amount scenario is solved.

[0080] Figure 2 is a structure diagram of a hash entry table and a data load block according to an exemplary embodiment. As shown in Figure 2 The original hash table is separated into a hash entry table and a data load block. The hash entry table is used to store hash entry data, and the data load block is used to store aggregated data. Specifically, the steps include:

[0081] The hash entry table is a list structure in a continuous memory layout mode. The hash entry table includes multiple rows of hash entries. Each row of hash entries includes:

[0082] The hash value prefix is used to store the high-order bytes of the hash value.

[0083] The block number is used to locate the target data load table in the data load block.

[0084] The in-block address is used to locate the data load entry in the target data load table.

[0085] Optionally, the hash entry table stores the hash prefix (the high byte of the complete hash value), the block number and the intra-block address in a three-column structure, directly locates the aggregated data (including the grouping column and the aggregation column) in the data load block through the block number and the intra-block address, avoids pointer jumping of the linked list method, and pre-stores the hash value of the grouping column in the data load block header, so that only the address mapping relationship of the hash entry table needs to be updated during expansion, without migrating data content, thereby reducing the performance loss of data copying and deconstruction. In addition, the hash prefix matching accelerates conflict filtering, and the prefetching characteristics of continuous memory layout further reduce invalid traversal and CPU stall time in the query stage. Compared with the traditional hash table, the performance of the aggregation operator is improved by one order of magnitude, and is suitable for high-frequency aggregation computing scenarios such as databases and graph databases.

[0086] The data load block is a list structure in a continuous memory layout mode, and the data load block includes a plurality of data load tables with consistent structures. Each data load table includes a plurality of data load entries, and each data load entry includes:

[0087] a hash value, used for storing the hash value of the grouping column;

[0088] a grouping column, used for storing grouping column data;

[0089] an aggregation column, used for storing the calculation result of the aggregation function.

[0090] Optionally, the data load block stores a plurality of data load tables (DPT) in a list structure in a continuous memory layout mode. Each data load table includes data load entries (DE) containing the complete hash value of the grouping column, the grouping column data and the aggregation column result. The beneficial effects are as follows: the continuous memory layout accesses data through pointer offset, avoids the non-continuous memory jump of the traditional linked list method, and reduces cache miss and CPU stall overhead; the design of the data load table with consistent structure makes memory allocation and release more efficient, and the fixed-size data load table also supports expansion of the number without migrating data content during dynamic expansion, thereby reducing the expansion overhead; the pre-stored hash value of the grouping column can be directly used for quick verification in hash conflict and address mapping update during expansion, thereby avoiding the performance loss of repeated hash value calculation; the separate storage of the grouping column and the aggregation column enables the aggregation calculation result (such as avg and count) to be directly updated in the aggregation column without the need to recombine the original data, thereby improving the writing efficiency; meanwhile, the prefetching characteristics of the continuous memory accelerate data access in the query stage and reduce invalid traversal. The performance bottleneck problem caused by hash conflict, expansion overhead and cache miss under large data volume is solved.

[0091] In an embodiment, in step S102, in response to receiving the write data instruction, a hash entry in the hash entry table is located according to the hash value of the write data. Specifically, the step includes:

[0092] calculating a hash value of the write data;

[0093] locating a row number of the hash entry table for the write data by taking the hash value modulo the number of rows of the hash entry table;

[0094] locating a hash entry in the hash entry table based on the row number.

[0095] Optionally, the hash value is mapped to a physical address space of the hash table by taking the hash value modulo the number of rows of the hash entry table, for example: achieving an initial locating efficiency of O(1) to avoid the non-continuous memory jump of the traditional linked list method; based on the row number locating the hash entry, when the linear probing method is used to handle the conflict, the adjacent entries in the continuous memory are sequentially accessed, and the prefetch mechanism of the CPU cache is used to reduce the cache miss and CPU stall time. The existing linked list method needs to traverse the non-continuous memory nodes, and the continuous access mode of the modulo locating and linear probing of the present scheme makes the conflict handling efficiency improved by one order of magnitude, and the design of pre-storing the hash value further speeds up the comparison process, and finally reduces the overall performance loss of the aggregation calculation.

[0096] Figure 3 is a flowchart of step S103 according to an exemplary embodiment, as shown in Figure 3 Step S103, based on the data storage state of the hash entry, stores the write data in the data load block through linear probing and updates the hash entry. Specifically, the following steps are included:

[0097] Step S1031, judging whether the hash entry stores data;

[0098] Step S1032, if the hash entry does not store data, storing the write data in the data load block and updating the data of the hash entry, and if the hash entry stores data, storing the write data in the data load block through linear probing.

[0099] Optionally, judging whether the hash entry stores data can quickly identify the validity of the initial location, and if it is a free entry, the write data is directly stored in the data load block and the hash table entry is updated to avoid unnecessary conflict handling; if there is a conflict, the write data is sequentially stored in the data load block by linear probing method, and combined with the cache prefetch characteristics of the continuous memory layout, the cache miss and CPU stall time caused by random access are reduced. The hash entry data load block avoids the pointer jump overhead of the linked list method, and finally realizes the high performance and low delay of the write operation in the aggregation calculation scenario.

[0100] It should be noted that the hash table entry will only be written with data including the hash prefix and the automatically allocated block number and block address after expansion or positioning when the corresponding entry row is empty; the probe is not empty, and the hash table entry does not need to be updated, that is, the data does not need to be stored in the hash table entry.

[0101] Further, by linear probing, the data to be written is stored in the data load block, which includes:

[0102] determining whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the data to be written;

[0103] If not, according to the continuous memory access rule, it is determined whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the data to be written;

[0104] If matched, the corresponding data load entry is read according to the block number and the block address in the hash entry table, it is determined whether the hash value in the data load entry is the same as the hash value of the data to be written, and whether the group column key value is the same as the key value of the data to be written, and in the case of being the same, the aggregate column data corresponding to the group column key value of the data load block is updated. In the case of being different, it is continued to determine whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the data to be written. Until in the case of being the same, the aggregate column data corresponding to the group column key value of the data load block is updated.

[0105] Optionally, the hash prefix matching is used as a first layer filtering condition, which quickly excludes non-target entries (such as conflict entries irrelevant to the data to be written) by intercepting the high byte of the hash value, reduces unnecessary complete hash value comparison and group column key value verification, thereby reducing the calculation overhead of invalid operations; the continuous memory access rule uses the prefetching characteristics of the CPU cache to sequentially access adjacent hash table entries, avoids the non-continuous memory jump of the linked list method, and further verifies whether the complete hash value and the group column key value are completely consistent after matching the entries with the same hash prefix, ensures the uniqueness of the data, and avoids redundant storage and calculation by directly updating the aggregate column instead of recombining the original data.

[0106] In an embodiment, after updating the data of the hash entry, when the number of rows of the hash entry table reaches the expansion threshold, the following expansion operation is performed:

[0107] The capacity of the hash entry table and the data load block is synchronously expanded to N times of the initial size; wherein the physical storage location in the data load block remains unchanged, and the data in the hash entry table is updated by the hash value.

[0108] Optionally, when the hash table load factor exceeds the threshold, the expansion operation needs to recalculate the hash values of all data and migrate the memory, involving a large amount of data copying and destruction operations, further exacerbating the write performance bottleneck. Therefore, when the number of rows of the hash entry table reaches the expansion threshold, the capacity of the synchronous expansion hash table and the data load block is expanded to the initial size N times, which can be 1-2 times in the embodiment of the application. Avoid the high overhead problem of full data migration and repeated hash calculation when expanding the traditional hash table. Specifically, the physical storage location of the data load block remains unchanged, and only the capacity of the data load table is expanded (such as appending a fixed size of data load table) to realize expansion, without migrating the original aggregated data content, thereby avoiding the data copying and destruction performance loss when expanding the traditional linked list or array. At the same time, the hash entry updates the address mapping through the pre-stored hash value, and directly calculates the row number of the new hash table using the complete hash value of the grouping column, avoiding the recalculation of the hash value of all entries, and significantly reducing the calculation overhead when expanding.

[0109] In an embodiment, step S104, in response to receiving the read data instruction, according to the hash value of the read data, the target aggregated data corresponding to the read data is obtained through linear probing. Specifically, it includes:

[0110] Calculate the hash value of the read data and linearly probe in the hash entry;

[0111] Determine whether the prefix of the hash value of the read data matches the hash prefix in the hash entry table;

[0112] If not, according to the continuous memory access rule, determine whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the read data, if it matches, read the corresponding data load entry according to the block number and block address in the hash entry table, determine whether the hash value in the data load entry is the same as the hash value of the read data, if they are the same, determine whether the grouping column key value is the same as the key value of the read data, in the case of the same, obtain the target aggregated data corresponding to the read data. In the case of not the same, continue to linearly probe in the hash entry until the case of the same, obtain the target aggregated data corresponding to the read data.

[0113] Optionally, when reading data, the hash prefix matching is used as the first layer filtering condition to quickly exclude non-target entries (such as conflict entries irrelevant to the read data), reduce unnecessary complete hash value comparison and group column key value verification, and thus reduce the calculation overhead of invalid operations; the linear probing combined with the continuous memory access rule sequentially accesses adjacent hash table entries, utilizes the prefetching characteristics of the CPU cache, reduces the cache miss and CPU stall time caused by random access. When the entry with the consistent hash prefix is matched, the complete hash value and the group column key value are further verified to ensure the data uniqueness, and the aggregated data is directly located by the block number and the intra-block address to avoid the pointer jump overhead of the linked list method.

[0114] In an embodiment, after obtaining the target aggregated data corresponding to the read data, the method further comprises:

[0115] According to the data type, the segmented aggregated hash table is optimized, wherein according to the fixed-length data type, the group column data of the data load entry is stored at the hash prefix position of the hash entry, and according to the variable-length string type, the group column is stored by the dictionary tree.

[0116] Optionally, for different types of aggregated data, different memory schemes can be designed for different data types due to the different memory spaces and data type characteristics. Fixed-length data type (integer optimization): taking int16 as an example, the group column only occupies 2 bytes, in order to avoid the non-continuous memory access from the hash entry table to the data load block, the int16 type group column data can be moved from the data load entry to the hash prefix position of the hash entry, and the key equal determination can be directly performed in the hash entry table during the linear probing. String type optimization: the string type belongs to the variable-length type, and the memory space occupied by the string type needs to be determined according to the string length, but the string data type characteristics are that all data is composed of enumerable characters, and the memory space can be saved by constructing the dictionary tree.

[0117] And / or, the segmented aggregated hash table is optimized by the parallel computing strategy, the input data is partitioned according to the hash range, each partition is processed by an independent thread, each independent thread processes the hash entry table and the data load block, and the local hash table is constructed by address reference, and the target aggregated data corresponding to the read data is obtained based on the local hash table.

[0118] Optionally, Figure 4 is a structural schematic diagram of the parallel segmented aggregated hash table according to an example embodiment, as Figure 4 shown, the input data received by the aggregation operator is first partitioned, different partition data is sent to different independent threads for processing, and each independent thread processes the hash entry table and the data load block to provide the user with a local hash table (such asFigure 4 The global hash table view is obtained based on the local hash table, and target aggregated data corresponding to the read data is obtained based on the global hash table view. The global hash table view does not copy actual data but only references the actual data by an address.

[0119] In summary, the embodiment of the present application provides a segmented aggregated hash table read-write method based on a graph database. The original hash table is separated into a hash entry table and a data load block. The hash entry table is used to store hash entry data, and the data load block is used to store aggregated data. The hash entry table and the data load block are decoupled, the memory access locality is improved, and the cache miss is reduced. In response to receiving a write data instruction, the hash entry in the hash entry table is located according to the hash value of the write data. The target entry is directly located by the hash value, the traversal process is avoided, and the calculation overhead is reduced. Based on the data storage state of the hash entry, the write data is stored in the data load block by linear probing, and the hash entry is updated. The conflict is solved in the continuous space of the hash entry table, and the probing process is accelerated. In response to receiving a read data instruction, the target aggregated data corresponding to the read data is obtained by linear probing according to the hash value of the read data. Invalid entries are quickly filtered by the hash prefix, the number of data load block access times is reduced, and the aggregated data read efficiency is improved. In the write process, the linear probing of the continuous memory layout reduces the cache invalidation, the hash value is pre-stored to avoid repeated calculation, in the query process, most of the non-matching items are skipped by hierarchical filtering through the hash prefix, only a small amount of data load block data needs to be accessed, the hash table read-write performance is improved, and the problem of low performance in the read-write process of the traditional hash table in the large data amount scenario is solved.

[0120] In a second aspect, the embodiment of the present application provides a segmented aggregated hash table read-write system based on a graph database. Figure 5 is a system structure block diagram of the segmented aggregated hash table read-write based on a graph database according to an exemplary embodiment. As shown in Figure 5 The system includes a separation module 510, a positioning module 520, a write module 530, and a read module 540; wherein:

[0121] The separation module 510 is configured to separate the original hash table into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregated data.

[0122] The positioning module 520 is configured to, in response to receiving a write data instruction, locate a hash entry in the hash entry table according to a hash value of the write data.

[0123] The write module 530 is configured to, based on a data storage state of the hash entry, store the write data in the data load block by linear probing, and update the hash entry.

[0124] The reading module 540 is configured to, in response to receiving a reading data instruction, acquire target aggregated data corresponding to the reading data according to a hash value of the reading data through linear probing.

[0125] To sum up, the application provides a segmented aggregated hash table reading and writing system based on a graph database, which comprises a separation module 510, a positioning module 520, a writing module 530 and a reading module 540. In writing, the linear probing of the continuous memory layout reduces cache invalidation, and the pre-stored hash value avoids repeated calculation. In query, the hash prefix hierarchical filtering is used to skip most non-matching items, and only a small amount of data load block data needs to be accessed, thereby improving the hash table reading and writing performance and solving the problem of low performance in the traditional hash table reading and writing process in a large data amount scenario.

[0126] It should be noted that the segmented aggregated hash table reading and writing system based on a graph database provided in the embodiment is used to implement the above-mentioned embodiments, and the description of which has been made. As used above, the terms "module", "unit", "sub-unit" and the like can be a combination of software and / or hardware that can implement a predetermined function. Although the above embodiment describes that the device is preferably implemented in software, the implementation of hardware or a combination of software and hardware is also possible and is conceived.

[0127] In a third aspect, the embodiments of the application provide an electronic device, Figure 6 is a block diagram of an electronic device according to an exemplary embodiment. As Figure 6 shown, the electronic device can include a processor 61 and a memory 62 having stored computer program instructions.

[0128] Specifically, the processor 61 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the application.

[0129] The memory 62 can include a mass storage for data or instructions. By way of example, and without limitation, the memory 62 can include a Hard Disk Drive (HDD), a floppy disk drive, a Solid State Drive (SSD), a flash drive, a Compact Disc Read Only Memory (CD-ROM), a Digital Versatile Disk (DVD), a Blu-Ray, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. The memory 62 can be removable and / or non-removable (or fixed) as appropriate. The memory 62 can be internal or external as appropriate. In certain embodiments, the memory 62 is a Non-Volatile Memory. In certain embodiments, the memory 62 includes a Read-Only Memory (ROM) and a Random-Access Memory (RAM). The ROM can be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), an Electrically Alterable ROM (EAROM), or a FLASH memory, or a combination of two or more of these, as appropriate. The RAM can be a Static Random-Access Memory (SRAM) or a Dynamic Random-Access Memory (DRAM), which can be a Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), an Extended Data Output Dynamic Random-Access Memory (EDODRAM), a Synchronous Dynamic Random-Access Memory (SDRAM), or the like, as appropriate.

[0130] The memory 62 can be used to store or buffer various data files required for processing and / or communication, and possible computer program instructions executed by the processor 61.

[0131] The processor 61 reads and executes the computer program instructions stored in the memory 62 to implement any of the above-mentioned embodiments of the graph database-based segmented aggregated hash table read-write method.

[0132] In an embodiment, the graph database-based segmented aggregated hash table read-write device can further include a communication interface 63 and a bus 60. As shown in the figure, the processor 61, the memory 62, and the communication interface 63 are connected through the bus 60 and complete communication with each other. Figure 6

[0133] The communication interface 63 is used to realize communication between various modules, devices, units, and / or equipment in the embodiments of the present application. The communication interface 63 can also realize data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.

[0134] ​Bus 60 includes hardware, software, or both, to couple components of the graph database based segmented aggregate hash table read-write device to each other and to couple components to other systems or devices. While bus 60 is shown for the sake of clarity as a single bus, bus 60 can include a combination of buses, including, for example, an Industry Standard Architecture (ISA), Video Electronics Standards Association (VESA), InfiniBand, Advanced Graphics Port (AGP), Personal Computer Memory Card International Association (PCMCIA), Small Computer Systems Interface (SCSI), PCI, PCI 2.2, Symbian, Bluetooth, and / or any other bus or interconnect, such as a bus to connect to a system across a wireless personal- area network. Bus 60 can be implemented using any suitable type, unidirectional and / or bidirectional, of connection including, for example, optical fiber, wire (including, for example, coaxial cable and the like), wireless media (including, for example, acoustic, RF, infrared, and / or microwave), and / or other connections.

[0135] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, having a program stored thereon, where the program, when executed by a processor, implements the graph database based segmented aggregate hash table read-write method provided in the first aspect.

[0136] More specifically, the computer readable storage medium can include, but is not limited to, portable discs, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0137] In possible implementation manners, the present application can also be implemented in the form of a program product, which comprises program codes for causing terminal equipment to execute steps of implementing the graph database-based segmented aggregation hash table reading and writing method provided by the first aspect when the program product is run on the terminal equipment.

[0138] Wherein, the program codes for executing the present application can be written in any combination of one or more programming languages, which can be executed entirely on the user equipment, partially on the user equipment, as an independent software package, partially on the user equipment and partially on a remote device, or entirely on a remote device.

[0139] The technical features of the above-described embodiments can be combined in any manner. In order to make the description concise, all possible combinations of the technical features in the above-described embodiments are not described, however, as long as the combinations of the technical features do not exist in contradiction, they should be considered as the scope of the present disclosure.

[0140] The above-described embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but should not be understood as a limitation on the scope of the patent. It should be pointed out that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A segmented aggregate hash table reading and writing method based on a graph database, characterized in that: The method comprises: Separate the original hash table into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregate data; the hash entry data includes a hash value prefix for storing the high-order byte of the hash value; a block number for locating a target data load table in the data load block; an intra-block address for locating a data load entry in the target data load table; the aggregate data includes a hash value for storing the hash value of a grouping column; a grouping column for storing grouping column data; and an aggregate column for storing the calculation result of an aggregate function; In response to receiving a write data instruction, locating a hash entry in the hash entry table according to a hash value of the write data; Based on the data storage state of the hash entry, storing the write data in the data payload block by linear probing, and updating the hash entry; In response to receiving a read data instruction, target aggregate data corresponding to the read data is acquired through linear probing according to a hash value of the read data.

2. A segmented aggregate hash table reading and writing method based on a graph database according to claim 1, characterized in that: The method of storing the write data in the data load block by linear probing based on the data storage state of the hash entry and updating the data of the hash entry comprises the following steps: Determining whether the hash entry stores data; If the hash entry does not store data, the write data is stored in the data payload block and the data of the hash entry is updated; if the hash entry stores data, the write data is stored in the data payload block through linear probing.

3. A segmented aggregate hash table reading and writing method based on a graph database according to claim 2, characterized in that: The step of storing the write data in the data load block by linear probing includes: Determine whether a hash prefix in the hash entry table of the next row matches a prefix of the hash value of the written data; If not, determining whether the hash prefix in the next row of the hash entry table matches the prefix of the hash value of the written data according to the continuous memory access rule; If there is a match, the corresponding data load entry is read according to the block number and the address within the block in the hash entry table, and it is determined whether the hash value in the data load entry is the same as the hash value of the written data, and whether the grouping column key value in the data load entry is the same as the key value of the written data. If they are the same, the aggregate column data corresponding to the grouping column key value of the data load block is updated.

4. A segmented aggregate hash table reading and writing method based on a graph database according to claim 2, characterized in that: After updating the data of the hash entry, when the number of rows in the hash entry table reaches an expansion threshold, the method further includes: The capacity of the hash entry table and the data payload block is synchronously expanded to N times of the initial size; wherein the physical storage location in the data payload block remains unchanged, and the data in the hash entry table updates the hash entry through the hash value.

5. The segmented aggregate hash table reading and writing method based on a graph database according to claim 1 is characterized in that: The method of acquiring target aggregate data corresponding to the read data by linear probing according to a hash value of the read data in response to receiving the read data instruction includes: Calculate the hash value of the read data and perform linear detection in the hash entry; Determining whether a prefix of the hash value of the read data matches a hash prefix in the hash entry table; If there is no match, then according to the continuous memory access rule, determine whether the hash prefix in the hash entry table of the next row matches the prefix of the hash value of the read data; if they match, read the corresponding data load entry according to the block number and the address within the block in the hash entry table, and determine whether the hash value in the data load entry is the same as the hash value of the read data; if they are the same, determine whether the grouping column key value in the data load entry is the same as the key value of the read data; if they are the same, obtain the target aggregate data corresponding to the read data.

6. A segmented aggregate hash table reading and writing method based on a graph database according to claim 1, characterized in that: The hash entry table is a list structure in a continuous memory layout manner, wherein the hash entry table includes multiple rows of hash entries, and each row of hash entries includes: Hash value prefix, used to store the high-order byte of the hash value; A block number, used to locate a target data load table in the data load block; An intra-block address, used to locate a data payload entry in the target data payload table.

7. A segmented aggregate hash table reading and writing method based on a graph database according to claim 1, characterized in that: The data load block is a list structure in a continuous memory layout mode. The data load block includes multiple data load tables with consistent structures. The data load tables include multiple rows of data load entries. Each row of the data load entries includes: Hash value, used to store the hash value of the grouping column, the hash value of the grouping column is calculated when writing data and pre-stored in the data payload entry; Grouping column, used to store grouping column data; Aggregate columns are used to store the calculation results of aggregate functions.

8. The segmented aggregate hash table reading and writing method based on a graph database according to claim 1 is characterized in that: The locating a hash entry in the hash entry table according to the hash value of the written data includes: Calculating a hash value of the written data; Locating the row number of the write data in the hash entry table by taking the modulo of the hash value and the number of rows in the hash entry table; Based on the row number, a hash entry in the hash entry table is located.

9. A segmented aggregate hash table reading and writing method based on a graph database according to any one of claims 1 to 8, characterized in that: After acquiring target aggregate data corresponding to the read data, the method further includes: Optimizing a segmented aggregate hash table based on data types, wherein, based on fixed-length data types, grouping column data of a data payload entry is stored at a hash prefix position of the hash entry, and based on variable-length string types, grouping columns are stored via a dictionary tree; And / or, optimizing the segmented aggregate hash table through a parallel computing strategy, partitioning the input data according to the hash range, processing each partition through an independent thread, each independent thread processing the hash entry table and the data load block, constructing a local hash table through address reference, and obtaining target aggregate data corresponding to the read data based on the local hash table.

10. A segmented aggregate hash table reading and writing system based on a graph database, characterized in that: The system includes a separation module, a positioning module, a writing module and a reading module; wherein: The separation module is used to separate the original hash table into a hash entry table and a data load block, wherein the hash entry table is used to store hash entry data, and the data load block is used to store aggregate data; the hash entry data includes a hash value prefix for storing the high-order byte of the hash value; a block number for locating a target data load table in the data load block; an intra-block address for locating a data load entry in the target data load table; the aggregate data includes a hash value for storing the hash value of a grouping column; a grouping column for storing grouping column data; and an aggregate column for storing a calculation result of an aggregate function; The positioning module is configured to, in response to receiving a write data instruction, locate a hash entry in the hash entry table according to a hash value of the write data; The writing module is configured to store the write data in the data payload block through linear probing based on the data storage state of the hash entry, and update the hash entry; The reading module is configured to, in response to receiving a data reading instruction, obtain target aggregated data corresponding to the read data through linear probing according to a hash value of the read data.

Citation Information

Patent Citations

  • Methods of organizing data and processing queries in a database system, and database system and software product for implementing such methods

    AU2008202360A1

  • Dynamic hash table operation method based on hybrid DRAM-NVM memory

    CN111459846A