Database processing system and method for offloading database operations
By designing a database offload engine, using hardware components such as vectorized adder and sum buffer to offload table scans and sum aggregation operations in the database, the high load problem of these operations on the host CPU is solved, and more efficient database processing and power consumption are achieved.
Patent Information
- Application Number
- CN201910770236.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-19
- Filing Date
- 2019-08-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-08-20
AI Technical Summary
When the table scan operation and sum aggregation operation are executed in the database, it will significantly consume the cycle and power of the host CPU, resulting in performance bottlenecks.
A database offload engine is designed, including vectorized adder, sum buffer, key address table and control circuit. By unloading these operations to special hardware, it reduces dependence on the host CPU.
Through the uninstall table scan and sum-aggregation operations, the overall processing speed of the database processing system is significantly improved and power consumption is reduced.
Smart Images

Figure CN110941600B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 735,688, filed on September 24, 2018, entitled “HIGHLY SCALABLE DATABASE OFFLOADING ENGINE FOR (K,V) AGGREGATION AND TABLE SCAN,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] One or more aspects of embodiments according to the present disclosure relate to database processing, and more particularly to a database offload engine. Background Art
[0004] Table scan operations and sum aggregation operations - when performed by the host CPU as part of query processing operations in a database - can significantly tax the CPU, consuming a significant portion of the CPU cycles and accounting for a significant portion of the power consumed by the CPU.
[0005] Therefore, there is a need for an improved system and method for performing table scan operations and sum aggregation operations in a database. Summary of the invention
[0006] According to an embodiment of the present invention, there is provided a database processing system including a database offload engine, wherein the database offload engine includes: a vectorized adder including a plurality of read-modify-write circuits; a plurality of sum buffers respectively connected to the read-modify-write circuits; a key address table; and a control circuit, wherein the control circuit is configured to: receive a first key and a corresponding first value; search the key address table for the first key; and in response to finding an address corresponding to the first key in the key address table, route the address and the first value to a read-modify-write circuit corresponding to the address among the plurality of read-modify-write circuits.
[0007] In some embodiments, the control circuit is further configured to: receive a second key and a corresponding second value; search the key address table for the second key; and in response to not finding an address corresponding to the second key in the key address table: select a new address, which does not exist in the key address table; store the second key and the new address in the key address table; and route the new address and the second value to a read-modify-write circuit corresponding to the new address among the multiple read-modify-write circuits.
[0008] In some embodiments, the database offload engine has a Non-Volatile Dual In-line Memory Module-P (NVDIMM-P) interface for establishing a connection with a host.
[0009] In some embodiments, the database offload engine has a Peripheral Component Interconnect express (PCIe) interface for establishing a connection with a host.
[0010] In some embodiments: the vectorized adder is a synchronous circuit within a clock domain, the clock domain is defined by a shared system clock, one of the multiple read-modify-write circuits is configured as a pipeline, the pipeline includes: a first stage for performing a read operation, a second stage for performing an addition operation, and a third stage for performing a write operation, and the pipeline is configured to receive an address and a corresponding value in each cycle of the shared system clock.
[0011] In some embodiments: the control circuit is a synchronous circuit within a clock domain, the clock domain is defined by a shared system clock, the control circuit includes a search circuit for searching the key address table for a key, the search circuit is configured to include a pipeline of multiple stages for searching the key address table, and the pipeline is configured to receive the key in each cycle of the shared system clock.
[0012] In some embodiments, the database processing system also includes a host connected to the database offload engine, the host includes a non-temporary storage medium storing database application instructions and driver layer instructions, the database application instructions include function calls, when the function calls are executed by the host, the function calls cause the host to execute driver layer instructions, and the driver layer instructions cause the host to control the database offload engine to perform sum aggregation operations.
[0013] In some embodiments, the database unloading engine also includes multiple table scanning circuits; one table scanning circuit among the multiple table scanning circuits includes a conditional test circuit that can be programmed with conditions, an input buffer and an output buffer, and the conditional test circuit is configured to: determine whether the first entry at the first address in the input buffer satisfies the condition, and in response to determining that the first entry satisfies the condition, write the corresponding result into the output buffer.
[0014] In some embodiments, the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write a corresponding element of the output vector into the output buffer.
[0015] In some embodiments, the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write the first address to a corresponding element of an output vector in the output buffer.
[0016] In some embodiments: the vectorized adder is a synchronous circuit within a clock domain, the clock domain is defined by a shared system clock, one of the multiple read-modify-write circuits is configured as a pipeline, the pipeline includes: a first stage for performing a read operation, a second stage for performing an addition operation, and a third stage for performing a write operation, and the pipeline is configured to receive an address and a corresponding value in each cycle of the system clock.
[0017] In some embodiments: the control circuit is a synchronous circuit within a clock domain, the clock domain is defined by a shared system clock, the control circuit includes a search circuit for searching the key address table for a key, the search circuit is configured to include a pipeline of multiple stages for searching the key address table, and the pipeline is configured to receive a key with each cycle of the system clock.
[0018] In some embodiments, the database offload engine has an NVDIMM-P interface for establishing a connection with a host.
[0019] According to an embodiment of the present invention, there is provided a database processing system including a database unloading engine, the database unloading engine including: a plurality of table scanning circuits; one table scanning circuit among the plurality of table scanning circuits includes a conditional test circuit, an input buffer and an output buffer which can be programmed with conditions, the conditional test circuit being configured to: determine whether a first entry at a first address in the input buffer satisfies the condition, and in response to determining that the first entry satisfies the condition, write a corresponding result into the output buffer.
[0020] In some embodiments, the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write a corresponding element of the output vector into the output buffer.
[0021] In some embodiments, the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write the first address to a corresponding element of an output vector in the output buffer.
[0022] In some embodiments, the database offload engine has an NVDIMM-P interface for establishing a connection with a host.
[0023] In some embodiments, the database offload engine has a PCIe interface for establishing a connection with a host.
[0024] According to an embodiment of the present invention, a method for offloading database operations from a host is provided, the method comprising: calling a driver function by an application running on the host to perform a sum aggregation operation; and performing the sum aggregation operation by a database offload engine, the database offload engine comprising: a vectorized adder, comprising a plurality of read-modify-write circuits; a plurality of sum buffers, respectively connected to the plurality of read-modify-write circuits; a key address table; and a control circuit, wherein performing the sum aggregation operation comprises: receiving a first key and a corresponding first value; searching the key address table for the first key; in response to finding an address corresponding to the first key in the key address table, routing the address and the first value to a read-modify-write circuit corresponding to the address in the plurality of read-modify-write circuits; receiving a second key and a corresponding second value; searching the key address table for the second key; in response to not finding an address corresponding to the second key in the key address table: selecting a new address that does not exist in the key address table; storing the key and the new address in the key address table; and routing the new address and the second value to a read-modify-write circuit corresponding to the new address in the plurality of read-modify-write circuits.
[0025] In some embodiments, the method also includes: the application calling a driver function to perform a table scan operation; and the database offload engine performing the table scan operation, wherein the performing of the table scan operation includes: a conditional testing circuit of the database offload engine determines whether a first entry at a first address in an input buffer of the database offload engine satisfies a condition, and in response to determining that the first entry in the input buffer satisfies the condition, writing a corresponding result into an output buffer of the database offload engine. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] These and other features and advantages of the present disclosure will be appreciated and understood with reference to the specification, claims and drawings, in which:
[0027] Figure 1 is a block diagram of a database processing system according to an embodiment of the present disclosure.
[0028] Figure 2 is a process flow diagram of a table scan operation according to an embodiment of the present disclosure.
[0029] Figure 3 is a process flow chart of the sequence of database processing operations according to an embodiment of the present disclosure.
[0030] Figure 4 is a block diagram of a database processing system according to an embodiment of the present disclosure.
[0031] Figure 5 is a block diagram of a database offload engine according to an embodiment of the present disclosure.
[0032] Figure 6 is a block diagram of a table scanning circuit according to an embodiment of the present disclosure.
[0033] Figure 7 is a block diagram of a sum aggregation circuit according to an embodiment of the present disclosure.
[0034] Figure 8 is a hardware-software block diagram of a database processing system according to an embodiment of the present disclosure.
[0035] Fig.9A is a process flow chart of a database processing operation according to an embodiment of the present disclosure.
[0036] Fig. 9B is a process flow chart of a database processing operation according to an embodiment of the present disclosure.
[0037] [Explanation of Symbols]
[0038] 105: host;
[0039] 110: Database uninstall engine;
[0040] 305: First step;
[0041] 310: Second step;
[0042] 315: The third step;
[0043] 320: Step 4;
[0044] 325: Subsequent steps;
[0045] 405: Database application;
[0046] 410: driver layer;
[0047] 415: Peripheral component interconnect fast interface;
[0048] 420: first parallelizer;
[0049] 425: Database unloading circuit / unloading circuit;
[0050] 430: second parallelizer;
[0051] 435: memory interface circuit;
[0052] 440: memory;
[0053] 505: a group of control registers;
[0054] 510: vector table scanning circuit;
[0055] 515: table scanning circuit;
[0056] 520: sum aggregation circuit;
[0057] 525: pre-fetch circuit;
[0058] 530: memory reading circuit;
[0059] 535: memory writing circuit;
[0060] 605: conditional test circuit;
[0061] 610: input buffer;
[0062] 615: output buffer;
[0063] 620: register;
[0064] 705: control circuit;
[0065] 710: vectorized adder;
[0066] 715: pre-fetch buffer;
[0067] 720: read-modify-write circuit;
[0068] 725: sum buffer;
[0069] 730: Find the circuit;
[0070] 735: key address table buffer;
[0071] 740: key address table / address table;
[0072] 745: first 8×8 crossbar switch;
[0073] 750: second 8×8 crossbar switch;
[0074] 805: Database application;
[0075] 810: operating system;
[0076] 815: host CPU;
[0077] 820: controller;
[0078] 825: Host main memory. DETAILED DESCRIPTION
[0079] The detailed description described below in conjunction with the accompanying drawings is intended as an illustration of an exemplary embodiment of the database unloading engine provided according to the present disclosure, and is not intended to represent the only form in which the present invention can be constructed or utilized. The description sets forth the features of the present invention in conjunction with the illustrated embodiments. However, it should be understood that the same or equivalent functions and structures may be implemented by different embodiments that are also intended to be encompassed within the scope of the present invention. As shown elsewhere herein, the same element numbers are intended to indicate the same elements or features.
[0080] Reference Figure 1 In some embodiments, the database processing system includes a host 105 and a database offload engine 110; the database offload engine 110 can be part of the host, or the database offload engine 110 can be connected to the host, as shown. For example, the host 105 can be a computer or server, including a central processing unit (CPU), main memory, and persistent storage (e.g., a hard disk drive or solid state drive). The database offload engine 110 can include or be connected to persistent storage. Persistent storage is a type of memory that balances speed, capacity, and persistence. One of the advantages of offloading aggregation operations and table scan operations closer to the data is that these operations are data intensive rather than compute intensive. The database offload engine 110 may include processing circuitry (discussed in further detail below) and memory. The database offload engine 110 can be connected to the host via any of a number of interfaces, including NVDIMM-p (via a memory channel) and PCIe. The host can perform various database processing operations, including executing queries.
[0081] An example of a database processing operation is a table scan operation. Figure 2 A table scan operation may include searching an entire column of a table of a database for entries that satisfy a condition. The condition may be part of a query being executed by the host or may be derived from the query. Figure 2In the example of , the table scan to be performed needs to identify each entry in the column corresponding to "New York", where the name "New York" corresponds to the integer 772 according to a dictionary (also stored in the database). The column to be searched can be stored in a compressed form, with each entry having 17 bits. To perform the table scan, the database processing system decompresses the column to be searched (converting each 17-bit number to a 32-bit number) and tests for each entry whether it satisfies a condition (in this case, whether the decompressed integer is equal to 772). The results of the scan can be represented in one of two formats, either (i) as a vector with the same number of elements as the column to be searched (in Figure 2 In some embodiments, the vector may be a vector of 257 elements in the example of 257 elements), containing one for each entry that satisfies the condition and zero for each entry that does not satisfy the condition, or (ii) a vector of indices (or addresses) in the column for each element in the column that satisfies the condition. In some embodiments, the vector instead contains the address of each element that does not satisfy the condition.
[0082] Another example of a database processing operation is a sum aggregation operation. This operation can be performed on a key value table, which can be a two-column table, one of which includes a set of keys (where, for example, each key is a 4-byte number) and the other includes a set of corresponding values (where, for example, each value is an 8-byte number). Some keys can be repeated (in any order); for example, key 23876 can appear 27 times in the table, with up to 27 different corresponding values. The sum aggregation operation generates a second key value table from the first key value table, in which each key appears exactly once (for example, in ascending order of arrangement), and the value corresponding to each key is the sum of all values corresponding to the key in the first key value table.
[0083] Table scan operations and sum aggregation operations and other database processing operations (e.g., GroupBy operations) can be performed by the host together. It should be understood that the example of the GroupBy operation is only an example, and in general the host can perform any operation on the data. When the host performs table scan operations and sum aggregation operations, the table scan operations and sum aggregation operations may consume a considerable portion of the host's processing cycles; therefore, if these operations are instead performed by (i.e., offloaded from the host to) a database offload engine, the overall processing speed of the database processing system can be significantly improved. In addition, for example, if the database offload engine uses dedicated hardware designed for, for example, table scan operations and sum aggregation operations, power consumption can be reduced, wherein the dedicated hardware may require less energy than the general-purpose hardware of the host CPU to perform a given operation.
[0084] Figure 3A process flow illustrating such offloading for performance improvement is shown. In a first step 305, the host CPU generates a query plan that includes a table scan operation and a sum aggregation operation. In a second step 310, the database offload engine (or "accelerator") performs the table scan operation. In a third step 315, the host CPU performs additional database operations including a GroupBy operation. In a fourth step 320, the database offload engine performs a sum aggregation operation. In a subsequent step 325 and possibly in other subsequent steps, the host CPU then uses the results of the table scan operation and / or the sum aggregation operation to perform additional database operations.
[0085] Figure 4 A block diagram of a database offload engine and two software layers interacting therewith according to one embodiment is shown. A database application 405 (e.g., a SAP High-Performance ANAlytic Appliance (HANA) database application) runs on a host and uses appropriate calls to a driver layer 410 to offload database operations (e.g., table scan operations and sum aggregation operations) to the database offload engine 110. The database offload engine 110 includes a PCIe interface 415 for communicating with the host, a CPU 416 for routing offloaded database operations to a plurality of (in Figure 4 4, a first parallelizer 420 for any one of two database offload circuits 425, and a second parallelizer 430 for establishing connections between the offload circuit 425 and a plurality of memory interface circuits 435 (e.g., mach interface generators (mig)). The database offload engine 110 may be connected to a plurality of memories 440 as shown, or, in some embodiments, the memory 440 may be part of the database offload engine 110. In some embodiments, the database offload engine 110 includes more than two offload circuits 425, for example, the database offload engine 110 may include 8 or more, 16 or more, or 32 or more such circuits.
[0086] Reference Figure 5, the offload circuit 425 may include a set of control registers 505, which the host may use to control the operation of the offload circuit 425. The offload circuit 425 may also include a vectorized table scan circuit 510, which includes a plurality of table scan circuits (vsearch) 515, a sum aggregation circuit (e.g., a Translation Lookaside Buffer (TLB), a vectorized adder (vadder), a scan) 520, and a prefetch circuit 525. The table scan circuit 515 and the sum aggregation circuit 520 may perform table scan operations and sum aggregation operations, respectively, as discussed in further detail below, and while database processing operations are being performed in the table scan circuit 515 and the sum aggregation circuit 520, the prefetch circuit 525 may fetch data from the memory 440 and store the fetched data in corresponding buffers in the table scan circuit 515 and the sum aggregation circuit 520. The vectorized table scan circuit 510 may receive compressed data and decompress it before further processing it. Prefetch circuitry 525 may fetch data from memory 440 using memory read circuitry 530 , and may write the results of the table scan operation and the sum aggregation operation to memory 440 using memory write circuitry 535 .
[0087] Reference Figure 6, each of the table scan circuits 515 may include a condition test circuit 605, an input buffer 610, an output buffer 615, and a set of registers 620. The registers 620 may include a pointer into the input buffer 610, a pointer into the output buffer 615, and one or more registers that specify the condition to be tested. The registers that specify the condition may include one or more value registers that specify a reference value, and one or more relationship registers that specify a relationship. For example, if the first relationship register contains a value corresponding to the relationship "equal to" and the corresponding value register contains the value 37, then the condition test circuit 605 may generate a one when the current input buffer value (the value at the address identified by the pointer into the input buffer 610 in the input buffer 610) is equal to 37. If the result of the scan being performed is to be formatted as a vector containing a one for each entry that satisfies the condition and a zero for each entry that does not satisfy the condition, then the condition test circuit 605 may then write the current result (e.g., a one) into the output buffer 615 at the address identified by the pointer into the output buffer 615 in the output buffer 615. Each time after the conditional test circuit 605 performs a test, both the pointer into the input buffer 610 and the pointer into the output buffer 615 may be incremented. In contrast, if the result of the scan being performed is to be formatted as a vector of the index (or address) within the column for each element in the column that satisfies the condition, then the conditional test circuit 605 may write to the output buffer 615 only when the current result is one (i.e., when the condition is satisfied), and in those cases the conditional test circuit 605 may write to the output buffer 615 the index (or address) of the current entry (the entry being tested) in the column being scanned. Each time after the conditional test circuit 605 performs a test, the pointer into the input buffer 610 may be incremented, and each time after the conditional test circuit 605 writes to the output buffer 615, the pointer into the output buffer 615 may be incremented. Some other possible relationships include "greater than" and "less than". Some or all of the registers 620 may be linked to or within the set of control registers 505 and may be set (i.e., written) by the host.
[0088] As described above, the vectorized table scan circuit 510 may include multiple table scan circuits 515. For example, if a table is to be scanned several times (for corresponding multiple conditions), if several tables are to be scanned, or if the scanning of the table is to be accelerated by dividing the table into multiple parts and having each of the table scan circuits 515 perform a table scan operation on the corresponding part, then these table scan circuits 515 may be employed to perform multiple table scan operations in parallel. The table scan circuits 515 may be pipelined so that one test is performed for each clock cycle, and they may be vectorized so that comparisons are performed in parallel.
[0089] Reference Figure 7 In some embodiments, the sum aggregation circuit 520 includes a control circuit 705, a vectorized adder 710, and a pre-fetch buffer 715. The sum aggregation circuit 520 can be a synchronous circuit in a single clock domain (with a single clock, referred to as the "system clock" of the circuit). The vectorized adder 710 includes a plurality of read-modify-write circuits 720, each of which can be connected to a corresponding sum buffer 725. In operation, the pre-fetch circuit 525 copies the key value table (or a portion of such a table) into the pre-fetch buffer 715. Each key-value pair in the pre-fetch buffer 715 can be converted into an address-value pair by an address conversion process described in further detail below, and the address-value pair can be sent to a corresponding one of the read-modify-write circuits 720. The read-modify-write circuit 720 then takes the current value sum from the address (of the address value pair) in the sum buffer 725, updates the value sum by adding the value from the address value pair to it, and saves the updated value sum back into the sum buffer 725 at the address of the address value pair (overwriting the value sum previously stored at this address). Each of the read-modify-write circuits 720 may be pipelined so that, for example, the first stage in the pipeline performs a read operation, the second stage in the pipeline performs an addition, and the third stage in the pipeline performs a write operation, and so that the read-modify-write circuit 720 may be able to receive a new address value pair during each cycle of the system clock. Once the entire key value table has been processed, a new key value table containing the sums may be formed by associating the key corresponding to the address where the sums are stored with each sum stored in the sum buffer 725.
[0090] The above-described address translation process may be advantageous because the key space corresponding to all possible 4-byte numbers (if each key is a 4-byte number) may be quite large, but any key value table may include only a small subset of possible 4-byte numbers. Therefore, the control circuit 705 may perform address translation to convert each key into a corresponding address, which forms a set of continuous addresses. The control circuit 705 may include a plurality of lookup circuits 730 and a plurality of key address table buffers 735 that together form a key address table 740. In operation, each lookup circuit 730 may receive a key value pair one at a time and (i) if an address has already been assigned, then look up the address of the key by searching the key address table buffer 735, or (ii) if an address has not yet been assigned to the key, then generate a new address and assign it to the key. The next address register (which may be located at the end of the set of control registers 505 ( Figure 5) may contain the next available address and may be incremented each time an address is assigned to a key that was not previously assigned an address. Each key address table buffer 735 may be associated with a subset of possible (4-byte) keys (e.g., based on the three least significant bits of the key, as discussed in further detail below), such that to search for a key in the key address table 740, only one of the key address table buffers 735 may need to be searched. The keys may be stored in each key address table buffer 735 in (increasing or decreasing) order so that a log search may be used to search the key address table buffers 735. The lookup circuit 730 may be a pipeline including multiple stages for searching the key address table, each stage corresponding to a step in the log search, such that the lookup circuit 730 may receive a key at each cycle of the system clock.
[0091] The address table 740 may include, for example, 8 key address table buffers 735, each of which may be used to store the address of the key based on the three least significant bits of the key. For example, the first key address table buffer ( Figure 7 A first 8×8 key address table buffer 745 may be used to store addresses of keys ending with 000 (i.e., having 000 as the three least significant bits), a second key address table buffer may be used to store addresses of keys ending with 001, a third key address table buffer may be used to store addresses of keys ending with 010, and so on. A first 8×8 cross-bar 745 may be used to enable each of the lookup circuits 730 to access all of the key address table buffers 735. For example, in operation, one of the lookup circuits 730 (e.g., the eighth) may receive a key-value pair having a key with 000 as the three least significant bits. It may then search PTLB0 for this key; if the key is in PTLB0, it may fetch the corresponding address and send the address-value pair to the assigned address read-modify-write circuit 720.
[0092] If there are eight read-modify-write circuits 720, then the address assignment to the read-modify-write circuits 720 may also be done based on the least significant bit (eg, based on the three least significant bits of the address), such as in Figure 7In an embodiment of . For example, if the address read from PTLB0 ends with 010, then the eighth lookup circuit may send the address and value to the third read-modify-write circuit in the read-modify-write circuit 720. The routing of the address and value may be completed by the second 8×8 crossbar switch 750. If the result of searching for the key in PTLB0 is that the key is not found, then the eighth lookup circuit may store the key in PTLB0 together with the address in the next address register, increment the next address register, and send the key and address to the read-modify-write circuit 720 to which the address is assigned. Whenever a new sum aggregation operation is started, the next address register may be initialized to zero. Conflicts within the crossbar switch may be resolved by arbitration at the competing outputs. The rate at which conflicts occur may depend on the temporal locality of the input key. Extending the sequential key into eight levels enables parallelism and reduces conflicts.
[0093] In some embodiments, the database offload engine has an NVDIMM-P (or memory channel) interface with the host (and the database offload engine can be packaged in an NVDIMM-P form factor). The host can then interact with the database offload engine through operating system calls that accommodate asynchronous access to memory. Such an asynchronous interface can facilitate performing operations in the database offload engine (which can introduce delays that may be unacceptable when using a synchronous memory interface). When performed in a hardware element that appears to the host as memory, such operations can be referred to as "function-in-memory (FIM)" processing. Reference Figure 8 In some such embodiments, a database application 805 (e.g., a SAP HANA database application) executes within an operating system 810 that includes a driver layer 410 that operates as a function-in-memory software interface. The host CPU 815 can communicate with the host main memory 825 and with the database offload engine through a controller 820 (if the database offload engine has an NVDIMM-P interface, the controller 820 can be a memory controller that supports both dynamic random access memory (DRAM) dual in-line memory module (DIMM) memory and NVDIMM-P memory, or if the database offload engine has a PCIe interface, the controller 820 can be a combination of a memory controller and a PCIe controller).
[0094] If the database offload engine has an NVDIMM-P interface, the database application 805 running on the host can use the memory of the database offload engine (e.g., memory 410 ( Figure 4)) to store database tables regardless of whether table scan operations or sum aggregation operations will be performed on them. Various database operations can be performed on tables and the results stored in other tables in the database offload engine's memory. When a table scan operation or sum aggregation operation is required, the host CPU can simply instruct the database offload engine to perform the operation on the tables already in the database offload engine's memory.
[0095] In contrast, if the database offload engine has a PCIe interface, then it may be inefficient to generally store the table in the memory of the database offload engine because the speed of performing host CPU operations on the data in the table may be greatly reduced due to the need to transfer data to and from the host CPU via the PCIe interface. Therefore, if the database offload engine has a PCIe interface, then the table of the database may generally be stored in the host main memory 825 and copied to the memory of the database offload engine as needed for performing table scan operations or sum aggregation operations in the database offload engine. Since the table needs to be copied to and from the database offload engine in such embodiments, perhaps the performance of an embodiment in which the database offload engine has an NVDIMM-P interface may generally be better than an embodiment in which the database offload engine has a PCIe interface.
[0096] Reference Fig.9AIn an embodiment where the database offload engine has an NVDIMM-p interface, temporary (or "time") data may be stored in the host main memory 825. The database main store (or "Hana Main Store") may be stored in the memory 440 of the database offload engine (or in the memory 440 connected to the database offload engine), and the host CPU and the database application 805 and one or more driver layers 410 running on the host CPU interface with the memory through a set of control registers 505. The database operation may include decompressing one or more tables in the database main store to form source data, and processing the source data (using database operations performed by the host CPU, or in the case of table scan operations or sum aggregation operations, database operations performed by the database offload engine) to form the destination data. In such an embodiment, the host may generate a query plan and call a function in the offload application programming interface (API) that causes the device driver to command the offload engine to perform a sum aggregation operation or a table scan operation. The database offload engine can then decompress the table data from the database main storage as needed, save the uncompressed data in the source area of the memory, perform the sum aggregation operation or table scan operation in a pipelined, vectorized manner, and store the results of the sum aggregation operation or table scan operation in the destination area of the memory. The host can then read the results from the destination area of the memory and perform additional database operations as needed.
[0097] Reference Fig. 9BIn an embodiment where the database offload engine has a PCIe interface, the database main storage may instead be located in the host main memory 825. To perform a table scan operation or a sum aggregation operation, the compressed data may be copied from the host main memory 825 to the database offload engine's memory 440 (or to a memory 440 connected to the database offload engine) using direct memory access (DMA) (e.g., direct memory access initiated by the database offload engine), decompressed in the database offload engine to form source data, and processed (using a table scan operation or a sum aggregation operation) to form destination data. In such an embodiment, the host may generate a query plan and call a function in the offload API that causes the device driver to command the offload engine to perform the sum aggregation operation or the table scan operation. The database offload engine may then copy the data from the database main storage in the host main memory 825 to the source area of the database offload engine's memory using direct memory access. The database offload engine can then decompress the table data from the source area of the memory as needed, store the uncompressed data in the source area of the memory, perform the sum aggregation operation or table scan operation in a pipelined, vectorized manner, and store the result of the sum aggregation operation or table scan operation in the destination area of the memory. The host can then read the result from the destination area of the memory through the PCIe interface and perform additional database operations as needed.
[0098] The term "processing circuit" is used herein to refer to any combination of hardware, firmware, and software used to process data or digital signals. The processing circuit hardware may include, for example, an application specific integrated circuit (ASIC), a general or dedicated central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), and a programmable logic device (such as a field programmable gate array (FPGA)). In the processing circuit used herein, each function is performed by hardware configured (i.e., hardwired) to perform the function, or by more general hardware (such as a CPU) configured to execute instructions stored in a non-temporary storage medium. The processing circuit may be fabricated on a single printed circuit board (PCB) or distributed on several interconnected PCBs. The processing circuit may include other processing circuits; for example, the processing circuit may include two processing circuits interconnected on a PCB, namely an FPGA and a CPU.
[0099] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or section from another element, component, region, layer or section. Therefore, without departing from the spirit and scope of the present inventive concept, the first element, first component, first region, first layer or first section described herein may be referred to as the second element, second component, second region, second layer or second section.
[0100] The terms used herein are only used to illustrate specific embodiments and are not intended to limit the inventive concept. As used herein, the terms "substantially", "about" and similar terms are used as approximate terms rather than as terms of degree, and are intended to take into account the inherent deviations of the measured or calculated values that will be recognized by a person of ordinary skill in the art. As used herein, the term "major component" refers to a component present in the composition or product in an amount greater than the amount of any other single component in the composition, polymer or product. In contrast, the term "primary component" refers to a component that constitutes at least 50% or more by weight of a composition, polymer or product. As used herein, the term "major portion" means at least half of each item when applied to multiple items.
[0101] As used herein, the singular forms "a and an" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "comprises and / or comprising" used in this specification indicate the presence of stated features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. When preceding a series of elements, expressions such as "at least one of..." modify the entire series of elements and do not modify individual elements in the series. In addition, "may" used when describing embodiments of the inventive concept refers to "one or more embodiments of the inventive concept." In addition, the term "exemplary" is intended to refer to an example or illustration. As used herein, the terms "use", "using" and "used" may be considered synonymous with the terms "utilize", "utilizing" and "utilized", respectively.
[0102] It should be understood that when an element or layer is referred to as being "on," "connected to," "coupled to," or "adjacent to" another element or layer, the element or layer may be directly on, directly connected to, directly coupled to, or directly adjacent to the other element or layer, or one or more intervening elements or layers may be present. In contrast, when an element or layer is referred to as being "directly on," "directly connected to," "directly coupled to," or "immediately adjacent to" another element or layer, there are no intervening elements or layers.
[0103] Any numerical range described herein is intended to include all sub-ranges of the same numerical precision within the described range. For example, the range of "1.0 to 10.0" is intended to include all sub-ranges between the described minimum value of 1.0 and the described maximum value of 10.0 (and including the described minimum value of 1.0 and the described maximum value of 10.0), that is, the minimum value is equal to or greater than 1.0 and the maximum value is equal to or less than 10.0, such as 2.4 to 7.6. Any maximum numerical limit described herein is intended to include all lower numerical limits contained therein, and any minimum numerical limit described in this specification is intended to include all higher numerical limits contained therein.
[0104] Although exemplary embodiments of the database unloading engine have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Therefore, it should be understood that the database unloading engine constructed according to the principles of the present disclosure may be implemented in ways other than those specifically described herein. The invention is further defined in the following claims and their equivalents.
Claims
1. A database processing system, comprising a database unloading engine, The database offload engine include: a vectorized adder including a plurality of read-modify-write circuits; a plurality of summing buffers, respectively connected to the plurality of read-modify-write circuits; Key address table; as well as control circuit, where The control circuit is configured to: receiving a first key and a first value corresponding to the first key; searching the key address table for the first key, and In response to finding a first address corresponding to the first key in the key address table, routing the first address and the first value to a read-modify-write circuit corresponding to the first address among the plurality of read-modify-write circuits based on a plurality of least significant bits of the first address, The read-modify-write circuit corresponding to the first address updates the value stored in the corresponding sum buffer among the plurality of sum buffers based on the first value.
2. The database processing system according to claim 1, wherein the control circuit is further configured to: receiving a second key and a second value corresponding to the second key; searching the key address table for the second key, and In response to not finding a second address corresponding to the second key in the key address table, selecting a new address that does not exist in the key address table, storing the second key and the new address in the key address table, and routing the new address and the second value to a read-modify-write circuit corresponding to the new address among the multiple read-modify-write circuits.
3. The database processing system according to claim 1, wherein the database offload engine has a non-volatile dual inline memory module-P interface for establishing a connection with a host.
4. The database processing system according to claim 1, wherein the database offload engine has a PCI Express interface for establishing a connection with a host.
5. The database processing system according to claim 1, in: The vectorized adder is a synchronous circuit within a clock domain defined by a shared system clock, One of the plurality of read-modify-write circuits is configured as a pipeline, the pipeline comprising: The first stage for performing read operations, a second stage for performing addition operations, and The third level is used to perform write operations, and The pipeline is configured to receive a third address and a third value corresponding to the third address at each cycle of the shared system clock.
6. The database processing system according to claim 1, in: The control circuit is a synchronous circuit within a clock domain defined by a shared system clock, The control circuit includes a search circuit for searching the key address table for a key, The lookup circuit is configured to include a pipeline of multiple stages for searching the key address table, The pipeline is configured to receive a key at each cycle of the shared system clock.
7. The database processing system according to claim 1, further comprising a host connected to the database unloading engine, wherein The host includes a non-transitory storage medium storing database application instructions and driver layer instructions, and The database application instruction includes a function call. When the function call is executed by the host, the function call causes the host to execute a driver layer instruction. The driver layer instruction causes the host to control the database offload engine to perform a sum aggregation operation.
8. The database processing system according to claim 1, wherein the database unloading engine further comprises a plurality of table scanning circuits, wherein A table scan circuit of the plurality of table scan circuits includes a conditional test circuit programmable with conditions, an input buffer, and an output buffer, and The conditional test circuit is configured as follows: determining whether the first entry at the fourth address in the input buffer satisfies the condition, and In response to determining that the first entry satisfies the condition, a corresponding result is written into the output buffer.
9. The database processing system according to claim 8, wherein the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write a corresponding element of the output vector into the output buffer. 10 . The database processing system of claim 8 , wherein the conditional testing circuit is configured to write the fourth address to a corresponding element of an output vector in the output buffer in response to determining that the first entry satisfies the condition.
11. The database processing system according to claim 8, in: The vectorized adder is a synchronous circuit within a clock domain defined by a shared system clock, A read-modify-write circuit among the plurality of read-modify-write circuits is configured as a pipeline including a first stage for performing a read operation, a second stage for performing an addition operation, and a third stage for performing a write operation, and The pipeline is configured to receive a fifth address and a fifth value corresponding to the fifth address at each cycle of the system clock.
12. The database processing system according to claim 8, in: The control circuit is a synchronous circuit within a clock domain defined by a shared system clock, The control circuit includes a search circuit for searching the key address table for a key, The lookup circuit is configured to include a pipeline of multiple stages for searching the key address table, The pipeline is configured to receive a key at each cycle of the shared system clock.
13. The database processing system according to claim 8, wherein the database offload engine has a non-volatile dual inline memory module-P interface for establishing a connection with a host.
14. A database processing system, comprising a database unloading engine, The database unloading engine includes a plurality of table scanning circuits, wherein A table scan circuit of the plurality of table scan circuits includes a conditional test circuit programmable with conditions, an input buffer, and an output buffer, and The conditional test circuit is configured as follows: determining whether a first entry at a first address in the input buffer satisfies the condition, and In response to determining that the first entry satisfies the condition, a corresponding result is written into the output buffer.
15. The database processing system of claim 14, wherein the conditional testing circuit is configured to, in response to determining that the first entry satisfies the condition, write an element of the output vector to the output buffer. 16 . The database processing system of claim 14 , wherein the conditional testing circuit is configured to write the first address to a corresponding element of an output vector in the output buffer in response to determining that the first entry satisfies the condition.
17. The database processing system according to claim 14, wherein the database offload engine has a non-volatile dual inline memory module-P interface for establishing a connection with a host.
18. The database processing system according to claim 14, wherein the database offload engine has a PCI Express interface for establishing a connection with a host.
19. A method for offloading database operations from a host, the method include: Calling a driver function by an application running on the host to perform a sum aggregation operation; as well as The sum aggregation operation is performed by the database unloading engine, wherein The database unloading engine comprises: a vectorized adder including a plurality of read-modify-write circuits; a plurality of summing buffers, respectively connected to the plurality of read-modify-write circuits; a key address table; and control circuit, and Executing the sum aggregation operation includes: receiving a first key and a first value corresponding to the first key; Search the key address table for the first key; In response to finding a first address corresponding to the first key in the key address table, routing the first address and the first value to a read-modify-write circuit corresponding to the first address among the plurality of read-modify-write circuits based on a plurality of least significant bits of the first address; receiving a second key and a second value corresponding to the second key; Search the key address table for the second key; In response to not finding an address corresponding to the second key in the key address table: selecting a new address that does not exist in the key address table; storing the second key and the new address in the key address table; and routing the new address and the second value to a read-modify-write circuit corresponding to the new address among the plurality of read-modify-write circuits, The read-modify-write circuit corresponding to the first address updates the value stored in the corresponding sum buffer among the plurality of sum buffers based on the first value.
20. The method according to claim 19, further comprising: include: The application calls a driver function to perform a table scan operation; as well as The table scan operation is performed by the database offload engine, wherein Executing the table scan operation includes: determining, by a condition test circuit of the database offload engine, whether a first entry at a second address in an input buffer of the database offload engine satisfies a condition, and In response to determining that the first entry in the input buffer satisfies the condition, a corresponding result is written into an output buffer of the database offload engine.
Citation Information
Patent Citations
Object storage system managing error-correction-code-related data in key-value mapping information
US20170255508A1