Direct memory access (DMA) for tabular-columnar data
The proposed method improves the efficiency of handling high-speed tabular-columnar data by using direct memory access channels to process and store data in a zero-copy format, addressing the inefficiencies of current DMA drivers and ensuring valid output data.
Patent Information
- Application Number
- PCT/IB2024/061874
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Current DMA drivers are inefficient in handling high-speed tabular-columnar data, particularly in writing processed data to memory, and they do not support zero-copy formats that are immediately usable by follow-on users/processors.
The method involves receiving an input stream of data with column identifications, determining corresponding direct memory access channels, buffering values for each channel, and outputting them as batches via the associated channels to different memory locations, while maintaining validity and performing calculations as needed.
This approach enables faster and more efficient processing and storage of columnar format data, reducing the need for additional processing and format conversion, and ensuring that output data is valid and immediately usable.
Smart Images

Figure IB2024061874_05062025_PF_FP_ABST
Abstract
Description
[0001] DIRECT MEMORY ACCESS (DMA) FOR TABULAR-COLUMNAR DATA
[0002] Cross-Reference to Related Application
[0003] This application claims the benefit of priority of United States Provisional Patent Application No. 63 / 603,425, filed on November 28, 2023. The foregoing application is incorporated herein by reference in its entirety.
[0004] Technical Field
[0005] The disclosed embodiments generally relate to computer memory and, in particular, the disclosed embodiments involve digital data processing, formatting, and handling.
[0006] Background Information
[0007] Computer memory is conventionally accessed using DMA (direct memory access) drivers that are designed and implemented to handle a wide range of data input and output types.
[0008] SUMMARY
[0009] In some embodiments, a method for columnar direct memory access processing includes: receiving an input stream of data including one or more values, each of the values including a column identification; determining, for each value, based on the column identification, a corresponding direct memory access channel; and processing each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location.
[0010] In some embodiments, the method includes: receiving an input stream of data including one or more values; determining, for each value, a corresponding category based on each respective value, and processing each of the values based on the corresponding category.
[0011] In some further embodiments, the input stream includes a plurality of portions of data, each portion including one or more of the values.
[0012] In some further embodiments, the values are tabular data. In some further embodiments, each of the values includes a column identification. In some further embodiments, a direct memory access channel is associated with each of the column identifications. In some further embodiments, values for each of the column identifiers are routed to the associated direct memory access channel. In some further embodiments, each of the associated direct memory access channels writes to a different associated memory location.
[0013] In some further embodiments, the determining includes determining the corresponding category based on the column identification. In some further embodiments, the processing includes routing each of the values based on the column identification. In some further embodiments, the routing is to one of a plurality of channel processing modules. In some further embodiments, each of the plurality of channel processing modules is associated with a direct memory access channel. In some further embodiments, a plurality of values are sent to each of the plurality of channel processing modules, buffered, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel.
[0014] In some further embodiments, the processing further includes modifying one or more values from a first format to output data in a second format. In some further embodiments, the second format is a standardized representation. In some further embodiments, the standardized representation is the Apache Arrow format.
[0015] In some further embodiments, the processing further includes calculations on the values for each corresponding category. In some further embodiments, each corresponding category is associated with a direct memory access channel, and the calculations include operations for the values of each of the associated direct memory access channels. In some further embodiments, the calculations include table-specific calculations. In some further embodiments, the calculations include operations selected from the group including: counting validity of input data; counting validity of values; aggregating validity data for each corresponding category; aggregating validity data from each portion of data (hunk) of the input stream, each portion of data including one or more of the values; calculating offsets for values in the input stream of data; calculating offsets for values of variable sized data in the input stream of data; and generating offsets for output data based on the input data stream values.
[0016] In some further embodiments, the processing further includes maintaining validity of the input stream values. In some further embodiments, maintaining validity includes processing each of the values to determine if the value has an associated validity indicator, maintaining association of the validity indicator with the value during the processing, and generating output data, the output data based on the value and having an output validity indicator.
[0017] In some embodiments, non-transitory computer-readable medium having stored thereon computer-readable instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving an input stream of data including one or more values, each of the values including a column identification; determining, for each value, based on the column identification, a corresponding direct memory access channel; and processing each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location. In some embodiments, a non-transitory computer-readable medium having stored thereon computer-readable instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving an input stream of data including one or more values; determining, for each value, a corresponding category based on each respective value, and processing each of the values based on the corresponding category.
[0018] In some embodiments, a system for processing data includes: an input module configured to receive an input stream of data including one or more values, each of the values including a column identification; a processing module including at least one processor including circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: determine, for each value, based on the column identification a corresponding direct memory access channel; and process each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location.
[0019] In some embodiments, a system for processing data includes: an input module configured to receive an input stream of data including one or more values; a processing module including at least one processor including circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: determine, for each value, a corresponding category based on each respective value and process each of the values based on the corresponding category.
[0020] In some further embodiments, the input stream includes a plurality of portions of data, each portion including one or more of the values.
[0021] In some further embodiments, the values are tabular data, each of the values includes a column identification, and a direct memory access channel is associated with each of the column identifications. In some further embodiments, the processing module is further configured for routing values for each of the column identifiers to the associated direct memory access channel. In some further embodiments, the processing module is further configured for writing each of the associated direct memory access channels to a different associated memory location.
[0022] In some further embodiments, the system further includes: a plurality of channel processing modules, the processing module is further configured for routing each of the values based on the column identification, to a corresponding channel processing module. In some further embodiments, the processing module is further configured for routing a plurality of values to each of the plurality of channel processing modules, and each of the plurality of channel processing modules is further configured to: receive the plurality of values, buffer the plurality of values, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel.
[0023] In some further embodiments, the processing module is further configured for modifying one or more values from a first format to output data in a second format. In some further embodiments, the second format is a standardized representation, and the standardized representation is the Apache Arrow format.
[0024] In some further embodiments, the processing module is further configured for performing calculations on the values for each corresponding category. In some further embodiments, each corresponding category is associated with a direct memory access channel, and the calculations include operations for the values of each of the associated direct memory access channels. In some further embodiments, the calculations include table-specific calculations.
[0025] In some further embodiments, the processing module is further configured for maintaining validity of the input stream values. In some further embodiments, maintaining validity includes processing each of the values to determine if the value has an associated validity indicator, maintaining association of the validity indicator with the value during the processing, and generating output data, the output data based on the value and having an output validity indicator.
[0026] Consistent with other disclosed embodiments, one or more non-transitory computer readable storage media may store program instructions, which are executed by at least one processing device and perform any of the methods described herein.
[0027] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims.
[0028] BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments. In the drawings:
[0030] FIG. 1 is an example of a computer (CPU) architecture.
[0031] FIG. 2 is an example of a graphics processing unit (GPU) architecture.
[0032] FIG. 3 is a diagrammatic representation of a computer memory with an error correction code (ECC) capability.
[0033] FIG. 4 is a diagrammatic representation of a process for writing data to the memory module.
[0034] FIG. 5 is a diagrammatic representation of a process for reading from memory.
[0035] FIG. 6 is a diagrammatic representation of an architecture including memory processing modules. FIG. 7 is a diagrammatic representation of an exemplary connection of a host to a memory appliance.
[0036] FIG. 8 is an example of implementations of processing systems.
[0037] FIG. 9 is an example of a high-level architecture for a data analytics accelerator.
[0038] FIG. 10 is an example of the software layer for the data analytics accelerator.
[0039] FIG. 11 is an example of the hardware layer for the data analytics accelerator.
[0040] FIG. 12 is an example of the storage layer and bridges for the data analytics accelerator.
[0041] FIG. 13 is an example of networking for the data analytics accelerator.
[0042] FIG. 14 is a high-level example of an architecture for a columnar DMA system.
[0043] FIG. 15 is a high-level general processing flow of the columnar DMA system.
[0044] FIG. 16 is a logical view of input data.
[0045] FIG. 17A is an example of values being input to the columnar DMA module.
[0046] FIG. 17B shows exemplary memory data views.
[0047] FIG. 18 is a flowchart of a method for columnar DMA processing.
[0048] FIG. 19 is a high-level partial block diagram of an exemplary system configured to implement the columnar DMA processing module, and other modules, of the disclosed embodiments.
[0049] DETAILED DESCRIPTION
[0050] The following detailed description refers to the accompanying drawings. Wherever convenient, the same reference numbers are used in the drawings and the following description to refer to the same or similar parts. While several illustrative embodiments are described herein, modifications, adaptations and other implementations are possible. For example, substitutions, additions, or modifications may be made to the components illustrated in the drawings, and the illustrative methods described herein may be modified by substituting, reordering, removing, or adding steps to the disclosed methods. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the proper scope is defined by the appended claims.
[0051] In one example, a disclosed embodiment is a system and method for processing, formatting, and handling data. The system may facilitate handling of higher speed data more efficiently than conventional systems and methods. The system and method may be used for columnar DMA (direct memory access) processing, however, this is not limiting, and other implementations are possible and consistent with the disclosed embodiments.
[0052] Example Architecture FIG. 1 is an example of a computer (CPU) architecture. A CPU 100 may comprise a processing unit 110 that includes one or more processor subunits, such as processor subunit 120a and processor subunit 120b. Although not depicted in the current figure, each processor subunit may comprise a plurality of processing elements. Moreover, the processing unit 110 may include one or more levels of on-chip cache. Such cache elements are generally formed on the same semiconductor die as processing unit 110 rather than being connected to processor subunits 120a and 120b via one or more buses formed in the substrate containing processor subunits 120a and 120b and the cache elements. An arrangement directly on the same die, rather than being connected via buses, may be used for both first-level (LI) and second-level (L2) caches in processors. Alternatively, in older processors, L2 caches were shared amongst processor subunits using back-side buses between the subunits and the L2 caches. Back-side buses are generally larger than front-side buses, described below. Accordingly, because cache is to be shared with all processor subunits on the die, cache 130 may be formed on the same die as processor subunits 120a and 120b or communicatively coupled to processor subunits 120a and 120b via one or more back-side buses. In both embodiments without buses (e.g., cache is formed directly on-die) as well as embodiments using back-side buses, the caches are shared between processor subunits of the CPU.
[0053] Moreover, processing unit 110 may communicate with shared memory 140a and memory 140b. For example, memories 140a and 140b may represent memory banks of shared dynamic random-access memory (DRAM). Although depicted with two banks, memory chips may include between eight and sixteen memory banks. Accordingly, processor subunits 120a and 120b may use shared memories 140a and 140b to store data that is then operated upon by processor subunits 120a and 120b. This arrangement, however, results in the buses between memories 140a and 140b and processing unit 110 acting as a bottleneck when the clock speeds of processing unit 110 exceed data transfer speeds of the buses. This is generally true for processors, resulting in lower effective processing speeds than the stated processing speeds based on clock rate and number of transistors.
[0054] FIG. 2 is an example of a graphics processing unit (GPU) architecture. Deficiencies of the CPU architecture similarly persist in GPUs. A GPU 200 may comprise a processing unit 210 that includes one or more processor subunits (e.g., subunits 220a, 220b, 220c, 220d, 220e, 220f, 220g, 220h, 220i, 220j, 220k, 2201, 220m, 220n, 220o, and 220p). Moreover, the processing unit 210 may include one or more levels of on-chip cache and / or register files. Such cache elements are generally formed on the same semiconductor die as processing unit 210. Indeed, in the example of the current figure, cache 210 is formed on the same die as processing unit 210 and shared amongst all of the processor subunits, while caches 230a, 230b, 230c, and 230d are formed on a subset of the processor subunits, respectively, and dedicated thereto.
[0055] Moreover, processing unit 210 communicates with shared memories 250a, 250b, 250c, and 250d. For example, memories 250a, 250b, 250c, and 250d may represent memory banks of shared DRAM. Accordingly, the processor subunits of processing unit 210 may use shared memories 250a, 250b, 250c, and 250d to store data that is then operated upon by the processor subunits. This arrangement, however, results in the buses between memories 250a, 250b, 250c, and 250d and processing unit 210 acting as a bottleneck, similar to the bottleneck described above for CPUs.
[0056] FIG. 3 is a diagrammatic representation of a computer memory with an error correction code (ECC) capability. As shown in the current figure, a memory module 301 includes an array of memory chips 300, shown as nine chips (i.e., chip-0, 100-0 through chip-8, 100-8, respectively). Each memory chip has respective memory arrays 302 (e.g., elements labelled 302- 0 through 302-8) and corresponding address selectors 306 (shown as respective selector-0 106-0 through selector-8 106-8). Controller 308 is shown as a DDR controller. The DDR controller 308 is operationally connected to CPU 100 (processing unit 110), receiving data from the CPU 100 for writing to memory, and retrieving data from the memory to send to the CPU 100. The DDR controller 308 also includes an error correction code (ECC) module that generates error correction codes that may be used in identifying and correcting errors in data transmissions between CPU 100 and components of memory module 301.
[0057] FIG. 4 is a diagrammatic representation of a process for writing data to the memory module 301. Specifically, the process 420 of writing to the memory module 301 can include writing data 422 in bursts, each burst including 8 bytes for each chip being written to (in the current example, 8 of the memory chips 300, including chip-0, 100-0 to chip-7, 100-7). In some implementations, an original error correction code (ECC) 424 may be calculated in the ECC module 312 in the DDR controller 308. The ECC 424 is calculated across each of the chip’s 8 bytes of data, resulting in an additional, original, 1-byte ECC for each byte of the burst across the 8 chips. The 8-byte (8xl-byte) ECC is written with the burst to a ninth memory chip serving as an ECC chip in the memory module 301, such as chip- 8, 100-8.
[0058] The memory module 301 can activate a cyclic redundancy check (CRC) check for each chip’s burst of data, to protect the chip interface. A cyclic redundancy check is an errordetecting code commonly used in digital networks and storage devices to detect accidental changes to raw data. Blocks of data get a short check value attached, based on the remainder of a polynomial division of the block’s contents. In this case, an original CRC 426 is calculated by the DDR controller 308 over the 8 bytes of data 422 in a chip’s burst (one row in the current figure) and sent with each data burst (each row / to a corresponding chip) as a ninth byte in the chip’s burst transmission. When each chip 300 receives data, each chip 300 calculates a new CRC over the data and compares the new CRC to the received original CRC. If the CRCs match, the received data is written to the chip’s memory 302. If the CRCs do not match, the received data is discarded, and an alert signal is activated. An alert signal may include an ALERT_N signal.
[0059] Additionally, when writing data to a memory module 301, an original parity 428 A is normally calculated over the (exemplary) transmitted command 428B and address 428C. Each chip 300 receives the command 428B and address 428C, calculates a new parity, and compares the original parity to the new parity. If the parities match, the received command 428B and address 428C are used to write the corresponding data 422 to the memory module 301. If the parities do not match, the received data 422 is discarded, and an alert signal (e.g., ALERT_N) is activated.
[0060] FIG. 5 is a diagrammatic representation of a process 530 for reading from memory. When reading from the memory module 301, the original ECC 424 is read from the memory and sent with the data 422 to the ECC module 312. The ECC module 312 calculates a new ECC across each of the chips’ 8 bytes of data. The new ECC is compared to the original ECC to determine (detect, correct) if an error has occurred in the data (transmission, storage). In addition, when reading data from memory module 301, an original parity 538A is normally calculated over the (exemplary) transmitted command 538B and address 538C (transmitted to the memory module 301 to tell the memory module 301 to read and from which address to read). Each chip 300 receives the command 538B and address 538C, calculates a new parity, and compares the original parity to the new parity. If the parities match, the received command 538B and address 538C are used to read the corresponding data 422 from the memory module 301. If the parities do not match, the received command 538B and address 538C are discarded and an alert signal (e.g., ALERT_N) is activated.
[0061] Overview of Memory Processing Modules and Associated Appliances
[0062] FIG. 6 is a diagrammatic representation of an architecture including memory processing modules. For example, a memory processing module (MPM) 610, as described above, may be implemented on a chip to include at least one processing element (e.g., a processor subunit) local to associated memory elements formed on the chip. In some cases, an MPM 610 may include a plurality of processing elements spatially distributed on a common substrate among their associated memory elements within the MPM 610.
[0063] In the example of FIG. 6, the memory processing module 610 includes a processing module 612 coupled with four, dedicated memory banks 600 (shown as respective bank-0, 600-0 through bank-3, 600-3). Each bank includes a corresponding memory array 602 (shown as respective memory array-0, 602-0 through memory array-3, 602-3) along with selectors 606 (shown as selector-0 606-0 to selector-3 606-3). The memory arrays 602 may include memory elements similar to those described above relative to memory arrays 302. Local processing, including arithmetic operations, other logic-based operations, etc. can be performed by processing module 612 (also referred to in the context of this document as a “processing subunit,” “processor subunit,” “logic,” “micro mind,” or “UMIND”) using data stored in the memory arrays 602, or provided from other sources, for example, from other of the processing modules 612. In some cases, one or more processing modules 612 of one or more MPMs 610 may include at least one arithmetic logic units (ALU). Processing module 612 is operationally connected to each of the memory banks 600.
[0064] A DDR controller 608 may also be operationally connected to each of the memory banks 600, e.g., via an MPM slave controller 623. Alternatively, and / or in addition to the DDR controller 608, a master controller 622 can be operationally connected to each of the memory banks 600, e.g., via the DDR controller 608 and memory controller 623. The DDR controller 608 and the master controller 622 may be implemented in an external element 620. Additionally, and / or alternatively, a second memory interface 618 may be provided for operational communication with the MPM 610.
[0065] While the MPM 610 of FIG. 6 pairs one processing module 612 with four, dedicated memory banks 600, more or fewer memory banks can be paired with a corresponding processing module to provide a memory processing module. For example, in some cases, the processing module 612 of MPM 610 may be paired with a single, dedicated memory bank 600. In other cases, the processing module 612 of MPM 610 may be paired with two or more dedicated memory banks 600, four or more dedicated memory banks 600, etc. Various MPMs 610, including those formed together on a common substrate or chip, may include different numbers of memory banks relative to one another. In some cases, an MPM 610 may include one memory bank 600. In other cases, an MPM may include two, four, eight, sixteen, or more memory banks 600. As a result, the number of memory banks 600 per processing module 612 may be the same throughout an entire MPM 610 or across MPMs. One or more MPMs 610 may be included in a chip. In a non-limiting example, included in an XRAM chip 624. Alternatively, at least one processing module 612 may control more memory banks 600 than another processing module 612 included within an MPM 610 or within an alternative or larger structure, such as the XRAM chip 624.
[0066] Each MPM 610 may include one processing module 612 or more than one processing module 610. In the example of the current figure, one processing module 612 is associated with four dedicated memory banks 600. In other cases, however, one or more memory banks of an MPM may be associated with two or more processing modules 612.
[0067] Each memory bank 600 may be configured with any suitable number of memory arrays 602. In some cases, a bank 600 may include only a single array. In other cases, a bank 600 may include two or more memory arrays 602, four or more memory arrays 602, etc. Each of the banks 600 may have the same number of memory arrays 602. Alternatively, different banks 600 may have different numbers of memory arrays 602.
[0068] Various numbers of MPMs 610 may be formed together on a single hardware chip. In some cases, a hardware chip may include just one MPM 610. In other cases, however, a single hardware chip may include two, four, eight, sixteen, 32, 64, etc. MPMs 610. In the particular non-limiting example represented in the current figure, 64 MPMs 610 are combined together on a common substrate of a hardware chip to provide the XRAM chip 624, which may also be referred to as a memory processing chip or a computational memory chip. In some embodiments, each MPM 610 may include a slave controller 613 (e.g., an extreme / Xele or XSC slave controller (SC)) configured to communicate with a DDR controller 608 (e.g., via MPM slave controller 623), and / or a master controller 622. Alternately, fewer than all of the MPMs onboard an XRAM chip 624 may include a slave controller 613. In some cases, multiple MPMs (e.g., 64 MPMs) 610 may share a single slave controller 613 disposed on XRAM chip 624. Slave controller 613 can communicate data, commands, information, etc. to one or more processing modules 612 on XRAM chip 624 to cause various operations to be performed by the one or more processing modules 612.
[0069] One or more XRAM chips 624, which may include a plurality of XRAM chips 624, such as sixteen XRAM chips 624, may be configured together to provide a dual in-line memory module (DIMM) 626. Traditional DIMMs may be referred to as a RAM stick, which may include eight or nine, etc., dynamic random-access memory chips (integrated circuits) constructed as / on a printed circuit board (PCB) and having a 64-bit data path. In contrast to traditional memory, the disclosed memory processing modules 610 include at least one computational component (e.g., processing module 612) coupled with local memory elements (e.g., memory banks 600). As multiple MPMs may be included on an XRAM chip 624, each XRAM chip 624 may include a plurality of processing modules 612 spatially distributed among associated memory banks 600. To acknowledge the inclusion of computational capabilities (together with memory) within the XRAM chip 624, each DIMM 626 including one or more XRAM chips (e.g., sixteen XRAM chips, as in the FIG. 6 example) on a single PCB may be referred to as an XDIMM (or eXtremeDIMM or XeleDIMM). Each XDIMM 626 may include any number of XRAM chips 624, and each XDIMM 624 may have the same or a different number of XRAM chips 624 as other XDIMMs 626. In the FIG. 6 example, each XDIMM 626 includes sixteen XRAM chips 624.
[0070] As shown in FIG. 6, the architecture may further include one or more memory processing units, such as an intense memory processing unit (IMPU) 628. Each IMPU 628 may include one or more XDIMMs 626. In the current figure example, each IMPU 628 includes four XDIMMs 626. In other cases, each IMPU 628 may include the same or a different number of XDIMMs as other IMPUs. The one or more XDIMMs included in IMPU 628 can be packaged together with or otherwise integrated with one or more DDR controllers 608 and / or one or more master controllers 622. For example, in some cases, each XDIMM included in IMPU 628 may include a dedicated DDR controller 608 and / or a dedicated master controller 622. In other cases, multiple XDIMMs included in IMPU 628 may share a DDR controller 608 and / or a master controller 622. In one particular example, IMPU 628 includes four XDIMMs 626 along with four master controllers 622 (each master controller 622 including a DDR controller 608), where each of the master controllers 622 is configured to control one associated XDIMM 626, including the MPMs 610 of the XRAM chips 624 included in the associated XDIMM 626.
[0071] The DDR controller 608 and the master controller 622 are examples of controllers in a controller domain 630. A higher-level domain 632 may contain one or more additional devices, user applications, host computers, other devices, protocol layer entities, and the like. The controller domain 630 and related features are described in the sections below. In a case where multiple controllers and / or multiple levels of controllers are used, the controller domain 630 may serve as at least a portion of a multi-layered module domain, which is also further described in the sections below.
[0072] In the architecture represented by FIG. 6, one or more IMPUs 628 may be used to provide a memory appliance 640, which may be referred to as an XIPHOS appliance. In the example of the current figure, memory appliance 640 includes four IMPUs 628.
[0073] The location of processing elements 612 among memory banks 600 within the XRAM chips 624 (which are incorporated into XDIMMs 626 that are incorporated into IMPUs 628 that are incorporated into memory appliance 640) may significantly relieve the bottlenecks associated with CPUs, GPUs, and other processors that operate using a shared memory. For example, a processor subunit 612 may be tasked to perform a series of instructions using data stored in memory banks 600. The proximity of the processing subunit 612 to the memory banks 600 can significantly reduce the time required to perform the prescribed instructions using the relevant data.
[0074] FIG. 7 is a diagrammatic representation of an exemplary connection of a host to a memory appliance. As shown in the current figure, a host 710 may provide instructions, data, and / or other input to memory appliance 640 and read output from the same. Rather than requiring the host to access a shared memory and perform calculations / functions relative to data retrieved from the shared memory, in the disclosed embodiments, the memory appliance 640 can perform the processing associated with a received input from host 710 within the memory appliance (e.g., within processing modules 612 of one or more MPMs 610 of one or more XRAM chips 624 of one or more XDIMMs 626 of one or more IMPUs). Such functionality is made possible by the distribution of processing modules 612 among and on the same hardware chips as the memory banks 600 where relevant data needed to perform various calculations / functions / etc. is stored.
[0075] The architecture described in FIG. 6 may be configured for execution of code. For example, each processor subunit 612 may individually execute code (defining a set of instructions) apart from other processor subunits in an XRAM chip 624 within memory appliance 640. Accordingly, rather than relying on an operating system to manage multithreading or using multitasking (which is concurrency rather than parallelism), the XRAM chips of the present disclosure may allow for processor subunits to operate fully in parallel.
[0076] In addition to a fully parallel implementation, at least some of the instructions assigned to each processor subunit may be overlapping. For example, a plurality of processor subunits 612 on an XRAM chip 624 (or within an XDIMM 626 or IMPU 628) may execute overlapping instructions as, for example, an implementation of an operating system or other management software, while executing non-overlapping instructions in order to perform parallel tasks within the context of the operating system or other management software.
[0077] For purposes of various structures discussed in this description, the Joint Electron Device Engineering Council (JEDEC) Standard No. 79-4C defines the DDR4 SDRAM specification, including features, functionalities, AC and DC characteristics, packages, and ball / signal assignments. The latest version at the time of this application is January 2020, available from JEDEC Solid State Technology Association, 3103 North 10th Street, Suite 240 South, Arlington, VA 22201-2107, www.jedec.org, and is incorporated by reference in its entirety herein.
[0078] Exemplary elements such as XRAM, XDIMM, XSC, and IMPU are available from NeuroBlade Ltd., Tel Aviv, Israel. Details of memory processing modules and related technologies can be found in PCT / IB2018 / 000995 filed 30-July-2018, PCT / IB2019 / 001005 filed 6-September-2019, PCT / IB2020 / 000665 filed 13-August-2020, and PCT / US2021 / 055472 filed 18-October-2021. Exemplary implementations using XRAM, XDIMM, XSC, IMPU, etc. elements are not limiting, and based on this description one skilled in the art will be able to design and implement configurations for a variety of applications using alternative elements.
[0079] Data Analytics Processor FIG. 8 is an example of implementations of processing systems and, in particular, processing systems for data analytics. Many modern applications are limited by data communication 820 between storage 800 and processing (shown as general -purpose compute 810). Current solutions include adding levels of data cache and re-layout of hardware components. For example, current solutions for data analytics applications have limitations including: (1) Network bandwidth (BW) between storage and processing, (2) network bandwidth between CPUs, (3) memory size of CPUs, (4) inefficient data processing methods, and (5) access rate to CPU memory.
[0080] In addition, data analytics solutions have significant challenges in scaling up. For example, when trying to add more processing power or memory, more processing nodes are required, therefore more network bandwidth between processors and between processors and storage is required, leading to network congestion.
[0081] FIG. 9 is an example of a high-level architecture for a data analytics accelerator. A data analytics accelerator 900 is configured between an external data storage 920 and an analytics engine (AE) 910 optionally followed by completion processing 912, for example, on the analytics engine 910. The external data storage 920 may be deployed external to the data analytics accelerator 900, with access via an external computer network. The analytics engine (AE) 910 may be deployed on a general-purpose computer. The accelerator may include a software layer 902, a hardware layer 904, a storage layer 906, and networking (not shown). Each layer may include modules such as software modules 922, hardware modules 924, and storage modules 926. The layers and modules are connected within, between, and external to each of the layers. Acceleration may be done at least in part by applying one or more innovative operations, data reduction, and partial processing operations between the external data storage 920 and the analytics engine 910 (or general-purpose compute 810). Implementations of our solutions may include, but are not limited to, features such as, in-line, high parallelism computation, and data reduction. In an alternative operation, (only) a portion of data is processed by the data analytics accelerator 900 and a portion of the data bypasses the data analytics accelerator 900.
[0082] The data analytics accelerator 900 may provide at least in part a streaming processor, and is particularly suited, but not limited to, accelerating data analytics. The data analytics accelerator 900 may drastically reduce (for example, by several orders of magnitude) the amount of data which is transferred over the network to the analytics engine 910 (and / or the general- purpose compute 810), reduces the workload of the CPU, and reduces the required memory which the CPU needs to use. The accelerator 900 may include one or more data analytics processing engines which are tailor-made for data analytics tasks, such as scan, join, filter, aggregate etc., doing these tasks much more efficiently than analytics engine 910 (and / or the general-purpose compute 810). An implementation of the data analytics accelerator 900 is the Hardware Enhanced Query System (HEQS), which may include a Xiphos Data Analytics Accelerator (available from NeuroBlade Ltd., Tel Aviv, Israel).
[0083] FIG. 10 is an example of the software layer for the data analytics accelerator. The software layer 902 may include, but is not limited to, two main components: a software development kit (SDK) 1000 and embedded software 1010. The SDK provides abstraction of the accelerator capabilities through well-defined and easy to use data-analytics oriented software APIs for the data analytics accelerator. A feature of the SDK is enabling users of the data analytics accelerator to maintain the users’ own DBMS, while adding the data analytics accelerator capabilities, for example, as part of the users’ DBMS’s planner optimization. The SDK may include modules such as:
[0084] A run-time environment 1002 may expose hardware capabilities to above layers. The run-time environment may manage the programming, execution, synchronization, and monitoring of underlying hardware engines and processing elements.
[0085] A Fast Data I / O providing an efficient API 1004 for injection of data into the data analytics accelerator hardware and storage layers, such as an NVMe array and memories, and for interaction with the data. The Fast Data I / O may also be responsible for forwarding data from the data analytics accelerator to another device (such as the analytics engine 910, an external host, or server) for processing and / or completion processing 912.
[0086] A manager 1006 (data analytics accelerator manager) may handle administration of the data analytics accelerator.
[0087] A toolchain may include development tools 1008, for example, to help developers enhance the performance of the data analytics accelerator, eliminate bottlenecks, and optimize query execution. The toolchain may include a simulator and profiler, as well as a LLVM compiler.
[0088] Embedded software component 1010 may include code running on the data analytics accelerator itself. Embedded software component 1010 may include firmware 1012 that controls the operation of the accelerator’s various components, as well as real-time software 1014 that runs on the processing elements. At least a portion of the embedded software component code may be generated, such as auto generated, by the (data analytics accelerator) SDK.
[0089] FIG. 11 is an example of the hardware layer for the data analytics accelerator. The hardware layer 904 includes one or more acceleration units 1100. Each acceleration unit 1100 includes one or more of a variety of elements (modules), which may include a selector module 1102, filter and projection module (FPE) 1103, Join-and-group-by (JaGB) module 1108, and bridges 1110. Each module may contain one or more sub-modules, for example, the FPE 1103 which may include a string engine (SE) 1104 and a filtering and aggregation engine (FAE) 1106.
[0090] In FIG. 11, a plurality of acceleration units 1100 are shown as first acceleration unit 1100-1 to nth acceleration unit 1100-N. In the context of this description, the element number suffix “-N”, where “N” is an integer, generally refers to an exemplary one of the elements, and the element number without a suffix refers to the element in general or the group of elements. One or more acceleration units 1100, individually or in combination, may be implemented using one or more individual or combination of FPGAs, ASICs, PCBs, and similar. Acceleration units 1100 may have the same or similar hardware configurations. However, this is not limiting, and modules may vary from one to another of the acceleration units 1100.
[0091] An example of element configuration will be used in this description. As noted above, element configuration may vary. Similarly, an example of networking and communication will be used. However, alternative, and additional connections between elements, feed forward, and feedback data may be used. Input and output from elements may include data and alternatively or additionally includes signaling and similar information.
[0092] The selector module 1102 is configured to receive input from any of the other acceleration elements, such as, for example, from at least from the bridges 1110 and the Join- and-group-by engine (JaGB) 1108 (shown in the current figure), and optionally / alternatively / in addition from the filtering and projection module (FPE) 1103, the string engine (SE) 1104, and the filtering and aggregation engine (FAE) 1106. Similarly, the selector module 1102 can be configured to output to any of the other acceleration elements, such as, for example, to the FPE 1103.
[0093] The FPE 1103 may include a variety of elements (sub-elements). Input and output from the FPE 1103 may be to the FPE 1103 for distribution to sub-elements, or directly to and from one or more of the sub-elements. The FPE 1103 is configured to receive input from any of the other acceleration elements, such as, for example, from the selector module 1102. FPE input may be communicated to one or more of the string engine 1104 and FAE 1106. Similarly, the FPE 1103 is configured to output from any of the sub-elements to any of the other acceleration elements, such as, for example, to the JaGB 1108.
[0094] The Join-and-group-by (JaGB) engine 1108 may be configured to receive input from any of the other acceleration elements, such as, for example, from the FPE 1103 and the bridges 1110. The JaGB 1108 may be configured to output to any of the acceleration unit elements, for example, to the selector module 1102 and the bridges 1110.
[0095] FIG. 12 is an example of the storage layer and bridges for the data analytics accelerator. The storage layer 906 may include one or more types of storage deployed locally, remotely, or distributed within and / or external to one or more of the acceleration units 1100 and one or more of the data analytics accelerators 900. The storage layer 906 may include non-volatile memory (such as local data storage 1208) and volatile memory (such as an accelerator memory 1200) deployed local to the hardware layer 904. Non-limiting examples of the local data storage 1208 include, but are not limited to solid state drives (SSD) deployed local and internal to the data analytics accelerator 900. Non-limiting examples of the accelerator memory 1200 include, but are not limited to FPGA memory (for example, of the hardware layer 904 implementation of the acceleration unit 1100 using an FPGA), processing in memory (PIM) 1202 memory for example, banks 600 of memory 602 in a memory processing module 610, and SRAM, DRAM, and HBM (for example, deployed on a PCB with the acceleration unit 1100). The storage layer 906 may also use and / or distribute memory and data via the bridges 1110 (such as, for example, the memory bridge 1114) via a fabric 1306 (described below in reference to FIG. 13), for example, to other acceleration units 1100 and / or other acceleration processors 900. In some embodiments, storage elements may be implemented by one or more elements or sub-elements.
[0096] One or more bridges 1110 provide interfaces to and from the hardware layer 904. Each of the bridges 1110 may send and / or receive data directly or indirectly to / from elements of the acceleration unit 1100. Bridges 1110 may include storage 1112, memory 1114, fabric 1116, and compute 1118.
[0097] Bridges configuration may include the storage bridge 1112 interfaces with the local data storage 1208. The memory bridge interfaces with memory elements, for example the PIM 1202, SRAM 1204, and DRAM / HBM 1206. The fabric bridge 116 interfaces with the fabric 1306. The compute bridge 1118 may interface with the external data storage 920 and the analytics engine 910. A data input bridge (not shown) may be configured to receive input from any of the other acceleration elements, including from other bridges, and to output to any of the acceleration unit elements, such as, for example, to the selector module 1102.
[0098] FIG. 13 is an example of networking for the data analytics accelerator. An interconnect 1300 may include an element deployed within each of the acceleration units 1100. The interconnect 1300 may be operationally connected to elements within the acceleration unit 1100, providing communications within the acceleration unit 1100 between elements. In FIG. 13, exemplary elements (1102, 1104, 1106, 1108, 1110) are shown connected to the interconnect 1300. The interconnect 1300 may be implemented using one or more sub-connection systems using one or more of a variety of networking connections and protocols between two or more of the elements, including, but not limited to, dedicated circuits and PCI switching. The interconnect 1300 may facilitate alternative and additional connections feed forward, and feedback between elements, including but not limited to looping, multi-pass processing, and bypassing one or more elements. The interconnect can be configured for communication of data, signaling, and other information.
[0099] Bridges 1110 may be deployed and configured to provide connectivity from the acceleration unit 1100-1 (from the interconnect 1300) to external layers and elements. For example, connectivity may be provided as described above via the memory bridge 1114 with the storage layer 906, via the fabric bridge 1116 with the fabric 1306, and via the compute bridge 1118 with the external data storage 920 and the analytics engine 910. Other bridges (not shown) may include NVME, PCIe, high-speed, low-speed, high-bandwidth, low-bandwidth, and so forth. The fabric 1306 may provide connectivity internal to the data analytics accelerator 900-1 and, for example, between layers like hardware 904 and storage 906, and between acceleration units, for example between a first acceleration unit 1100-1 to additional acceleration units 1100- N. The fabric 1306 may also provide external connectivity from the data analytics accelerator 900, for example between the first data analytics accelerator 900-1 to additional data analytics accelerators 900-N.
[0100] The data analytics accelerator 900 may use a columnar data structure. The columnar data structure can be provided as input and received as output from elements of the data analytics accelerator 900. In particular, elements of the acceleration units 1100 can be configured to receive input data in the columnar data structure format and generate output data in the columnar data structure format. For example, the selector module 1102 may generate output data in the columnar data structure format that is input by the FPE 1103. Similarly, the interconnect 1300 may receive and transfer columnar data between elements, and the fabric 1306 between acceleration units 1100 and accelerators 900.
[0101] Streaming processing avoids memory bounded operations which can limit communication bandwidth of memory mapped systems. The accelerator processing may include techniques such as columnar processing, that is, processing data while in columnar format to improve processing efficiency and reduce context switching as compared to row-based processing. The accelerator processing may also include techniques such as single instruction multiple data (SIMD) to apply the same processing on multiple data elements, increasing processing speed, facilitating "real-time” or “line-speed” processing of data. The fabric 1306 may facilitate large scale systems implementation.
[0102] Accelerator memory 1200, such as PIM 1202 and HBM 1206 may provide support for high bandwidth random access to memory. Partial processing may produce data output from the data analytics accelerator 900 that may be orders of magnitude less than the original data from storage 920. Thus, facilitating the completion of processing on analytics engine 910 or general- purpose compute with a significantly reduced data scale. Thus, computer performance is improved, for example, increasing processing speeds, decreasing latency, decreasing variation of latency, and reducing power consumption.
[0103] Consistent with the examples described in this disclosure, in some embodiments, a system includes a hardware based, programmable data analytics processor configured to reside between a data storage unit and one or more hosts, wherein the programmable data analytics processor includes: a selector module configured to input a first set of data and, based on a selection indicator, output a first subset of the first set of data; a filter and project module configured to input a second set of data and, based on a function, output an updated second set of data; a join and group module configured to combine data from one or more third data sets into a combined data set; and a communications fabric configured to transfer data between any of the selector module, the filter and project module, and the join and group module. The modules may correspond to the modules discussed above in connection with, for example, FIGS. 8-13.
[0104] In some embodiments, the first set of data has a columnar structure. For example, the first set of data may include one or more data tables. In some embodiments, the second set of data has a columnar structure. For example, the second set of data may include one or more data tables. In some embodiments, the one or more third data sets have a columnar structure. For example, the one or more data sets may include one or more data tables.
[0105] In some embodiments, the second set of data includes the first subset. In some embodiments, the one or more third data sets include the updated second set of data. In some embodiments, the first subset includes a number of values equal to or less than the number of values in the first set of data.
[0106] In some embodiments, the one more third data sets include structured data. For example, the structured data may include table data in column and row format. In some embodiments, the one or more third data sets include one or more tables, and the combined data set includes at least one table based on combining columns from the one or more tables. In some embodiments, the one or more third data sets include one or more tables, and the combined data set includes at least one table based on combining rows from the one or more tables.
[0107] In some embodiments, the selection indicator is based on a previous filter value. In some embodiments, the selection indicator may specify a memory address associated with at least a portion of the first set of data. In some embodiments, the selector module is configured to input the first set of data as a block of data in parallel and use SIMD processing of the block of data to generate the first subset.
[0108] In some embodiments, the filter and project module includes at least one function configured to modify the second set of data. In some embodiments, the filter and projection module is configured to input the second set of data as a block of data in parallel and execute a SIMD processing function of the block of data to generate the second set of data.
[0109] In some embodiments, the join and group module is configured to combine columns from one or more tables. In some embodiments, the join and group module is configured to combine rows from one or more tables. In some embodiments, the modules are configured for line rate processing.
[0110] In some embodiments, the communications fabric is configured to transfer data by streaming the data between modules. Streaming (or stream processing or distributed stream processing) of data may facilitate parallel processing of data transferred to / from any of the modules discussed herein.
[0111] In some embodiments, the programmable data analytics processor is configured to perform at least one of SIMD processing, context switching, and streaming processing. Context switching may include switching from one thread to another thread and may include storing the context of the current thread and restoring the context of another thread.
[0112] Example Implementation - Direct Memory Access (DMA) for Tabular-Columnar Data
[0113] Current DMA drivers are designed and implemented to handle a wide range of data input and output types.
[0114] Writing data, in particular table data, and more particular columnar format table data is typically done using conventional DMA (direct memory access) modules. There is a need to be able to more efficiently handle higher speed data that has been processed and that now needs to be written to memory. In addition, there is a desire to have the output data be in a “zero-copy” format that is immediately usable by follow-on users / processors without the need for further processing / data format conversion.
[0115] The disclosed systems and methods for columnar DMA may provide a relatively faster and more efficient solution for columnar format data, for example table data, as compared with current general-purpose DMA drivers. While the current description uses columnar DMA processing as an exemplary implementation, this is not limiting. The system and method may be used for other implementations such as, for example, determining for each value in the input stream a corresponding category based on each value, and using the determined category for processing the value. For example, the value may include data that is determined / categorized for a row identification (ID), format of data (data format, data structure), and / or data type (content of the data).
[0116] Referring now to the drawings, FIG. 14 is a high-level example of an architecture for a columnar DMA system in the data analytics architecture 900, consistent with the disclosed embodiments. Data analytics architecture 900 may be the same as, or similar to, the data analytics architecture described with respect to FIG. 9. Various additional details regarding the overall system architecture are provided there. The columnar DMA system and method may be implemented as a portion of one or more of the software layer 902, software modules 922, hardware layer 904 hardware modules 924, the storage layer 906, and the storage modules 926. Hardware and firmware implementations may be configured in the hardware layer 904.
[0117] The disclosed embodiments solve the problem of how to efficiently write from processing to memory. For example, data that has been processed by the acceleration unit 1100, possibly in a first format (for example, in a custom, internal format), streaming at high speeds, may need to be output via the bridges 1110 to the external data storage 920, in a second format, for example, a standardized representation, such as the Apache Arrow format.
[0118] FIG. 15 is a high-level general processing flow of the columnar DMA system, consistent with the disclosed embodiments. Input data 1500 may include one or more portions of data, shown as “chunks” or “hunks” designated Hn, where “n” is an integer representing one of the hunks. As noted elsewhere in this description, each hunk may include one or more values. Chronological hunks HO, Hl, H2, H3, H4, and H5 are shown in the current figure. The hunks Hn may be an input stream of data 1502 into the columnar DMA module 1504. The columnar DMA module may be implemented as a columnar DMA driver for various other modules and devices, such as a memory 1508. The columnar DMA module 1504 may output one or more DMA channels, and may output memory mapped output data 1506, shown as (logical) memory 1508 including rows Rn and columns Cn. The memory 1508 may be in a variety of locations, such as in the storage layer 906. In the examples of this description, the memory may be on a host, or analytics engine 910 for use by a user.
[0119] At a high level, to orient the reader, an example of a method for columnar DMA processing may include identifying the tabular column association in an input stream, modifying or encoding the column element data from the internal accelerator representation to a more standardized representation, buffering the modified / encoded data, copying the column data to a host memory, each column being copied to a dedicated location using multi-channel capabilities, and adding metadata per column.
[0120] A feature of some embodiments is that each input hunk of data may have validity (validity data / validity indicator for the corresponding values) and output data may maintain the validity. The columnar DMA module 1504 inputs the validity data in association with input values, maintains and processes the validity as required, and generates output data (memory mapped data 1506) including values and associated validity.
[0121] Another feature of some embodiments is that each DMA channel may be a column ID (column identification) / associated with a column ID. Input data includes an ID associated with each column. The columnar DMA module 1504 may process the column ID, handling each column as a DMA channel. Each channel may be output (written) to a different memory location.
[0122] Another feature of some embodiments is that each DMA channel may perform calculations, for example operations and table-specific calculations. For example, counting validity of input data and aggregating validity data from input data (hunks) to output data. Another example is calculating offsets for the input values, such as variable sized data, and generating corresponding offsets in the output data.
[0123] FIG. 16 is a logical view of input data, consistent with the disclosed embodiments. The input data 1610 is an example of FIG. 15 input data 1500, and includes exemplary fixed 1610-A and variable 1610-B values. Eogical views may be considered as views of a table with rows and columns of data. Several table rows (shown 0 to 7) each have both fixed (“FIXED”) and variable length (“V ARIABEE”) data. An exemplary implementation may include fixed values (DOO, D01, D02, D03, D04, D05, D06, D07) being 8 Bytes each. Exemplary variable length data may include D10 (24 Bytes D10-0, D10-1, D10-2), Dl l (8 bytes Dl l-0), D12 (8 Bytes D12-0), D13 (16 Bytes D13-0, D13-1), D14 (8 Bytes D14-0), D15 (8 Bytes D15-0), D16 (24 Bytes D16-0, D16-2, D16-2), and D17 (8 Bytes D17-0).
[0124] FIG. 17A is an example of values being input to the columnar DMA module 1504, consistent with the disclosed embodiments. Portions of the values 1610 may be taken as hunks of data values 1700 for streaming 1702 into the columnar DMA module 1504. In the current example, rows 0 and 1 are portioned into hunk-0 HO, rows 2 and 3 are portioned into hunk-1 Hl, rows 4 and 5 are portioned into hunk-2 H2, and rows 6 and 7 are portioned into hunk-3 H3. The streaming input data 1702 is input to the columnar DMA module 1504. The columnar DMA module 1504 may receive the streaming input data 1702, perform processing as described in relation to FIG. 18, and output memory mapped data 1706. The data values 1700, streaming input data 1702, and output memory mapped data 1706 may be respectively similar to the data values 1500, streaming input data 1502, and output memory mapped data 1506.
[0125] FIG. 17B shows exemplary memory data views, consistent with the disclosed embodiments. Examples of the memory 1508 include fixed sized (fixed length of data) memory data view 1712-A and variable sized (the length of each data and / or each value may vary from each other) memory data view 1712-B. In this case, the input data 1610 has been processed and output as exemplary fixed 1712-A and variable 1712-B data. In FIG. 17B, additional columns that are not part of the stored memory, are shown for simplicity and reference. A simplified address (ADDR.) column is shown, as will be apparent to one skilled in the art (for example, as compared to typical 32-bit or 64-bit hexadecimal addressing). A description (DESCR.) column may assist the reader with identifying the type of data being stored at each memory address. Descriptions shown include header data (HEADER), validity data (VALIDITY), offsets (OFFSET), and values (VALUE). Data stored in the memory is shown in the data (DATA) column. The exemplary data includes the following: “Hx” represents header data, with header data for the fixed data shown as “HA” and header data for the variable data shown as “HB”. Validity is represented by a “1” (valid) or “0” (invalid) for the corresponding data. Offsets are shown as “Fn” where the number “n” is the offset corresponding to the value “n”. Values are shown as “Dxn-y”, where “x” is a set of values (“0” is the fixed data and “1” is the variable length data), “n” is a value, and “-y” is a portion of the value.
[0126] FIG. 18 is a flowchart of a method for columnar DMA processing, consistent with the disclosed embodiments. In a non-limiting example, the previous figures’ exemplary configuration and data will be used in the current method description. When data, such as the input stream 1502 is ready to be processed, the method begins with initialization (initialize) 1802. After initialization 1802, processing may include both global operations 1810 and per channel operations 1820 to generate an output batch.
[0127] The global operations 1810 may be generally performed on the entire input stream 1502. Global operations 1810 may include (at block 1814) receiving the input stream (1502, 1702), in this case a stream of hunks (1500, 1700). The input stream of data (1502, 1702) may be generally processed to determine, for each value, a corresponding category based on each respective value. The processing (1810, 1820) may be done by the columnar DMA module 1504, and the channel processing modules (1822) may be configured internal (as a part of) or external (implemented separately from) the columnar DMA module 1504. Each value may be processed based on the corresponding category. For example, in some embodiments, all values for each category may be sent to the same channel processing module 1822.
[0128] In the current example, the stream of data may be processed to determine and select in block 1816 a corresponding DMA channel for each of the hunks or portions. The data may be tabular data, and the tabular data values may include a column ID. The DMA channel may be the column ID. That is, a DMA channel may be associated with each of the column IDs. Hence, this block may select the column ID corresponding (associated with) each value (or each hunk). Based on the corresponding DMA channel, each value, may be sent for per channel operations 1820.
[0129] A feature of some embodiments is that each DMA channel may be a column ID (column identification) / associated with a column ID. Input data includes an ID associated with each column. The columnar DMA module 1504 may process the column ID, handling each column as a DMA channel. Each channel may be output (written) to a different memory location. Per channel operations 1820 may be implemented by one or more channel processing modules 1822. In the current figure a single channel processing module (exemplary channel- 1 1822-1) is shown. Per channel processing may be done in parallel. Based on the corresponding DMA channel, each value (or each hunk) may be sent to a corresponding channel processing module 1822. For example, a value with associated column ID “1” is sent to channel-1 1822-1 processing module, and a value with associated column ID “2” is sent to channel-2 1822-2 processing module (not shown). Each channel processing module 1822 may maintain a channel state in block 1824 that can include information for the channel (for example, common for the entire channel), such as destination address(es), a data elements counter, and a null counter.
[0130] Each input hunk may be processed as appropriate for each hunk. Processing may include block 1826 accumulating one or more values from the hunk. The values may be accumulated by, for example, writing the values to a values buffer. For example, in the variable memory view 1712-B writing values Dln-y to a buffer (a values buffer). If the hunk has variable length values, then in block 1827, the process includes accumulating one or more offsets to a buffer (for example writing the offsets to an offsets buffer) with each offset corresponding to one or more variable length values. For example, the variable memory view 1712-B offsets Fn are buffered for corresponding values Din. In block 1828, accumulation is done of validity. The accumulation may be in a validity vector. For example, the variable memory view 1712-B validity data. For example, the validity (“1” one or “0” zero) at memory addresses 2 to 9 (8 validity indicators), corresponding to values Din (where n is “0” zero to “7” seven, 8 data values in memory addresses 18 to 30).
[0131] Another feature of some embodiments is that each input hunk of data may have validity (one or more validity data for the corresponding values) and output data may maintain the validity. The columnar DMA module 1504 may input the validity data in association with input values, maintain, accumulate, and process the validity as needed or requested, and generate output data (memory mapped data 1506) including values and associated validity.
[0132] In block 1829, counters are updated as appropriate. For example, block 1829 may include collecting metadata regarding the columnar DMA processing, collecting metadata regarding operation of the channel processing module, counting values, counting validity, etc.
[0133] The blocks (such as blocks 1826, 1827, and 1828) may be implemented as modules in software, firmware, or hardware. Blocks may accumulate continuously, accumulating and buffering data as the data is available from the channel processing and / or data stream. Each block / module may operate independently from the other blocks, accumulating and buffering data as appropriate. Each of the blocks may write data as necessary, for example, when a buffer is full then writing the contents of the buffer to the batch (to a memory). The update counters block 1829 may be updated only once when processing of a batch is finished, and then written to the batch.
[0134] Another feature of some embodiments is that each channel processing module 1822 (also referred to in the context of this description as each DMA channel) may perform table-specific calculations. For example, some embodiments may include counting validity of input data and aggregating validity data from input data (hunks) to output data. Another example is calculating offsets for the input values, such as variable sized data, and generating corresponding offsets in the output data. These table-specific calculations may be done by individual blocks, for example 1826, 1827, and / or 1828. Alternatively, a single block such as 1829 may be configured to perform one or more calculations.
[0135] In block 1830, a check is performed to determine if there are more hunks to process and if there is more space available in the current output batch. For example, the check may include checking whether there is more data 1700 and the batch is not filled. If there is more data to process, and there is more space available in the current output, then the method loops back (“YES”) to block 1814 to receive more data, and the method continue there. In general, in the current block, and optionally other blocks such as block 1838, pre-defined parameters may be used to determine if and when to buffer values, stop buffering values, write output data, etc. For example, each channel processing module 1822 may process and buffer values until a predefined amount (for example Bytes) of data has been buffered. For example, the pre-defined values may include an optimal amount, range, or maximum amount of data for output and handling by a DMA channel. In another example, the pre-defined values can be based on the category, for example, an optimal amount of data to save for a particular data format. In another example, the pre-defined values can be based on the operational parameters of the memory to be written to. The buffered values may be output as a batch via the associated DMA channel, and buffering continues with subsequently input values.
[0136] In block 1830, if there are not more hunks to process, or if there is not more space available in the current output, then the method continues (“NO”) to complete processing of the current batch. In block 1832, the accumulated validity is written. For example, if the accumulation was done in a validity vector, the validity vector is written to the memory for the batch. In block 1834, the column header is written. Other data is written as appropriate (values, offsets, etc.). For example, the variable memory view 1712-B header data HB. In block 1836, when writing of the batch has been completed, a signal can be initiated to the appropriate follow- on module that the batch is available. For completeness, in block 1840, the values are written, however, as noted elsewhere in this description, values may be written as necessary. For example, values may be written to the batch (for example from the accumulated values buffer) one or more times during processing and / or when block 1830 is “NO”.
[0137] After signaling a batch is available in block 1836, then in block 1838, if there is more input data (more hunks), the method can loop back and repeat at block 1814. If there is no more data to process (no more hunks), the method can finish 1842 by performing any remaining processing, closing the batches, performing any post-processing, cleanup, etc.
[0138] Although not shown in the current figure, optional and residual processing may be done at or between any of the processing stages.
[0139] FIG. 19 is a high-level partial block diagram of an exemplary system 1900 configured to implement the columnar DMA (direct memory access) processing module, and other modules, of the disclosed embodiments. The disclosed embodiments may include but are not limited to the columnar DMA module 1504 and the channel processing modules 1822.
[0140] For example, system (processing system) 1900 may include a processor 1902 (one or more) and four exemplary memory devices: a RAM 1904, a boot ROM 1906, a mass storage device (hard disk) 1908, and a flash memory 1910, all communicating via a common bus 1912. As is known in the art, processing and memory can include any computer readable medium storing software and / or firmware and / or any hardware element(s) including but not limited to field programmable logic array (FPLA) element(s), hard-wired logic element(s), field programmable gate array (FPGA) element(s), and application-specific integrated circuit (ASIC) element(s). Any instruction set architecture may be used in processor 1902 including but not limited to reduced instruction set computer (RISC) architecture and / or complex instruction set computer (CISC) architecture. A module (processing module) 1914 is shown on mass storage 1908, but as will be apparent to one skilled in the art, could be located on any of the memory devices.
[0141] Mass storage device 1908 is a non-limiting example of a non-transitory computer- readable storage medium bearing computer-readable code for implementing the data processing, formatting, and handling methodology described herein. Other examples of such computer- readable storage media include read-only memories such as CDs bearing or storing such code.
[0142] System 1900 may have an operating system stored on the memory devices, the ROM may include boot code for the system, and the processor may be configured for executing the boot code to load the operating system to RAM 1904, executing the operating system to copy computer-readable code to RAM 1904 and execute the code.
[0143] Network connection 1920 provides communications to and from system 1900. Typically, a single network connection provides one or more links, including virtual connections, to other devices on local and / or remote networks. Alternatively, system 1900 can include more than one network connection (not shown), each network connection providing one or more links to other devices and / or networks.
[0144] System 1900 can be implemented as a server or client respectively connected through a network to a client or server.
[0145] To the extent that the appended claims have been drafted without multiple dependencies, this has been done only to accommodate formal requirements in jurisdictions that do not allow such multiple dependencies. Note that all possible combinations of features that would be implied by rendering the claims multiply dependent are explicitly envisaged and should be considered part of the disclosed embodiments.
[0146] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
[0147] As used herein, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise.
[0148] The word “exemplary” is used herein to mean “serving as an example, instance or illustration”. Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments and / or to exclude the incorporation of features from other embodiments.
[0149] It is appreciated that certain features of the disclosed embodiments, which are, for clarity, described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the disclosed embodiments, which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable sub-combination or as suitable in any other described embodiment of the disclosed embodiments. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
[0150] The foregoing description has been presented for purposes of illustration. It is not exhaustive and is not limited to the precise forms or embodiments disclosed. Modifications and adaptations will be apparent to those skilled in the art from consideration of the specification and practice of the disclosed embodiments. Additionally, although aspects of the disclosed embodiments are described as being stored in memory, one skilled in the art will appreciate that these aspects can also be stored on other types of computer readable media, such as secondary storage devices, for example, hard disks or CD ROM, or other forms of RAM or ROM, USB media, DVD, Blu-ray, 4K Ultra HD Blu-ray, or other optical drive media.
[0151] Computer programs based on the written description and disclosed methods are within the skill of an experienced developer. The various programs or program modules can be created using any of the techniques known to one skilled in the art or can be designed in connection with existing software. For example, program sections or program modules can be designed in or by means of .Net Framework, .Net Compact Framework (and related languages, such as Visual Basic, C, etc.), Java, C++, Objective-C, HTML, HTML / AJAX combinations, XML, or HTML with included Java applets.
[0152] Moreover, while illustrative embodiments have been described herein, the scope of any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., of aspects across various embodiments), adaptations and / or alterations as would be appreciated by those skilled in the art based on the present disclosure. The limitations in the claims are to be interpreted broadly based on the language employed in the claims and not limited to examples described in the present specification or during the prosecution of the application. The examples are to be construed as non-exclusive. Furthermore, the steps of the disclosed methods may be modified in any manner, including by reordering steps and / or inserting or deleting steps. It is intended, therefore, that the specification and examples be considered as illustrative only, with a true scope and spirit being indicated by the following claims and their full scope of equivalents.
Claims
WHAT IS CLAIMED IS:
1. A method for columnar direct memory access processing, the method comprising: receiving an input stream of data including one or more values, each of the values including a column identification; determining, for each value, based on the column identification, a corresponding direct memory access channel; and processing each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location.
2. A method for processing data, the method comprising: receiving an input stream of data including one or more values; determining, for each value, a corresponding category based on each respective value, and processing each of the values based on the corresponding category.
3. The method of claim 2, wherein the input stream includes a plurality of portions of data, each portion including one or more of the values.
4. The method of claim 2, wherein the values are tabular data.
5. The method of claim 4, wherein each of the values includes a column identification.
6. The method of claim 5, wherein a direct memory access channel is associated with each of the column identifications.
7. The method of claim 6, wherein values for each of the column identifiers are routed to the associated direct memory access channel.
8. The method of claim 7, wherein each of the associated direct memory access channels writes to a different associated memory location.
9. The method of claim 5, wherein the determining includes determining the corresponding category based on the column identification.
10. The method of claim 5, wherein the processing includes routing each of the values based on the column identification.
11. The method of claim 10, wherein the routing is to one of a plurality of channel processing modules.
12. The method of claim 11, wherein each of the plurality of channel processing modules is associated with a direct memory access channel.
13. The method of claim 12 wherein a plurality of values are sent to each of the plurality of channel processing modules, buffered, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel.
14. The method of claim 2, wherein the processing further includes modifying one or more values from a first format to output data in a second format.
15. The method of claim 14, wherein the second format is a standardized representation.
16. The method of claim 15, wherein the standardized representation is the Apache Arrow format.
17. The method of claim 2, wherein the processing further includes calculations on the values for each corresponding category.
18. The method of claim 17, wherein each corresponding category is associated with a direct memory access channel, and the calculations include operations for the values of each of the associated direct memory access channels.
19. The method of claim 17, wherein the calculations include table- specific calculations.
20. The method of claim 17, wherein the calculations include operations selected from the group comprising of: counting validity of input data; counting validity of values; aggregating validity data for each corresponding category; aggregating validity data from each portion of data (hunk) of the input stream, each portion of data including one or more of the values; calculating offsets for values in the input stream of data; calculating offsets for values of variable sized data in the input stream of data; and generating offsets for output data based on the input data stream values.
21. The method of claim 2 wherein the processing further includes maintaining validity of the input stream values.
22. The method of claim 21, wherein maintaining validity includes processing each of the values to determine if the value has an associated validity indicator, maintaining association of the validity indicator with the value during the processing, and generating output data, the output data based on the value and having an output validity indicator.
23. A non-transitory computer-readable medium having stored thereon computer- readable instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: receiving an input stream of data including one or more values, each of the values including a column identification; determining, for each value, based on the column identification, a corresponding direct memory access channel; and processing each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location.
24. A non-transitory computer-readable medium having stored thereon computer- readable instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: receiving an input stream of data including one or more values; determining, for each value, a corresponding category based on each respective value, and processing each of the values based on the corresponding category.
25. A system for processing data, the system comprising: an input module configured to receive an input stream of data including one or more values, each of the values including a column identification; a processing module including at least one processor including circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: determine, for each value, based on the column identification a corresponding direct memory access channel; and process each of the values to buffer the values for each direct memory access channel, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel to a different associated memory location.
26. A system for processing data, the system comprising: an input module configured to receive an input stream of data including one or more values;a processing module including at least one processor including circuitry and a memory, wherein the memory includes instructions that when executed by the circuitry cause the at least one processor to: determine, for each value, a corresponding category based on each respective value and process each of the values based on the corresponding category.
27. The system of claim 26 wherein the input stream includes a plurality of portions of data, each portion including one or more of the values.
28. The system of claim 26 wherein the values are tabular data, each of the values includes a column identification, and a direct memory access channel is associated with each of the column identifications.
29. The system of claim 28 wherein the processing module is further configured for routing values for each of the column identifiers to the associated direct memory access channel.
30. The system of claim 29 wherein the processing module is further configured for writing each of the associated direct memory access channels to a different associated memory location.
31. The system of claim 28 further comprising: a plurality of channel processing modules, the processing module is further configured for routing each of the values based on the column identification, to a corresponding channel processing module.
32. The system of claim 31 wherein the processing module is further configured for routing a plurality of values to each of the plurality of channel processing modules, and each of the plurality of channel processing modules is further configured to: receive the plurality of values, buffer the plurality of values, and then based on pre-defined parameters the buffered values are output as a batch via the associated direct memory access channel.
33. The system of claim 26 wherein the processing module is further configured for modifying one or more values from a first format to output data in a second format.
34. The system of claim 33 wherein the second format is a standardized representation, and the standardized representation is the Apache Arrow format.
35. The system of claim 26 wherein the processing module is further configured for performing calculations on the values for each corresponding category.
36. The system of claim 35 wherein each corresponding category is associated with a direct memory access channel, and the calculations include operations for the values of each of the associated direct memory access channels.
37. The system of claim 35 wherein the calculations include table-specific calculations.
38. The system of claim 26 wherein the processing module is further configured for maintaining validity of the input stream values.
39. The system of claim 38 wherein maintaining validity includes processing each of the values to determine if the value has an associated validity indicator, maintaining association of the validity indicator with the value during the processing, and generating output data, the output data based on the value and having an output validity indicator.
Citation Information
Patent Citations
Flexible wheel for harmonic speed reducer and manufacturing method thereof
WO2018000995A1
Method and device for adjusting brightness
WO2019001005A1
Image processing method, device and apparatus, and storage medium
WO2020000665A1
Memory appliances for memory intensive operations
WO2022082115A1
Pattern matching across multiple input data streams
US20150161214A1