System and method for data analysis based on zero-extraction conversion loading
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-29
AI Technical Summary
Existing ETL data analysis processes suffer from low data transmission and processing efficiency, resulting in high resource requirements, increased total cost of ownership, and a lack of flexibility and parallelism.
By offloading ETL functionality from storage devices and utilizing devices such as compute storage drives and smart SSDs for data processing, combined with distributed clusters and RDMA technology, parallel transformation and processing of data at storage locations can be achieved, reducing data transfer, improving performance, and lowering resource requirements.
It improves the performance and flexibility of ETL processing, reduces the total cost of ownership, achieves greater parallelism and scalability, reduces data transfer steps, and improves processing efficiency.
Smart Images

Figure CN122122571A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 546,740, filed October 31, 2023, which is incorporated herein by reference for all purposes. Technical Field
[0003] This disclosure generally relates to memory systems, and more specifically to systems and methods for data analysis based on extract-transform load (ETL). Background Technology
[0004] This background section is intended to provide context only, and the disclosure of any concepts in this section does not constitute an admission that the concepts are prior art.
[0005] Data analytics is the process of using tools, techniques, and methods to analyze raw data to extract insights and make decisions. It encompasses a multidisciplinary field that uses a variety of tools and techniques, including those from mathematics, statistics, and computer science, to analyze data. Data analytics can be used to improve decision-making, streamline operations, and increase revenue. It can be used to enhance processes, optimize products and services, increase productivity, identify and leverage sources of competitive advantage, and mitigate risk. Data analytics can be used for a wide range of business purposes, including marketing, employee metrics, inventory metrics, delivery logistics, and more. Data analytics can involve collecting, cleaning, and analyzing datasets to solve problems. Summary of the Invention
[0006] In various embodiments, the systems and methods described herein include systems, methods, and apparatuses for data analysis based on zero-extract-transform-load (ETL). In some aspects, the technology described herein relates to a data analysis method comprising: generating a data request for at least a subset of data stored in a database; converting the data request into an object storage request; determining, based on parsing the object storage request, that the object storage request includes an extract-transform-load (ETL) request; creating an ETL command based on the object storage request; routing the ETL command to a storage device based on the storage device including at least one ETL function requested in the data request; and determining information based on result data received from the storage device in response to routing the ETL command to the storage device.
[0007] In some respects, the techniques described herein relate to a method in which the resulting data is generated based on a subset of data transformed at the storage device according to ETL functionality.
[0008] In some respects, the techniques described herein relate to a method in which a deterministic object storage request includes an ETL request, and the ETL command is configured as a representational state transfer (REST) command.
[0009] In some respects, the techniques described herein relate to a method in which an ETL command includes the name of a storage node containing a storage device.
[0010] In some respects, the techniques described herein relate to a method that further includes: determining that a data request is part of a batch data request; and combining the data request with at least the second request based on the fact that the data request and the second request are associated with the same host server.
[0011] In some respects, the technology described herein relates to a method that further includes: determining that a second storage device is associated with a processing load imbalance; and, based on the processing load imbalance, allocating at least one of an ETL function of the second storage device or a second ETL command assigned to the second storage device from the second storage device to the storage device.
[0012] In some respects, the techniques described herein relate to a method in which a storage device includes at least one of a storage server or a computing storage drive, the computing storage drive including one or more processors configured to perform a conversion function.
[0013] In some respects, the techniques described herein relate to a method in which a data request includes at least one of the following: key-value (KV) fetch or a filtering criterion for narrowing data in a database to a subset of data, and the database includes a KV database.
[0014] In some respects, the techniques described herein relate to a method in which a data request specifies an identifier for the data request.
[0015] In some respects, the techniques described herein relate to a method in which object storage requests include ETL identifiers.
[0016] In some aspects, the techniques described herein relate to a data analysis method comprising: receiving an Extract Transform Load (ETL) command at a storage node, the ETL command including a data request for a subset of data stored in a database; reading the subset of data from the database at the storage node; providing the subset of data to a storage device of the storage node for processing the subset of data at the storage device; processing the subset of data at the storage device based on the storage device including at least one ETL function requested in the data request; and calculating resulting data based on the processing.
[0017] In some aspects, the techniques described herein relate to a method that further includes at least one of providing result data to a client device, an application of the client device, a storage device for storing the result data, a different storage device, or a function for requesting result data, wherein the client device provides a Representational State Transfer (REST) command associated with an ETL command.
[0018] In some aspects, the techniques described herein relate to a method that further includes transmitting resulting data to the client device providing the ETL command, based on ETL commands including direct information instructing the client device to include RDMA data from a network interface card supporting Remote Direct Memory Access (RDMA).
[0019] In some aspects, the techniques described herein relate to a method that also includes identifying a second data request based on an ETL command being configured as a batch command, the ETL command including a second data request for a second subset of data stored in a database.
[0020] In some respects, the techniques described herein relate to a method that further includes processing a second subset of data at a second storage device, based on at least a second ETL function requested in a second data request.
[0021] In some respects, the techniques described herein relate to a method in which a storage device includes at least one of a storage server or a computing storage drive, the computing storage drive including one or more processors for processing a subset of data at the storage device.
[0022] In some respects, the techniques described herein relate to a method in which a data request includes key-value (KV) retrieval and the database includes a KV database.
[0023] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing code including instructions executable by a processor to: generate a data request for at least a subset of data stored in a database; convert the data request into an object storage request; determine, based on parsing the object storage request, that the object storage request includes an Extract Transform Load (ETL) request; create an ETL command based on the object storage request; route the ETL command to a storage device based on the storage device including at least one ETL function requested in the data request; and determine information based on result data received from the storage device in response to routing the ETL command to the storage device.
[0024] In some respects, the techniques described herein relate to a non-transitory computer-readable medium in which the resulting data is generated based on a subset of data transformed at a storage device according to an ETL function.
[0025] In some respects, the techniques described herein relate to a non-transitory computer-readable medium in which a deterministic object storage request includes an ETL request, and the ETL command is configured as a Representational State Transfer (REST) command.
[0026] A computer-readable medium is disclosed. This medium can store instructions that, when executed by a computer, cause the computer to perform operations substantially the same as or similar to those further disclosed herein. Similarly, non-transitory computer-readable media, devices, and systems for performing operations substantially the same as or similar to those described herein are further disclosed.
[0027] The systems and methods described herein offer several advantages and benefits. For example, they provide distributed cluster-level offloading. These systems and methods integrate inter-node communication, data distribution, and the aggregation of ETL functionality into a given framework. Based on the described systems and methods, including offloading ETL functionality to storage servers, there is no data transfer or minimal data transfer, leading to improved performance and reduced total cost of ownership due to lower overall resource requirements. Some solutions require reading the entire object into compute nodes before processing. The disclosed zero-ETL (zETL) systems and methods avoid reading the entire object into compute nodes by running on chunked data and processing near the location where the data is stored, and storing the results in the requested format, all of which avoids additional data copying. The zETL systems and methods disclosed herein can be extended to offload transformation functionality or any computation in an ETL pipeline to a compute-capable storage device (e.g., compute storage drives, Smart-SSDs, CXL-HC devices, etc.). The systems and methods can be configured to process ETL through data formatting, distributed processing to ensure parallel offloading, and can handle several data formats for storing results. Based on this system and method, computing resources can be automatically scaled. As more storage is added, the offloaded capacity can increase, thereby increasing overall computing power. The system and method disclosed herein can provide higher performance with greater parallelism and ETL capabilities. When data is distributed across storage clusters, offloading can occur in parallel, thus accelerating the ETL pipeline. Some publicly available solutions transfer the entire data to compute nodes and execute functions, while this system and method run functions in storage devices (e.g., compute storage drives, storage servers) without requiring data transfer. Because the zETL adapter can handle zETL-specific API communications, this system and method enable applications to use the zETL framework described herein without any changes to the pipeline code. Some ETL solutions may require the use of specific APIs to write to the pipeline without any user control over where data is processed. This system and method implement a hybrid model of ETL, where at least some functions are offloaded to storage servers, at least some functions are offloaded to storage devices capable of computation, and / or at least some computations are run on client servers, thus providing a flexible and customizable solution. Attached Figure Description
[0028] The foregoing and other aspects of this system and method will be better understood when this application is read with reference to the following accompanying drawings, wherein like numerals denote similar or identical elements. Furthermore, the drawings provided herein are for illustrative purposes only; other embodiments, which may not be explicitly shown, are not excluded from the scope of this disclosure.
[0029] These and other features and advantages of this disclosure will be appreciated and understood by referring to the specification, claims and drawings, wherein:
[0030] Figure 1 An example system according to one or more embodiments described herein is shown.
[0031] Figure 2 One or more embodiments described herein are shown. Figure 1 The details of the system.
[0032] Figure 3 An example system according to one or more embodiments described herein is shown.
[0033] Figure 4 An example system flow according to one or more embodiments described herein is shown.
[0034] Figure 5 An example system according to one or more embodiments described herein is shown.
[0035] Figure 6 An example system according to one or more embodiments described herein is shown.
[0036] Figure 7 An example system flow according to one or more embodiments described herein is shown.
[0037] Figure 8 An example system according to one or more embodiments described herein is shown.
[0038] Figure 9 An example system according to one or more embodiments described herein is shown.
[0039] Figure 10 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0040] Figure 11 A flowchart illustrating an example method associated with the disclosed system according to an example implementation described herein is depicted.
[0041] Figure 12 An example system according to one or more embodiments described herein is shown.
[0042] While the system and method are readily adaptable to various modifications and alternatives, specific embodiments thereof are illustrated by way of example in the accompanying drawings and will be described herein. The drawings may not be drawn to scale. However, it should be understood that the drawings and their detailed description are not intended to limit the system and method to the specific forms disclosed, but rather, the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the system and method as defined by the appended claims. Detailed Implementation
[0043] Details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims.
[0044] Various embodiments of this disclosure will now be described more fully below with reference to the accompanying drawings, which illustrate some, but not all, of the embodiments. In fact, this disclosure may be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided to enable this disclosure to meet applicable legal requirements. Unless otherwise stated, the term “or” is used herein in a substitute and connective sense. The terms “illustrative” and “example” are used for examples without an indication of quality level. The same reference numerals always denote the same elements. Arrows in each figure depict bidirectional data flow and / or bidirectional data flow capabilities. The terms “path,” “pathway,” and “route” are used interchangeably herein.
[0045] Embodiments of this disclosure can be implemented in various ways, including as an article of manufacture of a computer program. A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program components, scripts, source code, program code, object code, bytecode, compiled code, interpreted code, machine code, executable instructions, etc. (also referred to herein as executable instructions, instructions for execution, computer program product, program code, and / or similar terms used interchangeably herein). Such non-transitory computer-readable storage medium includes all computer-readable media (including volatile and non-volatile media).
[0046] In one embodiment, a non-volatile computer-readable storage medium may include a floppy disk, flexible disk, hard disk, solid-state storage device (SSS) (e.g., solid-state drive (SSD)), solid-state card (SSC), solid-state module (SSM), enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium. Non-volatile computer-readable storage media may include punched cards, paper tape, optical marking sheets (or any other physical medium having a perforated pattern or other optically identifiable markings), optical disc read-only memory (CD-ROM), rewritable optical disc (CD-RW), digital versatile optical disc (DVD), Blu-ray disc (BD), or any other non-transitory optical medium. Such non-volatile computer-readable storage media may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., serial, NAND, NOR, etc.), multimedia memory card (MMC), secure digital storage (SD) card, smart media card, compressed flash memory (CF) card, memory stick, etc. In addition, non-volatile computer-readable storage media may include conductive bridged random access memory (CBRAM), phase change random access memory (PRAM), ferroelectric random access memory (FeRAM), non-volatile random access memory (NVRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), silicon-oxide-nitride-oxide-silicon memory (SONOS), floating junction gate random access memory (FJG RAM), millipede memory, racetrack memory, etc.
[0047] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data output dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type 2 synchronous dynamic random access memory (DDR2 SDRAM), double data rate type 3 synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), dual transistor RAM (TTRAM), thyristor RAM (T-RAM), zero capacitor (Z-RAM), Rambus through-hole memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, etc. It should be understood that, where embodiments are described as using computer-readable storage media, other types of computer-readable storage media may be used in place of the aforementioned computer-readable storage media or in addition to the aforementioned computer-readable storage media.
[0048] It should be understood that various embodiments of this disclosure can be implemented as methods, apparatus, systems, computing devices, computing entities, etc. Thus, embodiments of this disclosure can take the form of apparatuses, systems, computing devices, computing entities, etc., that execute instructions stored on a computer-readable storage medium to perform certain steps or operations. Therefore, embodiments of this disclosure can take the form of entirely hardware embodiments, entirely computer program product embodiments, and / or embodiments including a combination of computer program products and hardware that perform certain steps or operations.
[0049] Embodiments of this disclosure are described below with reference to block diagrams and flowcharts. Therefore, it should be understood that each block in the block diagrams and flowcharts can be implemented as a computer program product, a complete hardware embodiment, a combination of hardware and computer program products, and / or an apparatus, system, computing device, computing entity, etc., which executes instructions, operations, steps, and interchangeable similar terms (e.g., executable instructions, instructions for execution, program code, etc.) on a computer-readable storage medium for execution. For example, code retrieval, loading, and execution can be performed sequentially, such that one instruction is retrieved, loaded, and executed at a time. In some example embodiments, retrieval, loading, and / or execution can be performed in parallel, such that multiple instructions are retrieved, loaded, and / or executed together. Thus, such embodiments can produce machines specifically configured to perform the steps or operations specified in the block diagrams and flowcharts. Therefore, the block diagrams and flowcharts support various combinations of embodiments for performing specified instructions, operations, or steps.
[0050] Throughout this specification, references to "an embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment may include in at least one embodiment disclosed herein. Therefore, the phrases "in one embodiment," "in an embodiment," or "according to an embodiment" (or other phrases with similar meanings) appearing in various places throughout this specification may not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" should not be construed as necessarily preferred or advantageous over other embodiments. Additionally, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Moreover, depending on the context discussed herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may occasionally be used interchangeably with their non-hyphenated versions (e.g., "two dimensional", "predetermined", "pixel specific", etc.), and uppercase entries (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with their non-uppercase versions (e.g., "counter clock", "row select", "pixout", etc.). This occasional interchangeability should not be considered inconsistent with each other.
[0051] Furthermore, depending on the context discussed herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. It should also be noted that the various figures (including component diagrams) shown and discussed herein are for illustrative purposes only and are not drawn to scale. Similarly, various waveforms and timing diagrams are shown for illustrative purposes only. For example, the dimensions of some components may be enlarged relative to others for clarity. Additionally, reference numerals are repeated in the figures where appropriate to indicate corresponding and / or similar components.
[0052] The terminology used herein is for the purpose of describing some exemplary embodiments only and is not intended to limit the claimed subject matter. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of the stated features, numbers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, numbers, steps, operations, elements, components, and / or combinations thereof.
[0053] It will be understood that when an element or layer is referred to as being on, "connected to," or "coupled to" another element or layer, that element or layer may be directly on, directly connected to, or directly coupled to the other element or layer, or there may be intermediate elements or layers present. Conversely, when an element is referred to as being "directly on," "directly connected to," or "directly coupled to" another element or layer, there are no intermediate elements or layers present. The same reference numerals always denote the same elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0054] As used herein, the terms “first,” “second,” etc., serve as labels for nouns that follow them and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such. Furthermore, the same reference numerals may be used in two or more figures to refer to components, parts, blocks, circuits, units, or modules having the same or similar functions. However, this usage is merely for simplification and ease of discussion; it does not imply that the construction or architectural details of these components or units are identical in all embodiments, or that these commonly referenced components / modules are the only way to implement some of the exemplary embodiments disclosed herein.
[0055] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the same meaning as their meaning in the context of the relevant field, and shall not be interpreted in an idealized or overly formal sense unless expressly defined herein.
[0056] As used herein, the term "module" means any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with modules. For example, software may be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any implementation described herein may, for example, individually or in any combination, include components, hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware storing instructions executed by programmable circuitry. Modules may be embodied collectively or individually as circuitry forming part of a larger system, such as, but not limited to, integrated circuits (ICs), system-on-a-chip (SoCs), components, etc.
[0057] The described systems and methods may be based on and / or include Extract-Transform-Load (ETL) processes. ETL can include the process of moving data from a source system to a target system (typically a data warehouse). This process may involve extracting data from the source system, transforming it to be more suitable for analysis, and loading it into the target system. ETL triggers can include scheduled triggers (e.g., automatically executing an ETL pipeline at a specific time each day); event triggers (e.g., monitoring for new data arrivals from a data source and triggering an ETL pipeline); data-driven triggers (e.g., triggering an ETL pipeline when data arrives at different partitions or folders based on certain criteria); manual triggers (e.g., manually starting an ETL pipeline for ad-hoc data processing tasks); lambda functions (e.g., subscribing to a Simple Queue Service (SQS) queue to start an ETL pipeline when new files are processed); and / or glue triggers (e.g., manually or automatically starting an ETL job or crawler). ETL can provide a comprehensive view of the data for in-depth analysis and reporting. Managing multiple datasets requires time and coordination and can lead to inefficiencies and delays. ETL can combine databases and various forms of data into a single, cohesive view. The systems and methods described may include performing ETL operations. However, instead of performing ETL operations separately from storage, the ETL operations described herein can be performed on storage devices (e.g., where ETL data may be stored).
[0058] In ETL, transformation steps prepare data for analysis by changing its structure or format to match a target system. Transformation functions can include cleaning (e.g., removing duplicates, filling in missing values); reshaping (e.g., converting currencies, pivot tables); calculation (e.g., calculating new dimensions and measures); filtering (e.g., selecting subsets of data based on criteria); sorting (e.g., sorting data); combining (e.g., combining data); aggregating (e.g., combining data); splitting (e.g., partitioning data); validation (e.g., ensuring data quality); authentication (e.g., ensuring data quality); auditing (e.g., ensuring data quality and compliance); masking, hashing, or removing (e.g., protecting data subject to regulations); formatting (e.g., matching the schema of a target data warehouse), and more. Transformation steps in ETL improve data integrity, standardize data, and prepare it for effective analysis and decision-making. Therefore, ETL can help companies gain deeper insights into their operations, processes, customers, and more.
[0059] ETL can be implemented in a variety of ways. For example, a financial institution might have customer information in several departments, and each department might list that customer's information in a different way. The membership department might list customers by name, while the accounting department might list them by number. ETL can bundle all these data elements and merge them into a unified representation, such as for storage in a database or data warehouse. Another way companies can use ETL is to move information to another application. For example, the new application might use a different database vendor and a different database schema. ETL can be used to transform the data into a format suitable for the new application. Another example could include expense and cost recovery systems used by accountants, consultants, etc. The data might eventually appear in time and billing systems, although some businesses might use the raw data for employee productivity reports in human resources and / or equipment usage reports in facilities management.
[0060] The described systems and methods may be based on and / or include block storage, object storage, file storage, etc. In the case of block storage, a block is a data block, and data blocks (chunks) can be combined to create a file. A block has an address, and an application retrieves a block by calling that address. Similar to file storage, object storage is used for unstructured data, while block storage can be used for structured data, such as information within a database. Object storage stores a file as an object by dividing it into multiple objects (e.g., objects of a set size) and storing the objects. Unlike files and file systems, objects can be stored in a flat structure. Each object includes data, metadata, and a unique identifier, which applications can use for easy access and retrieval. There are no folders or orientations in object storage, which makes data retrieval easy because an exact location is not required. From a pool of objects, a given object can be retrieved by presenting its object ID. Objects can be local or geographically distributed, but because they are in a flat address space, they are retrieved in the same way (e.g., by object ID). The described systems and methods may include storing data in data blocks (e.g., 1 MB data blocks).
[0061] The described systems and methods may be based on and / or include key-value (KV) databases or KV stores. KV databases may include data storage paradigms for storing KV pairs for high-performance read and write operations at the edge. KV databases may be designed to store, retrieve, and manage associative arrays. KV databases may use hash tables to store unique keys and their corresponding data values. KV pairs may include data structures that associate keys with values. Keys may include constants defining a dataset (e.g., species, color), and values may include variables belonging to the dataset (e.g., human, green). Thus, a KV pair may include "species = human" or "color = green". A KV request may include a request made to a KV database for data (e.g., KV pairs) stored in the KV database. The described systems and methods may include retrieving data comprising one or more KV pairs.
[0062] Requesting data from a database (e.g., a key-value database) can be based on one or more filters. Filtering data from a database can include criteria for narrowing down a large dataset to display only the most relevant information (e.g., a subset of data). Filters can be used to display specific records in a form, report, query, or data table. Filters can also be used to provide only certain records from a table, query, or report.
[0063] The described systems and methods may be based on and / or include Application Programming Interfaces (APIs). APIs may include a set of rules that allow software applications to communicate and exchange data. APIs can be used in multiple applications, including mobile applications, web applications, cloud services, etc. The described systems and methods may be based on and / or include JavaScript Object Notation (JSON). JSON may include a text-based format for storing and exchanging human-readable and machine-parseable data. JSON files may include keys as names and values containing related data. Data can be separated by commas, curly braces can hold objects, and square brackets can hold arrays. The described systems and methods may include performing ETL operations based on one or more APIs.
[0064] The described systems and methods may be based on and / or include Representational State Transfer (REST) commands. REST commands may include requests to a server to retrieve or modify data. REST may be based on a software architectural style that uses the Hypertext Transfer Protocol (HTTP) to deliver APIs. REST commands can be executed by invoking methods on a REST resource and passing parameters or requests in JSON format. REST commands may include HTTP verbs defining the operation to be performed; headers allowing clients to pass information about the request; a path to the resource (e.g., the location of the resource); and / or a message body including data. REST commands may return a response from a dictionary containing state and content. State may include an HTTP response code, and content may include a response body as JSON or text. REST commands may include methods that invoke REST resources and pass requests or parameters in JSON format. The described systems and methods may include performing ETL operations based on one or more REST commands (e.g., ETL commands converted to REST commands).
[0065] The described systems and methods may be based on and / or include Simple Storage Service (S3). S3 storage may include cloud-based object storage services that allow users to store, access, and manage data. S3 requests (e.g., object storage requests, S3 retrieval requests) may include requests to create, retrieve, update, or delete objects in S3 storage. The described systems and methods may include performing ETL operations based on the S3 protocol (e.g., S3-based interface storage).
[0066] The described systems and methods may be based on and / or include MinIO. MinIO may include a high-performance distributed object storage system that can be used for data storage, backup and recovery, and cloud-based applications. MinIO may include a software-defined, open-source object storage system running on industry-standard hardware. MinIO can be used for private and hybrid cloud object storage and can run on any cloud or on-premises infrastructure. MinIO is API-compatible, enabling integration with S3-based applications. MinIO provides robust data protection for AI storage datasets through several different features, including erase coding and site replication, ensuring data redundancy and fault tolerance to prevent hardware failure or data corruption. The described systems and methods may include performing ETL operations based on MinIO (e.g., performing ETL operations relative to an object storage system based on MinIO).
[0067] The described systems and methods may be based on and / or include a Disaggregated Storage Solution (DSS). A DSS may include a rack-scalable, high-read-bandwidth-optimized, S3-compatible object storage solution. It utilizes a separate architecture to enable independent scaling of storage and compute. It features an end-to-end KV semantic communication stack, completely eliminating the need for traditional software storage stacks. The described systems and methods may include performing ETL operations based on the DSS (e.g., interacting with object storage systems based on the DSS).
[0068] The described systems and methods may be based on and / or include Non-Volatile Memory Fast (NVMe®), Structure-based NVMe (NVMe over Fabrics (NVMe-oF)), and / or Remote Direct Memory Access (RDMA). Storage communication can utilize the NVMeOF-KV-RDMA protocol. NVMeOF-KV-RDMA achieves high end-to-end performance using zero-copy transfers. The DSS client stack may include a high-performance wrapper library for easier application integration. Applications utilizing the DSS client library can eliminate the need for bucket semantics, key distribution, and load balancing between server-side S3 endpoints. The described systems and methods may include performing ETL operations based on NVMe-oF (e.g., providing ETL computation results based on NVMe-oF).
[0069] The described systems and methods may be based on and / or include web-based business intelligence (BI) tools (e.g., COGNOS®, smartphones, tablets, etc.), which can be configured to enable data exploration (exploring and preparing data); data visualization (e.g., creating interactive dashboards and enterprise reports); predictive analytics (e.g., performing forecasting and decision trees); AI-assisted analytics (e.g., machine learning insights); data sharing (e.g., sharing data across different platforms or in the cloud); storytelling (e.g., combining charts to create stories based on data analytics); and / or enabling users to analyze data, create reports, and make informed decisions, including data preparation (e.g., uploading, connecting, linking, and modeling data).
[0070] The described systems and methods may be based on and / or include SVK plugins and / or Cognos plugins. Cognos plugins and / or SVK plugins may include a set of compute resources and task types that can extend the functionality of storage devices (e.g., storage servers, compute storage drives configured to offload the execution of ETL functions). Plugins may be used for specific systems or technologies. For example, Cognos / SVK plugins may be used to perform one or more ETL functions on a storage device; move or delete files on a remote server; retrieve files from a remote server and make local copies; return a directory list of a specific path on a remote server. The described systems and methods may include performing ETL operations based on Cognos (e.g., performing ETL operations on a storage device, at least in part, based on Cognos and / or the Cognos SVK plugin).
[0071] In some systems, data (e.g., structured, unstructured, and / or flat data) is transferred via a network (e.g., a structure) to computing resources (e.g., host CPUs, GPUs, NPUs, etc.). The computing resources can process the data (e.g., transforming the raw data for analysis by cleaning, reshaping, reformatting, improving data integrity, and / or computing new dimensions and metrics based on the data). The transformed data can then be written back to storage and made available for querying about the transformed data. Large datasets can be moved from various data sources (e.g., storage servers, storage clusters) to transformation functions (e.g., computing resources, CPUs, GPUs, accelerators) via structures (e.g., data networks). This data movement degrades the performance of the transformation phase of ETL (e.g., increased latency and reduced throughput) and increases the total cost of ownership (TCO) of a given system.
[0072] Based on the zero-ETL system and method described herein, data is retrieved from a storage device (e.g., a storage server, storage drive, storage cluster, compute storage), transformed within the storage device via one or more processors, and stored in the storage device so that the transformed data is available to client devices, applications, other functions, another database, etc. (e.g., based on queries or requests from clients, applications, etc.). Therefore, this system and method avoids or minimizes ETL steps (such as data analytics) in large applications, especially large data transfer steps, by dynamically offloading the transformation algorithm and performing computation near the storage device. This reduces data transfer size and system resource requirements, thereby improving overall system performance while reducing the cost of ownership. In the case of zero-ETL, data can be accessed in real-time or near real-time, further improving ETL processing efficiency.
[0073] Figure 1 An example system 100 according to one or more embodiments described herein is illustrated. Figure 1 The image shows machine 105, which can be referred to as a host, system, or server. Although Figure 1 Machine 105 is described as a tower computer, but embodiments of this disclosure can be extended to machines of any form factor or type. For example, machine 105 may be a rack server, blade server, desktop computer, tower computer, mini-tower computer, desktop server, laptop computer, notebook computer, tablet computer, etc.
[0074] Machine 105 may include processor 110, memory 115, and storage device 120. Processor 110 may be any type of processor. Note that, for ease of illustration, processor 110 and other components discussed below are shown as being external to the machine: embodiments of this disclosure may include these components within the machine. Although Figure 1 A single processor 110 is shown, but the machine 105 may include any number of processors, each of which may be a single-core or multi-core processor, each of which may implement a Reduced Instruction Set Computer (RISC) architecture or a Complex Instruction Set Computer (CISC) architecture (and other possibilities), and may be mixed in any desired combination.
[0075] Processor 110 may be coupled to memory 115. Memory 115 may be any type of memory, such as flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), persistent random access memory, ferroelectric random access memory (FRAM), or non-volatile random access memory (NVRAM), such as magnetoresistive random access memory (MRAM), phase-change memory (PCM), or resistive random access memory (ReRAM). Memory 115 may include volatile and / or non-volatile memory. Memory 115 may use any desired form factor: for example, single in-line memory module (SIMM), dual in-line memory module (DIMM), non-volatile DIMM (NVDIMM), etc. Memory 115 may be any desired combination of different memory types and may be managed by memory controller 125. Memory 115 may be used to store data that can be referred to as "short-term": that is, data that is not expected to be stored for a long time. Examples of short-term data may include temporary files, data used locally by the application (which may have been copied from other storage locations), etc.
[0076] Processor 110 and memory 115 can support various operating systems on which applications can run. These applications can issue requests (which may be referred to as commands) to read data from or write data to memory 115 or storage device 120. When storage device 120 is used to support applications that read or write data via a certain file system, device driver 130 can be used to access storage device 120. Although Figure 1 A storage device 120 is shown, but any number (one or more) of storage devices may be present in machine 105. Storage device 120 may include one or more compute storage devices (e.g., storage devices with compute resources, processing units, etc.). Storage device 120 may support any desired one or more protocols, including, for example, the Non-Volatile Memory Fast (NVMe®) protocol, the Serial Attached Small Computer System Interface (SCSI) (SAS) protocol, or the Serial AT Accessory (SATA) protocol. Storage device 120 may include any desired interface, including, for example, a Peripheral Component Interconnect Fast (PCIe®) interface or a Compute Fast Link (CXL®) interface. Storage device 120 may take any desired form factor, including, for example, U.2 form factor, U.3 form factor, M.2 form factor, Enterprise and Data Center Standard Form Factor (EDSFF) (including all its types, such as E1 Short, E1 Long, and E3 types), or Add-in Card (AIC).
[0077] Although Figure 1The term "storage device" is used, but embodiments of this disclosure may include any storage device format that can benefit from the use of a computing storage unit, examples of which may include hard disk drives, solid-state drives (SSDs), or persistent memory devices such as PCM, ReRAM, or MRAM. Any references to "storage device" or "SSD" below should be understood to include other embodiments of this disclosure and other kinds of storage devices. In some cases, the term "storage unit" may cover both storage device 120 and memory 115. Machine 105 may include a power supply 135. Power supply 135 may provide power to machine 105 and its components.
[0078] Machine 105 may include a transmitter 145 and a receiver 150. The transmitter 145 or receiver 150 may be used to transmit or receive data, respectively. In some cases, the transmitter 145 and / or receiver 150 may be used to communicate with memory 115 and / or storage device 120. The transmitter 145 may include write circuitry 160, which may be used to write data into a storage device (such as a register) within memory 115 and / or storage device 120. Similarly, the receiver 150 may include read circuitry 165, which may be used to read data from memory 115 and / or storage device 120 from a storage device such as a register. In the illustrated example, machine 105 may include an accelerator 155, which may be used to perform one or more of the operations described herein.
[0079] In one or more examples, machine 105 can be implemented using any type of device. Machine 105 can be configured as one or more servers (e.g., as its host), such as a computing server, storage server, storage node, network server, supercomputer, data center system, etc., or any combination thereof. Additionally or alternatively, machine 105 can be configured as one or more computers such as a workstation, personal computer, tablet computer, smartphone, etc., or any combination thereof (e.g., as its host). Machine 105 can be implemented using any type of device, which can be configured to include, for example, accelerator devices, storage devices, network devices, memory expansion and / or buffer devices, central processing unit (CPU), graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), optical processing unit (OPU), etc., or any combination thereof.
[0080] Any communication between devices including machine 105 (e.g., host, compute storage device, and / or any intermediate device) can occur through an interface that can be implemented using any type of wired and / or wireless communication medium, interface, protocol, etc., including PCIe, NVMe, Ethernet, NVMe-oF, Compute Fast Link (CXL), and / or coherence protocols (such as CXL.mem, CXL.cache, CXL.IO, etc.), Gen-Z, Open Coherent Accelerator Processor Interface (OpenCAPI), Cache Coherent Interconnect for Accelerators (CCIX), Advanced Extensible Interface (AXI), etc. or any combination thereof, Transmission Control Protocol / Internet Protocol (TCP / IP), Fibre Channel, Unlimited Bandwidth, Serial AT Accessory (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, any generation of wireless networks (including 2G, 3G, 4G, 5G, etc.), any generation of Wi-Fi, Bluetooth, Near Field Communication (NFC), etc., or any combination thereof. In some embodiments, the communication interface may include a communication structure, which includes one or more links, buses, switches, hubs, nodes, routers, converters, repeaters, etc. In some embodiments, system 100 may include one or more additional devices having one or more additional communication interfaces.
[0081] Any of the functions described herein (including host functions, device functions, extract-transfer-load (ETL) controller 140 functions, etc.) can be implemented in hardware, software, firmware, or any combination thereof, including, for example, hardware and / or software combinational logic, sequential logic, timers, counters, registers, state machines, volatile memory (such as at least one or any combination of dynamic random access memory (DRAM) and / or static random access memory (SRAM)), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory), memory with varying bulk resistance, phase-change memory (PCM), etc. and / or any combination thereof, complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), application-specific integrated circuit (ASIC) CPUs (including complex instruction set computer (CISC) processors (such as x86 processors) and / or reduced instruction set computer (RISC) processors (such as RISC-V and / or ARM processors)), GPUs, NPUs, TPUs, OpUs, etc., executing instructions stored in any type of memory. In some embodiments, one or more components of the ETL controller 140 may be implemented as a SoC.
[0082] In some examples, ETL controller 140 may include any one or a combination of logic (e.g., logic circuitry), hardware (e.g., processing units, memory, storage devices), software, firmware, etc. In some cases, ETL controller 140 may perform one or more functions in conjunction with processor 110. In some cases, at least a portion of ETL controller 140 may be implemented in or by processor 110 and / or memory 115. One or more logic circuitry of ETL controller 140 may include any one or a combination of multiplexers, registers, logic gates, arithmetic logic units (ALUs), caches, computer memory, microprocessors, processing units (processor 110, accelerator 155, CPU, GPU, NPU, and / or TPU), FPGAs, ASICs, etc., enabling ETL controller 140 to provide systems and methods for data analysis based on zero-extraction transform-load (ETL).
[0083] In one or more examples, ETL controller 140 can provide distributed cluster-level offloading for ETL processing. ETL controller 140 avoids or minimizes data transfers, resulting in improved performance and reduced total cost of ownership due to lower overall resource requirements. ETL controller 140 can offload transformation functions or any computation in the ETL pipeline to compute-capable storage devices (e.g., compute storage drives, smart SSDs, CXL-HC devices, etc.). ETL controller 140 runs ETL functions within storage devices (e.g., compute storage drives, storage servers), thereby avoiding or reducing the need for ETL data transfers. ETL controller 140 implements a hybrid ETL model where functionality can be offloaded to storage servers, compute-capable storage devices, and / or client servers, providing a flexible and customizable solution.
[0084] Figure 2 Examples based on the description in this article are shown. Figure 1Details of machine 105 are provided. In the illustrated example, machine 105 may include processor 110. Processor 110 may include one or more processors and / or one or more dies. Processor 110 may include memory controller 125 (e.g., one or more memory controllers) and clock 205 (e.g., one or more clocks), which can be used to coordinate the operation of the machine's components. Processor 110 may be coupled to memory 115 (e.g., one or more memory chips, stacked memory, etc.), and as an example, memory may include random access memory (RAM), read-only memory (ROM), or other state-saving media. Processor 110 may be coupled to storage device 120 (e.g., one or more storage devices) and to network connector 210, which may be, for example, an Ethernet connector or a wireless connector. Processor 110 may be connected to bus 215 (e.g., one or more buses), user interface 220 (e.g., one or more user interfaces), and I / O interface ports that can be managed using input / output (I / O) engine 225 (e.g., one or more I / O engines), as well as other components, may be attached to bus 215. As shown in the figure, processor 110 can be coupled to ETL controller 230, which can be... Figure 1 An example of an ETL controller 140. Additionally or alternatively, processor 110 may be connected to bus 215, and ETL controller 230 may be attached to bus 215. In some cases, one or more of the zETL operations described herein may be performed by processor 110 and / or in conjunction with processor 110.
[0085] Figure 3 An example system 300 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 300 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system 300 may be implemented by the machine 105, components of the machine 105, or any combination thereof, or may be implemented in conjunction with the machine 105, components of the machine 105, or any combination thereof.
[0086] As shown in the figure, system 300 may include application 305, zero-ETL (zETL) adapter 310, and storage adapter 315. Storage adapter 315 may include zETL API 320 (e.g., one or more zETL APIs). In some cases, storage adapter 315 may include or represent a split storage solution (DSS) client (e.g., a client of application 305). In some cases, application 305 may include at least one user application. In some cases, application 305 may include an application of a host (e.g., machine 105). In some cases, zETL adapter 310 may include an SSD Value Kit (SVK) zETL adapter. In some cases, SVK devices (e.g., SVK zETL adapter, SVK controller) may be part of an SVK framework that enables ETL computation functionality to be offloaded to a storage device. The SVK framework may include hardware, firmware, and / or software (e.g., APIs) that enable ETL computation functionality to be offloaded to a storage device. In some cases, application 305 may interface with zETL adapter 310 via an API (e.g., zETL API). In some cases, the zETL adapter 310 can interface with the storage adapter 315 via the client API.
[0087] As shown in the figure, storage adapter 315 can interface with storage server 325, which may include zETL API 330 (e.g., one or more zETL APIs). In some cases, the interface between storage adapter 315 and storage server 325 may include an object storage server interface (e.g., a Simple Storage Service (S3) storage interface). In the example shown, system 300 may include storage server 325 and storage controller 335. In some cases, storage controller 335 may include an SVK controller. The SVK controller (e.g., storage controller 335) can control one or more aspects of the SVK framework (e.g., controlling data requests, controlling the application of APIs, controlling aspects of zETL processing, etc.). Therefore, storage controller 335 can control one or more aspects of offloading ETL processing to storage devices.
[0088] In the example shown, storage server 325 can interface with storage controller 335 based on one or more APIs. For example, storage server 325 can interface with storage controller 335 based on a Representational State Transition (REST) API. REST APIs can include API types that allow two computer systems to exchange information over a network (e.g., via the Internet, a cloud network).
[0089] As shown in the figure, system 300 may include one or more node servers (e.g., a storage cluster including node servers 340 and 345). As shown, node server 340 may include zETL plugin 350, database (DB) adapter 355, and storage device 360. Node server 345 may include zETL plugin 365, DB adapter 370, and storage device 375. Storage device 360 and / or storage device 375 may include one or more storage media (e.g., disks, NAND flash memory, non-volatile memory, volatile memory, etc.). As shown, storage controller 335 may interface with node server 340 and / or node server 345 via a REST API. In some cases, the REST API between storage controller 335 and the storage cluster of node servers 340 and 345 may be based on remote procedure calls (RPC), such as Google Remote Procedure Call (gRPC), which may include an open-source, cross-platform, high-performance remote procedure call framework for implementing APIs using HTTP / 2 connection services and devices. In some cases, the offloading of ETL functionality to the storage device can be based on a REST API and / or gRPC. In some cases, zETL plugin 350 and / or zETL plugin 365 may include SVK plugins (e.g., SVKzETL feature plugins) and / or Cognos-based zETL feature plugins. In some cases, DB adapter 355 may include the DSS NKV library.
[0090] In some examples, application developers can identify and build ETL functions (e.g., ETL transformation functions, ETL APIs, ETL feature plugins, etc.). In some cases, a group of ETL functions can be chained (e.g., chained functions, linked functions). In some cases, ETL functions can be combined within a directed acyclic graph (DAG). For example, multiple ETL functions can be defined within the same DAG, using bitwise shift operators (>>) to set dependencies between them to specify the execution order, chaining functions together to create a logical workflow where the output of one task becomes the input of the next. In some cases, zETL adapter 310 (e.g., the Cognos zETL adapter) can use the zETL API (e.g., the Cognos zETL API) to offload these functions from host or client nodes to storage nodes or storage devices (e.g., data server nodes, storage servers, compute storage drives). In some cases, storage server 325 can manage the routing and distribution of ETL functions. In some cases, zETL plugin 350 and / or zETL plugin 365 can facilitate the execution of one or more zETL functions in or near a storage device. In some cases, ETL functions can be executed in parallel across two or more storage nodes (e.g., all storage nodes) where the data being processed is distributed. For example, one or more ETL functions can be executed on storage devices (e.g., SSDs, Smart SSDs, compute storage devices, storage drives with one or more processing units). In some cases, ETL functions can be executed on storage devices that store the ETL data (e.g., storage devices storing data extracted, transformed, and / or loaded in ETL operations). In some cases, DB adapter 355 and / or DB adapter 370 can format the results of ETL functions in the requested data format and provide the formatted data to the requested location (e.g., on the same storage device, loaded into an application, loaded into another ETL function, loaded into the next ETL function in a chained ETL function sequence, etc.).
[0091] In some examples, storage server nodes (e.g., node server 340, node server 345) may include multiple storage drives (e.g., 16 SSDs, 16 Smart SSDs, 16 compute storage drives), enabling a given cluster to run multiple parallel offloaded ETL functions (e.g., 16 parallel offloaded ETL functions). For example, an ETL function may be executed on one or more storage drives. The ETL function may run computations on different data ranges in the ETL pipeline and may write the results of multiple computations back to storage devices. For example, a first storage drive may compute computations on a first data range in the ETL pipeline, a second storage drive may compute computations on a second data range in the ETL pipeline, and so on. The results (e.g., results from one or more storage drives) may be returned to an application (e.g., application 305), may be queried later, and / or may be used for further computations in the ETL pipeline and / or other ETL pipelines. Because the data is local and does not move around, the response time of ETL processing is significantly improved. Data can be preprocessed and stored in storage devices, and an object server (e.g., storage server 325) can manage the layout and location of the data.
[0092] The zETL architecture of System 300 provides per-device offloading of ETL algorithms, which increases the number of parallel operations and improves overall system performance. Configuring data closer to processing (e.g., on storage processing) and avoiding data transfer over the system network improves overall system performance and reduces system costs.
[0093] In some cases, the SVK zETL solution stack can be divided into two components: a zETL adapter (e.g., zETL adapter 310) and a zETL plugin (e.g., zETL plugin 350, zETL plugin 365). The zETL adapter may include a user interface, and the zETL plugin can be configured as the main working element in the SVK framework for handling and / or performing the offloading of ETL functions.
[0094] System 300 can implement an SVK zero-ETL architecture, which includes several components and can be considered as comprising several layers. A user application (e.g., application 305), an SVK-zETL adapter (e.g., zETL adapter 310), and a storage adapter (e.g., storage adapter 315) can be three components of the SVK zero-ETL architecture, including hardware that executes the application running on the client's physical machine. This layer can interface with the storage server layer via the S3 API. The storage server layer can include storage servers (e.g., storage server 325) and an SVK controller (e.g., storage controller 335). The storage servers can communicate with the SVK controller to perform one or more ETL offloading operations. The SVK controller can communicate with a storage cluster, which can include one or more SVK storage node servers (e.g., node server 340, node server 345). SVK zero-ETL plugins (e.g., zETL plugin 350, zETL plugin 365) can run on these storage node servers, where the execution of the offloading ETL functionality occurs.
[0095] Figure 4 An example system flow 400 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system flow 400 may be derived by... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system process 400 may be implemented by machine 105, components of machine 105, or any combination thereof, or in conjunction with machine 105, components of machine 105, or any combination thereof.
[0096] In the illustrated example, system process 400 may include client 405 and storage device 410. As shown, system process 400 may include operations associated with client 405 and storage device 410. In some cases, client 405 may include an application (e.g., application 305), one or more zETL adapters (e.g., zETL adapter 310), and / or a storage adapter (e.g., storage adapter 315). In some examples, client 405 and / or storage device 410 may include one or more zETL APIs (e.g., zETL API 320, zETL API 330). In some cases, storage device 410 may include one or more storage devices (e.g., storage server, compute storage drive). Storage device 410 may include one or more processors (e.g., configured to perform ETL transformation functions).
[0097] At 415, client 405 may import one or more zETL adapters. A zETL adapter (e.g., zETL adapter 310) may include hardware, firmware, and / or software configured to perform one or more of the operations described herein. In some cases, the zETL adapter may enable client 405 to offload one or more ETL functions to a storage device (e.g., storage device 410).
[0098] At 420, client 405 can perform zETL writes. In some cases, zETL writes can distribute data (e.g., data being processed) across a storage cluster (e.g., storage device 410).
[0099] At 425, storage device 410 can store data (e.g., zETL-based writes) in data blocks (e.g., 1 MB data blocks). In some cases, data blocks may include comma-separated value (CSV) files (e.g., 1 MB CSV files).
[0100] As shown in the figure, one or more operations of system process 400 (e.g., 415, 420, and / or 420) may be part of preprocessing (e.g., offline preprocessing). In some cases, one or more operations of system process 400 (e.g., at least one operation from 430 to 490) may be part of online processing (e.g., online application).
[0101] At 430, client 405 may import one or more zETL adapters (e.g., zETL adapter 310), which may include hardware, firmware, and / or software configured to perform one or more of the operations described herein. In some cases, the zETL adapter may enable client 405 to offload one or more ETL functions to a storage device (e.g., storage device 410).
[0102] At 435, client 405 can initiate zETL and / or perform zETL initialization operations. For example, client 405 can initiate the offloading of ETL functionality to storage device 410.
[0103] At 440, storage device 410 may initiate the offloading of ETL functionality to one or more aspects of storage device 410. For example, storage device 410 may initiate one or more aspects of an SSD Value Kit (SVK), which may be associated with a storage cluster (e.g., the storage cluster of storage device 410).
[0104] At 445, client 405 may execute one or more aspects of zETL, including offloading the execution of ETL functions to storage device 410. In some cases, client 405 may load data associated with zETL unloading. In some cases, client 405 may load or invoke one or more APIs associated with zETL unloading. In some cases, client 405 may load or generate one or more ETL commands associated with zETL unloading.
[0105] At 450, storage device 410 can perform a replication operation relative to a storage controller (e.g., an SVK controller) associated with offloading the execution of ETL functions to storage device 410. In some cases, storage device 410 can replicate data to a storage controller. In some cases, storage device 410 can replicate data to the storage controller of storage device 410. In some cases, storage device 410 can replicate data to the storage controller of storage device 410. In some cases, storage device 410 can replicate one or more zETL commands to the storage controller of storage device 410. In some examples, storage device 410 can replicate ETL .SO files (e.g., etl.so, dynamically linked ETL shared object libraries, shared ETL objects, shared ETL libraries, or shared ETL object libraries). In some examples, storage device 410 can replicate ETL .JSON files (e.g., etl.json, a JavaScript Object Notation (JSON) file used to offload the execution of ETL functions to storage device 410). In some cases, storage device 410 can copy .SO files and / or .JSON ETL files via client 405 (e.g., a split storage solution (DSS) client).
[0106] At 455, client 405 can perform read operations to unload the execution of ETL functions to storage device 410. In some cases, client 405 can read files and / or commands associated with the ETL unloading operation.
[0107] At 460, storage device 410 may read data associated with the execution of ETL functions (e.g., read data stored on storage device 410). In some cases, storage device 410 may read or receive ETL commands and ETL APIs associated with performing ETL functions on storage device 410. Executing ETL functions offloaded to storage device 410 may include at least operations 460, 465, 470, 475, and / or 480 from system process 400.
[0108] At 465, storage device 410 can perform one or more load operations (e.g., load one or more ETL functions, load the executed ETL functions simultaneously or in parallel). In some cases, storage device 410 can load ETL.SO files (e.g., as SVK plugins). For example, storage device 410 can load ETL.SO files as SVK plugins on the SVK server of storage device 410 (e.g., the storage cluster server of storage device 410).
[0109] At 470, storage device 410 can perform one or more data cleanup operations (e.g., performing data cleanup operations of multiple ETL functions simultaneously or in parallel). For example, storage device 410 can receive ETL commands, read data stored at storage device 410 (e.g., at 460), and clean the data. Data cleanup may include identifying and correcting corrupt, inaccurate, or irrelevant records from datasets, tables, or databases. Data cleanup may include detecting incomplete, incorrect, or inaccurate portions of data and then replacing, modifying, and / or deleting the affected data.
[0110] At 475, storage device 410 can perform one or more data processing operations (e.g., performing multiple ETL functions simultaneously or in parallel). For example, storage device 410 can perform at least one ETL function on data stored at storage device 410 based on ETL commands received by storage device 410 from client 405 (e.g., applying at least one ETL transformation function). In some cases, storage device 410 can perform a series of combined ETL functions (e.g., ETL functions combined in a DAG).
[0111] At 480, storage device 410 may store the results of zETL data processing operations at 475. In some cases, storage device 410 may store the results to a storage drive of storage device 410 (e.g., in a data warehouse of storage device 410). In some cases, storage device 410 may provide the results to client 405 and / or load the results into an application on client 405. In some cases, storage device 410 may provide the results to another database (e.g., to a database physically separate from storage device 410). In some cases, storage device 410 may provide the results to a function (e.g., another function executed on or to be executed on storage device 410 or on another storage device). In some cases, storage device 410 may store the results in an SQL table (e.g., a database object in a relational database that stores data in row and column format). For example, storage device 410 may store the results in Apache Iceberg (APACHE ICEBERG®). Apache Iceberg may include table formats for large data volumes with asynchronous metadata writing processes.
[0112] At 485, client 405 can execute one or more queries. For example, client 405 can query storage device 410 (e.g., the data warehouse of storage device 410). In some cases, the query of client 405 may include a query for the results of the zETL data processing operation at 475.
[0113] At 490, client 405 can draw one or more charts. For example, client 405 can draw one or more charts based on the query at 485. In some cases, client 405 can generate one or more plots or graphs that visually represent one or more aspects of the results of the zETL data processing operation at 475.
[0114] Figure 5 An example system 500 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 500 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system 500 may be implemented by machine 105, components of machine 105, or any combination thereof, or in conjunction with machine 105, components of machine 105, or any combination thereof.
[0115] In the illustrated example, system 500 may include application 505, storage node 510 (e.g., node server, storage server), data warehouse 515, and REST server 520. In some cases, application 505 may be associated with a client (e.g., client 405). In some cases, storage node 510 and / or data warehouse 515 may be associated with a storage system (e.g., storage device 410). As shown, storage node 510 may include minIO 525, ETL offload 530, remote procedure call (RPC) server 535, worker node 540, and storage device 550. As shown, worker node 540 may include one or more ETL functions (e.g., ETL function 545). Storage device 550 may include one or more processors (e.g., processor 555).
[0116] In some cases, application 505 can initiate the offloading of ETL processing to storage node 510. For example, application 505 can generate ETL commands and / or send ETL commands to storage node 510. As shown, minIO 525 can process ETL commands and forward them to ETL offloading 530. As shown, ETL offloading 530 processes ETL commands via REST server 520. In some examples, ETL offloading 530 can transmit ETL commands to RPC server 535 via REST server 520. For example, ETL offloading 530 can send REST commands to REST server 520, and REST server 520 can transmit the REST commands to RPC server 535. In some cases, RPC server 535 can be configured as an SVK RPC server (e.g., a storage cluster server). In some cases, an application (e.g., application 505) can communicate with the SVK server via REST (e.g., via REST server 520).
[0117] In some examples, RPC server 535 can request data from data warehouse 515 based on REST commands. In some cases, RPC server 535 can transmit ETL commands to worker node 540 based on REST commands. In some cases, RPC server 535 can provide data from data warehouse 515 to worker node 540. Worker node 540 can include one or more ETL functions (e.g., ETL function 545). In some cases, worker node 540 can perform ETL function 545 based on ETL commands. In some cases, worker node 540 can perform ETL function 545 on data provided by RPC server 535 from data warehouse 515. In some cases, worker node 540 can include zETL modules and / or zETL plugins that worker node 540 can implement to perform ETL function 545. In some cases, worker node 540 can obtain the results of performing ETL function 545 on data. In some cases, worker node 540 can save the results to data warehouse 515 (e.g., via RPC server 535). In some cases, storage device 550 can store data, metadata, ETL processing results, etc., including files (e.g., comma-separated value (CSV) files). In some cases, one or more of these files can be divided into data blocks (e.g., 1 MB data blocks).
[0118] As shown in the figure, application 505 can generate and / or execute query 560. In some cases, query 560 can be sent to data warehouse 515. Application 505 can query data processed by ETL (e.g., results) and create plots or charts based on the ETL-processed data. In some cases, application 505 can invoke ETL functions within storage node 510 (e.g., ETL function 545), and ETL function 545 can provide the results to application 505.
[0119] System 500 provides a zETL framework for offloading ETL processing to a storage device (e.g., storage node 510). System 500 can be configured as a Cognos zETL system, enabling applications (e.g., application 505) to efficiently preprocess data (e.g., preprocess data from data warehouse 515 to storage device 550 before storage) and / or postprocess data (e.g., after local retrieval on storage node 510) using any storage solution. System 500 provides a framework where applications can offload ETL functionality to a storage server to avoid large data transfers. The framework of System 500 avoids data movement from the data source (storage device) to the transformation function processing (computation) by performing functions within a storage server (e.g., storage node 510). The framework of System 500 fine-tunes I / O based on ETL function requirements and back-end storage. The framework of System 500 provides the ability to replicate ETL functionality to achieve parallelism, thus providing flexibility in managing data distribution.
[0120] Based on the systems and methods described herein, data can be pre-formatted to improve data retrieval in distributed storage systems and reduce ETL turnaround time (e.g., formatting data from data warehouse 515 and storing the formatted data in storage 550). ETL algorithms (ETL function 545) that process the data are offloaded to the storage system (e.g., storage node 510), allowing ETL processing to be performed on the storage node where the data resides, thereby avoiding or minimizing data movement across the structure or network and improving system efficiency. In some cases, intermediate results can be stored back to storage devices (e.g., storage device 550, data warehouse 515) or processed inline, providing flexibility for the solution. System 500 can provide an SSD Value Kit (SVK) platform (e.g., RPC server 535) that provides a unified interface for configuring, managing, and interfacing with worker node 540, ETL function 545, storage device 550, and / or processor 555. In some cases, storage node 510 can implement an SVK cluster architecture (e.g., in storage device 550), enabling user applications (e.g., application 505) to interact with the SVKzETL API. These APIs may include a set of predefined interfaces for storage, ETL operations, and offloading zETL operations. The offloading schemes of the systems and methods described herein can be implemented using these APIs.
[0121] Based on the systems and methods described herein, system 500 enables the storage of results in various database type formats via data adapter plugins, thereby avoiding additional data copying for loading data into a data warehouse. ETL functions (e.g., ETL function 545) can be registered with the zero-ETL framework of system 500, and user applications can use the registered ETL functions in their pipelines without changing the pipeline code.
[0122] Figure 6 An example system 600 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 600 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system 600 may be implemented by the machine 105, components of the machine 105, or any combination thereof, or may be implemented in conjunction with the machine 105, components of the machine 105, or any combination thereof.
[0123] In the illustrated example, system 600 may include management node 605 and one or more hosts (e.g., host 610, host 615). In some cases, management application 620 may include and / or implement web-based business intelligence (BI) tools (e.g., Cognos). As shown, host 610 may include one or more host applications (e.g., host application 630), REST controller 635 (e.g., Cognos REST / controller service), node service 640 (e.g., Cognos node service), framework library 645 (e.g., Cognos framework library), zETL module 650 (e.g., SVK zETL module), and storage device 655. As shown, host 615 may include one or more host applications (e.g., host application 660), REST controller 665 (e.g., Cognos REST / controller service), node service 670 (e.g., Cognos node service), framework library 675 (e.g., Cognos framework library), zETL module 680 (e.g., SVK zETL module), and storage device 685. In some cases, storage device 655 and / or storage device 685 may include one or more CXL memory module-DRAM (CMM-D) devices; one or more CXL memory module-hybrid (CMM-H) devices; and / or one or more CXL memory module-box (CMM-B) memory pooling devices.
[0124] The described systems and methods may include unloadable modules (e.g., zETL modules) configured to conform to a specific format defined by SVK zETL and built by the user. In some examples, zETL module 650 may be configured as an unloadable module. These unloadable modules can be pre-built and unloaded into storage servers (e.g., storage server 325, node server 340, node server 345, storage node 510) using the SVK zETL API (e.g., zETL API 330). The capabilities of a given zETL module can be queried at runtime. Once the zETL module is unloaded, the runtime pipeline can utilize the functionality within the zETL module to process data within the storage server itself, instead of sending it to the application server, thus saving valuable bandwidth and reducing network latency. Furthermore, since ETL functionality can be applied independently to each object, multiple objects can be processed in parallel, and optimal throughput can be achieved when objects are distributed nearly evenly.
[0125] In some cases, unloadable modules can be added, modified, or removed as needed (e.g., zETL module 650, zETL module 680), thus providing the flexibility to seamlessly extend the SVK zETL system and methodology described herein. For example, the infrastructure used to build unloadable modules can be extended to support several different languages and / or have a common framework for writing unloadable modules or support converting them into a way that zETL can utilize modules.
[0126] In some examples, System 600 provides a framework for offloading ETL functionality to storage servers / devices at the cluster level. System 600 avoids data transfers within the storage system by performing pipelined functionality in storage servers and / or storage devices (e.g., storage device 655, storage device 685). System 600 provides a framework that enables ETL functionality to be offloaded and executed in parallel to achieve high throughput and reduced latency. System 600 avoids additional jumps by implementing ETL processing (e.g., custom data preprocessing) across different storage servers through storage adapters. System 600 provides a framework that enables the dynamic (re)distribution of ETL functionality as part of cluster rebalancing. For example, based on the systems and methods described herein, ETL functionality can be executed by one or more zETL modules (e.g., zETL module 650, zETL module 680, etc.). In some cases, one or more functions can be distributed to zETL module 650 (e.g., via management node 605). In some examples, one or more ETL functions can be exchanged between zETL modules. Alternatively or additionally, one or more functions may be distributed to zETL module 680 (e.g., via management node 605).
[0127] In some examples, at least one ETL function may be reallocated or moved from zETL module 650 to zETL module 680. Alternatively or additionally, one or more ETL functions may be reallocated or moved from zETL module 680 to zETL module 650. In some cases, management node 605 may identify an imbalance in the distribution of ETL functions. For example, management node 605 may determine that a zETL module is assigned a relatively high number of ETL functions. In some cases, management node 605 may determine that the latency or processing time of a zETL module exceeds a threshold (e.g., based on average processing time, based on expected processing time, etc.). Therefore, management node 605 may distribute and / or redistribute one or more ETL functions within a pool of zETL modules (e.g., zETL module 650, zETL module 680, etc.). In some cases, one or more ETL commands may be distributed and / or redistributed based on the processing load across zETL modules. For example, management node 605 may determine that the processing load of a zETL module exceeds a processing load threshold (e.g., based on the average processing load across zETL modules or the real-time average processing load). Based on the distribution and / or redistribution of ETL functions and / or ETL commands, ETL operations can be balanced across zETL modules of the storage cluster.
[0128] Figure 7 An example system flow 700 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system flow 700 may be derived by... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system process 700 may be implemented by machine 105, components of machine 105, or any combination thereof, or in conjunction with machine 105, components of machine 105, or any combination thereof.
[0129] In some examples, system flow 700 may depict one or more operations associated with offloading ETL processing to a storage device. In the illustrated example, system flow 700 may include preprocessing 705, data pipeline read 710, and data pipeline query 715. In some cases, preprocessing 705 may be asynchronous with the data pipeline. In some cases, preprocessing 705 may be performed before the data pipeline application runs.
[0130] At 720, preprocessing 705 may include loading data. For example, preprocessing 705 may include loading data from a data warehouse (e.g., data warehouse 515) to a local storage device (e.g., storage device 550). At 725, preprocessing 705 may include generating data blocks (e.g., 1 MB data blocks, CSV data blocks). In some cases, data blocks may be considered based on granularity and / or schema boundaries. At 730, preprocessing 705 may include storing data blocks (e.g., in storage nodes, storage device 360, storage device 410, storage device 550, storage device 655).
[0131] In some examples, data pipeline read 710 may include triggering a data pipeline application read that unloads the ETL system and method described herein. At 735, data pipeline read 710 may include data cleanup 735 (e.g., clean up data loaded from 720, clean up data blocks stored from 730). At 740, data cleanup may include removing empty data. At 745, data cleanup may include removing invalid data. At 750, data pipeline read 710 may include data processing (e.g., ETL processing, performing ETL functions on the data). At 755, data processing may include filtering data (e.g., based on filtering data criteria, based on data size, based on file type, etc.). For example, filtering data may include filtering hourly data based on rainfall and temperature range. At 760, data processing may include transforming data (e.g., converting temperature data from Celsius to Fahrenheit). At 765, data processing may include identifying information (e.g., calculating weather comfort rating scores, adding a column for rating scores, adding a column to a CSV file).
[0132] At 770, data pipeline reading 710 may include storing data. At 775, data storage may include storing processed data (e.g., the results of ETL data processing) to local storage devices and / or data warehouses. For example, processed data may be stored in storage devices for subsequent queries of the processed data and information derived from the processed data.
[0133] In some examples, data pipeline query 715 may include querying a data warehouse to obtain data pipeline applications that process data. At 780, data pipeline query 715 may include executing a query. At 785, the query may include querying the processed data (e.g., querying data processed by ETL). For example, an application (e.g., a top-level application) may query a data warehouse where the query retrieves at least a portion of the processed data. At 790, the query may include displaying the query results. For example, the results of the query may be displayed. In some cases, displaying the query results may include generating one or more charts or plots that provide a visual depiction of the information obtained from the data processed by ETL. For example, charts may be generated based on the query results. For example, time (e.g., days, hours) based on pleasure rating scores may be displayed in a defined order (e.g., arranged from highest to lowest rating).
[0134] Based on the systems and methods described herein, including system flow 700, data can be stored as data blocks to achieve better distribution and increased ETL processing parallelism. Based on these systems and methods, ETL functions (including query evaluation) can be offloaded and processed in parallel. Based on these systems and methods, ETL processing results can be stored in a database (e.g., local storage device, data warehouse). In some cases, data (e.g., preprocessed acquired data, ETL processed data) can be formatted according to a specified format (e.g., specified by the zETL adapter).
[0135] Figure 8 An example system 800 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 800 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is used for implementation or in combination with it. In some configurations, one or more aspects of the system 800 may be implemented by machine 105, components of machine 105, or any combination thereof, or may be implemented in combination with machine 105, components of machine 105, or any combination thereof.
[0136] In the example shown, system 800 may include a data warehouse management stack 805, a management REST API 810, a network 815, a REST controller 820, and one or more storage nodes (e.g., storage node 825, one or more Cognos node servers, etc.).
[0137] In the example shown, the data warehouse management stack 805 can call a management REST API (e.g., the Cognos management REST API) to download ETL binaries and register the ETL binaries with the zETL framework (e.g., register an ETL transformation module to process data warehouse data). Registration may include registering the feature name of the ETL function (e.g., feature_name="zETL") and / or registering the transformation module (TM) name (e.g., TM_name="filter_logic_1").
[0138] In some examples, requests (e.g., ETL requests) can be handled by REST controllers (e.g., the Cognos REST / controller service).
[0139] In some examples, requests can be validated (e.g., via REST controller 820) and routed to at least one host (e.g., to storage node 825, to a second storage node, etc.; to all hosts on which the Cognos node server is running).
[0140] At 830, storage node 825 can be configured to provide node services. In some examples, storage node 825 can be configured to parse requests and identify the zETL module to which the request is destined, and route the request to the corresponding zETL module (e.g., via the Cognos command framework). At 835, storage node 825 can be configured to invoke the command framework.
[0141] At 840, storage node 825 can be configured to send requests (e.g., ETL requests) to zETL modules (e.g., zETL module 650). In some cases, upon receiving a request, the zETL module can further parse the request to obtain subtasks. In this case, the subtasks may include downloading and registering dedicated ETL functions.
[0142] At 845, storage node 825 can be configured to register one or more ETL functions (e.g., register one or more ETL transformation module binaries). In some cases, functions can be registered to multiple nodes (e.g., to storage node 825, a second storage node, etc.). For example, a first function can be registered to a first node, a second function can be registered to a second node, and so on. Additionally or alternatively, a first function can be registered to both a first and a second node, a second function can be registered to both a first and a second node, and so on. In some cases, storage node 825 can load one or more ETL functions (e.g., to a location indicated in the request, or a location provided in the request).
[0143] In some cases, ETL functions can be defined and registered into the zETL framework as pipelines and / or as standalone functions. Depending on the application, ETL functions can run in parallel or simultaneously. In some cases, the zETL API can manage data read / write operations between local storage and ETL functions. In other cases, the zETL API can manage data read / write operations between ETL functions and user applications.
[0144] In some examples, application developers can write transformation modules (e.g., ETL functions) by implementing the APIs of the Cognos zETL framework. Application developers can identify desired functions (e.g., decryption, decompression, buffer parsing, etc.) and apply the transformation functions they want to implement. In some cases, application developers can return the resulting buffers to the zETL framework (e.g., optionally in a compressed / encrypted format). Application developers can create loadable binary formats, such as shared library formats (e.g., for Linux). Multiple transformation modules can be written for multiple ETL transformation functions to perform pipelined operations (e.g., filtering and then counting, filtering and concatenating two KV values, etc.). Application-specific ETL transformation modules can be uploaded / registered to the zETL framework based on the Cognos management API. During the execution path, the zETL framework can use storage solution-specific DB connector modules (e.g., the dss_db_connector module) to retrieve data from storage, pass buffers to the appropriate transformation modules, load the results into another bucket, save the results to a database, and / or load the results into the user application.
[0145] Figure 9 An example system 900 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 900 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is used for implementation or in combination with it. In some configurations, one or more aspects of the system 900 may be implemented by machine 105, components of machine 105, or any combination thereof, or in combination with machine 105, components of machine 105, or any combination thereof.
[0146] In the example shown, system 900 may include a data warehouse query engine 905, a management REST API 910, nodes 915 (e.g., storage nodes, server nodes, DSS storage nodes), and a REST controller service 920. As shown, node 915 may include a DSS S3 service 925, a DSS storage node 930, and a storage service 935. As shown, storage service 935 may include one or more storage drives (e.g., one or more SSDs, storage clusters, etc.).
[0147] In the example shown, the data warehouse query engine 905 (e.g., the SNOWFLAKE® query engine, a cloud-based data warehouse and analytics platform) can perform data retrieval (e.g., retrieving one or more key-value (KV) pairs (in batches)). In some cases, the data warehouse query engine 905 can specify filtering (transformation), logical names (e.g., registering with Cognos), and / or other aspects selected by the data warehouse query engine 905 for data retrieval.
[0148] In some examples, the management REST API 910 can send a get request to node 915. In some cases, an S3 client wrapper library from the DSS stack can translate a get request from the management REST API 910 (e.g., a non-standard ETL get request) into an S3 get request based on a private key. An example of a private key could be S3::Get{key = “cognos::zETL::<key_name> ", val =<rdd_info> +<max_size_allocated> ,subcmd = filter_x}, where rdd can represent RDMA data directly. When the DSS stack (e.g., on node 915) recognizes a network interface card (NIC) with RDMA capability on the client side, the DSS stack can directly transmit the resulting data to the client with zero overhead.
[0149] After receiving a request, DSS S3 service 925 can parse the request and determine that it includes zETL commands. DSS S3 service 925 can create a Cognos zETL REST command understood by the Cognos framework. In some cases, the zETL REST command may include the name of the host server storing the requested KV pair. If the command is part of a batch operation, commands associated with the same storage node (e.g., requests for KV pairs on the same storage node) can be batched into a single command (e.g., a batch command). At 940, DSS S3 service 925 can perform a fetch (e.g., an S3 fetch request). At 945, DSS S3 service 925 can form one or more ETL commands based on the fetch request.
[0150] In some examples, the DSS SE service 925 can send REST commands (e.g., Cognos zETL REST commands) to the REST controller service 920 (e.g., the Cognos REST / controller service), which can run on the same node (e.g., node 915) or on different nodes. In some cases, the REST controller service 920 may include a REST service and / or a controller service (e.g., the Cognos REST and controller service). In some cases, the REST controller service 920 may receive host server information (e.g., only receiving host server information for a given fetch request or ETL command). The REST controller service 920 may pass requests to the DSS storage node 930 (e.g., the Cognos node service running on a specified host server and on a specified node).
[0151] In some examples, after parsing the request, DSS storage node 930 can determine the zETL module to which the request should be directed and route the request to the specified zETL module (e.g., via the Cognos command framework). For example, at 950, DSS storage node 930 can invoke the command framework. At 955, DSS storage node 930 can send the request to the zETL module.
[0152] In some examples, DSS storage node 930 can determine that the request is for an ETL function (e.g., an ETL function request). In some cases, DSS storage node 930 can issue one or more local KV reads based on the identification of an ETL function request. In some cases, DSS storage node 930 can publish local KV reads to the DSS storage service indicating the stored KV in conjunction with the DSS DB connector module. For example, at 960, DSS storage node 930 can retrieve KV pairs from a DSS storage device (e.g., storage service 935). As shown, storage service 935 may include a DSS KV adapter 975 and a Storage Performance Development Kit (SPDK) service 980. In some cases, SPDK service 980 can provide a set of tools and libraries for writing high-performance, scalable, user-mode ETL storage applications.
[0153] In some examples, after retrieving the value buffer (e.g., a buffer of key-value pairs; via RDD zero-copy), the DSS storage node 930 can pass the value buffer (e.g., from storage service 935) to the requested zETL module for processing. For example, at 965, the DSS storage node 930 can execute filter transformation logic (e.g., perform one or more ETL functions on the DSS storage node 930). The processing at the zETL module can be pipelined for multi-stage processing with multiple ETL transformation modules.
[0154] In some examples, after processing, DSS storage node 930 can send the result buffer to the client (e.g., via RDD zero-copy), store the result buffer back to the DSS storage device (e.g., storage service 935, DSS storage node 930), or load the result buffer into the client application or a different database. At 970, DSS storage node 930 can send filtered data (e.g., the result buffer, ETH processing results) to the client (e.g., via the management REST API 910).
[0155] Figure 10 A flowchart illustrating an example method 1000 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, one or more aspects of method 1000 may be derived by... Figure 1 ETL controller 140 and / or Figure 2The method 1000 is implemented or implemented in conjunction with the ETL controller 230. In some configurations, one or more aspects of method 1000 may be implemented by machine 105, components of machine 105, or any combination thereof, or in combination with machine 105, components of machine 105, or any combination thereof. The depicted method 1000 is merely one implementation, and one or more operations of method 1000 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0156] In step 1005, method 1000 may include generating a data request. For example, a client or its application may generate a data request for at least a subset of data stored in a database.
[0157] In 1010, method 1000 may include converting a data request into an object storage request. For example, a client may convert a data request into an object storage request.
[0158] In 1015, method 1000 may include determining that the object storage request includes an Extract Transform Load (ETL) request. For example, a client or client application may determine that the object storage request includes an Extract Transform Load (ETL) request based on parsing the object storage request.
[0159] In 1020, method 1000 may include creating an ETL command based on an object storage request. For example, a client or client application may create an ETL command based on an object storage request.
[0160] In 1025, method 1000 may include routing ETL commands to a storage device. For example, a client or client application may route ETL commands to a storage device (e.g., a storage server, compute storage drive) based on the storage device including at least one ETL function requested in the data request.
[0161] In 1030, method 1000 may include determining information from the result data. For example, a client or client application may determine information based on result data received from the storage device in response to routing an ETL command to the storage device.
[0162] Figure 11 A flowchart illustrating an example method 1100 associated with the disclosed system according to an example implementation described herein is depicted. In some configurations, one or more aspects of method 1100 may be derived by... Figure 1 ETL controller 140 and / or Figure 2The method 1100 is implemented or implemented in conjunction with the ETL controller 230. In some configurations, one or more aspects of method 1100 may be implemented by machine 105, components of machine 105, or any combination thereof, or may be implemented in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 1100 is merely one implementation, and one or more operations of method 1100 may be rearranged, reordered, omitted, and / or otherwise modified to make other implementations possible and contemplated.
[0163] In 1105, method 1100 may include receiving an Extract Transform Load (ETL) command. For example, a storage node (e.g., a node server, storage server, compute storage drive, storage cluster) may receive an Extract Transform Load (ETL) command at the storage node, the ETL command including a data request for a subset of data stored in a database.
[0164] In 1110, method 1100 may include reading a subset of data from the database of the storage node. For example, the storage node may read a subset of data from its own database.
[0165] In step 1115, method 1100 may include providing a subset of data to a storage device of a storage node. For example, a storage node may provide a subset of data to a storage device of the storage node so that the subset of data can be processed at the storage device.
[0166] At 1120, method 1100 may include processing a subset of data at the storage device. For example, the storage node may process the subset of data at the storage device based on the storage device including at least one ETL function requested in the data request.
[0167] In 1125, method 1100 may include calculating the result data based on processing. For example, a storage node may calculate the result data based on processing.
[0168] Figure 12 An example system 1200 according to one or more embodiments described herein is illustrated. In some configurations, one or more aspects of system 1200 may be derived from... Figure 1 ETL controller 140 and / or Figure 2 The ETL controller 230 is implemented or implemented in conjunction with it. In some configurations, one or more aspects of the system 1200 may be implemented by the machine 105, components of the machine 105, or any combination thereof, or may be implemented in conjunction with the machine 105, components of the machine 105, or any combination thereof.
[0169] In the example shown, system 1200 may include an SSD Value Kit (SVK) framework 1205. In some cases, system 1200 may implement one or more aspects of the systems and methods for zero-extraction transformation loading (ETL) data analysis described herein. As shown, the SVK framework 1205 may include a management application 1210, a CXL management API 1215, a REST / controller service 1220, a node service 1225, a framework library 1230, zETL features 1235, CXL local storage management 1240, CXL memory module box (CMMB) management 1245, CXL memory sharing 1250, and CXL service level agreement (SLA) monitoring 1255.
[0170] In some examples, management application 1210 can provide a user interface and / or command-line interface for the zETL operations described herein. For example, management application 1210 can enable users or applications to define storage clusters for zETL operations, add devices (e.g., storage nodes) to the zETL system, reconfigure devices, etc. In some cases, CXL management API 1215 can manage APIs associated with the zETL operations described herein. In some cases, CXL management API 1215 can provide and / or manage one or more REST APIs for the zETL operations described herein. REST / controller service 1220 can provide REST services (e.g., associated with REST APIs). In some cases, REST / controller service 1220 can control one or more aspects of the REST services associated with the zETL operations described herein (e.g., enabling REST APIs, modifying REST APIs, disabling REST APIs, managing data and / or data movement based on REST APIs, etc.). In some examples, node service 1225 can manage one or more nodes associated with the zETL operations described herein (e.g., configuring, adding, removing, and / or updating storage nodes).
[0171] In some cases, framework library 1230 can be configured to communicate with zETL feature modules (e.g., zETL feature 1235). In some cases, zETL feature 1235 may be an example of ETL controller 140, ETL controller 230, zETL plugin 350, zETL plugin 365, worker node 540, zETL module 650, zETL 680, and / or data processing 750. In some cases, framework library 1230 may provide a library of features and configurations associated with the zETL operations described herein.
[0172] In some examples, zETL feature 1235 can be configured to perform one or more zETL functions via one or more processing units of a storage node (e.g., a storage server, compute storage drive, etc.). In some examples, CXL Local Memory Management 1240 can manage one or more aspects of the memory (e.g., CXL Local Memory) of a given storage node, including memory operations (e.g., read, write, allocate, deallocate, etc.). CMMB Management 1245 can be configured to provide memory pooling services based on the CXL protocol. CXL Memory Sharing 1250 can be configured to share memory among one or more entities (e.g., storage nodes, hosts, applications, etc.). CXL SLA Monitoring 1255 can be configured to manage how users, devices, and / or applications can use the zETL operations provided by zETL feature 1235.
[0173] In the examples described herein, the configurations and operations are example configurations and operations, and various additional configurations and operations may be involved that are not explicitly shown. In some examples, one or more aspects of the configurations and / or operations shown may be omitted. In some embodiments, one or more operations may be performed by components other than those shown herein. Additionally or alternatively, the order and / or timing of operations may be changed.
[0174] Specific embodiments may be implemented in one or a combination of hardware, firmware, and software. Other embodiments may be implemented as instructions stored on a computer-readable storage device that can be read and executed by at least one processor to perform the operations described herein. A computer-readable storage device may include any non-transitory memory structure for storing information in a machine-readable (e.g., computer) form. For example, a computer-readable storage device may include read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, and other storage devices and media.
[0175] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. As used herein, the terms "computing device," "user equipment," "communication station," "station," "handheld device," "mobile device," "wireless device," and "user equipment" (UE) refer to wired and / or wireless communication devices such as switches, routers, network interface controllers, cellular phones, smartphones, tablets, netbooks, wireless terminals, laptops, femtocells, high data rate (HDR) subscriber stations, access points, printers, point-of-sale equipment, access terminals, or other personal communication system (PCS) devices. Such devices can be wireless, wired, mobile, and / or fixed.
[0176] As used in this document, the term "communication" is intended to include sending, or receiving, or both. Similarly, when only the functionality of one device is claimed, bidirectional data exchange between two devices (where both devices send and receive during the exchange) can be described as "communication." The term "communication" as used herein with respect to wired and / or wireless communication signals includes sending and / or receiving wired and / or wireless communication signals. For example, a communication unit capable of transmitting wired and / or wireless communication signals may include a wired / wireless transmitter to send communication signals to at least one other communication unit, and / or a wired / wireless receiver to receive communication signals from at least one other communication unit.
[0177] Some embodiments can be used with a variety of devices and systems, such as personal computers (PCs), desktop computers, mobile computers, laptop computers, notebook computers, tablet computers, server computers, handheld computers, handheld devices, personal digital assistant (PDA) devices, handheld PDA devices, airborne devices, non-airborne devices, hybrid devices, in-vehicle devices, non-in-vehicle devices, mobile or portable devices, consumer devices, non-mobile or non-portable devices, wireless communication stations, wireless communication devices, wireless access points (APs), wired or wireless routers, wired or wireless modems, video devices, audio devices, audio-video (A / V) devices, wired or wireless networks, wireless local area networks, wireless video local area networks (WVANs), local area networks (LANs), wireless LANs (WLANs), personal area networks (PANs), wireless PANs (WPANs), etc.
[0178] Some embodiments may be used in combination with the following: one-way and / or two-way radio communication systems, cellular wireless telephone communication systems, mobile phones, cellular phones, wireless phones, personal communication system (PCS) devices, PDA devices containing wireless communication devices, mobile or portable global positioning system (GPS) devices, devices containing GPS receivers or transceivers or chips, devices containing RFID elements or chips, multiple-input multiple-output (MIMO) transceivers or devices, single-input multiple-output (SIMO) transceivers or devices, multiple-input single-output (MISO) transceivers or devices, devices having one or more internal antennas and / or external antennas, digital video broadcasting (DVB) devices or systems, multi-standard wireless devices or systems, wired or wireless handheld devices (e.g., smartphones), Wireless Application Protocol (WAP) devices, etc.
[0179] Some embodiments can be used in conjunction with one or more types of wireless communication signals and / or systems that conform to one or more wireless communication protocols, such as radio frequency (RF), infrared (IR), frequency division multiplexing (FDM), orthogonal FDM (OFDM), time division multiplexing (TDM), time division multiple access (TDMA), extended TDMA (E-TDMA), General Packet Radio Service (GPRS), extended GPRS, code division multiple access (CDMA), wideband CDMA (WCDMA), CDMA 2000, single-carrier CDMA, multi-carrier CDMA, multi-carrier modulation (MDM), discrete multi-tone (DMT), and Bluetooth. TM Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee TM Ultra-wideband (UWB), Global System for Mobile Communications (GSM), 2G, 2.5G, 3G, 3.5G, 4G, fifth-generation (5G) mobile networks, 3GPP, Long Term Evolution (LTE), LTE Advanced, Enhanced Data Rate GSM Evolution (EDGE), etc. Other embodiments can be used in a variety of other devices, systems, and / or networks.
[0180] Although an example processing system has been described above, embodiments of the subject matter and functional operation described herein can be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof.
[0181] The embodiments of the subject matter and operation described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs (i.e., one or more components of computer program instructions) encoded on a computer storage medium for execution by or control of an information / data processing device. Alternatively or additionally, program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information / data for transmission to a suitable receiver device for execution by the information / data processing device. The computer storage medium may be or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or combinations thereof. Furthermore, while the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in an artificially generated propagating signal. The computer storage medium may also be or be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0182] The operations described herein can be implemented as operations performed by an information / data processing device on information / data stored on one or more computer-readable storage devices or received from other sources.
[0183] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines used for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or a combination of the foregoing. The apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0184] A computer program (also known as a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or information / data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more components, subroutines, or code sections). A computer program can be deployed on a single computer or located at a single site or distributed across multiple sites and interconnected by a communication network.
[0185] The processes and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input information / data and generating output. As an example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and information / data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving or transferring information / data, or both, from one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer does not need to have such devices. Suitable devices for storing computer program instructions and information / data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM discs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0186] To provide interaction with the user, embodiments of the subject matter described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information / data to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, voice, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.
[0187] Embodiments of the subject matter described herein can be implemented in a computing system that includes backend components, such as an information / data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface or web browser through which a user can interact with embodiments of the subject matter described herein, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital information / data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (e.g., the Internet) and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).
[0188] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship is established by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends information / data (e.g., HTML pages) to the client device (e.g., to display information / data to a user interacting with the client device and to receive user input from the user interacting with the client device). Information / data generated at the client device (e.g., the result of user interaction) may be received at the server from the client device.
[0189] While this specification contains numerous details of specific embodiments, these should not be construed as limiting any embodiment or the scope of any possible claim, but rather as descriptions of features specific to particular embodiments. Certain features described herein in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0190] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0191] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
[0192] Many modifications and other examples described herein will come to mind for those skilled in the art upon benefiting from the teachings presented in the foregoing description and the accompanying drawings. Therefore, it should be understood that the embodiments are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terminology is used herein, it is used only in a general and descriptive sense and not for limiting purposes.
Claims
1. A data analysis method, comprising: Generate a data request for at least a subset of the data stored in the database; Convert the data request into an object storage request; Based on parsing the object storage request, it is determined that the object storage request includes an Extract Transform Load (ETL) request; Create an ETL command based on the object storage request; Based on the fact that the storage device includes at least one ETL function requested in the data request, the ETL command is routed to the storage device; as well as In response to routing the ETL command to the storage device, information is determined based on the result data received from the storage device.
2. The method according to claim 1, wherein, The resulting data is generated based on the subset of data transformed at the storage device according to the ETL function.
3. The method according to claim 1, wherein, Based on the determination that the object storage request includes the ETL request, the ETL command is configured as a Representational State Transfer (REST) command.
4. The method according to claim 3, wherein, The ETL command includes the name of the storage node, which includes the storage device.
5. The method according to claim 1, further comprising: It is determined that the data request is part of a batch data request; as well as Based on the fact that the data request and the second request are associated with the same host server, the data request is combined with at least the second request.
6. The method according to claim 1, further comprising: The second storage device was identified as being associated with an imbalance in processing load. as well as Based on the aforementioned processing load imbalance, at least one of the ETL function of the second storage device or the second ETL command assigned to the second storage device is distributed from the second storage device to the storage device.
7. The method according to claim 1, wherein, The storage device includes at least one of a storage server or a compute storage drive, the compute storage drive including one or more processors configured to perform the conversion function.
8. The method according to claim 1, wherein, The data request includes at least one of the following: Key-value (KV) retrieval, and the database includes a KV database, or Filtering criteria used to narrow down data in the database to the subset of data.
9. The method according to claim 1, wherein, The data request specifies the identifier of the data request.
10. The method according to claim 1, wherein, The object storage request includes an ETL identifier.
11. A data analysis method, comprising: At the storage node, an Extract Transform Load (ETL) command is received, the ETL command including a data request for a subset of data stored in the database; Read the subset of data from the database of the storage node; The data subset is provided to the storage device of the storage node for processing at the storage device; The data subset is processed at the storage device, which includes at least one ETL function requested in the data request. as well as The data is based on the processing and calculation results.
12. The method of claim 11, further comprising at least one of providing the result data to a client device, an application of the client device, a storage device for storing the result data, a different storage device, or a function for requesting the result data, wherein the client device provides a Representational State Transition (REST) command associated with the ETL command.
13. The method of claim 11, further comprising transmitting the result data to the client device providing the ETL command based on the ETL command including RDMA data direct information instructing the client device to include a network interface card supporting remote direct memory access (RDMA).
14. The method of claim 11, further comprising identifying a second data request based on the ETL command being configured as a batch command, the ETL command including the second data request for a second subset of data stored in the database.
15. The method of claim 14, further comprising processing the second data subset at the second storage device based on the fact that the second storage device includes at least the second ETL function requested in the second data request.
16. The method according to claim 11, wherein, The storage device includes at least one of a storage server or a compute storage drive, the compute storage drive including one or more processors for processing the subset of data at the storage device.
17. The method according to claim 11, wherein, The data request includes key-value (KV) retrieval, and the database includes a KV database.
18. A non-transitory computer-readable medium storing code, said code comprising instructions executable by a processor to perform the following: Generate a data request for at least a subset of the data stored in the database; Convert the data request into an object storage request; Based on parsing the object storage request, it is determined that the object storage request includes an Extract Transform Load (ETL) request; Create an ETL command based on the object storage request; Based on the fact that the storage device includes at least one ETL function requested in the data request, the ETL command is routed to the storage device; as well as In response to routing the ETL command to the storage device, information is determined based on the result data received from the storage device.
19. The non-transitory computer-readable medium according to claim 18, wherein, The resulting data is generated based on the subset of data transformed at the storage device according to the ETL function.
20. The non-transitory computer-readable medium according to claim 18, wherein, Based on the determination that the object storage request includes the ETL request, the ETL command is configured as a Representational State Transfer (REST) command.