Enhanced no-code ETL system for automated big data transformation and sharing

JP2025137474A5Pending Publication Date: 2026-04-21CLOUDBLUE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CLOUDBLUE LLC
Filing Date
2025-03-05
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing ETL systems for big data management are inefficient, requiring significant manual intervention, struggle with large data sets, and lack scalability and transparency, especially in scenarios involving thousands of subscriptions and complex data reconciliation.

Method used

A no-code ETL system that automates data processing and sharing, capable of handling large data sets in stream mode, integrating with various data formats, and providing a user-friendly interface for administrators to manage data transformations and validations without coding expertise.

Benefits of technology

Enables efficient, scalable, and transparent data management, reducing manual operations and ensuring data integrity and accuracy, with automated validation and audit capabilities, suitable for complex business environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a system, method and a computer-readable device, for providing a versatile, efficient, and user-friendly ETL system, to improve big data management.SOLUTION: A method includes: extracting data from multiple sources by a stream mode processing unit which segments the data into two or more manageable chunks, the sources being selected from at least one of a cloud storage, an external API and direct file upload; ensuring structural consistency of the extracted data and accuracy of content using a data consistency algorithm; and converting the extracted data using a conversion processing unit. The conversion includes applying predefined and custom transformation templates. The conversion processing unit integrates conversion logics defined by a user through API. The method includes formatting the converted data for loading into destination systems.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Data management, especially in the extract, transform, and load (ETL) of big data, faces significant challenges in efficiency and scalability. Existing ETL systems, such as Amazon Web Services (AWS), AIRFLOW APACHE, and APPACHE BEAM, rely primarily on code-based operations. This approach requires a significant number of programmers to manage data processing, making it impractical for scenarios that require simplified and automated management oversight, such as reconciliation of thousands of subscriptions.

[0002] Existing systems struggle to process large data sets in a timely and efficient manner. Data reconciliation, billing, and pricing processes often involve overly large Excel files and other formats, require significant manual intervention, and lead to prolonged operational periods, sometimes stretching into weeks. This inefficiency is exacerbated when dealing with tabular big data, as these systems reach their limits in handling files above a certain size without compromising performance.

[0003] Additionally, the lack of automation in the previous ETL process required a high degree of manual input, leading to scalability issues. As the company grew and the number of contractors increased globally, the need for a more robust and scalable system became apparent. Additionally, the current system suffers from weak audit capabilities due to the involvement of multiple systems and operators, and a reliance on email communication, which lacks transparency and traceability.

[0004] These shortcomings highlight the need for innovative solutions that can efficiently handle large-scale data processing, ensure scalability, and reduce reliance on manual operations, thereby improving overall operational efficiency and transparency in big data management. Summary of the Invention

[0005] The embodiments described herein improve big data management by providing a versatile, efficient, and user-friendly ETL system. Systems and methods are provided that address the manual operational, scalability, and auditability complexities of traditional ETL processes and facilitate a comprehensive solution to modern big data needs. Specifically, the disclosed embodiments provide a novel no-code ETL (extract, transform, load) system configured to automate and simplify the process of big data management, particularly targeting reconciliation, billing, and pricing operations in business environments. The system and method are uniquely configured to go beyond traditional programmers and enable administrators to automate data processing and sharing with minimal human intervention.

[0006] In some embodiments, systems and methods are provided for processing large data streams, including, for example, tabular big data, without the size limitations typical of conventional systems. In some embodiments, the systems and methods are scalable to handle files significantly larger than previous limits (e.g., greater than 1 GB), making the disclosed embodiments particularly well-suited for processing large volumes of subscription data in a medium suitable for data exchange (e.g., Microsoft Excel, etc.).

[0007] In some embodiments, the system and method can be adaptable to a variety of data formats, integrate with external CRM systems via APIs, and customize data presentation using adapters. Such flexibility can extend to the system's output capabilities, which are not limited to Excel but include other formats such as CSV and JSON.

[0008] In some embodiments, no-code operation is enabled and the system is configured with a simplified user interface, making it easy to upload data, configure transformations, and share results without requiring programming skills.

[0009] In some embodiments, comprehensive data processing supports a wide range of transformations, such as copying, renaming columns, mathematical operations, currency conversion using real-time rates, and data filling from various sources. Data processing can be enabled by an embedded programming language and the ability to extend functionality through low-code modification.

[0010] In some embodiments, for example, automated data validation and auditing can be performed before conversion begins, checking each data component for consistency, and conversion steps can be logged, improving transparency and accountability.

[0011] In some embodiments, the system can be scaled and adapted for different partners, marketplaces, and / or products. The system's modular design allows for replication of data streams for use with other partners, improving operational flexibility. In some embodiments, a test environment can be provided, allowing users to test configurations on sample data before applying them to production data, ensuring reliability and accuracy of data processing. A notification system incorporates email and / or internal notifications to alert users when transformations require human intervention, streamlining the process flow. In some embodiments, the system ensures data privacy by limiting visibility within an account and providing the option to publish results to specific partners under defined agreements. [Brief explanation of the drawings]

[0012] [Figure 1] 1 illustrates a system for automated no-code ETL, according to some embodiments. [Figure 2] 1 illustrates an example operating environment for an automated no-code ETL system, according to some embodiments. [Figure 3] 1 illustrates a data extraction component of a system for automated no-code ETL, according to some embodiments. [Figure 4] 1 illustrates a data transformation component of a system for automated no-code ETL, according to some embodiments. [Figure 5] 1 illustrates an automated process for no-code ETL, according to some embodiments. [Figure 6] 1 illustrates a data extraction process of a system for automated no-code ETL, according to some embodiments. [Figure 7] 1 illustrates a data transformation process of a system for automated no-code ETL, according to some embodiments. [Figure 8] 1 illustrates an exemplary computer system, according to some embodiments. [Figure 9] 1 illustrates a user interface (UI) for automated no-code ETL, according to some embodiments. [Figure 10] 1 illustrates a user interface (UI) for automated no-code ETL, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0013] The embodiments may be implemented in hardware, firmware, software, or any combination thereof. The embodiments may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and others. Furthermore, firmware, software, routines, and instructions may be described herein as performing particular actions. However, it should be understood that such description is merely for convenience and that such actions are actually the results obtained by a computing device, processor, controller, or other device executing the firmware, software, routines, instructions, etc.

[0014] It should be understood that the acts shown in the example methods are not exhaustive and that other acts may similarly occur before, after, or between any of the acts shown. In some embodiments of the present disclosure, acts may be performed in a different order and / or may differ.

[0015] FIG. 1 illustrates a system 100 for automating big data reconciliation, billing, and pricing processes. The system, referred to as a no-code ETL (extract, transform, load) system 100, is configured to facilitate the automated processing and publishing of data streams, particularly in tabular formats such as Excel files. With reference to FIG. 1 , a no-code ETL system 100 for automating and simplifying complex data manipulation is provided. In one non-limiting example, the no-code ETL system 100 for processing big data can include a data extraction module 102 and a data transformation module 104.

[0016] In some embodiments, the data extraction module 102 can include functionality for interfacing with various data sources, including cloud-based platforms and external APIs. For example, the data extraction module 102 can be configured to pull data from a client's CRM system, providing flexibility in sourcing the initial data. While the module focuses on Excel files here, it is capable of handling a variety of file types and is configured to read files in stream mode, thereby mitigating technical limitations on file size.

[0017] In some embodiments, stream mode processing, which allows for handling large data sets, especially when dealing with tabular big data, is implemented in the high-performance no-code ETL system of the present invention. Those skilled in the art will appreciate that stream mode processing techniques enable efficient reading and processing of data in manageable chunks, significantly reducing memory overhead and facilitating real-time data processing. By leveraging stream mode, the system can dynamically adjust to varying data size and complexity, ensuring optimal performance without sacrificing processing speed and accuracy. Stream mode processing facilitates operations requiring continuous data flow, such as real-time data transformation and loading, and provides a scalable solution for no-code ETL methodologies.

[0018] In one non-limiting example, the data extraction module 102 can use the .NET Framework and implement the FileStream.Read method to perform stream-mode data extraction from large files. The FileStream.Read method allows the module to read data in segments, enabling processing of large data sets without having to load the entire file into memory. By adopting this stream-based approach, the system maximizes performance with minimal memory consumption, facilitating a more scalable and responsive data extraction process.

[0019] In some embodiments, the data transformation module 104 can include a set of tools and functionality for processing the extracted data. This module allows users to configure a series of transformations and operations on the data, including mathematical operations, currency conversions using real-time rates, and data filling from various sources. This module supports a wide range of out-of-the-box transformation functions, such as manufacturer's suggested retail price (MSRP) and cost of goods sold (COGS) calculations, markup and margin calculations, and value-added tax (VAT) rate application. Additionally, the data transformation module 104 can support custom transformations via one or more application programming interfaces (APIs) (e.g., the CLOUDBLUE CONNECT Transformation Extension Module API), allowing for further customization and flexibility.

[0020] To improve the system's capabilities for data extraction, in some embodiments, the data extraction module 102 can be configured to interact with the data conversion module 104 by providing automatically queued data transitions for transformation (e.g., after extraction and structural and content validation). The interaction between the data extraction module 102 and the data conversion module 104 can be configured to ensure that the extracted data conforms to transformation criteria set and interpreted by the data conversion module 104, which can then apply the required operations without requiring manual intervention.

[0021] In some embodiments, the no-code ETL system 100 can include a UI module 106 configured for ease of use, allowing users to upload data files, configure transformations, and monitor the entire data processing workflow. The UI module 106 can also be configured to facilitate sharing processed data with partners in either an automated or semi-automated mode. Exemplary UIs are shown in Figures 9 and 10.

[0022] As shown in FIG. 9 , UI 900 is for a high-performance, no-code ETL system designed for transforming and sharing big data. It is configured to allow a wide range of transformation logic to be implemented without coding expertise, for example, by data transformation module 104. In a non-limiting example, a user can customize pricing streams as shown, implementing granular rules in a UI configured for ease of use, such as browsing products, searching suggested retail price (MSRP) and cost of goods sold (COGS) values, etc. FIG. 10 illustrates UI 1000 for a high-performance, no-code ETL system designed for transforming and sharing big data. As shown, UI 1000 includes a section for outbound pricing batch management, including status indicators 1010 and 1012. Status indicator 1010 indicates the overall progress of batch processing, displaying a percentage indicating how much of the batch is complete. Another status indicator 1012 displays the finalized or scheduled status of individual transformations using a similar percentage-completed metric. These indicators form a comprehensive dashboard that allows users to efficiently monitor and manage data transformation tasks, including viewing batch details, processing transformation requests, and tracking overall progress. UI900 and UI1000 are each designed to streamline operations for users without coding expertise, allowing them to configure, run, and monitor ETL processes through a simplified, intuitive interface.

[0023] The data conversion module 104 can be configured to initiate one or more pre-configured or custom conversions upon receiving data from the data extraction module 102. The data conversion module 104 can be configured to communicate with the UI module 106 to receive and apply user-defined conversion rules accordingly. In some embodiments, a bidirectional information flow between the data conversion module 104 and the UI module 106 allows a user to set conversion parameters and receive feedback regarding the conversion results.

[0024] In some embodiments, the system 100 may further include an automated data validation and audit module 108. This module is responsible for validating the data structure row by row against a predefined sample file structure and provides functionality for checking cell values ​​based on user-defined constraints, such as blank cell tolerance and decimal point precision.

[0025] In some embodiments, the system 100 may include a data publishing module 110. This module enables the distribution of processed data to various parties, e.g., vendors, distributors, resellers, etc. This module is particularly adept at handling complex computational chains where data must go through multiple processing and approval stages before it can be shared.

[0026] The UI module 106 can provide a central hub for user interaction with the system. For example, it can collect input to the data transformation module 104 and display processed data from the data publication module 110. This allows the UI module 106 to aggregate the transformation rules applied by the data transformation module 104 and the output prepared for publication, providing an informative overview of the data lifecycle within the system 100.

[0027] In some embodiments, the data publishing module 110 can receive processed and transformed data from the data transformation module 104 and format this data into publishable content for delivery to target users / entities. From a scalability and flexibility perspective, the no-code ETL system 100 can allow account administrators to replicate data streams for use with different partners or marketplaces. This feature ensures that the system can adapt and scale as the needs of the business grow and change.

[0028] Additionally, the system incorporates a notification module 112 that alerts users via email and an internal notification system when the conversion process requires human intervention. This module helps streamline operations and ensure timely human input when needed. The data publication module 110 can operatively interact with the notification module 112 to notify users of publication status and required actions, ensuring all users and entities have access to current data.

[0029] The system's architecture allows for the connection of data streams to partner accounts, marketplaces, and individual products or plans. This capability ensures that the processed data is relevant and can be effectively utilized for its intended purpose, such as pricing, billing, or reconciliation.

[0030] In some embodiments, system 100 can include a testing module 114, which allows administrators to test configurations on sample data before deploying the transformations to production data. This module is provided to ensure that the data processing workflow remains reliable and accurate. In some embodiments, testing module 114 can be configured to provide a controlled environment for testing the data transformations applied by data transformation module 104.

[0031] In some embodiments, the notification module 112 can be coordinated between the test module 114 and the UI module 106. The notification module 112 can alert a system administrator to conversion errors detected during testing and prompt necessary adjustments through the UI module 106 to ensure data integrity before publishing.

[0032] In some embodiments, the system 100 can include an internal database and logging module 116 that maintains a comprehensive log of all transformation changes, including the time of action, users involved, and associated comments. This module improves the transparency and traceability of the system, allowing users to effectively audit the transformation process. The testing module 114 can be integrated with the internal database and logging module 116 and configured to log all test results for audit and compliance purposes. A feedback loop between modules can ensure that only verified, accurate transformations are deployed to production.

[0033] In some embodiments, the data processing engine 118, which is part of the cloud platform EaaS module, handles the actual data processing tasks. It uses a multi-threaded approach to break down source data into parts that can be processed in parallel, ensuring efficiency and speed in handling large data sets. The data processing engine 118 executes data processing tasks and can integrate with the data extraction module 102 and the data transformation module 104 to ensure that computing resources are efficiently allocated and that the data flow between the extraction and transformation stages is optimized for performance. This configures the no-code ETL system 100 to enable the automated processing, transformation, and sharing of big data for reconciliation, billing, and pricing processes in business environments.

[0034] 2 illustrates a block diagram of an operating environment 200 in which a no-code ETL system 216 may be implemented, which may be one embodiment of the system 100 previously described in FIG. 1. The operating environment 200 may include various components that each contribute to the efficient functioning of the no-code ETL system. It is important to note that in other embodiments, the operating environment 200 may not have all of the listed components and / or may include other elements instead of or in addition to the elements listed herein. In this disclosure, a user of the system may be referred to interchangeably as a customer, client, or operator.

[0035] Operating environment 200 is the environment in which no-code ETL system 216 resides and operates effectively. User systems 212 can be machines or systems used by users to interact with no-code ETL system 216. For example, user systems 212 can include handheld computing devices, mobile phones, laptops, workstations, or a network of computing devices. As illustrated in FIG. 2 (and FIG. 1), user systems 212 interact with no-code ETL system 216 via network 214, which can include multiple components configured to perform aspects of a no-code ETL process.

[0036] The no-code ETL system 216 encompasses an integrated data processing and transformation platform and is specifically adapted to perform such processes in an automated manner, allowing complex operations to be performed by users with limited knowledge of the system's technical intricacies. Instead, users can leverage the system for various data manipulations, such as data extraction, transformation, and loading. The application platform 218 provides a framework that enables applications of the system 216, including hardware and software resources such as an operating system. The no-code ETL system 216 may include an application platform 218 that facilitates the creation, management, and execution of various applications developed by either the system provider, users, or third-party developers.

[0037] System 200 may include multiple processes for data manipulation and transformation. For example, it may identify and extract data from various sources, transform this data based on user-configured rules and processes, and load the transformed data into a specified destination. These functionalities are enabled by various components of the system, including no-code ETL data storage 222, system data storage 224, process space 228, and processor system 217.

[0038] Network 214 is a network or combination of networks of devices communicating with each other. It may be a LAN, a WAN, a telephone network, a wireless network, or any other suitable configuration. The most common type of network in use today is a TCP / IP network, e.g., the Internet, although it should be understood that network 214 is not limited to this configuration and may include other network types.

[0039] User system 212 may communicate with system 216 at a higher network level using TCP / IP, and may use other common Internet protocols such as HTTP, FTP, or WAP. In some implementations, user system 212 may include a "browser" for sending and receiving messages to and from system 216 via network 214. The interface between system 216 and network 214 may include load-sharing functionality, such as a round-robin request distributor, to balance the load and distribute incoming requests evenly across multiple servers.

[0040] 2 implements a no-code ETL platform. For example, system 216 may include an application server configured to implement and run no-code ETL software applications, and may also receive and provide associated data, code, forms, web pages, and other information from user systems 212, and store and retrieve associated data and objects from database systems.

[0041] 2 include conventional, well-known elements that are only briefly described herein. For example, each of user systems 212 may include a personal computer, laptop, PDA, cellular phone, or WAP-enabled device capable of interfacing directly or indirectly with the Internet or other network connection. Each user system 212 typically runs a browser, allowing a user to access, process, and view information and applications available from system 216 via network 214.

[0042] According to one embodiment, each of user systems 212, and all of its components, may be operator-configurable using an application such as a browser. Similarly, system 216, and all of its components, may be operator-configurable using an application that includes computer code executed by a processor, such as processor system 217. Program code 228 may include instructions that may be used to program a computer to perform any of the processes of the embodiments described herein.

[0043] One arrangement of elements of system 216 is shown in Figure 2 and includes network interface 220, application platform 218, no-code ETL data storage 222 for storing configuration and transformation data, system data storage 224 for storing operational data, program code 228 for implementing various functions of system 216, and process space 228 for executing service processes and system-specific processes. Process space 228 may be used to run applications and host services as part of the no-code ETL platform. In some embodiments, process space 228 may include a multi-threaded processing environment capable of executing different data transformation tasks in parallel, improving overall system efficiency.

[0044] The no-code ETL data storage 222 is a critical component to the functioning of the system 216. It stores the configurations, rules, and templates used in the data transformation process. This storage is configured to be highly flexible and scalable to address the evolving needs of various data processing tasks. It allows users to easily and efficiently store and retrieve transformation configurations, thereby facilitating the rapid deployment of data processing tasks.

[0045] System data storage 224 provides a repository for operational data for the system 216. This can include logs, user information, system status reports, and other data essential to the smooth operation and monitoring of the system. The storage is built to handle large amounts of data and is optimized for quick access and data retrieval, ensuring that system performance is maintained at optimal levels.

[0046] Program code 228 comprises the core software logic that drives the functionality of system 216. It may include algorithms and processing logic for extracting data from various sources, applying user-configured transformations, and loading the processed data to a desired destination. This code is provided to promote flexibility and customization, allowing users to define data processing logic without requiring extensive programming knowledge.

[0047] Application platform 218 is the backbone of system 216, providing the hardware and software infrastructure necessary for running program code 228 and storing and retrieving data from no-code ETL data storage 222 and system data storage 224. The platform is configured to be robust and scalable, capable of supporting large volumes of data processing while maintaining high performance and reliability.

[0048] The processor system 217 is responsible for executing the instructions of the program code 228. It comprises one or more processor units capable of handling the intensive computational tasks involved in extracting, transforming, and loading data. The processor system is selected and optimized to handle the specific requirements of the no-code ETL process, ensuring that data manipulations are performed quickly and efficiently.

[0049] The process space 228 is where the actual execution of data processing tasks takes place. It provides an isolated and secure environment for executing the various processes and applications that form part of the system 216. This space is configured to maximize processing efficiency and ensure the integrity and security of the data being processed.

[0050] Some of the elements of the system shown in FIG. 2 are well-known and conventional, and will only be briefly described herein. For example, each of user systems 212 may include a desktop personal computer, a workstation, a laptop, a PDA, a mobile phone, or a Wireless Access Protocol (WAP)-enabled device, or other computing device capable of interfacing directly or indirectly with the Internet or other network connection. Each of user systems 212 typically executes a browsing program, such as an HTTP client, e.g., Microsoft's Edge browser, Google's Chrome browser, or Opera browser, or a WAP-enabled browser in the case of a mobile phone, PDA, or other wireless device, to enable a customer of user system 212 to access, process, and view information, pages, and applications available from system 216 via network 214. Each of user systems 212 also typically includes one or more user interface devices, such as a keyboard, mouse, trackball, touchpad, touchscreen, pen, etc., for interacting with a graphical user interface (GUI) provided by a browser on a display (e.g., a monitor screen, LCD display, etc.) along with pages, forms, applications, and other information provided by system 216 or other systems or servers. For example, the user interface devices may be used to access data and applications hosted by the system 216, conduct searches of stored data, and otherwise enable a user to interact with various GUI pages that may be presented to the user. As noted above, embodiments are suitable for use with the Internet, which refers to a specific global internetwork of networks. However, it should be understood that other networks may be used in place of the Internet, such as, for example, an intranet, an extranet, a virtual private network (VPN), a non-TCP / IP-based network, any LAN or WAN, etc.

[0051] According to one embodiment, each of user systems 212, and all of its components, is operator-configurable using an application, such as a browser, that includes computer code executed using a central processing unit, such as an Intel Pentium® processor. Similarly, the system and all of its components may be operator-configurable using an application that includes computer code that executes using a central processing unit, such as processor system 217, which may include an Intel Pentium® processor, and / or multiple processor units. An embodiment of a computer program product may include a machine-readable storage medium having instructions stored thereon that can be used to program a computer to perform any of the processes of the embodiments described herein. The computer code for operating and configuring system 216 to interact with and process web pages, applications, and other data and media content described herein is downloaded and stored, for example, on a hard disk, although the entire program code, or portions thereof, may also be stored in other known volatile or non-volatile memory media or devices, such as ROM or RAM, or may be provided on media capable of storing program code, such as floppy disks, optical disks, digital versatile disks (DVDs), compact disks (CDs), microdrives, and magneto-optical disks, as well as magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of media or device suitable for storing instructions and / or data. Additionally, the entire program code, or portions thereof, may be transmitted and downloaded via a transmission medium, for example, via the Internet, from a software source or from another server, as is known, or transmitted via other conventional network connections, as is known (e.g., extranet, VPN, LAN, etc.), using known communication media and protocols (e.g., TCP / IP, HTTP, HTTPS, Ethernet, etc.).It will also be understood that computer code for implementing the embodiments may be implemented in a programming language that may be executed on client systems and / or servers or server systems, such as C, C++, HTML, other markup languages, Java, JavaScript, ActiveX, other scripting languages ​​such as VBScript, and many other well-known programming languages ​​(Java is a trademark of Sun Microsystems, Inc.).

[0052] This provides an operating environment 200 for the no-code ETL system 100. It enables automated processing, transformation, and sharing of big data for a variety of business applications. The operating environment is configured to be flexible, scalable, and user-friendly to meet the diverse needs of users involved in complex data manipulation.

[0053] 3 illustrates a data extraction module 300, which may be one embodiment of the data extraction module 102 in the no-code ETL system 100. This module is specifically configured to efficiently extract and process data from various sources. With reference to FIG. 3, a data extraction system 300 for handling the initial phase of data processing is provided. In one non-limiting example, the data extraction system 300 for streamlining data collection may include a stream mode processing unit 302 and a source compatibility interface 304.

[0054] The stream mode processing unit 302 can include advanced algorithms and techniques for reading and processing data in stream mode. This functionality makes it possible to handle large data files, such as massive Excel files, by processing the data in manageable chunks rather than loading the entire file into memory. This approach significantly improves the efficiency and performance of the data extraction process, especially when dealing with massive data sets.

[0055] The stream mode data processing capabilities of the data extraction module 300, which is part of the no-code ETL system 100, provide a technological approach to address the challenge of effectively handling large data files. The stream mode processing unit 302 can be designed to process data files and efficiently perform batch processing. In some embodiments, the stream mode processing unit 302 can be configured to read and process data in real time. In some embodiments, the stream mode processing unit 302 can be configured to read and process large data files by breaking them down into smaller, manageable chunks or batches. This allows the data extraction module 300 to handle data in a controlled and efficient manner that is suitable for processing large files.

[0056] Batches can be processed sequentially, ensuring that system memory usage is kept within manageable limits. This method significantly reduces the risk of memory overload, which is often associated with processing large files. By splitting the data, modules can treat each segment separately and apply any necessary extraction processes before moving on to the next. This batch processing approach also allows the system to pause, resume, or restart processing of data files as needed, adding an additional layer of flexibility and control.

[0057] The stream mode processing unit 302 can employ advanced algorithms to efficiently parse and process each batch. These algorithms are configured to optimize the reading, extraction, and initial processing of data, ensuring that each batch is processed promptly and accurately. This enables data-intensive applications to run while optimizing time and resources. Additionally, the data extraction module 300 can adjust the size of each batch based on the overall file size and the current system load. This scalability ensures that the module can adapt to fluctuating file sizes and system capacity, maintaining optimal performance regardless of data volume.

[0058] The stream mode processing unit 302 enables effective error handling and data integrity checking. Because the system processes data in successive batches, validation and error checking routines can be performed on each batch independently. This granularity in data processing identifies and addresses inconsistencies and errors within a particular batch without affecting the entire file. This allows the stream mode processing unit 302 to process data sequentially, working with a portion of data at a time in a controlled batch fashion, so that data is processed as it is read.

[0059] In this way, the data extraction module 300 can manage large data files using stream-mode processing. This method offers a balance of efficiency and control, allowing the system to handle large data files while reducing memory strain and improving processing speed. The module's algorithms are tailored to optimize batch processing, ensuring scalability, accurate data processing, and robust error management. While this approach differs from real-time streaming, it offers significant advantages in terms of resource management and operational flexibility, especially in scenarios involving large-scale data extraction.

[0060] The source compatibility interface 304 is configured to enable the data extraction module 300 to interface with a variety of data sources, including cloud storage platforms, CRM systems, external APIs, and direct user uploads. The interface includes a wide range of connectors and adapters, making it adaptable to different data environments and formats. This adaptability ensures that the module can effectively extract data regardless of its source or format, such as Excel, CSV, JSON, or other common data formats. The data extraction module 300 exhibits the technical capabilities to interface with a wide range of data sources. This versatility is achieved through the integration of multiple data connectors and adapters. These components are designed to establish a connection with a CRM system, access its database, and efficiently extract relevant data. In addition, the module includes interfaces for communicating with various external APIs, enabling it to access and retrieve data from a wide range of external systems and platforms. The module's adaptability extends to handling different file formats, including, but not limited to, Excel, CSV, and JSON. This functionality is based on advanced parsing algorithms that can interpret and process the structural nuances of these diverse formats, ensuring accurate data extraction and minimizing format-related errors.

[0061] Additionally, the data extraction module 300 incorporates data integrity mechanisms to ensure the accuracy and consistency of the extracted data. These mechanisms involve validation processes that check the structure and content of the data against predefined criteria and identify anomalies or inconsistencies. This feature is essential to maintaining the quality and reliability of data passing through subsequent processing stages of the ETL system.

[0062] Advanced data integrity assurance mechanisms systematically evaluate data against predefined criteria and identify inconsistencies such as missing values, format inconsistencies, and data corruption to ensure the accuracy of each data batch and maintain consistency across the dataset. This configures the data extraction module 300 to efficiently extract data from various sources while ensuring data integrity, setting a solid foundation for the effective operation of the system 100 in a variety of data processing scenarios.

[0063] 4 illustrates an embodiment of a data transformation module 400, which may be a specific embodiment of the data transformation module 104 in the no-code ETL system 100. This module facilitates advanced data transformations required for various business applications. Referring to FIG. 4, the data transformation system 400 may include a transformation processing unit 402 and a custom transformation integration unit 404 configured to handle complex data manipulations.

[0064] The transformation processing unit 402 employs algorithms to perform a wide range of data transformations. These transformations cover mathematical operations, currency conversions, and advanced data structuring. The unit uses programming languages ​​such as JQ to enable complex and flexible data manipulation. This capability allows users to tailor data transformations to their specific requirements.

[0065] The custom transformation integration unit 404 allows users to add custom transformations to the data transformation module 400. This unit integrates with the CloudBlue Connect transformation extension module API to provide a platform for incorporating user-defined transformations. Users write custom code according to the system's guidelines, which the system then integrates to extend the module's functionality.

[0066] In some embodiments, the data transformation module 400 can be operatively connected to an automation platform (e.g., CLOUDBLUE CONNECT) via an API or the like, enabling the integration of complex distribution and supply chain applications. This integration facilitates interactions with various entities in the distribution chain, including vendors, distributors, and resellers. The module can integrate an API gateway to enable users to create custom transformations tailored to specific supply chain and distribution scenarios. This can include automating interactions between different entities or integrating various distribution chain data formats into a unified processing system.

[0067] Integration allows for the inclusion of transformation logic tailored to specific user needs. Users can develop transformations that may include unique business logic or data manipulation routines that are not available in the standard set of transformations. This allows the system to incorporate user-defined transformations into a no-code operational framework. This capability is based on a dynamic linking mechanism, allowing the system to recognize and execute these custom transformations as if they were native components of the module. This makes the system more flexible and adaptable to niche and evolving data processing requirements.

[0068] The data transformation module 400 also features a user-friendly configuration interface. This interface streamlines the setup and application of transformations, making the process accessible even to users with limited programming expertise. It allows for straightforward selection, customization, and testing of transformations, ensuring accurate data processing according to the user's specifications. In accordance with the system's no-code operating principle, the interface presents users with a visual representation of the data flow and transformation process, allowing them to understand and configure transformations without having to write or understand complex code. The interface can also include features such as drag-and-drop capabilities, pre-built transformation templates, and interactive guides, which help simplify the configuration process. This user-centered design approach enables a wider range of users, including those with minimal technical background, to leverage advanced data processing capabilities, thereby fostering more inclusive use of the system across various organizational roles. This allows the data transformation module 400 to be configured to handle a variety of data transformations, from standard to custom processes, significantly contributing to the versatility of the system 100 in various data transformation contexts.

[0069] It should be understood that the acts shown in the example methods are not exhaustive and that other acts may similarly occur before, after, or between any of the acts shown. In some embodiments of the present disclosure, acts may be performed in a different order and / or may differ.

[0070] 5 is a flow diagram of a method 500, one embodiment for performing no-code ETL processes within system 100. The method provides an efficient approach to data handling and processing, encompassing data extraction, transformation, and loading. Method 500 is configured to efficiently manage data workflow, ensuring accuracy, scalability, and adaptability to various data environments. Based on the disclosure herein, the operations in method 100 may be performed in a different order and / or may be varied.

[0071] In operation 502, the computing device may perform data source identification. This step involves identifying the origin of the data, which may be a CRM system, an external API, cloud storage, or direct upload. The process may include evaluating the expected data format and structure from these sources to ensure that subsequent extraction processes are tailored to handle the data effectively.

[0072] In operation 504, the computing device can begin the data extraction process. This step involves the data extraction module 300 employing a stream-mode data processing technique. The module processes data in manageable batches, reducing memory load and improving processing efficiency. This step is important for working with large data files, such as massive Excel documents, where loading the entire file into memory is not feasible.

[0073] In operation 506, the computing device may perform a data integrity check. This validation process may include examining the data for structural accuracy, content consistency, and identifying anomalies or inconsistencies. The system employs a series of algorithms configured to detect and address issues such as missing values, format inconsistencies, and data corruption. This step is critical to ensuring the extracted data is reliable and suitable for further processing.

[0074] In operation 508, the computing device can determine whether the extraction and validation were successful. If successful, the system transitions to the data transformation phase. In this step, the user interacts with the user-friendly configuration interface of the data transformation module 400. The user can select from a wide range of pre-built transformation templates or configure a custom transformation. This process can include defining transformation logic, such as mathematical operations, currency conversions, and data structuring, tailored to the specific needs of the data processing task.

[0075] In operation 510, the computing device can perform optional specialized data manipulation. If required by the user, the system provides the ability to integrate custom transformations. Users can write and integrate transformation code via an API (e.g., the CLOUDBLUE CONNECT API). Operation 510 can include applying custom business logic or data manipulation routines, improving the adaptability of the system to specific user requirements.

[0076] In operation 512, the set configuration may cause the computing device to perform data transformation. Operation 512 may include applying defined transformation logic to the extracted data. The transformation processing unit 402 of module 400 may be configured to perform operation 512, such that each data batch is processed according to the configured rules and conditions. Operation 512 may include converting the extracted raw data into a format that is meaningful and useful to the end user.

[0077] In operation 514, the computing device may load the processed data into a designated destination system, which may include a database, data warehouse, or other storage system. The loading process is configured to be efficient and ensures that the transformed data is accurately and completely integrated into the target system.

[0078] In operation 516, the computing device can continuously monitor and manage the data flow. This can include tracking the progress of data extraction, transformation, and loading, as well as identifying and addressing any issues that may arise during the process. The system provides tools and interfaces for users to oversee the ETL workflow, providing insight into each step and the ability to intervene if necessary.

[0079] In operation 518, the computing device can collect feedback and iterate the process. Based on the performance of the ETL workflow and feedback from the user, the system can adjust and refine the process. This may involve fine-tuning transformation configurations, optimizing extraction methods, or modifying data loading techniques. This iterative approach ensures continuous improvement of the ETL process, adapting to changing data requirements and user needs.

[0080] Method 500 thereby provides a no-code ETL process performed by a computing device such as system 100. The no-code ETL process manages complex data workflows and enables diverse data processing requirements in various business environments.

[0081] 6 illustrates a method 600 for performing a data extraction process within the no-code ETL system 100. The method 600 ensures a reliable extraction stage of the ETL process.

[0082] In operation 602, the computing device identifies and evaluates data sources. This step involves a detailed analysis of the origins of the data, including CRM systems, external APIs, cloud storage platforms, or direct uploads. The format, structure, and specific characteristics of the data from these sources are evaluated to set the stage for optimal extraction.

[0083] In operation 604, depending on the nature of the data source, the computing device selects the most suitable extraction method. This may require a direct API call in the case of a cloud-based source, an SQL query in the case of a database system, or a parsing algorithm in the case of a file-based source. The selection is based on efficiency and compatibility of the method with the data source.

[0084] In operation 606, the computing device configures the necessary data connectors and adapters. This step ensures integration with the data source and allows the system to efficiently access and retrieve data. The connectors and adapters are tailored to handle the specific data protocols and formats native to the source.

[0085] In operation 608, the computing device initiates an extraction process to pull data from the source based on the configured methods and connectors. This process is performed while maintaining data integrity, ensuring that data is accurately captured from the source without loss or corruption.

[0086] For particularly large data files, in operation 610, the computing device employs a stream-mode data processing approach, which involves reading and processing data in manageable chunks, effectively reducing memory load and increasing processing speed. The system dynamically adjusts the size of these chunks based on file size and system capacity.

[0087] In operation 612, the computing device performs a series of data integrity checks and validation procedures, which may include verifying data format, ensuring structural accuracy, and detecting anomalies and inconsistencies. These checks are important to ensure the quality and reliability of the extracted data.

[0088] In operation 614, the computing device checks whether errors or problems occurred during the extraction and, if so, invokes one or more mechanisms to address and resolve them. This may include retrying the extraction process, adjusting extraction parameters, or flagging the problem for user intervention. In this way, operation 614 verifies and corrects to maintain the continuity and efficiency of the extraction process.

[0089] In operation 616, the computing device formats and standardizes the extracted data. This step ensures that the data adheres to a consistent structure and format, facilitating integration with subsequent transformation processes in the ETL workflow.

[0090] If the extraction and standardization are successful, then at operation 618 the computing device transitions to a data transformation phase, where the extracted data is ready for further manipulation and processing defined in subsequent stages of the ETL system. Method 600 thereby provides an efficient approach to data extraction within a no-code ETL system.

[0091] 7 illustrates a method 700 for performing data transformation processes within the no-code ETL system 100. The method 700 provides an efficient approach for transforming extracted data to ensure its relevance, accuracy, and suitability for the intended use.

[0092] Operation 702 begins with a computing device receiving data from a data extraction module. This data arrives in raw form directly from various data sources and requires transformation to achieve the desired structure and content.

[0093] In operation 704, the computing device analyzes the structure of the received data and identifies transformation requirements. This step involves understanding the end uses of the data, which can range from analytical processing to report generation, and defining the transformation logic necessary to meet those objectives.

[0094] In operation 706, the computing device displays pre-built transformation templates suitable for common data scenarios. In this step, the user selects an appropriate template that matches the data transformation goals. These templates simplify the transformation process, especially for users without significant technical expertise.

[0095] In operation 708, the computing device optionally enables scenarios requiring specialized data manipulation and allows the user to configure custom transformations, which may involve writing transformation logic or scripts, optionally using a programming language such as JQ, to define specific data manipulation routines.

[0096] In operation 710, the computing device uses the API to integrate any custom transformations into the transformation process. This step allows the system to process these user-defined transformations in parallel with standard transformations, increasing the flexibility and power of the data transformation module.

[0097] In operation 712, the computing device performs a conversion process, which applies selected or custom conversion logic to the raw data to convert it into the required format for its intended use.

[0098] In operation 714, the computing device validates the transformed data to ensure it meets predefined criteria, checking the data for consistency, accuracy, and alignment with what it was transformed to.

[0099] In operation 716, the computing device identifies errors or inconsistencies in the transformed data and takes necessary corrective action, which may include reapplying the transformation with adjusted parameters or flagging the issue for manual review and intervention.

[0100] In operation 718, the computing device formats the validated and converted data into a final structure, preparing it for loading into a target system or for direct use. This step ensures that the data is not only accurate and relevant, but also presented in a way that is accessible and understandable to the end user.

[0101] In operation 720, upon successful conversion and formatting, the computing device outputs the data to be loaded into a specified destination system, such as a database or data warehouse. This completes the data conversion process and transitions to the final phase of the ETL workflow. This enables method 700 to enable a data conversion process within a no-code ETL system. While the data involved in the process undergoes the necessary transformations, the process maintains the integrity and specific requirements of the end use.

[0102] 8 is a block diagram of example components of device 800. One or more computer systems 800 may be used, for example, to implement any of the embodiments described herein, as well as combinations and subcombinations thereof. Computer system 800 may include one or more processors (also referred to as central processing units, or CPUs), such as processor 804. Processor 804 may be connected to a communications infrastructure or bus 806.

[0103] The computer system 800 may also include user input / output devices 803 , such as a monitor, keyboard, pointing device, etc., which may communicate with a communications infrastructure 806 through a user input / output interface 802 .

[0104] One or more of the processors 804 may be a graphics processing unit (GPU). In one embodiment, a GPU may be a processor that is a specialized electronic circuit designed to process mathematically intensive applications. A GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common in computer graphics applications, images, video, etc.

[0105] Computer system 800 may also include a main or primary memory 808, such as random access memory (RAM). Main memory 808 may include one or more levels of cache. Main memory 808 may have control logic (i.e., computer software) and / or data stored therein.

[0106] Computer system 800 may also include one or more secondary storage devices or memory 810. Secondary memory 810 may include, for example, a hard disk drive 812 and / or a removable storage device or drive 814.

[0107] The removable storage drive 814 may interact with a removable storage unit 818. The removable storage unit 818 may include a computer-usable or readable storage device having computer software (control logic) and / or data stored thereon. The removable storage unit 818 may be a program cartridge and cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or other removable storage unit and associated interface. The removable storage drive 814 may read from and / or write to the removable storage unit 818.

[0108] Secondary memory 810 may include other means, devices, components, intermediaries, or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 800. Such means, devices, components, intermediaries, or other approaches may include, for example, removable storage unit 822 and interface 820. Examples of removable storage unit 822 and interface 820 may include a program cartridge and cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or other removable storage unit and associated interface.

[0109] Computer system 800 may further include a communications or network interface 824. Communications interface 824 may enable computer system 800 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referred to by reference number 828). For example, communications interface 824 may enable computer system 800 to communicate with external or remote devices 828 via communications path 826, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to or from computer system 800 via communications path 826.

[0110] Additionally, computer system 800 may be a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smartwatch or other wearable, an appliance, part of the Internet of Things, and / or an embedded system, or any combination thereof, to name a few non-limiting examples.

[0111] The computer system 800 may be a client or server that accesses or hosts applications and / or data through a delivery paradigm, including, but not limited to, remote or distributed cloud computing solutions, local or on-premise software ("on-premise" cloud-based solutions), "as a service" models (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS)), and / or hybrid models including combinations of the foregoing examples or other service or delivery paradigms.

[0112] Applicable data structures, file formats, and schemas in computer system 800 may be derived from standards including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or other functionally similar representations, alone or in combination. Alternatively, proprietary data structures, formats, or schemas may be used, either exclusively or in combination with known or open standards.

[0113] In some embodiments, a tangible, non-transitory apparatus or article of manufacture including a tangible, non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This may include, but is not limited to, computer system 800, main memory 808, secondary memory 810, removable storage units 818 and 822, and tangible articles of manufacture embodying any combination of the foregoing. Such control logic, when executed by one or more data processing devices (such as computer system 800), may cause such data processing devices to operate as described herein.

[0114] It is understood that the Detailed Description section, and not the Abstract section, is intended to be used to interpret the claims. The Abstract section may describe one or more, but not all, example embodiments of the invention as contemplated by the inventors, and thus is not intended to limit the invention and the appended claims in any way.

[0115] The present invention has been described above with the aid of functional building blocks illustrating the implementation of certain functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the certain functions and relationships thereof are properly performed.

[0116] The foregoing description of specific embodiments fully discloses the general nature of the present invention, and by applying the knowledge of those skilled in the art, such specific embodiments can be readily modified and / or adapted for various uses without undue experimentation and without departing from the general concept of the present invention. Such adaptations and modifications are therefore intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. It should be understood that the phraseology or terminology used herein is for the purpose of description and not of limitation, and should therefore be interpreted in light of the teaching and guidance provided by those skilled in the art.

[0117] The breadth and scope of the present invention should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A data processing system for automating the Extract, Transform, and Load (ETL) process, A server coupled to a processor, which can perform the following instructions, namely, Extracting data from multiple sources, wherein the sources are selected from one or more of cloud storage, external APIs, and direct file uploads, and the extraction is performed by a stream-mode processing unit configured to segment the data into two or more manageable chunks. The structural integrity and content accuracy of the extracted data are verified using a data integrity algorithm, wherein the verification is defined by the data integrity algorithm based on at least one of formal consistency checks, anomaly detection, and data corruption identification. The extracted data is transformed via a transformation processing unit, wherein the transformation includes the application of predefined and custom transformation templates, and the transformation processing unit, This involves integrating user-defined transformation logic through an API, wherein the integration enables customization of data transformations. A system comprising a server configured to perform conversion and format the converted data for loading into a destination system.

2. The system according to claim 1, wherein the stream mode processing unit dynamically adjusts the size of data chunks based on the size of the data file and the system capacity.

3. The system according to claim 1, wherein the data integrity algorithm comprises an error correction mechanism configured to address data inconsistencies identified during the extraction process.

4. The instructions configure the processor to perform the no-code implementation of the automation, and the no-code implementation comprises instructions. The conversion processing unit utilizes an embedded programming language for defining custom conversion logic, and the custom conversion logic enables users to define custom extensions in a low-code environment and to define rules related to custom operations. The system according to claim 1, wherein the stream mode processing unit enables the user to manage data chunk size and processing parameters without programming expertise.

5. The system according to claim 1, further comprising a no-code user interface (UI) configured to allow a user to select and configure transformations from a set of pre-built templates without programming expertise.

6. The system according to claim 1, wherein the API for integrating user-defined conversion logic is compatible with a wide range of external programming environments.

7. The system according to claim 1, wherein the server comprises a data loading module configured to load the converted data into two or more different destination systems, the destination systems being selected from databases and data warehouses.

8. The system according to claim 1, wherein the server is further configured to execute commands for monitoring the data flow throughout the entire ETL process, and further comprises tracking progress and performing real-time error correction.

9. The system according to claim 1, further comprising a logging module configured to record one or more data change logs for storing changes to data, including a timestamp, the nature of the change, and a user identifier associated with manual changes, wherein the logging module enables comprehensive auditing and / or traceability of each change.

10. The system according to claim 9, wherein the server comprises an authorization management module configured to coordinate access to the data change log and ensure traceability control.

11. A computer implementation method, Identifying a data source for extraction, wherein the data source is selected from one or more of the following: cloud storage, external API, and direct file upload. Using a server coupled to the processor, the process of extracting data from the identified data source is performed, To ensure structural accuracy and content consistency, a data integrity check is performed on the extracted data using a data integrity algorithm. Configuring data transformations based on validated data, comprising selecting from pre-built transformation templates and defining custom transformations, Integrating custom transformation logic into the data transformation process via API, Applying the configured data transformation to the extracted data, To ensure compliance with predefined standards, the converted data is verified, A method comprising: formatting the verified and converted data for loading into a target system.

12. The method according to claim 11, wherein the data extraction process comprises stream mode processing, and the stream mode processing comprises dividing the data into dynamically adjusted, manageable chunks based on file size and / or system capacity.

13. The method according to claim 11, wherein the data integrity check comprises one or more error correction processes for addressing inconsistencies identified during extraction.

14. The method according to claim 11, wherein configuring the data transformation comprises implementing an embedded programming language for custom transformation logic.

15. The method according to claim 11, further comprising receiving input regarding the selection and / or configuration of one or more intended transformations from a pre-built template via a user interface.

16. The method according to claim 11, wherein integrating custom conversion logic via an API is further comprising integrating custom conversion logic from one or more external programming environments via the API.

17. The method according to claim 11, wherein the method comprises loading the converted data into two or more different destination systems selected from a database and a data warehouse.

18. The method according to claim 11, wherein the method comprises monitoring the data flow throughout the entire ETL process, tracking progress, and performing real-time error correction.

19. A non-temporary, tangible, computer-readable device that, when executed by a computing device, Identifying a data source for extraction, wherein the data source is selected from one or more of the following: cloud storage, external API, and direct file upload. The computer performs the data extraction process from the identified data source, To ensure structural accuracy and content consistency, a data integrity check is performed on the extracted data using a data integrity algorithm. Configuring data transformations based on the verified data, comprising selecting from pre-built transformation templates and defining custom transformations, Integrating custom transformation logic into the data transformation process via API, Applying the configured data transformation to the extracted data, To ensure compliance with predefined standards, the converted data is verified, A computer-readable device having stored instructions that causes the computing device to perform an operation comprising formatting the verified and converted data for loading into a target system.

20. The computer-readable device according to claim 19, wherein the data extraction process comprises performing stream-mode processing, the data transformation comprises utilizing an embedded programming language for defining custom transformation logic, the custom transformation logic enables a user to define custom extensions in a low-code environment and to define rules relating to custom operations, and the stream-mode processing comprises dividing the data into dynamically adjusted, manageable chunks based on file size and / or system capacity.