Systems and methods for data distribution
The data distribution system addresses scalability and flexibility issues by enabling customizable data persistence and computation across multiple platforms using extensible components, ensuring efficient and secure data management.
Patent Information
- Application Number
- JP2025520824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-11
- Filing Date
- 2023-10-09
- Publication Date
- 2025-10-28
AI Technical Summary
Existing data distribution systems are not scalable, agnostic to multiple platforms, and lack flexibility in modifying behaviors, leading to inefficiencies in data persistence and computation across different storage platforms.
A data distribution system that allows users to extend data persistence, transport, and distributed computation, enabling scalability and customization through extensible components like an interpreted script integrator (ISI) and an extensibility framework (ExFrame), which support multiple scripting languages and allow users to define heuristics and behaviors independently of the system's underlying code.
Enables optimized, scalable data persistence and computation across various platforms, allowing users to customize system behavior without sharing proprietary data or algorithms outside their domain, and supports 'plug-and-play' uncoupled components for efficient data management.
Smart Images

Figure 2025535746000001_ABST
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to the field of data distribution systems. [Background technology]
[0002] Several types of data delivery platforms exist for customers to store and retrieve their data. For example, third-party data persistence and distributed computing solutions are designed that require customers to acquire a known, remotely managed solution. Platforms such as Amazon Web Services™ (AWS™), Microsoft™ Azure™ Cloud, and Google Cloud Services (GCS) all adopt this model.
[0003] As another example, third-party data persistence and distributed computing solutions exist in the open source domain. However, these solutions do not attempt to agnosticize the adoption of any combination of third-party licensed solutions, such as AWS, Azure, and GCS. These solutions rely on internal or external networked computers known to the management system. For example, these systems are agnostic to data persistence, computation mechanisms, etc. Furthermore, these systems are not scalable to support any number of persistence and computation domain models, such as implementing both the AWS platform and the Azure platform simultaneously.
[0004] However, these solutions do not allow the ability to modify any number of previous modifications to any behavior using any of several public or proprietary scripting solutions. In other words, these solutions provide a single layer of modification that itself uses a single scripting solution.
[0005] What is needed is a system that enables customers to persist their data across storage platforms in any number of different locations and then perform co-located, distributed computations on this data so that data persistence, data transfer, and data computation can be done in an optimized, scalable manner and agnostic to scientific applications and instruments that may generate large amounts of data. Summary of the Invention
[0006]
[0009] In one aspect, the method may include receiving, by at least one worker node (WN) of a data distribution system from a job manager, a job definition including a set of processes to be executed for a set of data requests, by the at least one WN to an extended binary access manager (BAMEx), a request for the set of data, by the at least one WN to an extended binary access manager (BAMEx), receiving, by the at least one WN, the set of data based on the request for the set of data from the BAMEx, executing, by the at least one WN, one or more processes according to the job definition, generating, by the at least one WN, a set of result files including results of at least one of the executed one or more processes, and transmitting, by the at least one WN, the set of result files to the BAMEx for storage.
[0007] Additionally, each subcomponent or module of the described methods and systems is extensible by an interpreted script integrator (ISI), which can be customized using the proprietary scripting language of the extension framework (ExFrame), which can extend or modify any of the methods and systems described above. [Brief explanation of the drawings]
[0008] For the purpose of illustrating the invention, there is shown in the drawings a form which is presently preferred, it being understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown. [Figure 1A] 1 illustrates a system for data distribution according to the present disclosure. [Figure 1B] 1 illustrates a system for data distribution according to the present disclosure. [Figure 2] 1 illustrates a system for data distribution according to the present disclosure. [Figure 3] 1 illustrates a system for data distribution according to the present disclosure. [Figure 4A] 1 illustrates a system for data distribution according to the present disclosure. [Figure 4B] 1 illustrates a system for data distribution according to the present disclosure. [Figure 5] 1 illustrates a system for data distribution according to the present disclosure. [Figure 6] 1 illustrates a system for data distribution according to the present disclosure. [Figure 7] 1 illustrates a system for data distribution according to the present disclosure. [Figure 8] 1 illustrates a system for data distribution according to the present disclosure. [Figure 9] 1 illustrates a system for data distribution according to the present disclosure. [Figure 10] 1 illustrates a system for data distribution according to the present disclosure. [Figure 11] 1 illustrates a process for data distribution according to the present disclosure. [Figure 12] 1 illustrates a computing device for performing aspects of data distribution according to the present disclosure. [Figure 13] 1 illustrates a computing device for performing aspects of data distribution according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] The present disclosure may be understood more readily by reference to the following detailed description, taken in conjunction with the accompanying drawings and examples, which form a part of this disclosure. It should be understood that the present invention is not limited to the specific devices, methods, applications, conditions, or parameters described and / or illustrated herein, and that the terminology used herein is for the purpose of describing particular embodiments by way of example only, and is not intended to limit the claimed invention. Also, as used in the specification, including the appended claims, the singular forms "a," "an," and "the" include the plural, and reference to a particular value includes at least that particular value unless the context clearly dictates otherwise. As used herein, the term "plurality" means two or more. When ranges of values are expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another embodiment. It should be understood that all ranges are inclusive and combinable, and that steps may be performed in any order. All documents cited herein are incorporated by reference in their entirety for all purposes.
[0010] Furthermore, the phrase "based on," as used herein, unless otherwise specified, should be understood to mean "based at least in part on."
[0011] It should also be understood that certain features of the invention, which are, for clarity, described herein in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any subcombination. Furthermore, references to values stated in ranges include each and every value within that range. Furthermore, the term "comprising" has its standard open-ended meaning, but should also be understood to encompass "consisting of." For example, a device comprising parts A and B may include parts in addition to parts A and B, or may be formed solely from parts A and B.
[0012] For example, systems in which microscopy data is generated, such as in the field of flow cytometry, typically export the data generated within the system into a standard format (e.g., .fcs format) and then manually copy and / or move it to a new location for further analysis. As an example, this copying / moving of data may occur via a thumb drive from the computer that generated or initially received the data, which may then be entered into another storage container, such as another computer. However, when large amounts of data are generated, such as when a dataset includes images, traditional copying and storage of the generated data can be very difficult to accomplish (e.g., a user may be able to move a small subset of the data at a given time).
[0013] Furthermore, the internal IT security within a customer's system (e.g., a microscopy lab system) may be unknown to the data distribution system. Therefore, traditional storage systems, such as traditional cloud computing, are cumbersome to implement because integrating a third-party data distribution system also typically requires relaxing the security measures of the customer's system. There is a need for a data storage system that enables universally reachable, mass data persistence where users are agnostic to the data storage system.
[0014] Systems and methods for data distribution and persistence are described herein. By implementing a data distribution system, any user can extend data persistence, transport, and distributed computation so that the solution is fully scalable and authorable. Any number of initiating, contributing, or consuming client applications are agnostic to the internal persistence, colocation, or distribution of data. Thus, any number of client applications and / or servers can instantiate, contribute to, and / or consume persisted or generated data managed or generated by the system.
[0015] Additionally, any user can define their own heuristics and behaviors that can add to, replace, or remove existing system behavior. Users of the system can write scripts that add, derive, override, or remove behavior independent of the system's underlying code implementation. Any number of public or proprietary scripting languages may be employed by customers to modify the system's behavior so that it performs the custom work required by the customer.
[0016] The deployed system may also be unaware of any other systems or persistence mechanisms implemented in conjunction with the deployed system. Furthermore, proprietary data, algorithms, or configurations are not shared outside of the user's (e.g., customer's) domain. There is no obligation to connect the user's domain to any third-party servers, including commercial data distribution systems, to persist or capture data.
[0017] The systems described herein may include various uncoupled components that can operate in concert with one another (e.g., "plug and play" type systems). Thus, it may be possible to employ a single component without having to employ other components of the system. Each component of the disclosed system is provided below.
[0018] FIG. 1 illustrates a system 100 for data distribution according to the present disclosure. The system may include a device server (DS) 105. The DS may be a RESTful server application that sits on top of a web server (e.g., a lightweight HTTP(s) server). For example, LibWebSocket may be implemented as a web server. However, the DS may be an abstraction, and the underlying server may be agnostic to the end client application. Additionally, the DS 105 may be instantiated within an application to introduce the DS into the system. In some cases, communication to the DS 105 relies on the POST method in REST and may take the form of a JSON-formatted payload.
[0019] The DS 105 can provide API contract declaration, introspection, and validation for the system. This allows the DS to enforce client / server API contracts and also provide a mechanism for applications to retrieve this API and declarations. This allows API validation to be dynamically provided at runtime. The integration of a callback system that defines the work of an endpoint client can utilize the payload subcomponent to enforce the contract between the endpoint declaration (in the DS) and the application that defines the work of the endpoint (in the application). Furthermore, when the DS is connected on both sides of a microservices architecture, the DS can provide push notifications so that one server can natively communicate with another server via a message center (described below).
[0020] The system may also include a message center (MC) 135. The MC may be a common component for providing "push" technology to receiving subscribers. Subscribers may be agnostic to the described system. For example, a subscriber may subscribe to the system's JSMS (e.g., via a third-party tool) for specific messages. Within the described system, both the sending microserver and the receiving microserver may engage the DS as a listener, and the MC 135 may function as a complement to them. In some cases, the MC may be part of another component of the system. Thus, the MC 135 may provide some degree of communication optimization between these two related components. The MC 135 may also be implemented for any other component within the described system. The MC 135 may utilize a payload to bind an application to the MC. The MC 135 may convert the payload into a JSON-formatted POST request via REST.
[0021] The system also includes an interpreted script integrator. The ISI 125 may include an Interactive Script Integrator (ISI) 125. The ISI 125 may support an extension framework that integrates with many other components, such as BAMEx, DS, WN, and Job Scheduler Microservice (JSMS).
[0022] The ISI 125 can provide a cascading callback solution that relies on the declaration of ISI scopes and their dynamic injection into the ISI's scope stack, which can be a literal stack of callback resolution domains that declare callback scripts in an associated interpreted scripting language. (Note that the interpreted scripting language can include, for example, a strongly typed language (e.g., C#) bound to the ISI as an interpreted scripting language.)
[0023] Callback resolution within ISI 125 can bind keywords to script callbacks. ISI can use a scope stack to resolve script callbacks based on the cascading order of ranges pushed on the range stack. Payloads can be utilized to bind interpreted scripting languages to applications (e.g., via the ISI scope stack).
[0024] The ISI 125 can be a subcomponent of an application. The ISI 125 can be instantiated, but it is also possible to have multiple instances within the same application. This allows an application designer to dedicate an instance of the ISI 125 to perform some types of work, while another instance of the ISI 125 within the same application might be responsible for another type of work. For example, the ISI 125 might be instantiated in an application that itself instantiates the DS 105, which, by its nature, already instantiates an instance of the ISI 125. In this case, the application designer may choose to have a single ISI instance serve both the application and the DS. Or, conversely, the designer may choose to have two separate ISI instances manage each of the two use cases.
[0025] The system can also include a worker node base (WN) 110. The WN can establish a base worker so that it can execute processes on a dataset in one of two ways. First, work can be executed via the ISI 125 through the extension metaphor established by the ISI. Second, work can be executed as a derived worker node that natively extends a virtualized API in the base WN 110. This can be referred to as subclassing the base WN. A subclassed WN can also employ the first method and extend its functionality via the ISI extension framework.
[0026] A worker node can declare itself with, for example, a "name:address" and a "type." This information allows a Job Scheduler Microservice (JSMS) 115, which can support enqueuing many different WNs across any number of logical domains, or a calling client application, to submit jobs that can be consumed by worker nodes of a particular WN type. A collection of worker nodes of a particular type within the same logical domain can create a "worker farm" of the specified type within the specified domain. JSMS 115 is described later in this disclosure.
[0027] The WN 110 can include a DS and an MC, as described herein. The JSMS 115 can register with the MC of a worker node to receive events (e.g., via a JSON-formatted REST POST request). A job, represented as a payload, can be bound to a worker node through a native (C++) payload subcomponent. Once represented as a payload, the job is passed to a derived worker node (e.g., CAFWrapper, etc.) to perform the work (e.g., process) dictated by the definition of the derived worker node. Examples of work that a worker can perform can include, but are not limited to, clustering data, supervised or unsupervised learning or training with data, decomposing data into datasets, reconstructing data into larger datasets, performing computations on data (e.g., single-stage computations, multi-stage computations, distributed computations, etc.), and querying data for data locations via query parameters.
[0028] The system may also include a JSMS 115. The JSMS 115 may support enqueuing many different WNs across any number of logical domains. A logical domain may be a construct that allows persisted data to be associated (e.g., colocated) with a specific type of WN. In some cases, a domain may support any number of WN types. Any number of calling clients may enqueue any number of jobs that are assigned to available WNs of the required type within a specified job domain by the JSMS 115. Thus, data located within one domain may be processed on a worker farm of a specified type within that same domain. As an example, one or more jobs may be submitted to the JSMS 115 requesting image processing of data residing in an S3 bucket in a region of a cloud system (e.g., AWS). These jobs may then be assigned to an existing (or dynamically instantiated) worker farm. This allows WNs to process data within the same domain, thereby reducing the impact and monetary cost of data transmission.
[0029] JSMS 115 can also be responsible for centralizing and managing distributed work across any number of worker farms and across any number of logical domains. JSMS 115 can also be responsible for failure conditions such as dropped worker nodes or timed-out job executions.
[0030] JSMS 115 can also utilize DS and MC to manage and manipulate queued jobs. A job can be a payload passed to JSMS (e.g., as a JSON-formatted REST POST request). The underlying job object can be marshaled into a payload within JSMS and managed in maps and lists maintained by JSMS (e.g., as native C++ objects).
[0031] The system may also include a job marshaller (JM) 120. The JM 120 may be a "super object" that allows a single embedded DS to route messages from external clients (e.g., JSMS or native applications) to any number of embedded WNs. These embedded WNs may not contain embedded DSs or MCs, but may natively reference the JM's embedded DS. The JM 120 can easily fully saturate any particular machine (computer). Thus, all resources may be shared and fully utilized by the collection of sub-WNs contained in the JM 120.
[0032] The system can also include an extensibility framework (ExFrame) 130. ExFrame 130 can include a proprietary interpreted scripting solution that can provide a clean, agnostic mechanism for users to develop custom behavior by overriding nodes in a declarative node graph. An ExFrame node graph can be developed so that users can individually override nodes in the graph to implement their own desired behavior. This provides a declarative, verified software development kit (SDK) solution.
[0033] In essence, ExFrame 130 can provide an interpreter capable of registering declarative, verified node graphs. Any number of node graphs can be declared and, similar to traditional ISI scopes, these graphs can be bound to keyword callbacks. ISI 130 can then resolve these keywords to their scoped scripts. When such resolution identifies an ExFrame node graph, the ExFrame interpreter can execute on the node graph, at which point ISI 130 has resolved the invocation callback keywords. Execution of nodes in the graph can resolve against other callbacks also registered with ISI in other scopes on the scope stack. Furthermore, similar to ISI 130, any number of ExFrames can be instantiated within an application.
[0034] The system can include an Extended Binary Access Manager (BAMEx) 140. BAMEx 140 can centralize access to data as binary blobs so that the persistence location of the data is obfuscated from invoking client applications. BAMEx 140 can also manage the lifecycle and maintenance of the binary data. For example, BAMEx 140 can relocate binary blobs based on overridable heuristics (e.g., see ISI 125, described below), thereby periodically relocating less frequently accessed data further from consuming worker nodes (e.g., see WN, described below) and client applications.
[0035] BAMEx 140 can be an instantiated component that can be embedded in another application. BAMEx 140 can utilize an extensibility model that allows for various data store types (e.g., local file system, AWS S3, AWS EBS, AWS EFS, AWS Glacier, Azure Blobs, etc.). BAMEx 140 can manage these data stores and the data contained within each data store through a universally reachable database (UDB). Some behaviors of BAMEx 140 can be customized using the ISI override subsystem. For example, a backup strategy may be written that overrides the built-in backup heuristics provided by the BAMEx base.
[0036] The system may also include a payload structure (Payload) 145 (shown in FIG. 10). The Payload allows the described system to bind transport protocols and scripting languages to native code (e.g., strongly typed C++). This dynamic binding allows for agnostic abstraction of bound objects to components of the described system (e.g., DS, JSMS, WN, BAMEx, ExFrame, etc.).
[0037] Payload 145 can be used to seamlessly exchange information between components. Payload 145 can utilize a polymorphic inheritance architecture that separates first-class citizens (which can be freely added to the hierarchy) from the underlying serialization technology (e.g., HTTP(S), GRPC, TCP / IP, JSON, XML, CSV, etc.). By providing this abstraction, the component architecture and code design of the described system can consume payload objects but then forward these objects to other components and protocols that do not depend on the managing component code. Payloads can be native code linked to an application. Embedding system components or subcomponents can be linked in payload modules. Thus, if an application utilizes any of the system components or subcomponents, the system's payload is linked to those components and subcomponents.
[0038] The system also uses a Common Error Manager (CEM). The error management component may include a Error Manager (CEM) 150 (disclosed in FIG. 11). Error management within this collection of uncoupled components of the described system may be implemented in CEM 150. This component may be used in all or some of the components of the described system. CEM 150 may provide a mechanism to register error codes (e.g., integers) into error records. The error records may include a human-readable error string, an error description, and supplemental error information. CEM 150 may be a singleton instantiated upon initial invocation of the application.
[0039] Each component employing CEM150 can register a set of errors with a CEM150 instance. This registration system guarantees the uniqueness of both the error code and the error record. In this way, components of the system are free to register any errors they need, regardless of which other components have previously registered them. CEM150 solves the subtle problems caused by the "plug-and-play" concept of the system's decoupled component architecture.
[0040] The system may also include a Universal Logger (ULog) 155 (shown in FIG. 11). ULog 155 can work with CEM 150 to provide a consistent logging solution that natively consumes CEM error codes and error records. ULog can also generate consistent, formatted (e.g., JSON-based) log records that can be sorted, presented, and mined through a proprietary log viewer.
[0041] Components and subcomponents of the system can be set up or deployed in several ways. For example, components such as DS, MC, JSMS, WN, JM, etc. can be instantiated. In some cases, components can be deployed into the system via a singleton pattern within an application such as CEM, ISI, and ExFrame. In some cases, components are deployed via native module (e.g., C++) module inclusion for payloads, etc. In the case of instantiation, the instantiated objects (e.g., DS, MC, JSMS, WN, JM, etc.) can be managed at the native application level.
[0042] In particular, FIG. 1 illustrates a system 100 in which two clients interact with a universally reachable JSMS 115. The clients can enqueue jobs to the JSMS 115. The JSMS 115 can manage the execution of these jobs across multiple worker nodes (WNs) 110. The system 100 has two types of worker nodes 110: type A and type B. Similarly, a job requiring a WN 110 of type A is sent to an available WN 110 of type A. Similarly, a job requiring a WN 110 of type B is sent to an available WN 110 of type B. The system 100 also includes a JM 120, which can include both WNs 110 of type A and WNs of type B. The JSMS 115 can send jobs to the JM 120 when the associated WN type becomes available within this JM 120.
[0043] System 100 shows an exploded WN 110 of Type A. In this exploded view of WN 110, it can be seen that WN 110 uses custom scripts to extend its default behavior via ExFrame 130 and ISI 125. WN 110 also communicates with external processes developed independently of system 100.
[0044] System 100 also shows that derived work of WN 110 overrides behavior with custom scripts via ExFrame 130 and ISI 125. In addition, derived work instantiates a local BAMEx 140 instance (within the WN application) that interacts with a universally reachable SQL DB (UDB). Data is managed by BAMEx 140, which persists and retrieves data from two data stores located in two separate domains, Domain 1 and Domain 2.
[0045] Some components of the described system can be used in a completely cohesive manner (e.g., no coupling) within other components. For example, Payload can be used ubiquitously as a common commodity that binds native (e.g., C++) code to both external communication protocols (e.g., HTTP(s), TCP / IP, GRPC, etc.) and interpreted scripting languages (e.g., Python, R, ExFrame, JavaScript, C#, etc.). Similarly, CEM and ULog can be implemented across all components within a solution domain. However, these subcomponents can also be cohesively consumed by a managing component. Therefore, there can be no coupling of these components. Coupling that occurs within the system can include singleton instances of CEM and ULog.
[0046] 2 illustrates a system 200 for data distribution according to the present disclosure. The system 200 may include a DMS topology including a JSMS and a WN (e.g., without derived subclasses for unique Work). In this case, the WN implements ExFrame 130 to extend the WN's basic functionality to invoke existing external processes. This particular system setup demonstrates that ExFrame 130 and ISI 125 enable a WN to invoke complex legacy systems without code changes to the legacy systems.
[0047] 3 illustrates a system 300 for data distribution according to the present disclosure. System 300 illustrates a WN that has been subclassed (e.g., via C++) to extend functionality (e.g., perform work). Derived WNs can be plugged into a base JSSMS. For example, system 300 can be used to perform complex machine learning / artificial intelligence (ML / AI) image processing on large volumes of images.
[0048] 4 illustrates a system 400 for data distribution according to the present disclosure. System 400 can include a JM for utilizing and sharing resources on the same box (e.g., a single JM can optimally deploy any number of WNs and share resources such as GPUs, memory, file systems, etc.). This allows the JM to optimally utilize high-end boxes (e.g., having many CPU or central processing unit cores, high memory, high-end GPU(s), SSDs, etc.).
[0049] For example, a single box can deploy a WN utilizing a GPU and four CPU cores, along with four WNs utilizing four CPUs. All five WNs can optimally "share" memory and IO resources, thus reducing the need for inter-process communication or complex data serialization (http, gRPC, etc.).
[0050] The JM can also "emulate" a virtual WN. This concept allows transient processes to be deployed independently of the rest of the DMS-deployed topology. For example, any number of compute instances (e.g., AWS Lambda) can be instantiated, invoked, and / or destroyed by the JM. For example, the JMS simply "talks" to the JM, which delegates these commands to the transient compute instances.
[0051] 5 illustrates a system 500 for data distribution according to the present disclosure. System 500 includes, as a single instance, potentially many "types" of WNs and many JMs that include many WNs of many types. In system 500, WNs and JMs can optionally be spread across many "domains." JSMS can be "universally reachable" within the IT universe provided by end users.
[0052] FIG. 6 illustrates a system 600 for data distribution according to the present disclosure. The system 600 can include a single, universally reachable JSMS that communicates with any number of deployed and subclassed WNs. These derived WNs can be specifically coded to process batches of images (e.g., batches containing more than 100,000 images). In some cases, these derived WNs utilize a CLR bridge to connect the WN (e.g., C++-based) to the derived "work" (e.g., C#-based). The system 600 can include a single, universally reachable BAMEx that can be used for optimized data sharing across one or more domains. By way of example, the domains can include AWS, Azure, LAN, etc.
[0053] FIG. 7 illustrates a system 700 for data distribution according to the present disclosure. System 700 illustrates how ISI 125 can override any number of code callbacks in any number of application subsystems. As illustrated, ExFrame 130 is simply an additional declarative node graph (similar to an ISI script). Thus, ISI 125 can manage a scope stack of callback resolution domains that declare callback scripts using their associated interpreted scripting languages across various components 705-a through 705-c of system 700, such as, but not limited to, DS, WN, and JSMS.
[0054] FIG. 8 illustrates a system 800 for data distribution according to the present disclosure. It shows how a DS 105 can be added to an application (e.g., the application can become a server). The DS 105 includes an ISI 125, which allows endpoints (e.g., clients 1-3) to be declared and bound to ISI scripts AND / OR for the ISI 125, overriding the default native (C++) behavior so that the ISI 125 can resolve this behavior using its cascading scope stack. Finally, the diagram shows how the DS 105 natively (C++) calls back to application components 805-a, 805-b. The components 805-a, 805-b can be components of the system, such as WN, JSMS, JM, etc. While the DS 105 can communicate with the components 805-a, 805-b via native callbacks, communication between the client and the DS 105 can occur via RESTful POST requests. In a non-limiting example, the DS 105 receives a RESTful POST request from a client. The ISI 125 can check the endpoint script to determine whether an override script for the DS function corresponding to the RESTful POST request exists. The DS 105 can communicate with component 805-a (e.g., JSMS), 805-b (e.g., another DS), or both via native callback functions according to the RESTful POST request (e.g., which may include a payload) and the endpoint script.
[0055] Figure 9 illustrates a system 900 for data distribution according to the present disclosure. Figure 9 illustrates BAMEx 140 instantiated within an application. Two application components 905-a, 905-b (e.g., DS, JSMS, JM, WN, etc.) access BAMEx 140 to persist and retrieve binary objects. Data management is maintained in a universal database (UDB). At least two data stores containing binaries are established. BAMEx 140 includes ISI 125, which allows default BAMEx behavior to be overridden by ISI scripts.
[0056] Figure 10 illustrates a system 1000 for data distribution according to the present disclosure. Figure 10 illustrates three independent (decoupled) components 1005-a through 1005-c, with several applications / servers / libraries, etc., communicating through the dynamic nature of payloads. This configuration allows one component to send complex objects between components, and such data schemas can be loosely typed within the strongly typed paradigm of C++.
[0057] FIG. 11 illustrates a system 1100 for data distribution according to the present disclosure. FIG. 11 shows how CEM can be used as a stand-alone component within an application. System 1100 can include three application components 1105-a through 1105-c (e.g., DS, JM, JSMS, WN, etc.) that use CEM to manage their error codes in a manner that prevents error code enumeration conflicts. FIG. 11 also illustrates that CEM can log errors to an application logging component, which typically uses an injectable logging implementation. In this case, a universal logger (ULog) is shown.
[0058] 12 illustrates a process 1200 for data distribution according to the present disclosure. Steps 1205-1270 of process 1200 may be implemented by a data distribution system, such as the systems described in FIGS. 1-11, without any limitation.
[0059] In step 1205, the data can be transferred to BAMEx for storage accessible by the system. For example, a client application can transfer a file from local storage (e.g., a single board computer (SBC)) to BAMEx. BAMEx can then transfer the file to a local file system and link the data to a database (e.g., a unified database (UDB)). For example, heuristics maintained by the system can determine the location of a long-term persistent storage container (e.g., an AWS S3 bucket, AWS Glacier, etc.). In some cases, the system's ISI can allow a user (e.g., a customer) to override or modify the system's heuristics. In some cases, BAMEx can store files in a local cache in addition to, or instead of, storing files in the local file system.
[0060] In step 1210, the client application can construct a job definition payload. In some cases, the job definition payload can be constructed as a file format such as JSON. In some cases, the client application can include an experiment analysis application such as Attune. The job definition payload can be sent from the client application to JSMS (e.g., for enqueuing). The job definition payload can include, for example, a job ID, job metadata, and job payload. The job ID can be used for WN routing by JSMS. The job metadata can define specific requests for specific worker nodes of a specific type. Additionally, the payload can pass any necessary data to the worker nodes to perform the underlying job function.
[0061] In step 1215, the JSMS can determine one or more WNs to receive the job definition. For example, the JSMS can determine a list of idle WNs. The WNs can be idled from working on any job, or the WNs are not queued to work on a job. Additionally, the JSMS can determine the WN type. For example, the WNs can be configured to perform a particular job (e.g., a particular image processing function). The JSMS can determine one or more WNs to receive the job based on the WN's availability, WN type, etc. In some cases, this determination can also be based on the colocation of the job's data (e.g., the data's domain). In some cases, the heuristics used by the JSMS can be overridden by ISI / ExFrame. Thus, control over how the JSMS determines where to send the job definition and to which WNs and domains to send it can ultimately be managed by the user (e.g., customer) and the user's domain.
[0062] The JSMS may transmit the job definition to the determined one or more WNs in step 1220. The job definition may be transmitted via an application layer protocol such as HTTP.
[0063] In step 1225, the WN may receive the job definition and open the job definition. The job definition may include one or more job parameters associated with the job. For example, the job parameters may be captured in a job metadata component of the job definition. Additionally, associated job data may be captured in a job payload component of the job definition. The job definition may be declarative and may be validated via the DS.
[0064] In some cases, the behavior of a predefined WN type can be overridden by ISI / ExFrame. This allows users to change the behavior of a given type of WN. Furthermore, in some cases, ISI / ExFrame can override a WN type and declare the WN as a different type before the WN is registered with JSMS. In this way, a user can modify a WN type similar to what they want and register it as a different WN type written by the user in a scripting language. Furthermore, the scripting language may not be required to recompile the original source. Therefore, these types of data-driven changes to subcomponents can be made independently of the distribution system.
[0065] In step 1230, the selected WN can request data from BAMEx based on the job definition. BAMEx can further send a request (e.g., via TCP) to a database (e.g., UDB) to access the data based on the ID of the job definition (e.g., the ID of the data to be retrieved). In response, the database can send information to BAMEx on how to access the data. For example, the database can provide BAMEx information about the location of the data (e.g., a specific database, a specific domain, etc.). BAMEx can then fetch the data and relay it to the WN for the job. In some cases, BAMEx can relay the file path of the data to the WN if the data is cached in the local file system (as opposed to the data being stored externally). The WN can then fetch the data based on the file path. In some cases, the BAMEx cache behavior can be overridden in ISI / ExFrame, so that BAMEx can follow user preferences when relying on the cache.
[0066] In step 1235, the WN can initiate the job function provided in the job definition. For example, the WN's derived work can perform any work on the data as needed, but cannot modify the original data. It can copy the original data and create modified or derived data based on that data. In some cases, the job function can include feeding the data to a common analysis framework (CAF). For example, the data can be fed over a CLR bridge that can connect the WN's native language (e.g., C++) to the CAF library's language (e.g., C#). If CAF is implemented, the WN can also send the job parameters included in the job definition to CAF to execute the job.
[0067] In step 1240, the job can be executed by the WN. In the case of a CAF implementation, the CAF can perform data analysis. For example, if the data includes images, the CAF can perform image data analysis and generate a file (e.g., a Masks & Extension File) containing analysis results (e.g., cell morphology measurements). The images can be run through an AI / ML model to generate masks and extension attributes. The result file can then be sent to the WN. In some cases, the data processing can be further pre- and post-manipulated using ISI / ExFrame so that the underlying CAF and CLR bridge remain unchanged, but users can pre-process images and then post-process the masks and extension data before it is sent to BAMEx for persistence. In some cases, if the CAF is not involved, the WN can execute the job functions provided in the job definition, which enable the generation of the result file.
[0068] The WN may send the results file to BAMEx in step 1245. The results file may also include information that associates the results file with the data fetched for the job, identification information such as a job identification.
[0069] In step 1250, BAMEx can send the result file to storage (e.g., UDB and data store). In some cases, the storage can be selected by ExFrame / ISI within BAMEx.
[0070] In step 1255, the WN can notify the JS that the job is complete. In some cases, any errors encountered during job execution can also be sent to the JS. The WN can then set their status to "idle."
[0071] In step 1260, the client application may send a job status request to the JS. In some cases, the request may be sent over an application layer protocol such as HTTP. The JS may send a response to the JS indicating the job status (e.g., based on information provided by the WN). The response may also include information on how to retrieve the result file from BAMEx (e.g., identification of the result file).
[0072] In step 1265, the client application can send a request to BAMEx for the results file. The request can include identification information for the results file. BAMEx can request the results file from storage (e.g., by first identifying a storage location). The results file can be sent to BAMEx. BAMEx can cache the results file, which can then be sent to the client application.
[0073] In step 1270, the client application can process the result file for user viewing. For example, in a scenario where the result file includes a mask and an extension file, the client application can separate the mask file and the extension file and provide user viewing of the mask file and the extensions separate from each other. In some cases, ISI / ExFrame can also be implemented with the client application to provide controlled user extension disclosure.
[0074] In at least some embodiments, an entity implementing some or all of one or more of the techniques described herein may include a general-purpose computer system that includes or is configured to access one or more computer-accessible media. Figure 13 illustrates a general-purpose computer system that includes or is configured to access one or more computer-accessible media. The example computer system of Figure 7 may be configured to implement one or more of the service platforms of Figures 1-11, DS105, WN110, JSMS115, JM120, ISI125, ExFrame130, MC135, BAMEx140, or combinations thereof.
[0075] In the illustrated embodiment, computing device 1300 includes one or more processors 1310-a, 1310-b, and / or 1310-n (sometimes referred to herein as "processor 1310" in the singular or plural) coupled to system memory 1320 via an input / output (I / O) interface 1330. Computing device 1310 further includes a network interface 1340 coupled to I / O interface 1330.
[0076] In various embodiments, computing device 1300 may be a uniprocessor system including one processor 1310, or a multiprocessor system including multiple processors 1310 (e.g., two, four, eight, or another suitable number). Processor 1310 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1310 may be a general-purpose or embedded processor implementing any of a variety of instruction set architectures (ISAs), such as the x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multiprocessor system, each of processors 1310 may generally, but not necessarily, implement the same ISA.
[0077] System memory 1320 may be configured to store instructions and data accessible by processor(s) 1310. In various embodiments, system memory 1320 may be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / Flash type memory, or any other type of memory. In the illustrated embodiment, program instructions and data implementing one or more desired functions, such as the methods, techniques, and data described above, are shown stored in system memory 1320 as code 1325 and data 1326.
[0078] In one embodiment, I / O interface 1330 may be configured to coordinate I / O traffic between processor 1310, system memory 1320, and any peripherals within the device, including network interface 1340 or other peripheral interfaces. In some embodiments, I / O interface 1330 may perform any necessary protocol, timing, or other data conversion to convert data signals from one component (e.g., system memory 1320) into a format suitable for use by another component (e.g., processor 1310). In some embodiments, I / O interface 1330 may include support for devices connected through various types of peripheral buses, such as, for example, variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, the functionality of I / O interface 1330 may be split into two or more separate components, such as, for example, a northbridge and a southbridge. Additionally, in some embodiments, some or all of the functionality of I / O interface 1330 , such as the interface to system memory 1320 , may be incorporated directly into processor 1310 .
[0079] Network interface 1340 may be configured to enable data exchange between computing device 1300 and one or more other devices 1360 connected to one or more networks 1350, such as, for example, other computer systems or devices. In various embodiments, network interface 1340 may support communications over any suitable wired or wireless general-purpose data network, such as, for example, a type of Ethernet network. Furthermore, network interface 1340 may support communications over a telecommunications / telephone network, such as an analog voice network or a digital fiber communications network, over a storage area network, such as a Fibre Channel storage area network (SAN), or over any other suitable type of network and / or protocol.
[0080] In some embodiments, system memory 1320 may be one embodiment of a computer-accessible medium configured to store program instructions and data, such as those described above, for implementing corresponding method and apparatus embodiments. However, in other embodiments, program instructions and / or data may be received, sent, or stored on different types of computer-accessible media. Generally speaking, computer-accessible media may include non-transitory storage or memory media, such as magnetic or optical media, e.g., disks or DVDs / CDs, coupled to computing device 1300 via I / O interface 1330. Non-transitory computer-accessible storage media may also include any volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM (read-only memory), etc., that may be included in some embodiments of computing device 1300 as system memory 1320 or another type of memory. Furthermore, computer-accessible media may include transmission media or signals, such as electrical, electromagnetic, or digital signals, carried over communications media, such as networks and / or wireless links, such as those that may be implemented via network interface 1340. The functionality described in various embodiments may be implemented using some or all of multiple computing devices, such as those shown in FIG. 13. For example, software components executing on various different devices and servers may cooperate to provide functionality. In some embodiments, some of the described functionality may be implemented using storage devices, network devices, or special-purpose computer systems in addition to, or instead of, being implemented using a general-purpose computer system. The term "computing device," as used herein, refers to at least all of these types of devices, but is not limited to these types of devices.
[0081] Compute nodes, which may also be referred to as computing nodes, may be implemented on a wide variety of computing environments, such as commodity hardware computers, virtual machines, web servers, computing clusters, and computing appliances. Any of these computing devices or environments may be conveniently described as a compute node.
[0082] Each of the processes, methods, and algorithms described in the previous sections may be embodied in and fully or partially automated by code modules executed by one or more computers or computer processors. The code modules may be stored on any type of non-transitory computer-readable medium or computer storage device, such as a hard drive, solid-state memory, optical disk, etc. The processes and algorithms may be partially or wholly implemented in application-specific circuitry. The results of the disclosed processes and process steps may be stored, permanently or otherwise, in any type of non-transitory computer storage, such as, for example, volatile or non-volatile storage.
[0083] Illustrative Embodiments The following embodiments are illustrative only and do not limit the scope of the present disclosure as defined in the appended claims. It should be understood that any part of any one or more embodiments may be combined with any part of any other one or more embodiments.
[0084] Embodiment 1 A method for data distribution, comprising: receiving, by at least one of a plurality of worker nodes (WNs) of a data distribution system, from a job manager, a job definition including a set of processes to be executed for a set of data requests; sending, by the at least one WN, a request for the set of data to an extended binary access manager (BAMEx); receiving, by the at least one WN, the set of data from the BAMEx based on the request for the set of data; executing, by the at least one WN, one or more processes according to the job definition; generating, by the at least one WN, a set of result files including results of at least one of the executed one or more processes; and sending, by the at least one WN, the set of result files to the BAMEx for storage.
[0085] Embodiment 2 2. The method of embodiment 1, wherein the job manager is a job scheduler microservice (JSMS).
[0086] Embodiment 3 The method of embodiment 1 or 2, further comprising: receiving, by JSMS, a job definition from a client source external to the data delivery system, the job definition including a request for a set of processes to be executed on a set of data; and identifying, by JSMS, a WN of the data delivery system based on the job definition.
[0087] Embodiment 4 The method of any of embodiments 1 to 3, wherein the set of data is cached locally at BAMEx or stored externally to the data distribution system.
[0088] Embodiment 5 The method of any of embodiments 1 to 4, wherein the set of data includes a set of images.
[0089] Embodiment 6 The method of any of embodiments 1 to 5, further comprising: determining, by BAMEx, a location of the set of data outside the data distribution system; and sending, by BAMEx, a request for retrieval of the set of data to a storage location based on the determined location, wherein the storage location is outside the data distribution system.
[0090] Embodiment 7 7. The method of any of embodiments 1 to 6, further comprising determining, by the JSMS, a status for each of a plurality of WNs in the data distribution system, wherein the WNs are identified based on the status.
[0091] Embodiment 8 The method of any one of embodiments 1 to 7, wherein the status includes a busy status or an idle status.
[0092] Embodiment 9 The method of any of embodiments 1 to 8, further comprising determining, by the JSMS, a type for each of a plurality of WNs in the data distribution system, wherein the WNs are identified based on the type.
[0093] Embodiment 10 The method of any of embodiments 1 to 9, wherein the type includes a data analyzer WN or a data derivation WN.
[0094] Embodiment 11 The method of any one of embodiments 1 to 10, further comprising: storing, by BAMEx, the set of result files in a local cache; and transmitting, by BAMEx, the set of result files to storage external to the distribution system.
[0095] Embodiment 12 12. The method of any of embodiments 1 to 11, further comprising: receiving, by BAMEx, a set of data from a client application; determining a storage location for the set of data based on heuristics maintained by the data distribution system; and transmitting the set of data to the storage location, wherein the storage location is external to the data distribution system.
[0096] Embodiment 13 13. The method of any of embodiments 1-12, further comprising: sending, by the WN to the JSMS, a job definition to the WN and then a notification of the busy status of the WN.
[0097] Embodiment 14 14. The method of any of embodiments 1-13, wherein the WN is grouped with at least one other WN of the plurality of WNs to comprise a single entity from the perspective of the JSMS.
[0098] Embodiment 15 15. The method of any of embodiments 1-14, further comprising grouping, by a job marshaller (JM) of the data distribution system, the WN and at least one other WN to comprise a single entity.
[0099] Embodiment 16 The method of any one of embodiments 1 to 15, wherein the WN comprises a central processing unit (CPU).
[0100] Embodiment 17 17. The method of any of embodiments 1 to 16, further comprising: receiving, by BAMEx, a request for a set of result files from a client application; determining, by BAMEx, a storage location for the set of result files according to heuristics maintained by the data distribution system; retrieving, by BAMEx, the set of result files from the storage location; and sending, by BAMEx, the set of result files to the client application.
[0101] Embodiment 18 18. The method of any of embodiments 1 to 17, wherein the storage location for the set of data and the storage location for the set of result files are unknown to the client application.
[0102] Embodiment 19 19. The method of any of embodiments 1-18, wherein the client application refrains from communicating with the storage system for the set of data and the storage location for the set of result files.
[0103] Embodiment 20 20. The method of any of embodiments 1-19, further comprising transmitting the set of data by the WN to a Common Analysis Framework (CAF), wherein the execution of the one or more job functions is performed by the CAF.
[0104] Embodiment 21 21. The method of any of embodiments 1-20, wherein the job definition originates from a client source external to the data delivery system.
Claims
1. 1. A method for data distribution, comprising: receiving, by at least one of a plurality of worker nodes (WNs) of a data distribution system, from a job manager, a job definition including a set of processes to be executed for a set of data requests; sending a request for a set of data to an Extended Binary Access Manager (BAMEx) by said at least one WN; receiving, by the at least one WN, from the BAMEx, the set of data based on the request for the set of data; executing, by said at least one WN, one or more processes in accordance with said job definition; generating, by said at least one WN, a set of result files containing results of at least one of said one or more executed processes; transmitting, by said at least one WN, said set of result files to said BAMEx for storage.
2. The method of claim 1 , wherein the job manager is a Job Scheduler Microservice (JSMS).
3. receiving, via the JSMS, from a client source external to the data distribution system, the job definition including a request for a set of processes to be performed on a set of data; The method of claim 2 , further comprising: identifying, by the JSMS, the WN of the data delivery system based on the job definition.
4. The method of claim 1 , wherein the set of data is cached locally in the BAMEx or stored externally to the data distribution system.
5. The method of claim 1 , wherein the set of data comprises a set of images.
6. determining, by the BAMEx, a location of the set of data outside of the data distribution system; 2. The method of claim 1, further comprising: sending, by the BAMEx, a request for retrieval of the set of data to a storage location based on the determined location, the storage location being external to the data distribution system.
7. The method of claim 2 , further comprising determining, by the JSMS, a status for each of the plurality of WNs of the data distribution system, the WNs being identified based on the status.
8. The method of claim 7 , wherein the status comprises a busy status or an idle status.
9. The method of claim 2 , further comprising determining, by the JSMS, a type for each of the plurality of WNs of the data distribution system, the WNs being identified based on the type.
10. The method of claim 9 , wherein the type includes a data analyzer WN or a data derivative WN.
11. storing, by the BAMEx, the set of results files in a local cache; The method of claim 1 , further comprising: transmitting, via the BAMEx, the set of result files to storage external to the data distribution system.
12. receiving, by the BAMEx, the set of data from a client application; determining a storage location for the set of data based on heuristics maintained by the data distribution system; and The method of claim 1 , further comprising: transmitting the set of data to the storage location, the storage location being external to the data distribution system.
13. 3. The method of claim 2, further comprising sending, by the WN to the JSMS, a notification of a busy status of the WN after sending the job definition to the WN.
14. The method of claim 2 , wherein the WN is grouped with at least one other WN of the plurality of WNs to comprise a single entity from the perspective of the JSMS.
15. 15. The method of claim 14, further comprising grouping, by a job marshaller (JM) of the data distribution system, the WN and the at least one other WN to comprise the single entity.
16. The method of claim 1 , wherein the WN comprises a central processing unit (CPU).
17. receiving, by said BAMEx, a request from a client application for said set of results files; determining, by the BAMEx, storage locations for the set of resulting files according to heuristics maintained by the data distribution system; retrieving the set of result files from the storage location by the BAMEx; The method of claim 1 , further comprising: sending, via the BAMEx, the set of result files to the client application.
18. 18. The method of claim 17, wherein the storage location for the set of data and the storage location for the set of results files are unknown to the client application.
19. 20. The method of claim 18, wherein the client application refrains from communicating with a storage system for the set of data and the storage location for the set of result files.
20. The method of claim 1 , further comprising transmitting, by the WN, the set of data to a Common Analysis Framework (CAF), wherein the performance of the one or more job functions is performed by the CAF.
21. The method of claim 1 , wherein the job definition originates from a client source external to the data delivery system.