System and method for data distribution

By designing a multi-working node data distribution system, using an extended binary access manager and interpreted script integrator, the problem that existing systems cannot persist data across domains is solved, and the persistence, transmission and calculation of data is independently optimized, and a large amount of data generated by scientific applications and instruments is supported, and users can customize system behavior.

CN119998794APending Publication Date: 2025-05-13LIFE TECHNOLOGIES CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071134.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing data distribution systems cannot persist data across domains across different storage platforms and cannot independently optimize data persistence, transmission and computing, especially when scientific applications and instruments generate large amounts of data.

Method used

A data distribution system is designed to receive job definitions from the job manager through multiple work nodes, send data requests to the extended binary access manager, and execute the process according to the job definition, generate the result file and store it in the extended binary access manager. The system utilizes an interpreted script integrator and an extended framework that allows users to define and customize system behavior, supporting any number of public or proprietary scripting languages.

Benefits of technology

It realizes data persistence and distributed computing across different storage platforms, independently optimizes data persistence, transmission and computing, supports large amounts of data generated by scientific applications and instruments, and the system behavior can be customized according to user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998794A_ABST
    Figure CN119998794A_ABST
Patent Text Reader

Abstract

Methods and systems for data distribution are described herein. In one aspect, a method may include receiving, by at least one working node (WN 110) of a plurality of WNs of a data distribution system, a job definition from a job manager, where the job definition includes a set of processes to be performed on a request for a set of data; sending, by the WN 110, a request for data to an extended binary access manager (BAMEx) 140; receiving, by the WN 110, the data from the BAMEx 140 based on the request for the data; executing, by the WN 110, one or more processes according to the job definition; generating, by the WN 110, a set of result files, the set of result files comprising a result of at least one of the one or more executed processes; and sending, by the WN 110, the set of result files to the BAMEx 140.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed technology relates to the field of data distribution systems. Background Art

[0002] There are many types of data distribution platforms for customers to store and retrieve their data. For example, third-party data persistence and distributed computing solutions have been designed that require customers to license known and remotely managed solutions. TM (AWS TM ), Microsoft TM Azure TM Platforms such as Cloud and Google Cloud Service (GCS) adopt this model.

[0003] As another example, third-party data persistence and distributed computing solutions also exist in the open source domain. However, these solutions do not attempt to allow any combination of third-party licensed solutions (such as AWS, Azure, and GCS) in an indiscriminate manner. These solutions rely on internal or external networked computers known to the management and control system. For example, these systems are not independent of data persistence, computing mechanisms, etc. In addition, these systems cannot be extended to support any number of persistence and computing domain models. For example, implementing AWS and Azure platforms at the same time.

[0004] However, these solutions cannot modify any number of previous modifications of any behavior using any number of public or proprietary scripting solutions. In other words, these solutions provide a single modification layer that itself uses a single scripting solution.

[0005] What is needed is a system that allows customers to persist their data across any number of differently located storage platforms and then perform collocated distributed computations on the data, such that data persistence, data transfer, and data computation can be performed in an optimized and scalable manner independent of scientific applications and instruments that can generate large amounts of data. Summary of the invention

[0006] Methods and systems for data distribution are described herein. In one aspect, a method may include: receiving a job definition from a job manager by at least one of a plurality of work nodes (WNs) of a data distribution system, wherein the job definition includes a set of processes to be executed for a request for a set of data; sending a request for the set of data to an extended binary access manager (BAMEx) by at least one WN; receiving the set of data from BAMEx based on the request for the set of data by at least one WN; executing one or more processes according to the job definition by at least one WN; generating a set of result files by at least one WN, the set of result files including a result of at least one of the one or more processes executed; and sending the set of result files to BAMEx by at least one WN for storage.

[0007] In addition, each subcomponent or module of the described methods and systems can be extended through an Interpreted Script Integrator (ISI). The ISI can be customized using the Extensible Framework (ExFrame) proprietary scripting language. ExFrame can extend or modify any of the aforementioned methods and systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] For the purpose of illustrating the invention, there is shown in the drawings a form that is presently preferred; it being understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown.

[0009] Figures 1 to 10 A system for data distribution according to the present disclosure is depicted.

[0010] Fig.11 A process for data distribution according to the present disclosure is depicted.

[0011] Fig.12 A computing device for performing aspects of data distribution in accordance with the present disclosure is depicted. DETAILED DESCRIPTION

[0012] By referring to the following specific embodiments of the drawings and examples that form a part of this disclosure, the present disclosure can be more easily understood. It should be understood that the present invention is not limited to the specific devices, methods, applications, conditions or parameters described and / or shown herein, and the terms used herein are only for the purpose of describing specific embodiments by example, and are not intended to limit the claimed invention. In addition, as used in the specification including the attached claims, the singular forms "one", "one" and "the" include plural numbers, and unless the context clearly stipulates otherwise, the reference to a specific numerical value at least includes the specific value. As used in this article, the term "multiple" means more than one. When expressing a series of values, another embodiment includes from one specific value and / or to other specific values. Similarly, when a value is expressed as an approximate value by using the antecedent "about", it will be understood that a specific value forms another embodiment. All ranges are inclusive and combinable, and it should be understood that steps can be performed in any order. For any and all purposes, any document cited herein is incorporated by reference as a whole.

[0013] Further, as used herein, the phrase "based on" should be understood to mean "based, at least in part, on" unless otherwise specified.

[0014] It should be understood that certain features of the invention described herein in separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for the sake of brevity may also be provided individually or in any sub-combination. In addition, references to values ​​within a range include each value within that range. In addition, the term "comprising" should be understood to have its standard, open-ended meaning, but also encompasses "consisting of". For example, a device comprising part A and part B may include parts other than part A and part B, but may also be formed by only part A and part B.

[0015] For example, systems that generate microscopy data in fields such as flow cytometry typically export the data generated in the system to a standard format (e.g., .fcs format) and then manually copy and / or move it to a new location for additional analysis. As one example, such copying / moving of data may occur via a thumb drive from a computer that generated or initially received the data, which may then be entered into another storage container (such as another computer). However, where large amounts of data are generated, such as where the data set includes images, conventional copying and storage of the generated data is very difficult to implement (e.g., a user may only be able to move a small portion of the data at a given time).

[0016] Additionally, the internal IT security within the customer's system (e.g., a microscope lab system) may not be known to the data distribution system. Therefore, conventional storage systems (such as conventional cloud computing) are cumbersome to implement because the integration of a third-party data distribution system typically also requires the relaxation of the security measures of the customer's system. A data storage system is needed that allows universally accessible persistence of large amounts of data without the user having to be aware of the data storage system.

[0017] This article describes systems and methods for data distribution and persistence. By implementing a data distribution system, any user can extend data persistence, transmission and distributed computing, making the solution fully scalable and authorizable. Any number of initiating, contributing or consuming client applications are independent of the internal persistence, collocation or distribution of data. Therefore, any number of client applications and / or servers can instantiate, contribute and / or consume persistent or generated data managed or generated by the system.

[0018] In addition, any user can define their own heuristics and behaviors, which can augment, replace, or remove existing system behaviors. Users of the system can write scripts that augment, derive, override, or remove behaviors that are unrelated to the underlying code implementation of the system. Customers can employ any number of public or proprietary scripting languages ​​to modify the behavior of the system so that the system's behavior performs the custom work required by the customer.

[0019] The deployed system may also be unaware of any other system or persistence mechanism implemented in combination with the deployed system. Furthermore, no proprietary data, algorithms, or configurations are shared outside of the user (e.g., customer) domain. There is no obligation to connect the user domain to any third-party server to persist or capture data, including commercial data distribution systems.

[0020] The systems described herein may include a variety of uncoupled components that can work in conjunction with each other (eg, a "plug and play" type system). Thus, a single component may be employed without employing other components of the system. Each component of the disclosed system is provided below.

[0021] Figure 1 A system 100 for data distribution according to the present disclosure is depicted. The system may include a device server (DS) 105. The DS may be a RESTful server application located on top of a web server (e.g., a lightweight HTTP server). For example, LibWebSockets may be implemented as a web server. However, the DS may be an abstract concept, and the underlying server may not need to be aware of the final client application. In addition, the DS 105 may be instantiated within an application to introduce the DS into the system. In some cases, communication with the DS 105 relies on a POST method within REST and may take the form of a JSON formatted payload.

[0022] DS 105 can provide API contract declaration, introspection and validation for the system. This can allow DS to enforce client / server API contracts and can further provide a mechanism for applications to retrieve the API and declaration, which can dynamically provide API validation at runtime. The integration of a callback system that defines the work of the endpoint client can utilize the payload subcomponent to enforce the contract between the endpoint declaration (within DS) and the application that defines the work of the endpoint (within the application). In addition, in the case where DS is attached on both sides of the microservice architecture, DS can provide push notifications so that one server can talk to another server in a native way via the message center (described below).

[0023] The system may also include a message center (MC) 135. MC can be a public component that provides "push" technology to receiving subscribers. Subscribers may not need to know the described system; for example, a subscriber may subscribe (e.g., via a third-party tool) to the system's JSMS to obtain specific messages. Within the described system, both the sending and receiving microservices may view the DS as a listener, and the MC 135 135 may complement them. In some cases, the MC may be part of other components of the system. Therefore, the MC 135 may provide a certain amount of communication optimization between these two related components. The MC 135 may also be implemented for any other component within the described system. The MC 135 may bind an application to the MC using a payload. The MC 135 may convert the payload into a JSON formatted POST request via REST.

[0024] The system may also include an interpreted script integrator (ISI) 125. The ISI 125 may support an extension framework in many components that are integrated into other components, such as BAMEx, DS, WN, and Job Scheduler Microservices (JSMS).

[0025] ISI 125 may provide a cascading callback solution that relies on the declaration and dynamic injection of ISI scopes onto an ISI scope stack. The scope stack may be a literal stack of callback resolution domains for callback scripts declared in an associated interpreted scripting language. (It should be noted that the interpreted scripting language may include, for example, a strictly typed language (e.g., C#) that is bound to the ISI as an interpreted scripting language).

[0026] Callback resolution within the ISI 125 may bind keywords to script callbacks. The ISI may use a scope stack to resolve script callbacks based on the cascading nature of the order of scopes pushed onto the scope stack. The payload may be used to bind an interpreted scripting language to an application (e.g., via the ISI scope stack).

[0027] ISI 125 can be a subcomponent of an application. ISI 125 can be instantiated, but it can also have multiple instances within the same application. This allows the application designer to set aside an instance of ISI 125 to perform certain types of work, while another instance of ISI 125, also in the same application, may be responsible for another type of work. For example, ISI 125 can be instantiated with an application that itself instantiates DS105. DS105 has essentially instantiated an instance of ISI 125. In this case, the application designer can choose to have a single ISI instance serve both the application and the DS. Or conversely, the designer can choose to have two separate ISI instances govern each of the two use cases.

[0028] The system may also include a Worker Node Library (WN) 110. The WN may establish a base Worker Node from which processing may be performed on a data set in one of two ways. First, work may be performed via the ISI 125 through an extended metaphor established by the ISI. Second, work may be performed as a derived Worker Node that extends a virtualization API within the base WN 110 in a native manner. This may be referred to as subclassing the base WN. The subclassed WN may also extend its functionality using the first approach and via the ISI extension framework.

[0029] A worker node may declare itself, for example, with a "name:address" and a "type". This information may allow a job scheduler microservice (JSMS) 115 to issue jobs that may be consumed by worker nodes of a particular WN type, which may support queuing or invoking client applications on many different WNs on any number of logical domains. A collection of worker nodes of a particular type within the same logical domain may create a "worker node cluster" of a specified type within a specified domain. JSMS 115 will be described later in this disclosure.

[0030] WN 110 may include a DS and MC as described herein. JSMS 115 may register with the MC of a worker node to receive events (e.g., via JSON formatted REST POST requests). Jobs represented as payloads may be bound to worker nodes via native (C++) payload subcomponents. Once represented as payloads, jobs are passed to derived worker nodes to perform work (e.g., processes) indicated by the definition of the derived worker node (e.g., CAFWrapper, etc.). Examples of work that a worker node may perform include, but is not limited to: cluster analysis of data; supervised or unsupervised learning or training based on the data; decomposing data into data sets; reorganizing data into larger data sets; computations performed on data (e.g., single-level computations, multi-level computations, distributed computations, etc.); and data queries on data locations based on query parameters.

[0031] The system may also include JSMS115. JSMS115 may support queuing many different WNs on any number of logical domains. A logical domain may be a construct that allows persistent data to be associated (e.g., co-located) with a particular type of WN. In some cases, a domain may support any number of WN types. Any number of calling clients may queue any number of jobs assigned by JSMS115 to available WNs of the necessary type within a specified job domain. Therefore, data located within one domain may be processed for a cluster of worker nodes of a specified type within the same domain. As an example, one or more jobs may be issued to JSMS115 requesting image processing of data within an S3 bucket residing in a region of a cloud system (e.g., AWS). These jobs may then be assigned to a pre-existing (or dynamically instantiated) cluster of worker nodes. This allows the WNs to process data within the same domain, thereby reducing data transfer impact and monetary costs.

[0032] JSMS 115 can also be responsible for centralizing and managing distributed work across any number of worker node clusters and across any number of logical domains. JSMS 115 can also be responsible for failure conditions, such as worker nodes going offline or job execution timeouts.

[0033] JSMS 115 can also manage and manipulate queued jobs using DS and MC. A job can be a payload passed to JSMS (e.g., as a JSON formatted REST POST request). The underlying job object can be marshaled into the payload within JSMS and can be managed within maps and lists maintained by JSMS (e.g., as native C++ objects).

[0034] The system may also include a job marshaler (JM) 120. The JM 120 may be a "super object" that allows a single embedded DS to route messages from an external client (e.g., JSMS or native application) to any number of embedded WNs. These embedded WNs may not include an embedded DS or MC, and are referred to as JM embedded DSs in terms of native. The JM 120 may facilitate full saturation of any particular machine (computer). Therefore, all resources may be shared and fully utilized by the collection of child WNs that the JM 120 includes.

[0035] The system may also include an extension framework (ExFrame) 130. ExFrame 130 may include a proprietary interpreted scripting solution that may provide a clean and versatile mechanism for users to develop custom behaviors by overriding nodes within a declarative node graph. An ExFrame node graph may be developed whereby nodes within the graph may be individually overridden by users to perform their own desired behaviors. This provides a declarative and proven software development kit (SDK) solution.

[0036] In short, ExFrame 130 can provide an interpreter that can register declarative and verified node graphs. Any number of node graphs can be declared, and like traditional ISI scopes, these graphs can be bound to keyword callbacks. ISI 130 can then parse these keywords into their scope scripts. When such parsing identifies an ExFrame node graph, the ExFrame interpreter can be executed on the node graph, and the ISI 130 has parsed the call callback keywords. The execution of the nodes in the graph can be parsed for other callbacks of ISIs in other scopes that are also registered on the scope stack. In addition, similar to ISI 130, any number of ExFrames can be instantiated within an application.

[0037] The system may include an extended binary access manager (BAMEx) 140. BAMEx 140 may centrally access data in the form of binary chunks so that the persistent location of the data is obfuscated and cannot be discovered by calling client applications. BAMEx 140 may also manage the lifecycle and maintenance of binary data. For example, BAMEx 140 may relocate binary chunks based on overridable heuristics (see, e.g., ISI 125 described below) so that less frequently accessed data may be periodically relocated away from consuming worker nodes (see, e.g., WN described below) and client applications.

[0038] BAMEx 140 can be an instantiated component that can be embedded in another application. BAMEx 140 can utilize an extension model that allows for various data store types (e.g., local file system, AWS S3, AWS EBS, AWS EFS, AWS Glacier, Azure Blob, etc.). BAMEx 140 can manage these data stores and the data included within the corresponding data stores via a universally reachable database (UDB). Some behaviors of BAMEx 140 can be customized using the ISI override subsystem. For example, a backup policy can be written to override the built-in backup heuristics provided by the BAMEx library.

[0039] The system may also include a payload configuration (payload) 145 (in Fig.10 The payload allows the described system to bind transport protocols and scripting languages ​​to native code (e.g., strictly typed C++). This dynamic binding can abstract the binding object to the components of the described system (e.g., DS, JSMS, WN, BAMEx, ExFrame, etc.) in an undifferentiated manner.

[0040] Payload 145 can be used to seamlessly exchange information between components. Payload 145 can utilize a polymorphic inheritance architecture that separates first-class citizens (which can be added to the hierarchy at will) from underlying serialization technologies (e.g., HTTP, GRPC, TCP / IP, JSON, XML, CSV, etc.). By providing this abstraction, the component architecture and code design of the described system can consume payload objects, but then these objects can be transmitted to other components and protocols that are unrelated to the management component code. The payload can be native code linked to the application. Embedded system components or subcomponents can be linked in the payload module. Therefore, if an application utilizes any of the system components or subcomponents, the system's payload is linked to these components and subcomponents.

[0041] The system may also include a common error manager (CEM) 150 (in Fig.11 ). Error management within this uncoupled set of components of the described system may be implemented in CEM 150. This component may be used for all or some of the components of the described system. CEM 150 may provide a mechanism to register an error code (e.g., an integer) to an error record. The error record may include a human-readable error string, an error description, and supplemental error information. CEM 150 may be a singleton that is instantiated upon initial invocation of an application.

[0042] Each component that employs CEM 150 can register a list of errors with a CEM 150 instance. This registration system guarantees uniqueness of error codes and error records. In this way, components of the system are free to record their own necessary errors, regardless of what other components have previously recorded. CEM 150 solves subtle problems caused by the "plug and play" concept of the system's uncoupled component architecture.

[0043] The system may also include a universal logger (ULog) 155 (in Fig.11 ). ULog 155 can work with CEM 150 to provide a consistent logging solution that consumes CEM error codes and error records in a native manner. ULog can also generate consistent formatted (e.g., JSON-based) log records that can be sorted, presented, and mined through a proprietary log viewer.

[0044] The components and subcomponents of the system can be set up or introduced in a variety of ways. For example, components can be instantiated, such as for DS, MC, JSMS, WN, JM, WN, JM, etc. In some cases, such as for CEM, ISI, and ExFrame, components can be introduced into the system via a singleton pattern within the application. In some cases, such as for payloads, components are introduced via native module (e.g., C++) module includes. For instantiation, instantiated objects (e.g., DS, MC, JSMS, WN, JM, etc.) can be managed at the native application level.

[0045] In particular, Figure 1 A system 100 is depicted in which two clients interact with a generally reachable JSMS 115. The clients can queue jobs for the JSMS 115. The JSMS 115 can manage the execution of these jobs on multiple work nodes (WNs) 110. In the system 100, there are two flavors of work nodes 110, namely type A and type B. Similarly, jobs that require a WN 110 of type A are sent to available WNs 110 of type A. Similarly, jobs that require a WN 110 of type B are sent to available WNs 110 of type B. The system 100 also includes a JM 120, which can include WNs 110 of type A and WNs of type B. The JSMS 115 can send jobs to the JM 120 when the associated WN type becomes available within the JM 120.

[0046] Depicted in system 100 is a decomposed WN 110 of type A. In this decomposed view of WN 110, it can be seen that WN 110 extends its default behavior using custom scripts via ExFrame 130 and ISI 125. WN 110 also communicates with external processes that were developed without regard to system 100.

[0047] The system 100 also shows that the derived work of WN 110 overrides the behavior with custom scripts via ExFrame 130 and ISI 125. In addition, the derived work instantiates a local BAMEx 140 instance (within the WN application) that interacts with a universally reachable SQL DB (UDB). Data is managed by BAMEx 140 to persist and retrieve data from two data stores located in two separate domains (Domain 1 and Domain 2).

[0048] Certain components of the described system can be used in other components in a completely cohesive manner (e.g., without coupling). For example, payloads can be widely used as a general commodity that binds native (e.g., C++) code to both external communication protocols (e.g., HTTP, TCP / IP, GRPC, etc.) and interpreted scripting languages ​​​​(e.g., Python, R, ExFrame, JavaScript, C#, etc.). Similarly, CEM and ULog can be implemented in all components within the solution domain. However, these subcomponents can also be cohesively absorbed by the management and control components. Therefore, there may be no coupling of these components. The coupling that does occur within the system may include singleton instantiations of CEM and ULog.

[0049] Figure 2 A system 200 for data distribution according to the present disclosure is depicted. The system 200 may include a DMS topology that includes a JSMS and a WN (e.g., for unique work, no derived subclasses). In this case, the WN is implementing an ExFrame 130 to extend the basic functionality of the WN to call pre-existing external processes. This particular system setup demonstrates that with the help of the ExFrame 130 and the ISI 125, the WN can call complex legacy systems without code modifications to the legacy systems.

[0050] Figure 3 A system 300 for data distribution according to the present disclosure is depicted. System 300 shows that WNs are subclassed (e.g., via C++) to extend functionality (e.g., perform work). Derived WNs can be attached to a base JSMS. For example, system 300 can be used to perform complex machine learning / artificial intelligence (ML / AI) image processing on large numbers of images.

[0051] Figure 4 A system 400 for data distribution according to the present disclosure is depicted. The system 400 may include JMs to utilize and share resources on the same box (e.g., a single JM may optimally deploy any number of WNs and share resources such as GPUs, memory, file systems, etc.). This allows high-end boxes (e.g., with a large number of CPU or central processing unit cores, high memory, high-end GPUs, SSDs, etc.) to be optimally utilized by the JM.

[0052] For example, a single box can deploy one WN utilizing a GPU and 4 CPU cores and 4 WNs utilizing 4 CPUs. All 5 WNs can then optimally "share" memory and IO resources. Thus, inter-process communication or complex data serialization (http, gRPC, etc.) can be alleviated.

[0053] The JM can also "emulate" a virtual WN. This concept allows for the deployment of transient processes independent of the rest of the DMS deployment topology. For example, the JM can instantiate, invoke, and / or destroy any number of compute instances (e.g., AWS Lambda). For example, JSMS only "talks" to the JM which delegates these commands to the transient compute instances.

[0054] Figure 5 A system 500 for data distribution according to the present disclosure is depicted. The system 500 includes WNs of potentially many "types" as a single instance and many JMs including many WNs of many types. In the system 500, the WNs and JMs may optionally be deployed on many "domains". The JSMS may be "universally reachable" within the IT domain provided by the end user.

[0055] Figure 6 A system 600 for data distribution according to the present disclosure is depicted. The system 600 may include a single, universally reachable JSMS that communicates with any number of deployed subclassed WNs. These derived WNs may be specifically coded to handle batches of images (e.g., where a batch includes greater than 100K images). In some cases, these derived WNs utilize a CLR bridge to connect the WN (e.g., C++ based) to derived "jobs" (e.g., C# based). The system 600 may include a single, universally reachable BAMEx that may be used for optimized data sharing across one or more domains. As an example, domains may include: AWS, Azure, LAN, etc.

[0056] Figure 7 Depicted is a system 700 for data distribution according to the present disclosure. System 700 shows how ISI 125 can cover any number of code callbacks within any number of application subsystems. As shown, ExFrame 130 is simply an additional declarative node graph (similar to an ISI script). Therefore, ISI 125 can manage a scope stack of callback resolution domains that declare callback scripts in an associated interpreted scripting language across various components 705-a-705-c (such as but not limited to DS, WN, JSMS, etc.) of system 700.

[0057] Figure 8 A system 800 for data distribution according to the present disclosure is depicted. Figure 8It shows how DS105 can be added to an application (e.g., this can make the application a server). DS105 includes ISI125, which can allow endpoints (e.g., clients 1-3) to be declared and bound to ISI scripts, and / or allow ISI 125 to override the behavior of default native (C++) operations so that ISI 125 uses its cascading scope stack to resolve the behavior. Finally, the figure shows how DS105 calls back to application components 805-a, 805-b in a native manner (C++), where components 805-a, 805-b can be components of the system, such as WN, JSMS, JM, etc. DS105 can communicate with components 805-a, 805-b via native callbacks, and communication between the client and DS105 can be carried out via RESTful POST requests. In a non-limiting example, DS105 receives a RESTful POST request from a client. ISI125 can check the endpoint script to determine whether there is an overriding script for the DS function corresponding to the RESTful POST request. DS 105 may communicate with component 805 - a (eg, JSMS), 805 - b (eg, another DS), or both via native callback functions based on a RESTful POST request (eg, which may include a payload) and an endpoint script.

[0058] Fig. 9 A system 900 for data distribution in accordance with the present disclosure is depicted. Fig. 9 An example BAMEx 140 within an application is shown. Two application components 905-a, 905-b (e.g., DS, JSMS, JM, WN, etc.) access BAMEx 140 to persist and retrieve binary objects. Management of the data is maintained within a universal database (UDB). At least two data stores including binary values ​​have been established. BAMEx 140 includes an ISI 125 that allows default BAMEx behavior to be overridden by ISI scripts.

[0059] Fig.10 A system 1000 for data distribution according to the present disclosure is depicted. Fig.10 Shown are three independent (decoupled) components 1005-a-1005-c with some application / server / library / etc communicating through the dynamic nature of the payload. This configuration allows one component to send complex objects between components, where such data patterns can be loosely typed within the C++ strongly typed paradigm.

[0060] Fig.11 Depicted is a system 1100 for data distribution in accordance with the present disclosure. Fig.11It shows how CEM can be used as a standalone component within an application.System 1100 may include three application components 1105-a-1105-c (eg, DS, JM, JSMS, WN, etc.) that use CEM to manage their error codes in a manner that prevents error code enumeration conflicts. Fig.11 It is also shown that the CEM can log errors to an application logging component, which is typically implemented using an injectable logging implementation. In this case, a Universal Logger (ULog) is shown.

[0061] Fig.12 A process 1200 for data distribution according to the present disclosure is depicted. Steps 1205-1270 of process 1200 may be implemented by a data distribution system such as Figures 1 to 11 The system described in (a) is implemented.

[0062] At step 1205, data may be transferred to BAMEx for storage in a storage device accessible by the system. For example, a client application may transfer a file from a local storage device (e.g., a single board computer (SBC)) to BAMEx. BAMEx may then transfer the file to a local file system and link the data to a database (e.g., a unified database UDB). For example, heuristics maintained by the system may determine the location of a long-term persistent storage container (e.g., AWS S3 bucket, AWS Glacier, etc.). In some cases, the system's ISI may allow a user (e.g., a customer) to override or modify the system's heuristics. In some cases, in addition to or instead of storing the file in the local file system, BAMEx may also store the file in a local cache.

[0063] At step 1210, the client application may construct a job definition payload. In some cases, the job definition payload may be constructed as a file format, such as JSON. In some cases, the client application may include an experimental analysis application, such as Attune. The job definition payload may be sent from the client application to the JSMS (e.g., for queuing). The job definition payload may include, for example, a job ID, job metadata, and a job payload. The job ID may be used for WN routing performed by the JSMS. The job metadata may define a specific request for a specific work node of a specific type. In addition, the payload may pass any necessary data to the work node to perform the underlying job function.

[0064] In step 1215, the JSMS may determine one or more WNs for receiving the job definition. For example, the JSMS may determine a list of idle WNs. The WN may be idle without performing any job, or may not be queued for a job. In addition, the JSMS may determine the WN type. For example, the WN may be configured to perform a specific job (a specific image processing function, etc.). The JSMS may determine one or more WNs for receiving the job based on the availability of the WN, the WN type, etc. In some cases, the determination may also be made based on the juxtaposition of the data for the job (e.g., the domain of the data). In some cases, the heuristics used by the JSMS may be overridden by ISI / ExFrame. Therefore, control over how the JSMS decides where to send the job definition and to which WN and domain it is sent may ultimately be managed by the user (e.g., the customer) and the user's domain.

[0065] The JSMS may send the job definition to the determined one or more WNs at step 1220. The job definition may be sent via an application layer protocol, such as HTTP.

[0066] In step 1225, the WN may receive a job definition and open the job definition. The job definition may include one or more job parameters associated with the job. For example, the job parameters may be captured in a job metadata component of the job definition. In addition, the associated job data may be captured in a job payload component of the job definition. The job definition may be declarative and may be validated via the DS.

[0067] In some cases, the behavior of predefined WN types can be overridden by ISI / ExFrame. This allows the user to change the behavior of a certain type of WN. In addition, in some cases, ISI / ExFrame can override the WN type and declare the WN as a different type before the WN is registered with JSMS. In this way, a WN type similar to what the user expects can be modified and registered as a different WN type that the user has written in a scripting language. In addition, a scripting language may not be required to recompile the original source. Therefore, these types of data-driven modifications to subcomponents can be made independently of the distributed system.

[0068] At step 1230, the selected WN may request data from BAMEx based on the job definition. Based on the identification of the job definition (e.g., identification of the data to be retrieved), BAMEx may further send a request to access the data to a database (e.g., UDB) (e.g., via TCP). In response, the database may send BAMEx information about how to access the data. For example, the database may provide BAMEx information about the location of the data (e.g., a specific database, a specific domain, etc.). BAMEx may then retrieve the data and relay the data to the WN for the job. In some cases, if the data is cached within the local file system (as opposed to the case where the data is stored externally), BAMEx may relay the file path of the data to the WN. The WN may then retrieve the data based on the file path. In some cases, the behavior of the BAMEx cache may be overridden by the ISI / ExFrame so that when BAMEx relies on the cache may be based on user preferences.

[0069] In step 1235, the WN may initiate a job function as provided in the job definition. For example, a derivative of the WN may perform any work on the data if necessary. However, the original data cannot be changed. The original data may be copied and modified, or derived data may be created based on the data. In some cases, the job function may include feeding data into a common analysis framework (CAF). The data feed may occur, for example, over a CLR bridge that may connect the native language of the WN (e.g., C++) to the language of the CAF library (e.g., C#). In the case of implementing CAF, the WN may also send the job parameters included in the job definition to the CAF to execute the job.

[0070] At step 1240, the job may be executed by the WN. In the case of a CAF specific implementation, the CAF may perform data analysis. For example, where the data includes images, the CAF may perform image data analysis and generate files (e.g., mask and extension files) including the results of the analysis (e.g., cell morphology measurements). The image may be run through an AI / ML model to generate mask and extension attributes. The result file may then be sent to the WN. In some cases, ISI / ExFrame may be used to further pre- and post-manipulate the digital processing so that the underlying CAF and CLR bridge remain unchanged, but the user may pre-process the image and then post-process the mask and extension data before sending it to BAMEx for persistent storage. In some cases where CAF is not involved, the WN may execute the job function as provided in the job definition, which may generate a result file.

[0071] The WN may send the result file to BAMEx at step 1245. The result file may also include identifying information, such as information associating the result file with the data retrieved for the job, the job identification, and the like.

[0072] BAMEx may send the result file to a storage device (eg, UDB and data storage) at step 1250. In some cases, the storage device may be selected by ExFrame / ISI within BAMEx.

[0073] In step 1255, the WN may notify the JS of the job completion. In some cases, any errors encountered during job execution may also be sent to the JS. The WNs may then set their status to "idle".

[0074] At step 1260, the client application may send a job status request to the JS. In some cases, the request may be sent via an application layer protocol such as HTTP. The JS may send a response to the JS indicating the job status (e.g., based on information provided by the WN). The response may also include information on how to retrieve the result file from BAMEx (e.g., identification information of the result file).

[0075] At step 1265, the client application may send a request for the result file to BAMEx. The request may include identification information of the result file. BAMEx may request the result file from a storage device (e.g., by first identifying the storage location). The result file may be sent to BAMEx, which may cache the result file and then send it to the client application.

[0076] At step 1270, the client application may process the result file for viewing by the user. For example, in a scenario where the result file includes a mask and an extension file, the client application separates the mask file from the extension file and provides the user with the mask file and the extension separated from each other for viewing. In some cases, ISI / ExFrame may also be implemented with the client application to provide controlled user exposure to the extension.

[0077] In at least some embodiments, an entity implementing part or all of one or more of the techniques described herein may include a general-purpose computer system that includes or is configured to access one or more computer-accessible media. Fig.13 A general purpose computer system including or configured to access one or more computer-accessible media is depicted. Figure 7 An example computer system may be configured to implement Figures 1 to 11 One or more of the service platform, DS105, WN 110, JSMS115, JM 120, ISI 125, ExFrame 130, MC 135, BAMEx 140, or a combination thereof.

[0078] In the illustrated embodiment, computing device 1300 includes one or more processors 1310-a, 1310-b, and / or 1310-n (which may be referred to herein as "a processor 1310" in the singular or as "processors 1310" in the plural) coupled to system memory 1320 via an input / output (I / O) interface 1330. Computing device 1310 also includes a network interface 1340 coupled to I / O interface 1330.

[0079] In various embodiments, computing device 1300 may be a uniprocessor system including one processor 1310 or a multiprocessor system including multiple processors 1310 (e.g., two, four, eight, or another suitable number). Processor 1310 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1310 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as an x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multiprocessor system, each of processors 1310 may typically (but not necessarily) implement the same ISA.

[0080] System memory 1320 may be configured to store instructions and data accessible by processor 1310. In various embodiments, system memory 1320 may be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / Type memory or any other type of memory. In the illustrated embodiment, program instructions and data that implement one or more desired functions (such as those methods, techniques, and data described above) are shown as stored in system memory 1320 as code 1325 and data 1326.

[0081] In one embodiment, the I / O interface 1330 may be configured to coordinate I / O traffic between the processor 1310, the system memory 1320, and any peripheral devices in the device, including the network interface 1340 or other peripheral interfaces. In some embodiments, the I / O interface 1330 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., the system memory 1320) into a format suitable for use by another component (e.g., the processor 1310). In some embodiments, the I / O interface 1330 may include support for devices attached through various types of peripheral buses (such as, for example, a variant of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard). In some embodiments, the functionality of the I / O interface 1330 may be divided into two or more separate components, such as, for example, a north bridge and a south bridge. In addition, in some embodiments, some or all of the functionality of the I / O interface 1330 (such as an interface to the system memory 1320) may be directly incorporated into the processor 1310.

[0082] The network interface 1340 may be configured to allow data to be exchanged between the computing device 1300 and one or more other devices 1360 (such as, for example, other computer systems or devices) attached to one or more networks 1350. In various embodiments, the network interface 1340 may support communication via any suitable wired or wireless general-purpose data network (such as, for example, an Ethernet type). In addition, the network interface 1340 may support communication via a telecommunications / telephone network (such as an analog voice network or a digital fiber optic communication network), via a storage area network (such as a Fibre Channel SAN (Storage Area Network)), or via any other suitable type of network and / or protocol.

[0083] In some embodiments, system memory 1320 can be an embodiment of a computer-accessible medium that is configured to store program instructions and data for implementing embodiments of the corresponding methods and devices as described above. However, in other embodiments, program instructions and / or data may be received, sent, or stored on different types of computer-accessible media. In general, computer-accessible media may include non-transitory storage media or memory media, such as magnetic or optical media, such as a disk or DVD / CD coupled to computing device 1300 via I / O interface 1330. Non-transitory computer-accessible storage media may also include any volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM (read-only memory), etc., which may be included in some embodiments of computing device 1300 as system memory 1320 or another type of memory. In addition, computer-accessible media may include transmission media or signals, such as electrical signals, electromagnetic signals, or digital signals carried via communication media such as a network and / or wireless links (such as those that may be implemented via network interface 1340). Multiple computing devices (such as Fig.13 Some or all of the computing devices (including those shown in the figure) can be used to implement the functions described in various embodiments; for example, software components running on various different devices and servers can cooperate to provide functions. In some embodiments, in addition to or instead of using a general-purpose computer system to implement, a storage device, a network device, or a special-purpose computer system can be used to implement part of the described functions. The term "computing device" as used herein refers to at least all of these types of devices and is not limited to these types of devices.

[0084] A compute node (which may also be referred to as a computing node) may be implemented on a variety of computing environments, such as commodity hardware computers, virtual machines, web services, computing clusters, and computing appliances. For convenience, any of these computing devices or environments may be described as a computing node.

[0085] Each of the processes, methods, and algorithms described in the foregoing sections may be embodied in code modules executed by one or more computers or computer processors and fully or partially automated by these code modules. The code modules may be stored on any type of non-transitory computer-readable medium or computer storage device, such as a hard drive, solid-state memory, optical disk, etc. The processes and algorithms may be implemented in part or in whole in dedicated circuits. The results of the disclosed processes and process steps may be persistently or otherwise stored in any type of non-transitory computer storage device, such as, for example, a volatile or non-volatile storage device.

[0086] Exemplary embodiments

[0087] The following embodiments are exemplary only and are not intended to limit the scope of the present disclosure or the appended claims. It should be understood that any part of any one or more embodiments may be combined with any part of any other one or more embodiments.

[0088] Implementation 1

[0089] A method for data distribution, the method comprising: receiving a job definition from a job manager by at least one of a plurality of working nodes (WNs) of a data distribution system, wherein the job definition comprises a set of processes to be executed for a request for a set of data; sending a request for the set of data to an extended binary access manager (BAMEx) by the at least one WN; receiving the set of data from the BAMEx based on the request for the set of data by the at least one WN; executing one or more processes according to the job definition by the at least one WN; generating a set of result files by the at least one WN, the set of result files comprising a result of at least one of the one or more processes executed; and sending the set of result files to the BAMEx for storage by the at least one WN.

[0090] Implementation 2

[0091] A method according to embodiment 1, wherein the job manager is a job scheduler microservice (JSMS).

[0092] Implementation 3

[0093] According to the method described in any one of embodiments 1 to 2, the method also includes: the JSMS receiving the job definition from a client source outside the data distribution system, the job definition including a request for a set of processes to be executed on a set of data; and the JSMS identifying the WN of the data distribution system based on the job definition.

[0094] Implementation 4

[0095] A method according to any one of embodiments 1 to 3, wherein the set of data is cached locally at the BAMEx, or stored externally to the data distribution system.

[0096] Implementation 5

[0097] A method according to any one of embodiments 1 to 4, wherein the set of data includes a set of images.

[0098] Implementation 6

[0099] According to the method described in any one of embodiments 1 to 5, the method also includes: determining, by the BAMEx, a location of the set of data outside the data distribution system; and sending, by the BAMEx, a request to retrieve the set of data to a storage location based on the determined location, wherein the storage location is outside the data distribution system.

[0100] Implementation Plan 7

[0101] According to the method described in any one of embodiments 1 to 6, the method further includes: determining, by the JSMS, a status of each of the multiple WNs of the data distribution system, wherein the WN is identified based on the status.

[0102] Implementation 8

[0103] A method according to any one of embodiments 1 to 7, wherein the state includes a busy state or an idle state.

[0104] Implementation Plan 9

[0105] According to the method described in any one of embodiments 1 to 8, the method further includes: determining, by the JSMS, a type of each of the multiple WNs of the data distribution system, wherein the WN is identified based on the type.

[0106] Implementation 10

[0107] A method according to any one of embodiments 1 to 9, wherein the type includes a data analyzer WN or a data derivation WN.

[0108] Implementation Plan 11

[0109] According to the method described in any one of embodiments 1 to 10, the method further includes: storing, by the BAMEx, the set of result files in a local cache; and sending, by the BAMEx, the set of result files to a storage device outside the distribution system.

[0110] Implementation Plan 12

[0111] According to the method described in any one of embodiments 1 to 11, the method also includes: receiving the set of data from the client application by the BAMEx; determining the storage location of the set of data based on the heuristic maintained by the data distribution system; and sending the set of data to the storage location, wherein the storage location is external to the data distribution system.

[0112] Implementation 13

[0113] According to the method described in any one of embodiments 1 to 12, the method further includes: after sending the job definition to the WN, the WN sends a notification of the busy status of the WN to the JSMS.

[0114] Implementation Plan 14

[0115] A method according to any one of embodiments 1 to 13, wherein the WN is grouped with at least one other WN among the multiple WNs to constitute a single entity as observed by the JSMS.

[0116] Implementation Plan 15

[0117] According to the method described in any one of embodiments 1 to 14, the method further includes: grouping the WN and the at least one other WN by a job marshaller (JM) of the data distribution system to form the single entity.

[0118] Implementation Plan 16

[0119] A method according to any one of embodiments 1 to 15, wherein the WN includes a central processing unit (CPU).

[0120] Implementation Plan 17

[0121] According to the method described in any one of embodiments 1 to 16, the method also includes: receiving, by the BAMEx, a request for the set of result files from the client application; determining, by the BAMEx, a storage location of the set of result files based on a heuristic maintained by the data distribution system; retrieving, by the BAMEx, the set of result files from the storage location; and sending, by the BAMEx, the set of result files to the client application.

[0122] Implementation Plan 18

[0123] A method according to any one of embodiments 1 to 17, wherein the storage location of the set of data and the storage location of the set of result files are unknown to the client application.

[0124] Implementation Plan 19

[0125] A method according to any one of embodiments 1 to 18, wherein the client application avoids communicating with the storage system of the set of data and the storage location of the set of result files.

[0126] Implementation Plan 20

[0127] According to the method described in any one of embodiments 1 to 19, the method further includes: sending the set of data by the WN to a common analysis framework (CAF); wherein executing one or more job functions is performed by the CAF.

[0128] Implementation Plan 21

[0129] A method according to any one of embodiments 1 to 20, wherein the job definition originates from a client source external to the data distribution system.

Claims

1. A method for data distribution, the method comprising: receiving, by at least one of a plurality of worker nodes (WNs) of a data distribution system, a job definition from a job manager, wherein the job definition includes a set of processes to be performed on a request for a set of data; Sending, by the at least one WN, a request for the set of data to an extended binary access manager (BAMEx); receiving, by the at least one WN, the set of data from the BAMEx based on the request for the set of data; executing, by the at least one WN, one or more processes according to the job definition; generating, by the at least one WN, a set of result files including results of at least one of the one or more processes executed; as well as The set of result files is sent by the at least one WN to the BAMEx for storage.

2. The method of claim 1, wherein the job manager is a Job Scheduler Micro Service (JSMS).

3. The method according to claim 1, further comprising: receiving, by the JSMS, from a client source external to the data distribution system, the job definition comprising a request for a set of processes to be performed on a set of data; as well as The WN of the data distribution system is identified by the JSMS based on the job definition.

4. The method of claim 1, wherein the set of data is cached locally at the BAMEx, or stored externally to the data distribution system. The method of claim 1 , wherein the set of data comprises a set of images.

6. The method according to claim 1, further comprising: determining, by the BAMEx, a location of the set of data external to the data distribution system; as well as A request to retrieve the set of data is sent by the BAMEx to a storage location based on the determined location, wherein the storage location is external to the data distribution system.

7. The method according to claim 1, further comprising: A status of each of the plurality of WNs of the data distribution system is determined by the JSMS, wherein the WN is identified based on the status. The method according to claim 7 , wherein the state comprises a busy state or an idle state.

9. The method according to claim 1, further comprising: A type of each of the plurality of WNs of the data distribution system is determined by the JSMS, wherein the WN is identified based on the type.

10. The method of claim 9, wherein the type comprises a data analyzer WN or a data derivation WN.

11. The method according to claim 1, further comprising: storing, by said BAMEx, said set of result files in a local cache; as well as The set of result files is sent by the BAMEx to a storage device external to the distribution system.

12. The method according to claim 1, further comprising: receiving, by said BAMEx, said set of data from a client application; determining a storage location for the set of data based on a heuristic maintained by the data distribution system; as well as The set of data is sent to the storage location, wherein the storage location is external to the data distribution system.

13. The method according to claim 1, further comprising: After sending the job definition to the WN, the WN sends a notification of the busy status of the WN to the JSMS.

14. The method of claim 1, wherein the WN is grouped with at least one other WN of the plurality of WNs to constitute a single entity as observed by the JSMS.

15. The method according to claim 14, further comprising: The WN and the at least one other WN are grouped by a job marshaler (JM) of the data distribution system to form the single entity.

16. The method of claim 1, wherein the WN comprises a central processing unit (CPU).

17. The method according to claim 1, further comprising: receiving, by the BAMEx, a request for the set of result files from the client application; determining, by said BAMEx, a storage location for said set of result files based on a heuristic maintained by said data distribution system; retrieving, by said BAMEx, said set of result files from said storage location; as well as The set of result files is sent by the BAMEx to the client application.

18. The method of claim 1, wherein the storage location of the set of data and the storage location of the set of result files are unknown to the client application.

19. The method of claim 1, wherein the client application avoids communicating with a storage system of the set of data and the storage location of the set of result files.

20. The method according to claim 1, further comprising: Sending the set of data to a common analysis framework (CAF) by the WN; The execution of one or more job functions is performed by the CAF.

21. The method of claim 1, wherein the job definition originates from a client source external to the data distribution system.